UK IT leaders are still buying training-grade GPU clusters for a workload that has already moved on. The UK government's own compute annex puts total UK AI compute demand at 1.8 GW in 2024, rising to 9.6 GW by 2035 in its main scenario — and, critically, says Gen-AI inference could account for up to 62% of that future demand. That's not a training story. It's an agents-and-chatbots-running-all-day story: copilots answering tickets, agents processing invoices, chat interfaces serving customers, all day, every day. Yet procurement teams keep specifying frontier-scale training rigs "just in case," when the UK demand curve is inference-shaped. This piece sets out what that means for buying inference server hardware in 2026 — and why getting the category wrong is an expensive mistake.
View the data behind this chart
| 2024 baseline | 2035 main scenario | 2035 high-demand scenario | |
|---|---|---|---|
| UK AI compute demand | GW1.8 | GW9.6 | GW13.6 |
The Market Has Already Voted for Inference
The starting point for any 2026 hardware decision is scale, and the UK government's own numbers are stark. Its compute evidence annex puts UK AI compute demand at 1.8 GW in 2024, rising to 9.6 GW by 2035 in the main scenario — and as high as 13.6 GW in a high-demand scenario that the annex attributes to McKinsey, a 7.5x increase on the 2024 baseline. That is a lot of new capacity to plan for, whichever scenario lands.
What matters more for a hardware buyer is the shape of that demand, not just its size. The same annex states that Gen-AI inference could account for up to 62% of future UK AI compute demand by 2035, and explicitly flags inference as expected to overtake training as the dominant driver of UK AI compute in some scenarios. In other words: the growth is coming from running models, not building them.

Inference vs Training: Different Jobs, Different Machines
Training and inference are different engineering problems wearing the same 'AI hardware' label. Training a foundation model is a rare, batch-oriented, throughput-maximising exercise typically run on a small number of large, tightly-interconnected clusters. Inference is the opposite: it is continuous, latency-sensitive, and runs every time an employee opens a copilot, an agent processes a document, or a customer starts a chat session. Buy for the wrong one and you either bottleneck your users or bankroll idle capacity.
The spend data already tells this story. Soldo's research found UK business AI spending rose 449% in Q1 2024 versus the same quarter a year earlier, with 63% of that expenditure going to ChatGPT — a hosted, inference-only product, not an in-house training run. Meanwhile Helium 42's 2026 benchmark found a striking adoption gap: only 28% of UK organisations meet the Department for Science, Innovation and Technology's definition of strategic AI deployment, yet 78% report broader AI-tool usage. That gap is inference happening informally, via chat tools and APIs, well ahead of any formal infrastructure plan — exactly the workload pattern that on-prem inference hardware needs to be sized against.
Beyond FLOPS: The Metrics That Actually Decide Inference Buys
Peak FLOPS is the wrong headline number for an inference buy. What decides whether an inference server hardware purchase pays off is a cluster of practical metrics: VRAM capacity (can the model and its KV cache fit at all), memory bandwidth (how fast tokens get generated once it does fit), latency under concurrent load, power draw per unit of throughput, and — the number that actually ties back to budget — cost per token, per query, or per served user.
That last framing matters specifically for UK buyers because it's the one procurement and finance teams can actually hold a supplier to. A GPU estate that wins on raw specification sheets can still lose badly on cost-per-served-user once concurrency, batching efficiency and idle time are factored in. Before specifying hardware, model the actual query volume and concurrency you expect — a tool built to calculate your GPU server needs exists for exactly this reason, and it forces the conversation onto workload shape rather than GPU count.
The 2026 Hardware Categories: Matching Silicon to the Job
For inference specifically, 2026 buyers are choosing between three broad hardware categories: general-purpose GPU accelerators, purpose-built inference ASICs, and CPU-based inference for smaller or less latency-critical workloads. Each trades flexibility against specialisation. GPU accelerators run almost any model architecture and benefit from the widest software support, which is why they remain the default choice for firms serving multiple model types or sizes. Purpose-built ASICs can offer a better cost-per-token at scale for a fixed, high-volume workload — but that specialisation cuts both ways, since swapping model architectures later is harder. CPU-based inference remains a legitimate option for smaller models, batch workloads, and edge deployments where GPU procurement and power overheads aren't justified.
The comparison below sets out the trade-off buyers actually face, category by category, rather than model by model — because the category decision (flexible vs specialised vs simple) has to happen before any specific vendor conversation. For most UK enterprises running a mix of internal copilots, customer-facing agents and document processing, general-purpose high-performance GPU accelerators remain the safer starting point; ASIC specialisation only pays off once a single workload is running at genuinely high, predictable volume.
Architecture: From Single Box to Distributed Serving Estate
Architecture decisions compound the hardware choice. A single well-specified inference node can serve a surprising amount of enterprise traffic, but agentic workloads — where one user request triggers multiple chained model calls — push firms toward multi-node serving with proper orchestration well before raw GPU count would suggest it's necessary. Interconnect bandwidth between nodes, not just within them, starts to matter once a workload spans more than one server.
This is also where UK-specific planning constraints bite. The government's own demand curve — from 1.8 GW to a 9.6 GW or 13.6 GW UK-wide figure by 2035 — is a proxy for how much extra power and cooling capacity data centres and enterprise server rooms will need to absorb over the same period. Buyers specifying inference server hardware today should assume rack power density keeps rising and plan cooling and electrical headroom accordingly, rather than sizing for today's draw only.
Software Stack: The Multiplier Most Buyers Ignore
Software determines how much of any given piece of hardware you actually get to use. Inference engines and serving stacks — vLLM, TGI, llama.cpp, TensorRT-LLM, OpenVINO and ROCm among them — each optimise for different hardware and workload combinations, and achievable throughput on identical silicon can vary enormously depending on which stack sits on top of it. A hardware decision made without a matching software decision is only half a decision.
Practically, this means UK buyers should treat the software stack as part of procurement, not an afterthought bolted on post-delivery. Ask any supplier which serving frameworks are validated on the proposed hardware, what the upgrade path looks like as models change, and whether the stack supports the orchestration tooling your own platform team already runs.
View the data behind this chart
| Layer | Detail |
|---|---|
| Gen-AI inference workloads | Up to 62% of UK AI compute demand by 2035 |
| Training and other AI compute | The remaining share of 2035 UK AI compute demand |
TCO and the UK Budget Reality
Budgets are moving fast enough that a mismatched purchase compounds quickly. Helium 42's 2026 benchmark put average annual UK AI spend at £15.94 million across organisations of all sizes, with large enterprises reporting £45 million to £78 million annually, and 85% to 91% of UK organisations increasing AI budgets year-on-year. Separately, Helium 42 sized the UK AI market's total addressable spend at £18.2 billion in 2025, rising to £31.4 billion by 2028; IMARC Group put the broader UK AI market at USD 4.0 billion in 2025, projected to reach USD 23.1 billion by 2034 at a 20.79% CAGR. Whichever estimate a buyer anchors to, the direction is the same: spend keeps growing, so a hardware category mismatch made now gets more expensive to unwind each year it's left in place.
For most UK firms, that argues for comparing the full cost of running inference on-prem against buying it as a managed service before committing capital — particularly while workload volumes and concurrency patterns are still settling. It's worth working through that comparison properly, including power, cooling, and refresh cycles, rather than defaulting to either extreme; a structured look at how to compare self-hosting vs cloud GPU costs is the right next step before signing a purchase order.
Data Residency, UK GDPR and the ICO's Expectations
Hardware choice for inference in the UK isn't purely a performance-and-cost decision — it also carries a governance dimension. Whichever category of inference server hardware a buyer lands on, procurement teams need to budget for electricity, rack density, and cooling constraints alongside aligning the purchase to UK data-residency and governance requirements under the UK GDPR and the ICO's AI and data-protection expectations. That governance layer sits on top of the workload-sizing exercise, not apart from it: an on-premise inference estate keeps the question of where inference data physically sits and how it's processed directly within the buyer's own control, which is one more reason the on-prem-versus-managed-service comparison above needs to weigh residency and compliance obligations, not just raw pounds-per-token.
In practice, this means data-residency and ICO-facing governance requirements should be scoped alongside the hardware specification — not bolted on afterwards — so that whichever inference architecture is chosen, the compliance conversation and the hardware conversation are resolved together.
Choosing Your Inference Hardware: The Practical Recommendation
The practical decision rule for 2026 is straightforward: size for the inference and agentic workload you can actually forecast over the next 12–24 months, not for a hypothetical future training run. Start from expected query volume, concurrency, and model size; choose the hardware category (GPU, ASIC, or CPU) that matches that workload's flexibility needs; validate the software stack alongside the hardware; and build in power and cooling headroom given how fast UK compute demand is projected to grow.
Because AI budgets are rising across the board — and because getting the category wrong is expensive to reverse — it's worth structuring the purchase properly rather than treating it as a one-off capital spend. That's true whether the answer ends up being an on-prem inference estate or a hybrid arrangement; either way, it's worth reviewing how to finance your AI server purchase so the commercial structure matches the workload-led sizing, rather than the other way round.
Sources
Every figure in this article traces to the sources below.
- •UK Government — compute evidence annex (UK AI compute demand scenarios, inference share)
- •IT Brief — Soldo research on UK business AI spending growth
- •Helium 42 — 2026 UK AI adoption and spend benchmark
- •IMARC Group — UK artificial intelligence market size and forecast
View the data behind this chart
| Best Fit Workloa… | Flexibility | Key Trade-off | |
|---|---|---|---|
| GPU accelerators | LLM & multimodal serving | High - wide support | Higher power draw |
| Inference-specific… | Fixed high-volume model | Low - narrow support | Best cost per token |
| CPU-based inference | Small models, batch | High - runs anything | Lower throughput |
