UK’s trusted IT infrastructure partner since 2003
Servnet
FinanceToolsConfiguratorGet in Touch
Networking

Right-Sizing Your Network for On-Prem AI in 2026

Servnet Editorial · IT infrastructure analysis8 min read
Share

GPU budgets get the headlines, but repatriated AI projects stall or starve for a quieter reason: the east-west network wasn't sized for the traffic pattern. Synchronous training collectives need a 1:1 non-blocking fabric, while inference can tolerate 2:1 to 3:1 oversubscription — treat them the same and you either overspend badly or leave GPUs idle waiting on the wire. In the UK, that network decision now has to be made alongside a grid-connection queue that hit 125 GW by mid-2025, up from 41 GW barely six months earlier. This is a buyer's analysis of what to actually provision, and in what order.

Ethernet Rate By Fabric Tier, 2026
1600120080040008-128 GPU cluster…ToR-to-spine (800G…Dominant 2026…Early AI-scale…Fabric tierLine rate (Gb/s)Typical line rate
View the data behind this chart
Ethernet Rate By Fabric Tier, 2026
Line rate (Gb/s)8-128 GPU cluster…ToR-to-spine (800G…Dominant 2026…Early AI-scale…
Typical line rate4008008001600

Why Traditional Data-Centre Networks Break Under Repatriated AI

Most enterprise networks were designed for north-south traffic: clients talking to a handful of servers, with predictable, bursty demand that tolerates a reasonably oversubscribed access-to-core ratio. Repatriating AI training onto that same fabric is where the plan usually falls apart, because the dominant traffic pattern flips entirely.

Synchronous pre-training and large fine-tuning runs are driven by all-reduce and all-gather collectives — every GPU in a job periodically needs to exchange gradients with every other GPU, simultaneously, on a tight schedule. That pattern needs a 1:1 non-blocking east-west fabric; any blocking point turns into a synchronisation barrier that every GPU in the job waits on, regardless of how much compute you bought.

This is the core reason on-prem AI refreshes disappoint: teams size the GPUs correctly and then treat the switch fabric as an afterthought inherited from the last server refresh. The network isn't a supporting cast member here — for training workloads specifically, it's the schedule.

Illustration: Right-Sizing Your Network for On-Prem AI in 2026

Bandwidth And Latency By Workload: Training, Inference And RAG

Not every AI workload needs the same fabric, and provisioning them identically wastes either capacity or performance. The starting point is understanding how each workload's traffic actually behaves.

Synchronous training traffic is the strict case: because collectives are tightly coupled across all participating GPUs, the recommended east-west oversubscription target is 1:1 — a genuinely non-blocking fabric. Inference is different: its flows are much more statistically independent of one another, which is why current guidance allows 2:1 to 3:1 oversubscription at the upper tiers without meaningfully hurting throughput.

Retrieval-augmented generation sits closer to the inference profile in network terms, because retrieval and generation for a given query are largely independent of what every other query is doing at the same moment — it doesn't create the synchronised, all-to-all pattern that defines training. The practical implication: don't build your RAG or inference tier to the same non-blocking standard as your training tier, and don't assume a single fabric design serves both well.

Reference Sizing: Small, Medium And Large On-Prem Clusters

For smaller on-prem deployments — the 8 to 128 GPU range that covers most UK mid-market pilots and single-team training clusters — a realistic fabric envelope in 2026 is 100G to 400G Ethernet. This is the practical sweet spot: enough east-west headroom for a non-blocking design at modest scale without paying for capacity you won't use for years.

At the top end, platforms aimed at million-GPU AI factories are now running at 800 Gb/s per port, with Ethernet-based options such as Spectrum-X and InfiniBand-class Quantum-X built specifically for that scale. Very few UK on-prem buyers sit in this bracket today, but it defines where the roadmap is heading and what your cabling plant needs to tolerate later.

A useful worked reference for mid-size leaf-spine design: 400G per port across 64 spine switches yields roughly 25.6 Tbps of bisectional bandwidth. That's the kind of arithmetic worth doing before you commit to a spine switch count — it tells you directly whether your planned topology actually delivers the non-blocking ratio your training workload needs, rather than assuming it does.

Ethernet vs InfiniBand vs RoCE: The 2026 Decision

IEEE 802.3dj, expected to complete by July 2026, standardises Ethernet at 200 Gb/s, 400 Gb/s, 800 Gb/s and 1.6 Tb/s — a single Ethernet roadmap now covers the full range that used to be InfiniBand's exclusive advantage. In practice, 100G, 200G, 400G and 800G are the dominant server- and fabric-facing speeds through 2026, with 1.6T only beginning to appear on early AI-scale spine and inter-cluster links.

This matters for the Ethernet-versus-InfiniBand question that every on-prem AI buyer eventually faces. Ethernet-based AI fabrics (RoCE-driven) have closed much of the historical gap by riding the same 802.3dj rate progression that mainstream data-centre gear uses, which simplifies operations, staffing and spares versus a parallel InfiniBand estate. InfiniBand and InfiniBand-class platforms still target the very largest, most latency-sensitive training-at-scale deployments. For most UK on-prem projects outside the million-GPU tier, the practical question isn't which fabric is faster in isolation — it's whether your team can operate a second specialist network alongside the Ethernet estate you already run. Read the debate between Infiniband and Ethernet for AI fabrics in more depth before committing capital.

Building The Fabric: Switches, NICs, Optics And Cabling

Physical layer choices need to match the reach and the oversubscription target you've committed to, not the other way round. For top-of-rack to spine links, 800G QSFP-DD in DR8 or FR4 variants is the current fit, with the choice driven by your physical plant and the reach the run actually needs.

Reach discipline matters more than it looks on a spreadsheet: inter-pod or inter-row links beyond 500 metres should move to coherent optics or 800G FR4 variants, and anything at 2 km or longer single-mode reach is treated as a genuinely different design class with different cost and power characteristics. Buyers who plan a single short-reach optics strategy for an entire campus tend to discover this the expensive way, mid-build.

Host-side, NIC selection increasingly means deciding how much network processing you offload from the CPU — this is where it's worth taking the time to understand DPUs and SmartNICs before specifying nodes. And don't assume your GPU interconnect figures cover storage: the 800G and 400G numbers above describe host and spine interconnect, not necessarily every storage tier, which is exactly why it's worth reviewing how to optimise storage networking with NVMe over Fabrics as a separate design decision rather than an inherited afterthought.

Reference On-Prem AI Fabric Layout
GPU Leaf Switches100G-400G host uplinksSpine Switches800G QSFP-DD DR8/FR4Storage FabricNVMe-oF, sized separatel…Inter-Row LinkCoherent optics beyond…

Securing East-West Traffic Without Starving The Fabric

AI clusters invert the usual security assumption: the volume of traffic that matters most now moves rack-to-rack and switch-to-switch, not client-to-server. A perimeter-only security posture leaves that entire traffic class effectively unmonitored, which is a real gap given how much of an AI cluster's value — and its intellectual property exposure — sits in gradients and model weights moving across that fabric.

The practical response is to build segmentation and inspection into the fabric design itself rather than bolting it on afterwards, and to size any inline security capacity against the same non-blocking targets you set for training traffic — a security appliance that reintroduces a blocking point defeats the point of the 1:1 fabric you just paid for. Review your approach to network security as part of the fabric design, not as a separate project that follows it.

The UK Planning Risk: Grid Queue And Curtailment

In the UK, network sizing decisions increasingly have to be made against a genuinely uncertain power backdrop, and that changes the calculus for how aggressively to overbuild. The transmission connection queue in Great Britain reached 125 GW by mid-2025, up sharply from 41 GW in late 2024. Data centers account for approximately 73 GW of this queue, which is swollen substantially by speculative demand rather than projects that will actually build.

Government and regulator response is moving in real time: a GOV.UK consultation aimed at tackling speculative grid-connection requests closed on 15 April 2026. Ofgem, in a consultation published in July 2026, is proposing a Data Centre Commitment Fee and additional queue management milestones to tackle speculative grid connections, while also exploring mandatory curtailment for future data centres during periods of system stress as a backstop measure, with further consultation expected in autumn 2026. Alongside this, Ofgem's planned Curate, Plan, and Connect framework with the National Energy System Operator is intended to stop viable projects being blocked behind speculative ones in the queue.

For network buyers, the implication is direct: an underutilised, power-hungry non-blocking fabric built well ahead of confirmed grid capacity is harder to justify in this environment, and curtailment risk is a real argument for phasing your fabric build to match your actual connection date rather than provisioning the full end-state topology on day one.

Future-Proofing And The TCO Verdict

With IEEE 802.3dj due to complete by July 2026 and covering 200G through to 1.6T on a single Ethernet standard, the sensible future-proofing move is optical and cabling plant that supports a reach upgrade path — not a switch estate locked into short-reach-only optics that can't carry you to the next tier without a full re-cable.

On total cost of ownership, the lever that actually moves the number is the oversubscription ratio you choose per tier, not the brand on the chassis: a 1:1 non-blocking design costs meaningfully more in ports and optics than a 2:1 or 3:1 design, so applying the strict ratio only where the workload genuinely needs it — training — while relaxing it for inference and RAG tiers is the single biggest optimisation available to a UK buyer in 2026.

Our recommendation: size the training tier to 1:1 non-blocking within the 100G–400G envelope appropriate to your GPU count, relax the inference/RAG tier to 2:1–3:1, plan optics reach against your actual campus layout rather than a generic template, and sequence the whole build against your confirmed grid connection date rather than the GPU delivery date. Use calculate your AI GPU requirements and explore financing options for network equipment as the next two concrete steps.

Sources

Every figure in this article traces to the sources below.

  • AI Data Centers — east-west oversubscription targets for training vs inference
  • Network Devices Inc. — realistic Ethernet fabric envelope and top-end port speeds
  • IEEE ComSoc Technology Blog — 802.3dj rates and dominant 2026 Ethernet speeds
  • HytoptoDevice — 800G optics, spine bisectional bandwidth example, reach thresholds
  • GOV.UK — consultation on speculative grid-connection demand reform
  • Bloomberg — Ofgem exploring mandatory data centre curtailment
  • AI Energy Intelligence — GB transmission connection queue growth
  • CMS — Ofgem's Curate, Plan, and Connect grid reform framework
Share
Key takeaways
  • Training traffic needs a 1:1 non-blocking east-west fabric; inference and RAG can run at 2:1–3:1 oversubscription — don't size both tiers the same way.
  • For 8–128 GPU clusters, 100G–400G Ethernet is the realistic 2026 fabric envelope; 800 Gb/s per port is reserved for million-GPU-class deployments.
  • Do the bisectional bandwidth arithmetic before committing spine count — 400G per port across 64 spines yields ~25.6 Tbps, and that number either meets your non-blocking target or it doesn't.
  • IEEE 802.3dj (expected complete July 2026) puts 200G–1.6T on one Ethernet roadmap, narrowing the historical InfiniBand advantage for most on-prem tiers.
  • Plan optics reach deliberately: 800G QSFP-DD DR8/FR4 for ToR-to-spine, coherent optics or FR4 beyond 500m, and treat 2km+ single-mode as a separate design and cost class.
  • UK builds carry grid risk on top of network risk — the GB connection queue hit 125 GW by mid-2025 (from 41 GW in late 2024, with data centres accounting for approximately 73 GW), and Ofgem's July 2026 consultation proposes a Data Centre Commitment Fee alongside possible mandatory curtailment, so phase your fabric build to your actual connection date.
Frequently asked

FAQs — Right-Sizing Your Network for On-Prem AI in 2026

What oversubscription ratio should I use for an on-prem AI training network?

Target a 1:1 non-blocking east-west fabric for synchronous training, because all-reduce and all-gather collectives create tightly coupled, simultaneous traffic across every GPU in a job. Any blocking point becomes a synchronisation barrier the whole job waits on, so this is the tier where oversubscription genuinely costs performance, not just headroom.

Can my inference or RAG network be oversubscribed to save cost?

Yes — inference workloads tolerate 2:1 to 3:1 oversubscription at upper tiers because their flows are far more statistically independent than training collectives. RAG behaves similarly since retrieval and generation for one query don't depend on every other query happening at the same time, so building this tier to the same 1:1 standard as training is usually unnecessary spend.

Should I choose Ethernet or InfiniBand for on-prem AI in 2026?

IEEE 802.3dj brings 200G to 1.6T onto a single Ethernet standard, so RoCE-based Ethernet now covers most on-prem tiers without the operational overhead of a separate fabric. InfiniBand and InfiniBand-class platforms remain aimed at the very largest, most latency-sensitive training-at-scale builds — most UK buyers outside that tier are better served staying on Ethernet.

How does the UK grid situation affect my on-prem AI network build timeline?

The GB transmission connection queue reached 125 GW by mid-2025, up from 41 GW in late 2024, with data centres accounting for approximately 73 GW of that total. Ofgem's July 2026 consultation proposes a Data Centre Commitment Fee and additional queue management milestones, while also exploring mandatory curtailment for data centres during system stress as a backstop measure, with further consultation expected in autumn 2026. Sequence your network build against your confirmed connection date rather than building the full non-blocking fabric before power capacity is certain.

What cabling reach should I plan for between rows or pods?

Inter-pod or inter-row links beyond 500 metres should use coherent optics or 800G FR4 variants; anything at 2km or longer single-mode reach is a distinct design and cost class. Plan this at the campus layout stage, not after the switch order is placed.

Does my storage network need the same bandwidth as my GPU interconnect?

Not necessarily. The commonly cited 800G and 400G figures describe GPU host and spine interconnect specifically, not every storage tier. Size your storage fabric — including NVMe over Fabrics paths — as a distinct decision so training nodes aren't starved by an undersized storage link hiding behind a well-sized compute fabric.

Related

Continue reading

More in Networking

Got a question this article didn't answer?

One conversation with an engineer who's done this before. No sales script.

Talk to Servnet →

Talk to a UK specialist

Get expert advice or a no-obligation quote — servers, storage, networking, maintenance, finance and cloud. We reply the same working day.

or call 0800 987 4111