UK’s trusted IT infrastructure partner since 2003
Servnet
FinanceToolsConfiguratorGet in Touch
AI Infrastructure

Nvidia Rubin GPU Specs 2026: The UK Buyer's Roadmap Tracker

Servnet Editorial · IT infrastructure analysis7 min read
Share

Nvidia's next GPU generation, Vera Rubin, is on track for the second half of 2026, with Rubin Ultra following in H2 2027. The number UK infrastructure teams should sit with first isn't a performance figure — it's power. Independent analysis published in March 2026 puts Rubin at 1,800-2,300 W per GPU, with no air-cooled configuration available. That single fact reshapes the procurement conversation for any UK site without liquid-cooling headroom already built in, regardless of how the FP4 throughput numbers compare.

HBM Memory Capacity by GPU Generation
1030 GB773 GB515 GB258 GB0 GB288 GBBlackwell Ultra288 GBRubin1024 GBRubin UltraHBM capacity
View the data behind this chart
HBM Memory Capacity by GPU Generation
Blackwell UltraRubinRubin Ultra
HBM capacityGB288GB288GB1024

Nvidia's Roadmap in Mid-2026: Where UK Buyers Actually Stand

As of July 2026, Nvidia's publicly stated cadence runs Blackwell Ultra NVL72 (shipped H2 2025), Vera Rubin (targeted H2 2026), then Rubin Ultra (targeted H2 2027). A March 2026 report said Rubin had already taped out and was targeted for enterprise deployment in the second half of this year, alongside a separate end-of-2026 availability target for Rubin CPX, a companion part in the same platform generation.

This tracker exists because the three most-cited pieces of coverage on Rubin — Nvidia's own architecture write-up, US-focused system-integrator roadmaps, and dense TCO modelling aimed at hyperscale buyers — none of them frame the decision the way a UK IT or procurement lead needs it framed: what ships when, what it costs to run, and whether waiting is the financially smarter move.

The short version: Blackwell Ultra is the nearer-term, already-shipping option; Rubin is a second-half-2026 proposition realistically reaching UK deployments through OEM servers and cloud instances rather than direct GPU purchase; Rubin Ultra is a 2027 refresh-cycle conversation, not a 2026 one.

Illustration: Nvidia Rubin GPU Specs 2026: The UK Buyer's Roadmap Tracker

Rubin vs Blackwell Ultra vs Rubin Ultra: The Core Numbers

Nvidia and independent outlets have converged on a consistent spec picture. Rubin is positioned at up to 50 PFLOPS of FP4 inference performance with up to 288 GB of HBM4 memory — matching Blackwell Ultra's reported 288 GB ceiling on capacity but moving to a newer HBM generation. Tom's Hardware separately reported Rubin's training performance at around 17 PFLOPS in INT8, a distinct metric class from the FP4 inference figure and not directly comparable to it.

Rubin Ultra, due in H2 2027, is the bigger jump: up to 100 PFLOPS FP4 and 1 TB of HBM4E memory, delivered across 12 stacks of HBM4E per package according to Tom's Hardware. That's a doubling of the inference throughput ceiling and a significant increase in memory capacity (over 3.5x) versus Rubin, in a single generational step.

Buyers evaluating this against current-generation hardware should also weigh a comparison of H100, H200, B100, and B200 GPUs already in production UK deployments, since Blackwell-class parts remain the primary option for immediate, direct procurement in mid-2026, though Rubin is now in full production with H2 2026 partner and cloud availability.

The Vera Rubin Platform: Extreme Co-Design Explained

Rubin isn't sold as a standalone chip decision — it's a platform generation pairing the Rubin GPU with Nvidia's Vera CPU, next-generation NVLink Switch fabric, DPU, and NIC, all co-designed as a single system rather than assembled from discrete parts. This is the same 'extreme co-design' logic behind Blackwell's NVL72 rack architecture, extended a generation further.

For UK enterprises, the practical implication is that Rubin's memory and interconnect gains only fully materialise at the rack or pod level, not the single-GPU level. Understanding how High Bandwidth Memory (HBM) capacity interacts with interconnect bandwidth is relevant here, because the platform's value proposition rests on keeping ever-larger model states resident close to compute rather than shuttling data across slower links.

Rubin CPX, targeted for end-of-2026 availability per the March 2026 StorageReview report, sits alongside the main Rubin GPU in this platform generation — evidence that Nvidia is now shipping companion silicon on a near-identical timeline to the flagship part, rather than staggering releases by a year or more as in prior cycles.

Agentic AI Claims Under Scrutiny

Nvidia markets Rubin heavily around 'agentic AI' — multi-step reasoning, tool use, and long-running autonomous workflows. The verifiable architectural basis for that claim is narrower than the marketing suggests: the confirmed specifics are a high memory capacity (up to 288 GB HBM4, matching Blackwell Ultra's capacity but using a newer HBM generation), a stated inference-throughput ceiling (50 PFLOPS FP4), and a separate, lower training-throughput figure (around 17 PFLOPS INT8) that Nvidia's own messaging tends not to foreground.

That gap between an inference metric and a training metric matters for agentic workloads specifically, because agentic pipelines typically run repeated inference calls with growing context, not fresh training runs — meaning the 50 PFLOPS FP4 figure is the more directly relevant number for most UK deployments considering Rubin for reasoning-heavy applications, not the training figure some coverage conflates it with.

UK technical teams evaluating vendor claims should treat the FP4/INT8 distinction as a hard boundary in any procurement conversation, and push suppliers for workload-specific benchmarks rather than accepting headline PFLOPS figures as a proxy for agentic-AI readiness.

Power and Cooling: What 1,800-2,300W Per GPU Means for UK Data Centres

The most operationally significant data point in this whole roadmap is Rubin's reported 1,800-2,300 W per GPU, with no air-cooled configuration available according to March 2026 analysis. That's a substantial step up in electrical and thermal demand per accelerator, and it lands directly on UK sites already managing constrained grid connections in several regions.

For a UK data centre operator, this changes the calculus around PUE, rack density, and cooling retrofit spend well before any performance benefit is realised. Sites without existing liquid-cooling infrastructure should treat Rubin as requiring an infrastructure upgrade project in its own right, not a drop-in GPU refresh — a dynamic explored further in the impact of HBM on AI server costs.

Energy contract terms and cooling-retrofit lead times should therefore sit alongside GPU allocation queues as a binding constraint on when a UK organisation can realistically deploy Rubin-based infrastructure — the power envelope alone may push a practical deployment date later than Nvidia's H2 2026 platform target.

Rubin vs Blackwell Ultra vs Rubin Ultra: Core Specs
Blackwell UltraRubinRubin UltraMemory capacity288 GB288 GB HBM41 TB HBM4EFP4 inference performa…Not disclosed50 PFLOPS100 PFLOPSPower envelope per GPUNot disclosed1,800-2,300WNot disclosedShipment windowH2 2025H2 2026H2 2027
View the data behind this chart
Rubin vs Blackwell Ultra vs Rubin Ultra: Core Specs
Blackwell UltraRubinRubin Ultra
Memory capacity288 GB288 GB HBM41 TB HBM4E
FP4 inference performa…Not disclosed50 PFLOPS100 PFLOPS
Power envelope per GPUNot disclosed1,800-2,300WNot disclosed
Shipment windowH2 2025H2 2026H2 2027

Cost of Ownership and the UK Buying Calculus

No official GBP list pricing exists for Rubin at this point, and none of the sourced reporting behind this tracker includes retail or system-level pricing in sterling. That absence is itself informative: Rubin and Rubin Ultra are aimed squarely at hyperscale and large-enterprise AI deployments, meaning UK buyers will almost always encounter them through OEM server configurations, cloud provider instances, or colocation-integrated systems — not standalone card purchases.

Early Rubin supply is also described as prioritised toward hyperscalers and frontier AI labs, which means UK enterprises outside that tier should expect allocation queues rather than open-market availability in the initial H2 2026 window, even once shipments begin.

The pragmatic UK procurement strategy that follows from this: use our AI GPU calculator to model current-generation Blackwell-class fleet economics now, and treat Rubin as a late-2026-into-2027 refresh-cycle decision rather than an immediate purchase, unless capacity is needed this quarter.

Rubin vs the Competitive Landscape — and Should You Upgrade?

Rival accelerators expected in the 2026-27 window from other silicon vendors are widely anticipated in the market, but independently verified specifications comparable in detail to Nvidia's Rubin disclosures were not available in the sources underpinning this tracker at time of writing. UK buyers should treat any competitive comparison claims circulating for that period as provisional until vendors publish equivalent architecture-level detail.

What is verifiable is the internal Nvidia trade-off: Blackwell Ultra is shipping now with a known 288 GB memory ceiling; Rubin maintains Blackwell Ultra's 288 GB memory capacity but shifts to HBM4 and a new platform generation, and while Blackwell Ultra's FP4 throughput is not officially disclosed, Rubin's 50 PFLOPS FP4 represents a significant generational leap over prior Blackwell-class parts like the B200; Rubin Ultra a year later does deliver the larger step-change, doubling FP4 throughput and significantly increasing memory capacity (over 3.5x) again.

For UK businesses already invested in Blackwell-class fleets, the decision is less about specs and more about timing and infrastructure readiness — a trade-off covered in more depth when deciding between Blackwell and Rubin. As a working rule: buy Blackwell Ultra now for immediate capacity needs; plan Rubin for late-2026/2027 refresh cycles where liquid-cooling readiness, energy contracts, and allocation timing all align; treat Rubin Ultra purely as a 2027-28 multi-year roadmap input.

Methodology

This tracker compiles GPU-generation specifications, shipment-window statements, and infrastructure-impact figures published between March 2025 and March 2026 by technology and business press covering Nvidia's roadmap disclosures, including outlets specialising in enterprise hardware analysis, semiconductor reporting, and infrastructure-focused technical coverage.

Figures were cross-checked against the original source publication for each data point and kept scoped exactly as reported — FP4 inference throughput, INT8 training throughput, HBM memory capacity by generation, and per-GPU power estimates are treated as distinct metric classes rather than merged into single comparative statements. Where no verified figure existed for a given generation or metric (for example, Blackwell Ultra's FP4 throughput or Rubin Ultra's power envelope), this is stated explicitly as not disclosed rather than estimated.

Shipment windows (H2 2025, H2 2026, H2 2027) are reported as roadmap targets rather than confirmed customer delivery dates, and early-supply prioritisation toward hyperscalers and frontier labs is flagged wherever source reporting stated it. This tracker will be updated as Nvidia and independent analysts publish further verified detail ahead of Rubin's H2 2026 shipment window.

Sources

Every figure in this article traces to the sources below.

  • Network World — Nvidia's stated shipment windows for Blackwell Ultra NVL72 and Vera Rubin
  • TechCrunch — Rubin's 50 PFLOPS FP4 inference and 288 GB memory figures from GTC 2025
  • The Next Platform — Rubin Ultra's H2 2027 target, 100 PFLOPS FP4 and 1 TB HBM4E
  • Tom's Hardware — Rubin Ultra's 12-stack HBM4E configuration and Rubin's INT8 training performance
  • StorageReview — March 2026 report on Rubin tape-out and Rubin CPX end-of-2026 target
  • Tech-Insider — Rubin's 1,800-2,300 W power envelope and cooling analysis
  • CNBC — Corroboration of Nvidia's H2 2026 next-generation GPU shipment timing
Nvidia GPU Shipment Windows, H2 2025 to H2 2027
W0W22W44W66W88W110W130Blackwell Ultra26wVera Rubin26wRubin CPX13wRubin Ultra26wTotal: 130 weeks end-to-end
View the data behind this chart
Nvidia GPU Shipment Windows, H2 2025 to H2 2027
PhaseStarts (week)Duration (weeks)
Blackwell Ultra026
Vera Rubin5226
Rubin CPX6513
Rubin Ultra10426
Open data

The 9 verified data points behind this study are free to download and reuse with attribution (CC BY 4.0).

Cite as: Servnet Research, “Nvidia Rubin GPU Specs 2026: The UK Buyer's Roadmap Tracker”, servnetuk.com, 2026.

Share
Key takeaways
  • Vera Rubin is targeted for H2 2026 shipment; Rubin Ultra follows in H2 2027 — treat these as roadmap targets, not confirmed delivery dates.
  • Rubin: up to 50 PFLOPS FP4 inference and up to 288 GB HBM4. Rubin Ultra: up to 100 PFLOPS FP4 and 1 TB HBM4E across 12 memory stacks.
  • Rubin's reported 1,800-2,300 W per-GPU power envelope has no air-cooled configuration — plan liquid-cooling infrastructure before GPU procurement, not alongside it.
  • No GBP pricing exists yet; UK buyers will access Rubin via OEM servers, cloud instances, or colocation, with early supply prioritised toward hyperscalers.
  • Blackwell Ultra remains the most readily available option for immediate, direct procurement; Rubin is available via hyperscalers and partners in H2 2026, and belongs in late-2026/2027 refresh planning for broader enterprise deployments.
  • Rubin's 50 PFLOPS figure is an inference metric, not a training metric — don't compare it directly against training-throughput numbers from other generations.
Frequently asked

FAQs — Nvidia Rubin GPU Specs 2026

When will Nvidia Rubin GPUs actually ship?

Nvidia's roadmap targets the second half of 2026 for Vera Rubin, with a March 2026 report confirming the chip had taped out and was on track for enterprise deployment in that window. Early supply is expected to prioritise hyperscalers and frontier AI labs before wider availability.

What is Rubin Ultra and when does it arrive?

Rubin Ultra is the following Nvidia GPU generation, targeted for H2 2027. It's projected at up to 100 PFLOPS of FP4 inference performance and 1 TB of HBM4E memory across 12 memory stacks — roughly double Rubin's throughput and significantly higher (over 3.5x) memory ceilings.

How does Rubin compare to Blackwell Ultra on memory?

Both are reported at up to 288 GB of memory capacity, but Rubin moves to the newer HBM4 standard versus Blackwell Ultra's memory configuration. The bigger memory jump doesn't arrive until Rubin Ultra's 1 TB HBM4E in 2027.

Can UK businesses cool a Rubin-based system with air cooling?

No. March 2026 analysis reported Rubin at 1,800-2,300 W per GPU with no air-cooled configuration available, meaning UK deployments will require liquid-cooling infrastructure — a significant consideration for sites without existing retrofit capability.

Should UK enterprises buy Blackwell Ultra now or wait for Rubin?

For immediate capacity needs, Blackwell Ultra is the most widely available option. Rubin is available via hyperscalers and cloud instances in H2 2026, but for direct enterprise procurement, it is more realistically a late-2026-into-2027 refresh-cycle purchase due to allocation prioritisation and cooling infrastructure demands.

What is Rubin CPX?

Rubin CPX is a companion part in the same Vera Rubin platform generation, targeted by Nvidia for end-of-2026 availability according to March 2026 reporting — shipping on a near-identical timeline to the main Rubin GPU rather than a year or more later.

Related

Continue reading

More in Research

Got a question this study didn't answer?

One conversation with an engineer who's done this before. No sales script.

Talk to Servnet →

Talk to a UK specialist

Get expert advice or a no-obligation quote — servers, storage, networking, maintenance, finance and cloud. We reply the same working day.

or call 0800 987 4111