Most AI accelerator comparisons quietly mix a single GPU's spec sheet with an 8-GPU system's totals — and the gap is enormous. Nvidia's own datasheet-style materials put a single B200 GPU at 180 GB of HBM3e and 7.7 TB/s of bandwidth, while the DGX B200 system page lists 1,440 GB and 64 TB/s at the 8-GPU node level — an eightfold jump that has nothing to do with a faster chip. This database rebuilds the Nvidia, Intel and AMD comparison from vendor documents only, labels every figure by its exact hardware scope, and — where the underlying data doesn't support a claim — says so rather than inventing a number.
View the data behind this chart
| HBM capacity | HBM bandwidth | Dense FP8 throug… | |
|---|---|---|---|
| Nvidia B200 (per GPU) | 180 GB HBM3e | 7.7 TB/s | 4,500 TFLOPS |
| Nvidia DGX B200 (node) | 1,440 GB total | 64 TB/s total | Not published |
| Intel Gaudi 3 | 8x HBM2e modules | 3.67 TB/s | 3,670 TFLOPS |
Why AI accelerator spec sheets keep disagreeing
Every vendor publishes accelerator specifications at a different level of the stack: some datasheets describe a single GPU die, others describe the module it ships in, and others describe the eight-GPU rack system it's sold as. When these get pasted into a single comparison table without labels, the resulting numbers look contradictory even when every source is individually correct.
Nvidia's Blackwell B200 is the clearest reference point available in 2026, because both a per-GPU datasheet-style document and Nvidia's own DGX B200 system page are publicly retrievable — which makes it possible to show, side by side, exactly where GPU-level figures end and node-level totals begin. Intel's Gaudi 3 material is well documented but has one confirmed gap. AMD's MI350-series public data set, as retrieved for this database, is too thin to support a full comparable row without guessing — so it isn't force-fitted into the table below.

Nvidia Blackwell B200: the per-GPU numbers
Datasheet-style material for the B200 (also sold in HGX B200 configurations) puts the single-GPU memory footprint at 180 GB of HBM3e with 7.7 TB/s of bandwidth. Interconnect is via NVLink 5, rated at 1.8 TB/s per GPU — not a PCIe-only design, which matters for anyone estimating multi-GPU scaling. The SXM module carries a 1,000 W TGP.
- •180 GB HBM3e per GPU, 7.7 TB/s bandwidth per GPU
- •NVLink 5 interconnect: 1.8 TB/s per GPU
- •1,000 W TGP per GPU (SXM module)
- •Dense FP8: ~4,500 TFLOPS per GPU; dense FP4: ~9,000 TFLOPS per GPU
Why the DGX B200's 1,440 GB isn't a GPU spec
Nvidia's own DGX B200 system page states 1,440 GB of total HBM3e and 64 TB/s of total HBM bandwidth — but this is the aggregate across the 8-GPU system, not a per-GPU figure. Anyone citing '1,440 GB' or '64 TB/s' as a single-accelerator number has conflated a whole NVIDIA DGX systems node with a GPU, and the resulting per-GPU price-performance or memory-per-dollar calculation will be off by roughly an order of magnitude.
This distinction is exactly what UK procurement teams need to hold onto when comparing quotes: a vendor bid that quietly quotes node-level memory against a competitor's per-GPU memory figure isn't offering a better spec — it's offering a mismatched comparison.
The conflicting secondary figure UK buyers should flag
A separate B200 market explainer in wide circulation states 192 GB of HBM3e and 8 TB/s of bandwidth for the same GPU — figures that conflict with the 180 GB / 7.7 TB/s numbers in the datasheet-style material used above. Rather than averaging or picking whichever number is more convenient, this database treats the two as distinct, conflicting data points and uses only the datasheet-style figures in its comparison table.
If a UK reseller or system integrator quotes you 192 GB per B200 GPU, ask which document that figure comes from — it may be tracing back to the same secondary report rather than Nvidia's own materials.
Intel Gaudi 3: solid figures with one open gap
Intel's Gaudi 3 material is consistent across sources: Hot Chips technical slides and Intel's own white paper describe an accelerator built around 8 integrated HBM2e devices, delivering 3.67 TB/s of HBM bandwidth and 3,670 FP8 TFLOPS. A 2026 coverage page corroborates the same 3,700 GB/s bandwidth and 3,670 FP8 TFLOPS figures, giving reasonable confidence in those two numbers specifically.
What the currently retrieved set does not include is a fresh, dated datasheet line for Gaudi 3's interconnect bandwidth and TDP. Those figures exist in Intel's broader materials, but this database won't publish a number it can't trace to a specific, checked document — UK buyers should request the current Intel datasheet directly for those two fields before finalising a Gaudi 3 procurement decision.
View the data behind this chart
| Nvidia B200 FP8 | Nvidia B200 FP4 | Intel Gaudi 3 FP8 | |
|---|---|---|---|
| Dense TFLOPS | TFLOPS4500 | TFLOPS9000 | TFLOPS3670 |
AMD MI350-series: the row we're not publishing
AMD's MI350X datasheet is hosted by NEC and is a genuine vendor document, but the portion retrieved for this database exposes only a single figure: a 160 GB/s bidirectional baseboard interconnect line. It does not expose HBM capacity, HBM bandwidth, TDP, or FP8/FP4 throughput at a level that could be safely compared to the Nvidia and Intel figures above.
Rather than filling those gaps with figures pulled from press coverage of unclear precision — which is precisely the mixed-marketing problem this database exists to avoid — MI350X is deliberately left out of the normalised comparison table below. Any table you see elsewhere claiming a full MI350X vs B200 spec-for-spec comparison should be checked against AMD's own datasheet before you rely on it for a purchasing decision.
Reading this table for UK procurement
The practical risk for UK IT buyers isn't a lack of data — it's blending scopes. A node-level DGX figure compared against a competitor's single-accelerator figure will distort cost-per-token and rack-density planning long before you get to a purchase order. Before sizing a cluster with tools like an AI GPU sizing calculator, confirm whether every input figure is per-GPU, per-module, or per-node.
Vendor list pricing for these accelerators was not available in the source material used for this database, so any UK total-cost-of-ownership modelling needs pricing sourced separately and mapped against the GPU-level specs above, not the node-level totals. For regulated or public-sector workloads, accelerator choice should also be checked against UK GDPR and security requirements before committing to an onshore colocation or sovereign AI build — see our broader AI servers data study for the wider system-level context around server configuration decisions.
Methodology
This database was compiled in August 2026 from vendor and vendor-adjacent primary sources: Nvidia's own DGX B200 system page, a datasheet-style HGX B200 reference document, Intel's Gaudi 3 white paper and Hot Chips conference material, and an AMD MI350X datasheet hosted by NEC. Secondary market-explainer and benchmark-coverage pages were used only to corroborate figures already present in primary material, and were explicitly excluded from the comparison table wherever they conflicted with a vendor document.
Every figure was checked against the exact hardware level it describes — single GPU/module, multi-GPU baseboard, or full system node — and labelled accordingly rather than normalised into a single 'per accelerator' number. Where a source provided only a partial or ambiguous figure (as with AMD's MI350X interconnect line, or Intel Gaudi 3's interconnect and TDP), that gap is stated directly rather than filled with an estimate. Readers using this data for procurement should still request current, dated vendor datasheets for any field this article flags as unverified.
Sources
Every figure in this article traces to the sources below.
- •NVIDIA — DGX B200 system page (node-level HBM totals)
- •GPU Smith — HGX B200 datasheet PDF (per-GPU HBM, bandwidth, NVLink)
- •GPU Smith — B200 hardware page (TDP, FP8/FP4 throughput)
- •Yobitel — B200 market explainer (conflicting secondary figures)
- •Hot Chips — Intel Gaudi 3 conference presentation
- •Intel — Gaudi 3 AI accelerator white paper
- •InferenceBench — Gaudi 3 corroborating coverage
- •NEC — AMD Instinct MI350X GPU datasheet
The 7 verified data points behind this study are free to download and reuse with attribution (CC BY 4.0).
Cite as: Servnet Research, “AI Accelerator Comparison 2026: Nvidia vs AMD vs Intel”, servnetuk.com, 2026.
