AMD's Instinct MI355X posted 100,282 tokens per second on the Llama2-70b-99.9 Server test in MLPerf Inference v6.0 — a 3.1x jump over the MI325X result AMD had previously submitted, published in AMD's own April 2026 benchmark blog. That is the strongest evidence yet that AMD's roadmap is delivering. But UK buyers weighing an AMD Instinct vs Nvidia decision this year face a sharper filter than raw throughput: whether their exact Linux distribution, ROCm release and framework version sit inside AMD's published support matrix on deployment day. This piece works from AMD's own compatibility documents and MLPerf submissions to give a workload-by-workload verdict versus Nvidia H200 vs H100 vs B100/B200 deployments.
View the data behind this chart
| MI325X (MLPerf v5.1… | MI355X (MLPerf v6.0) | |
|---|---|---|
| Tokens per second | tokens/s…32028 | tokens/s…100282 |
The AI accelerator landscape in mid-2026: a shifting duopoly
As of August 2026, AMD's Instinct line is no longer a paper alternative confined to press releases. Both MI325X and MI355X appear as currently supported hardware in AMD's live ROCm system-requirements matrix — MI355X on the CDNA4 gfx950 architecture, MI325X on the CDNA3 gfx942 architecture. That distinction matters immediately: these are two different software targets, not one product with a speed bump, and AMD's own documentation treats them as such.
Nvidia still holds the edge on procurement familiarity and ecosystem certainty that UK IT leaders have built years of process around. What has changed is that AMD now has a published, dated body of evidence — MLPerf Inference v6.0 results tagged as released on 1 April 2026 — that lets buyers evaluate the hardware on submitted numbers rather than marketing slides. The question for a UK deployment in 2026 is no longer purely "can AMD run my models" but "does my exact stack land inside AMD's current support window."

MI325X vs MI355X: the generation gap AMD itself documents
The clearest head-to-head data point in AMD's own materials is inside its Instinct line, not against Nvidia. In MLPerf Inference v6.0, AMD's MI355X delivered 100,282 tokens per second on the Llama2-70b-99.9 Server scenario, which AMD describes as 3.1x the throughput of the MI325X result it had previously submitted. A separate summary from Principled Technologies puts concrete numbers either side of that claim: MI355X on v6.0 at 100,282.36 tokens/second against MI325X on the earlier v5.1 round at 32,027.57 tokens/second. Both figures are real and AMD-sourced, but they span two different MLPerf rounds — v6.0 and v5.1 — so this is a generational trend, not a same-round bake-off, and should be quoted that way.
Behind the number sits a genuine architectural jump. AMD's MLPerf submission page confirms MI355X uses FP4 precision support alongside 288 GB of HBM3 memory and 8 TB/s of memory bandwidth. For inference workloads that need to hold very large model weights on a single card, that memory ceiling is the more durable fact than any single tokens-per-second figure — it changes how many GPUs a cluster needs before bandwidth or capacity becomes the bottleneck. AMD also reports exceeding 1 million tokens per second in multi-node inference, evidence that the throughput story extends beyond a single card when the deployment is built at that scale.
For direct throughput and efficiency comparisons against Nvidia's H200, B100 and B200 parts, AMD's own MLPerf submissions in this piece don't include Nvidia's equivalent figures — see the dedicated Nvidia H200 vs H100 vs B100/B200 breakdown for that side of the comparison before drawing a cross-vendor conclusion.
Training at scale: what AMD's MLPerf v6.0 submissions actually cover
Training is where buyers most often assume one number applies across an entire product family — AMD's own submission round argues against that. Its 2026 MLPerf Training v6.0 round included entries spanning MI325X, MI350X and MI355X Instinct GPUs, plus a separate 512-GPU MI300X submission from Oracle Cloud Infrastructure working with AMD. That is cards spanning different architectural generations, submitted under different configurations, not one homogeneous result.
A supplemental MLCommons discussion document adds a specific scale reference: Flux.1 FP8 multi-node data-parallel training ran on 64 MI325X GPUs. For a UK buyer sizing a training cluster, the practical lesson is to match the submitted configuration — GPU model, precision, node count — to the workload being planned, rather than assuming that any Instinct card will replicate a headline number found for a different SKU or scale.
ROCm vs CUDA: quantifying the version-pinning tax
The real "CUDA tax" for AMD in 2026 isn't a rewrite-your-stack problem — ROCm supports the mainstream frameworks. It is a documentation-discipline problem, and AMD's own matrices show why. MI355X requires ROCm 7.0.1 or higher; MI325X requires ROCm 6.3.1 or higher. These are not interchangeable stacks: a team standardised on the older ROCm baseline for MI325X cannot assume MI355X will simply drop in.
AMD's ROCm LLM Extension compatibility matrix gives buyers two currently viable pairings for version 26.04 of that tooling: MI355X, MI325X and MI300X together with ROCm 7.2.0 on Ubuntu 24.04, or the same three cards with ROCm 7.1.0 on Ubuntu 22.04. That is genuinely useful for teams wanting one LLM tooling baseline across mixed-generation fleets — but it only holds for those two exact pairings, not for arbitrary ROCm-and-OS combinations.
Virtualisation exposes the sharpest worked example of the pinning tax. AMD's hardware-support matrix lists MI355X KVM passthrough on Ubuntu 24.04 and RHEL 9.6, but SR-IOV support only on Ubuntu 24.04 — so a RHEL 9.6 shop that wants SR-IOV virtualisation, not just KVM passthrough, needs to add an Ubuntu 24.04 estate for that card. MI325X flips the pattern: KVM passthrough spans Ubuntu 24.04, Ubuntu 22.04, RHEL 9.6 and RHEL 9.4, but SR-IOV is validated only on Ubuntu 22.04. Get the OS-to-feature pairing wrong and a procurement decision that looked fine on paper stalls in the lab.
Distro breadth still favours the older card for now: ROCm 7.14.0 release notes list MI325X support across RHEL 10.0, SLES 15 SP7, Debian 13, Debian 12, Oracle Linux 10 and Oracle Linux 9 — six distributions, which is a wide net for UK regulated estates running varied Linux baselines. The same 7.14.0 notes don't extend that explicit list to MI355X, which is consistent with MI355X's more recent, narrower current support footprint. None of this is a permanent gap — these are version-specific snapshots, not fixed limitations — but it is the gap as documented today.
Beyond the version-and-distro matrices AMD publishes, a full ecosystem comparison would also weigh developer tooling, debugging workflows and the depth of third-party framework integration on each platform. AMD's published documentation for this piece is limited to compatibility matrices and MLPerf submissions rather than that broader tooling picture, so buyers evaluating day-to-day developer experience should treat ROCm's documented version support as a necessary but not complete picture of ecosystem readiness.
Total cost of ownership: where the public numbers run out
This is the point where an honest analysis has to stop rather than guess. No verified UK GBP list pricing for current-generation Instinct or Nvidia cards was found in AMD's published materials, which means any TCO comparison a buyer sees quoted in a vendor deck or blog post needs to be checked against a live quote, not treated as fact. What is verifiable is the hardware-density lever: MI355X's 288 GB of HBM3 memory and 8 TB/s of bandwidth can change how many GPUs are needed to hold a given large model, which is a real infrastructure and rack-density factor even without a published price per card.
Before requesting quotes from either vendor, size the actual cluster requirement to the workload — model size, precision, expected concurrency — using a tool such as the AI GPU Calculator, then ask both AMD and Nvidia channel partners for GBP-denominated, like-for-like pricing against that sizing rather than accepting a generic per-card list figure.
View the data behind this chart
| MI325X | MI355X | Buyer action | |
|---|---|---|---|
| Min ROCm version | 6.3.1+ | 7.0.1+ | Pin exact version |
| Architecture | gfx942 (CDNA3) | gfx950 (CDNA4) | Not interchangeable |
| Distro coverage | 6 distros (v7.14.0) | Not listed in v7.14.0 | MI325X wider net |
| SR-IOV OS support | Ubuntu 22.04 only | Ubuntu 24.04 only | Match SR-IOV OS first |
| Llama2-70B server | 32,028 tok/s (v5.1) | 100,282 tok/s (v6.0) | 3.1x AMD's own claim |
Workload-by-workload verdict
Pulling the compatibility matrices and MLPerf submissions together gives a defensible, non-speculative set of calls for where each card earns consideration in mid-2026:
- •Very large single-GPU inference where model weights approach the 288 GB HBM3 ceiling: MI355X is the stronger candidate, provided the buyer can commit to ROCm 7.0.1+ on the gfx950-supported stack.
- •Regulated or legacy Linux estates needing the broadest documented compatibility today: MI325X is the safer near-term pick, given its six-distro coverage in ROCm 7.14.0 and wider KVM passthrough footprint.
- •Extreme-scale multi-node inference: AMD's own published result of over 1 million tokens per second is real evidence, but only for teams operating at, and willing to replicate, that submission's scale and configuration.
- •Mixed-generation training clusters: AMD's MLPerf Training v6.0 round shows MI325X, MI350X and MI355X submitted together, plus a 512-GPU MI300X run — match the exact GPU, precision and node count to your workload rather than assuming parity across the family.
- •Deployments where Nvidia lead times are the binding constraint: AMD becomes genuinely credible as a fallback, but only for buyers prepared to verify the ROCm, distro and framework triple before signing procurement paperwork.
UK procurement reality: support, lead times and the verification step
For UK IT leaders, especially in public sector or other regulated environments, the practical decision hinges less on which vendor's card is faster in a vacuum and more on whether the target ROCm release, Linux distribution and framework version sit inside AMD's current, dated support window — because that documentation trail is what underpins predictable enterprise support later. Buyers who can standardise around ROCm 7.x on Ubuntu 24.04 or RHEL 9.6-class environments get access to MI355X's stronger compute profile; those who need the broadest documented compatibility today should treat MI325X as the safer commitment.
AMD's published evidence in 2026 is strongest on performance transparency and compatibility documentation, weaker on transparent list pricing, so UK procurement teams should still run parallel GBP quotes against Nvidia alternatives rather than pricing purely off benchmark claims. In practice, AMD is most credible in the UK where Nvidia availability or lead times are constraining a project timeline — though this piece does not have verified specific lead-time figures for either vendor's UK channel, so buyers should confirm current lead times directly with suppliers as part of the same quoting exercise. For teams ready to move to physical procurement, UK-available systems such as Supermicro AMD GPU Servers already ship configured for Instinct deployment, which shortens the path from compatibility-matrix verification to a working cluster.
Verdict: is AMD Instinct a real alternative in 2026?
Conditionally, yes. AMD's MI355X and MI325X are documented, MLPerf-benchmarked, currently supported hardware — not a roadmap promise. The 3.1x generational throughput jump and the 288 GB/8 TB/s memory profile on MI355X are real, sourced claims that change what a single card can hold and serve. But the deployment risk has moved from "does the silicon work" to "does my exact software stack match AMD's published support matrix on the day I need it in production" — and that verification step is non-negotiable, not optional due diligence.
Our recommendation: if you can standardise on a ROCm-supported configuration and validate benchmark parity for your specific model and precision, both MI355X and MI325X are legitimate options today, particularly where Nvidia lead times are the actual constraint. If you need the broadest packaged certainty right now with minimal in-house verification effort, Nvidia still wins on procurement familiarity. For the next generation of this argument, see how AMD MI455x vs Nvidia Rubin is shaping 2026 purchasing decisions before you commit multi-year budget to either roadmap.
Sources
Every figure in this article traces to the sources below.
- •AMD — ROCm system requirements matrix (MI355X/MI325X architecture support)
- •AMD — Instinct hardware support matrix (ROCm version requirements)
- •AMD — ROCm LLM Extension compatibility matrix
- •AMD — MLPerf Inference v6.0 results blog
- •AMD — MLPerf Inference v6.0 submission details
- •AMD — ROCm blog on MLPerf Training v6.0 submissions
- •MLCommons — MLPerf Training v6.0 supplemental discussion (Flux.1 FP8 on 64 MI325X GPUs)
- •AMD — ROCm 7.14.0 release notes (distro support)
- •AMD — ROCm simulation compatibility matrix
- •AMD — Cluster documentation hardware support (virtualisation)
View the data behind this chart
| Layer | Detail |
|---|---|
| FP4 precision support | Basis for AMD's MLPerf Inference v6.0 submission |
| 288 GB HBM3 memory | Per-GPU capacity for very large LLM weights |
| 8 TB/s memory bandwidth | Drives sustained tokens-per-second throughput |
| Over 1 million tokens/sec | AMD's published multi-node inference result |
