When Meta trained Llama 3 405B on 16,384 NVIDIA H100 GPUs over 54 days, the cluster logged 419 unexpected interruptions - roughly one every three hours - with GPU faults behind 30.1% of them and HBM3 memory faults behind another 17.2%. That is the most granular public failure data available for H100-class hardware, and it exposes a gap in most GPU procurement conversations: buyers get quoted a warranty term, not an interruption rate. This piece works from Meta's disclosure and three follow-up 2026 fleet papers to show what current-generation GPUs actually do at scale, why the popular "1-3 year lifespan" line doesn't answer a budgeting question, and how a UK buyer should size spares, SLAs and support cover around GPU-and-HBM failures rather than generic hardware warranty.
View the data behind this chart
| Llama3 GPU faults | Llama3 HBM3 faults | 100k-GPU HW total | 100k-GPU GPU-only | |
|---|---|---|---|---|
| Share of interruptions | %30.1 | %17.2 | %78 | %7.4 |
Lifespan is the wrong question: physical, economic and accounting realities
Three different "lifespans" get conflated in almost every GPU procurement conversation. Physical lifespan is how long the silicon and its memory package can keep operating before it genuinely wears out. Economic lifespan is how long it stays worth running versus replacing it with a faster, more efficient generation. Accounting lifespan is whatever depreciation schedule finance has written down, regardless of whether the card still works or still makes economic sense. A vendor quoting one of these numbers is not necessarily quoting the others, and a widely repeated "1-3 year lifespan" figure for AI GPUs is frequently a mix of all three, stated as if it were a single hard fact.
No verifiable public dataset reviewed for this piece quantifies enterprise GPU physical lifespan in years. What does exist, and what is far more useful for planning, is Meta's disclosed interruption data from a real, sustained training run - a measured failure cadence rather than a theoretical expiry date. If you're deciding whether to run an existing fleet longer or refresh it, that residual-value question deserves its own analysis; our GPU residual value study and the vocabulary in our server warranty and maintenance SLA guide are the better starting points than any generic lifespan claim.

Hard numbers: what Meta's Llama 3 run and later fleet papers actually show
Meta trained Llama 3 405B on 16,384 H100 80GB GPUs over 54 days and recorded 419 unexpected interruptions - about one every three hours. Of those, 148 interruptions (30.1%) were attributed to faulty GPUs and 72 (17.2%) to GPU HBM3 memory faults. This is the most granular fleet-scale reliability disclosure publicly available for H100-class hardware, and it is the baseline UK buyers should reference instead of anecdotal lifespan claims.
Independent reviews of that same disclosure read the "hardware" share differently depending on what they bundle together. An EDN summary of Meta's findings put hardware-related causes - GPU and host-component issues combined - at 78% of interruptions. A separate 2026 ZenML analysis of Meta's reliability framework put failures across SRAMs, HBMs, processing grids and network-switch hardware at over 66%. Neither figure is wrong; they are counting different combinations of components, which is exactly why buyers should ask any supplier "hardware share of what?" before comparing two failure statistics as if they were the same metric.
A different, larger dataset adds the most important forward-looking signal. A 2026 arXiv paper on fault-tolerant training across more than 100,000 GPUs found hardware-related failures still caused 78% of interruptions, but the mix had shifted: faulty GPU compute had fallen to just 7.4% of interruptions and was only the fourth-largest cause, while HBM issues had become the top cause in that newer system. A separate 2026 OpenReview paper on large training jobs likewise found hardware causes - GPU and host-component issues plus unplanned maintenance events - responsible for over 70% of interruptions. Read together, across different clusters and different years, the pattern holds: hardware, not software, drives most interruptions, and as raw GPU compute gets more reliable, memory and host components take over as the leading failure source.
- •Meta Llama 3 (16,384 H100 GPUs, 54 days): 419 interruptions - GPU 30.1%, HBM3 17.2%
- •EDN review of Meta's disclosure: 78% hardware-related (GPU + host component)
- •ZenML 2026 review of the same framework: over 66% hardware-related (SRAM, HBM, processing grid, network switch)
- •arXiv 100,000+-GPU HSDP paper (2026): 78% hardware-related; GPU-compute-only share fallen to 7.4%, HBM now the top cause
- •OpenReview 2026 paper on large training jobs: hardware causes (GPU/host/maintenance) over 70% of interruptions
Beyond the chip: workload, utilisation and environment
Meta's numbers come from a genuinely worst-case reliability workload: continuous, synchronous training across tens of thousands of GPUs running near full power draw with no idle windows, for 54 days straight. That is a categorically different stress profile from bursty inference, where accelerators cycle between load and idle and rarely sustain peak thermal and power delivery for weeks at a time. Sustained, full-utilisation training is what surfaces GPU and HBM faults at the cadence Meta recorded; it is also exactly the workload where the industry has invested most in fault handling, which is why Meta built purpose-designed silent-data-corruption detection and an automated error-handling stack around the Llama 3 pipeline rather than relying on standard warranty support.
Cooling and power quality sit underneath all of this. Sustained full-load training racks put continuous thermal and electrical stress on the GPU package and its HBM stack, and UK operators weighing air versus liquid cooling for new GPU halls should factor this reliability angle in alongside the energy-efficiency case - see our companion piece on how to optimise AI server cooling for dense GPU racks. The practical read-across for spares planning: any UK estate running sustained, synchronous multi-GPU training should expect interruption rates closer to Meta's disclosed cadence than to lighter, intermittent inference deployments, and should size spare pools accordingly.
Vendor and hyperscaler perspectives
Nvidia, AMD and the major cloud hyperscalers rarely publish raw component MTBF or interruption-level reliability figures at fleet scale. Instead, vendor and hyperscaler public disclosures focus on software mitigation layers and diagnostic tooling: Nvidia promotes Data Center GPU Manager (DCGM) health checks and Base Command automated node remediation, while hyperscalers like AWS and Microsoft Azure frame enterprise reliability around infrastructure availability SLAs, predictive node draining, and VM-level automated recovery rather than bare-metal component survival rates. Meta remains the most transparent public discloser of GPU-and-HBM-specific fleet failure data, which is why its raw interruption counts have become the de facto reference point even for on-premises enterprise buyers.
This vendor framing confirms an operational reality: hyperscalers and OEMs treat hardware interruptions as inevitable at scale, engineering around them with telemetry and automation rather than pretending silicon never fails. That combination - continuous health detection plus automated node evacuation and advance hardware swap - is the template UK buyers should look for in any OEM or third-party maintenance package, aligning service levels with how datacentre hardware actually behaves in production.
What UK compliance frameworks mean for GPU failure risk
No verified UK GBP failure-cost or GPU-price figure accompanies the reliability data reviewed here, so treat any specific per-incident cost claim you encounter elsewhere with scepticism unless it cites a dated, UK-specific source. What UK buyers can act on directly is the compliance and procurement framing. Operators in scope of the UK's Network and Information Systems Regulations 2018 must take appropriate and proportionate technical and organisational measures to manage risk to their network and information systems - a GPU fleet's measured interruption rate is precisely the kind of operational risk that framing is designed to capture. Large UK undertakings in scope of the Energy Savings Opportunity Scheme (ESOS) also face energy audits every four years, a natural checkpoint to re-cost cooling, power quality and spares strategy for GPU halls at the same time.
On procurement, UK buyers can source NVIDIA datacentre GPUs and support through UK resellers such as Scan.co.uk, but list prices for datacentre-class cards are typically quote-based rather than published, so budget conversations need to happen configuration by configuration rather than off a public price list. The decision that matters more than any single unit price is the one between OEM warranty cover and independent support: our server warranty and maintenance SLA comparison sets out exactly what each model does and doesn't cover once a GPU fleet is out of its first year.
View the data behind this chart
| Cluster / scope | What was measured | Hardware-related share | |
|---|---|---|---|
| Meta Llama 3 (2024) | 16,384 H100, 54 days | GPU vs HBM3 faults | GPU 30.1%, HBM3 17.2% |
| 100k-GPU HSDP (2026) | 100,000+ GPU training | All interruption causes | 78% HW; GPU-only 7.4% |
| Meta stack review (2025) | Llama 3 training stack | GPU/host component issues | 78% hardware-related |
Strategies for maximising GPU uptime and extending useful life
Translate Meta's cause breakdown directly into a spares and support spec. Because GPU and HBM faults dominate the disclosed interruptions, cover should be written for the GPU and its HBM stack explicitly, not just "the server" as a unit - many standard hardware warranties are written at board or chassis level and don't guarantee GPU/HBM-specific advance replacement. UK buyers should also insist on UK-based replacement logistics and stock, not offshore depot turnaround, given how frequently these faults recur at fleet scale.
- •Hold spare GPUs and HBM-adjacent components sized to your actual interruption cadence, not a generic "5% spares" rule of thumb (for example, scaling Meta's 148 GPU faults over 54 days for a 1,000-GPU estate running sustained training implies roughly 5 GPU replacement events per month at similar utilisation - scale from your own MTBF data)
- •Specify advance-replacement and on-site swap in the SLA, not just next-business-day depot repair
- •Push for the same silent-data-corruption detection and automated fault-handling Meta built into its own stack, from the OEM or via third-party maintenance providers who monitor at that resolution
- •Revisit buy-new-vs-support-existing each budget cycle using a residual-value view via our GPU residual value study, rather than a fixed depreciation assumption
Where reliability goes next
A suggestive pattern in the newer data is that GPU-compute reliability may be improving faster than memory reliability. In Meta's Llama 3 run, GPU faults were the single largest recorded cause at 30.1%. In the 2026 100,000-GPU HSDP paper, GPU-compute faults had fallen to just 7.4% and ranked only fourth, while HBM issues had become the leading cause. However, because these data points reflect different fleet sizes, study years, and potentially different GPU architectures, this shift is an indicator rather than a proven causal trend across identical hardware. For spares planning, that means the old assumption - "stock more GPUs" - is already dated for the newest fleets; the growth item in a 2026 spares pool is HBM-adjacent replacement capacity and memory-level silent-data-corruption tooling, not just GPU cards.
Conclusion: plan for interruptions, not a mythical lifespan
Buyers searching for "GPU failure rate" are usually trying to answer a budgeting question - how many spares, what SLA, when to refresh - and the generic "1-3 year lifespan" framing doesn't answer any of it, because no verifiable public dataset reviewed here quantifies GPU physical lifespan in years. What Meta's disclosure and the follow-up 2026 fleet papers do answer, precisely, is how often GPUs and their HBM stacks generate unplanned interruptions at real operating scale, and where that failure locus is moving. Size your spares pool and your support SLA around that evidence, specify GPU-and-HBM-level cover explicitly, and revisit the assumption every budget cycle as the underlying failure mix keeps shifting.
Sources
Every figure in this article traces to the sources below.
- •Data Center Dynamics — Meta Llama 3 interruption counts, GPU and HBM3 fault shares
- •Meta — official statement on hardware reliability and automated error-handling stack
- •ZenML — 2026 analysis of Meta's hardware reliability framework
- •EDN — summary of Meta's hardware failure and silent-data-corruption findings
- •OpenReview — 2026 paper on hardware-driven interruptions in large training jobs
- •arXiv — 2026 paper on fault-tolerant HSDP training across 100,000+ GPUs
- •UK Government — Energy Savings Opportunity Scheme (ESOS) guidance
- •UK Government — Network and Information Systems Regulations 2018
- •Scan.co.uk — UK reseller market for NVIDIA datacentre GPU procurement
View the data behind this chart
| Layer | Detail |
|---|---|
| GPU compute faults | 30.1% of Meta's Llama 3 interruptions (2024) |
| HBM3 memory faults | 17.2% of Meta's Llama 3 interruptions (2024) |
| Host, network & other hardware | Remaining causes flagged across Meta and later fleet papers |
