A routine fibre maintenance job knocked Azure's West US region offline for five hours, and Microsoft's own post-incident review shows why multi-region cloud alone isn't the safety net many UK buyers assume it is. Here's what actually broke, and what it should change in your disaster recovery thinking.
View the data behind this chart
| Feb 2026 | Jul 2026 | |
|---|---|---|
| Outage duration | hrs10 | hrs5 |
What Microsoft says actually happened
According to Microsoft's Preliminary Post Incident Review, connectivity to its West US cloud region was lost for five hours on 23 July, between 14:44 UTC and 19:41 UTC. The trigger was mundane: engineers were isolating a device for routine maintenance and had confirmed at least one of the two redundant network paths would stay operational before starting work.
The failure came from automation, not from the fibre itself. A bug in the request conversion system incorrectly pulled additional devices into the maintenance perimeter, and IP routes sitting between the datacentre and the wide-area network were removed from more equipment than intended. Microsoft first saw this as large-scale route churn on its WAN before correlating it with what it called recent fibre maintenance activity.
Crucially, workloads running entirely inside the West US region kept working. It was any traffic entering or leaving the facility — the exact traffic pattern that multi-region failover, hybrid connectivity and customer access all depend on — that broke.
A five-hour outage, and the second big one this year
The Register reports the incident affected 27 services before full recovery at 19:41 UTC. Engineers spotted the issue within the first hour and began reconnecting services, but full restoration still took most of a working day for affected customers.
This is not an isolated event. In February, Microsoft disclosed a separate 10-hour disruption across US West and US East after a transformer electrical failure caused loss of utility power to a datacentre, with a cascading control-system fault stopping automatic transfer to generators until UPS batteries ran down. Two significant West US-linked outages inside a single year is a pattern, not a one-off.
Why multi-region failover didn't fully save the day
Microsoft's standard advice after an incident like this is to design mission-critical workloads across multiple regions. That's sound guidance, but this outage exposes its limit: the fault sat in the WAN path connecting the datacentre to the outside world, meaning any failover traffic trying to leave the affected region was hit by the same problem as ordinary customer traffic.
In other words, multi-region architecture protects you from a regional compute or power failure, but it does not automatically protect you from a shared connectivity chokepoint, a misapplied maintenance script, or an automation bug that misclassifies which devices are in scope. UK teams that have ticked the box on 'multi-region resilience' should ask a harder question: does our failover path share any physical fibre, router, or automation system with the primary path? If the answer is yes, the redundancy is theoretical.

This is a pattern across Microsoft's cloud, not a freak event
Widen the lens and the West US incident sits alongside a run of very different Microsoft cloud failures over the past two years: a Western Europe outage traced to a misconfigured network device, a global Azure/Microsoft 365 disruption rolled back after a recent WAN change, a separate incident blamed on a DDoS attack combined with a defence-implementation error, and a March 2024 South Africa outage linked to multiple subsea fibre cuts on the African coast. Microsoft also warned in September 2025 that Red Sea subsea cable damage was forcing Azure traffic onto alternate, higher-latency routes.
The common thread isn't fibre, or software, or power — it's that Microsoft's failure modes are varied and recurring: routing mistakes, device misconfiguration, power-system cascades, DDoS mitigation errors, and physical cable damage have all taken down customer-facing services in the last two years. Any DR strategy built on the assumption that a hyperscaler has 'solved' this class of risk is working from an outdated premise.
What this means for UK disaster recovery planning
For infrastructure buyers, the practical takeaway isn't to abandon Azure or any single hyperscaler — it's to stop treating cloud redundancy as a substitute for tested, independent recovery capability. A few concrete steps follow directly from this incident.
First, quantify what an outage like this actually costs your organisation before you decide how much resilience to buy; you can calculate the true cost of downtime in pounds per hour rather than working from assumption. Second, revisit your RTO and RPO targets and test whether they still hold if your 'secondary' region shares WAN infrastructure with your primary — this is exactly where teams should understand disaster recovery RTO and RPO in plain terms rather than vendor marketing language.
Third, weigh whether some workloads belong back closer to home. Organisations that kept an on-premise or colocated backup path were able to keep operating during this window regardless of what was happening to Azure's WAN; it's worth reading how firms are approaching this via cloud repatriation and running the numbers to compare cloud and on-premise TCO. Fourth, check your Microsoft 365 data protection separately from infrastructure DR — an Azure networking failure and a Microsoft 365 data-loss event are different risks requiring different insurance, and it's worth reviewing a Microsoft 365 backup comparison alongside your broader backup & disaster recovery posture.
View the data behind this chart
| Single-Region… | Multi-Region… | Hybrid + On-Prem | |
|---|---|---|---|
| Failover risk | High | Med-High | Low-Med |
| WAN dependency | Total | Partial | Reduced |
| Recovery control | Vendor-only | Vendor-only | In-house option |
| Cost profile | Lowest | Higher | Highest upfront |
| Data location | Cloud region | Multi-region | UK-controlled |
Building resilience that doesn't rely on one vendor's WAN
None of this argues against cloud adoption; it argues for architecture that assumes any single connectivity path, including a hyperscaler's, can fail. That means genuinely independent failover routes, hardware maintained on a schedule you control rather than one that depends on a third party's automation, and a tested plan for operating during extended connectivity loss rather than just data loss.
If ageing or single-sourced network hardware is part of your exposure, it's worth understanding how third-party maintenance options can extend support and reduce dependency on a single supplier's maintenance schedule, and how a Veeam-based backup layer sits alongside cloud-native protection. For organisations still building their overall approach, it's a good moment to choose a UK disaster recovery provider that can demonstrate independence from any one hyperscaler's network path.
- 01Network World — Microsoft explains why its West US Azure and cloud services failed · 24 July 2026
- 02The Register — Microsoft fiber foul-up cut off Azure California for almost five hours · 24 July 2026
- 03DatacenterDynamics — Microsoft suffered power interruption at West US cloud region in February · 8 February 2026
- 04DatacenterDynamics — Microsoft Azure's March outage in South Africa due to subsea cable cuts · 15 March 2024
- 05The Register — Red Sea submarine cable outage slows Microsoft cloud · 8 September 2025
- 06DatacenterDynamics — Microsoft misconfigured network device led to Azure outage · 30 January 2024
- 07BleepingComputer — Microsoft says massive Azure outage was caused by DDoS attack · 1 August 2024
