A routine maintenance job on Microsoft's West US network turned into a near five-hour Azure blackout on 23 July, taking 27 services offline. For UK infrastructure buyers, the incident is another reminder to strengthen your backup and disaster recovery strategy before the next fault, not after it.
View the data behind this chart
| Phase | Starts (week) | Duration (weeks) |
|---|---|---|
| Maintenance begins… | 0 | 5 |
| Root cause traced to fibre… | 5 | 7 |
| Rollback of route changes… | 12 | 3 |
| WAN restored, services… | 15 | 5 |
What actually broke in West US
According to Microsoft's preliminary post-incident review, the trouble began at 14:44 UTC on 23 July when engineers kicked off what the company called routine device maintenance in its West US region, which spans northern and central California including the Bay Area, San Jose, Los Angeles and Santa Clara.
Within a minute, Microsoft's networking and incident-response teams were already chasing traffic anomalies and packet loss. The root cause, Microsoft confessed, was a bug in the request-conversion system that isolates network paths for maintenance: it incorrectly marked extra devices as part of the job and stripped IP routes from more equipment than intended, cutting traffic entering and leaving the region.
Microsoft's team initially saw the fault as large-scale route churn across its wide-area network before tracing it back to a datacenter in West US, correlating it with recent fibre maintenance activity. Rollback started at 17:45 UTC, the WAN was stable by 18:26 UTC, and full recovery across all affected services landed at 19:41 UTC — a disruption lasting almost five hours from a single mismarked maintenance request.
Why a California network fault matters to UK buyers
It's tempting to file this under 'not our region, not our problem'. That would be a mistake. The Cyber Monitoring Centre has estimated that a 24-hour outage affecting major AWS or Azure regions serving the UK, Ireland, Europe or the eastern US could cost UK organisations somewhere between £650 million and £1 billion in direct revenue alone, before any downstream costs are counted.
This West US incident lasted a fraction of that window, but it demonstrates the mechanism that would produce those losses: a single maintenance bug removing routes from far more devices than intended, with 27 dependent services degrading almost simultaneously. If your UK workloads, backups, or failover targets sit behind Azure's networking or identity layers anywhere in the world, that blast radius is relevant to you.
This is exactly the kind of scenario that should feed into how you calculate the real cost of IT downtime for your own estate — not as an abstract exercise, but using the actual duration and service counts Microsoft has now published.
The pattern behind the outage
This is not an isolated West US event. Microsoft's own incident history shows a February 2026 power interruption in the same region, caused by an electrical failure in an on-site transformer, where utility power was actually still functioning but the datacenter lost it anyway. Separately, a Managed Identity issue in February 2026 knocked out both East US and West US for almost six hours, a clear example of a single shared dependency crossing supposedly independent regions.
Older incidents follow the same logic: a 2018 lightning strike caused power and cooling failures in South Central US that hit around 40 Azure services, with knock-on effects on Azure Active Directory and Azure Resource Manager well outside the affected region. A separate West Europe outage traced to a misconfigured network device disrupted one cluster for roughly two and a half hours. And in October 2025, a global Azure Front Door outage was serious enough that Microsoft advised customers to redirect traffic via Azure Traffic Manager while it reverted to a last-known-good configuration.
The common thread across all of these: control planes, identity systems and networking layers are shared far more widely than most DR diagrams admit.

What to audit in your DR plan this quarter
Microsoft's own Azure Site Recovery guidance is explicit that disaster recovery should use a secondary region, a Recovery Services vault, defined replication policies, a secondary virtual network, and regular test failovers — not a same-region or same-zone replica dressed up as DR.
Start your audit by confirming your replica genuinely sits in a different Azure region, not just a different availability zone within the same one. Then check whether that secondary region shares an identity service, a control plane, or a networking fabric with your primary — the February 2026 Managed Identity incident shows that 'different region' doesn't always mean 'different failure domain'.
Test failover matters just as much as design. Untested DR plans routinely fail under live conditions, and Azure Site Recovery's own documentation stresses validating failover without touching production. If you haven't run a full test failover this year, that's the first gap to close — and a good moment to understand RTO and RPO in disaster recovery terms your board can actually interrogate.
- •Confirm replicas sit in a genuinely separate Azure region, not just a separate zone
- •Map shared dependencies: identity, control plane, and WAN routing across your primary and secondary regions
- •Schedule a test failover this quarter rather than assuming last year's design still holds
- •Quantify your exposure using real incident durations, not worst-case guesses
View the data behind this chart
| Single-Region DR | Multi-Region ASR | Tested Multi-Reg… | |
|---|---|---|---|
| Regional fibre fault | Fails | Can survive | Can survive |
| Shared identity risk | Fails | Still exposed | Reduced risk |
| Untested failover | High risk | High risk | Low risk |
| RTO/RPO confidence | Unverified | Assumed | Verified |
| UK outage cost risk | £650m-£1bn/day | £650m-£1bn/day | Lower exposure |
Beyond the cloud provider: building your own resilience layer
No amount of Azure region diversity removes the value of an independent backup layer that doesn't depend on the same cloud vendor's control plane at all. UK buyers with Microsoft 365 workloads in particular should compare Microsoft 365 backup options rather than assuming platform-native retention covers a regional networking fault.
It's also worth reviewing whether your immutable backup approach can survive a scenario where the primary cloud region is unreachable for hours rather than minutes — this is where organisations look to implement immutable backup architectures that sit outside the affected provider's blast radius entirely.
Finally, if your recovery strategy still relies on ageing on-site hardware to bridge a failover window, it's worth weighing whether to compare cloud vs. on-premise TCO for critical workloads, and whether extending the life of existing kit through third-party maintenance for your hardware buys you time to redesign properly rather than reacting after the next outage.
- 01The Register — Microsoft fiber foul-up cut off Azure California for almost five hours · 24 July 2026
- 02The Register — EU and UK organizations ponder resilience after Azure outage · 30 October 2025
- 03The Register — Azure outages ripple across multiple dependent services · 3 February 2026
- 04DataCenterDynamics — Microsoft suffered power interruption at West US cloud region in February · 1 February 2026
- 05DataCenterDynamics — Microsoft publishes preliminary analysis of major cloud outage · 30 October 2025
- 06ComputerWeekly — UK's largest businesses dangerously exposed to cloud outages · 1 November 2025
- 07DataCenterDynamics — Microsoft misconfigured network device led to Azure outage · 1 January 2023
- 08DataCenterDynamics — Microsoft Azure suffers outage after cooling issue · 1 September 2018
- 09IBM — How to set up a multi-region DR architecture on Azure with Site Recovery · 1 January 2024
- 10The Register — Microsoft Azure challenges AWS for downtime crown · 29 October 2025
- 11DataCenterDynamics — Isolated power event causes small Azure outage in West Central US cloud region · 1 February 2026
