UK’s trusted IT infrastructure partner since 2003
Servnet
FinanceToolsConfiguratorGet in Touch
Backup & DR

Azure West US Outage 2026: What UK Buyers Must Fix

London · Servnet News Desk · IT infrastructure analysis5 min read
Share

A routine fibre maintenance job knocked Azure's West US region offline for five hours, and Microsoft's own post-incident review shows why multi-region cloud alone isn't the safety net many UK buyers assume it is. Here's what actually broke, and what it should change in your disaster recovery thinking.

Microsoft's West US-linked outages in 2026
10 hrs8 hrs5 hrs3 hrs0 hrs10 hrsFeb 20265 hrsJul 2026Outage duration
View the data behind this chart
Microsoft's West US-linked outages in 2026
Feb 2026Jul 2026
Outage durationhrs10hrs5

What Microsoft says actually happened

According to Microsoft's Preliminary Post Incident Review, connectivity to its West US cloud region was lost for five hours on 23 July, between 14:44 UTC and 19:41 UTC. The trigger was mundane: engineers were isolating a device for routine maintenance and had confirmed at least one of the two redundant network paths would stay operational before starting work.

The failure came from automation, not from the fibre itself. A bug in the request conversion system incorrectly pulled additional devices into the maintenance perimeter, and IP routes sitting between the datacentre and the wide-area network were removed from more equipment than intended. Microsoft first saw this as large-scale route churn on its WAN before correlating it with what it called recent fibre maintenance activity.

Crucially, workloads running entirely inside the West US region kept working. It was any traffic entering or leaving the facility — the exact traffic pattern that multi-region failover, hybrid connectivity and customer access all depend on — that broke.

A five-hour outage, and the second big one this year

The Register reports the incident affected 27 services before full recovery at 19:41 UTC. Engineers spotted the issue within the first hour and began reconnecting services, but full restoration still took most of a working day for affected customers.

This is not an isolated event. In February, Microsoft disclosed a separate 10-hour disruption across US West and US East after a transformer electrical failure caused loss of utility power to a datacentre, with a cascading control-system fault stopping automatic transfer to generators until UPS batteries ran down. Two significant West US-linked outages inside a single year is a pattern, not a one-off.

Why multi-region failover didn't fully save the day

Microsoft's standard advice after an incident like this is to design mission-critical workloads across multiple regions. That's sound guidance, but this outage exposes its limit: the fault sat in the WAN path connecting the datacentre to the outside world, meaning any failover traffic trying to leave the affected region was hit by the same problem as ordinary customer traffic.

In other words, multi-region architecture protects you from a regional compute or power failure, but it does not automatically protect you from a shared connectivity chokepoint, a misapplied maintenance script, or an automation bug that misclassifies which devices are in scope. UK teams that have ticked the box on 'multi-region resilience' should ask a harder question: does our failover path share any physical fibre, router, or automation system with the primary path? If the answer is yes, the redundancy is theoretical.

Illustration: Azure West US Outage 2026: What UK Buyers Must Fix

This is a pattern across Microsoft's cloud, not a freak event

Widen the lens and the West US incident sits alongside a run of very different Microsoft cloud failures over the past two years: a Western Europe outage traced to a misconfigured network device, a global Azure/Microsoft 365 disruption rolled back after a recent WAN change, a separate incident blamed on a DDoS attack combined with a defence-implementation error, and a March 2024 South Africa outage linked to multiple subsea fibre cuts on the African coast. Microsoft also warned in September 2025 that Red Sea subsea cable damage was forcing Azure traffic onto alternate, higher-latency routes.

The common thread isn't fibre, or software, or power — it's that Microsoft's failure modes are varied and recurring: routing mistakes, device misconfiguration, power-system cascades, DDoS mitigation errors, and physical cable damage have all taken down customer-facing services in the last two years. Any DR strategy built on the assumption that a hyperscaler has 'solved' this class of risk is working from an outdated premise.

What this means for UK disaster recovery planning

For infrastructure buyers, the practical takeaway isn't to abandon Azure or any single hyperscaler — it's to stop treating cloud redundancy as a substitute for tested, independent recovery capability. A few concrete steps follow directly from this incident.

First, quantify what an outage like this actually costs your organisation before you decide how much resilience to buy; you can calculate the true cost of downtime in pounds per hour rather than working from assumption. Second, revisit your RTO and RPO targets and test whether they still hold if your 'secondary' region shares WAN infrastructure with your primary — this is exactly where teams should understand disaster recovery RTO and RPO in plain terms rather than vendor marketing language.

Third, weigh whether some workloads belong back closer to home. Organisations that kept an on-premise or colocated backup path were able to keep operating during this window regardless of what was happening to Azure's WAN; it's worth reading how firms are approaching this via cloud repatriation and running the numbers to compare cloud and on-premise TCO. Fourth, check your Microsoft 365 data protection separately from infrastructure DR — an Azure networking failure and a Microsoft 365 data-loss event are different risks requiring different insurance, and it's worth reviewing a Microsoft 365 backup comparison alongside your broader backup & disaster recovery posture.

DR approaches after a shared-path cloud failure
Single-Region…Multi-Region…Hybrid + On-PremFailover riskHighMed-HighLow-MedWAN dependencyTotalPartialReducedRecovery controlVendor-onlyVendor-onlyIn-house optionCost profileLowestHigherHighest upfrontData locationCloud regionMulti-regionUK-controlled
View the data behind this chart
DR approaches after a shared-path cloud failure
Single-Region…Multi-Region…Hybrid + On-Prem
Failover riskHighMed-HighLow-Med
WAN dependencyTotalPartialReduced
Recovery controlVendor-onlyVendor-onlyIn-house option
Cost profileLowestHigherHighest upfront
Data locationCloud regionMulti-regionUK-controlled

Building resilience that doesn't rely on one vendor's WAN

None of this argues against cloud adoption; it argues for architecture that assumes any single connectivity path, including a hyperscaler's, can fail. That means genuinely independent failover routes, hardware maintained on a schedule you control rather than one that depends on a third party's automation, and a tested plan for operating during extended connectivity loss rather than just data loss.

If ageing or single-sourced network hardware is part of your exposure, it's worth understanding how third-party maintenance options can extend support and reduce dependency on a single supplier's maintenance schedule, and how a Veeam-based backup layer sits alongside cloud-native protection. For organisations still building their overall approach, it's a good moment to choose a UK disaster recovery provider that can demonstrate independence from any one hyperscaler's network path.

Share
Key takeaways
  • Microsoft's PIR confirms a five-hour West US outage (14:44–19:41 UTC, 23 July) was caused by an automation bug removing IP routes during routine fibre maintenance, not a hardware failure alone.
  • The outage hit any traffic entering or leaving the region — exactly the path multi-region failover depends on — showing shared WAN infrastructure can defeat regional redundancy.
  • This is the second major Microsoft cloud outage in 2026 after a 10-hour February disruption, and part of a wider pattern including misconfiguration, DDoS and subsea fibre incidents since 2024.
  • UK buyers should test whether their 'secondary region' truly avoids shared network paths, and quantify downtime cost before deciding how much independent, non-cloud recovery capability to keep.
Frequently asked

FAQs — Azure West US Outage 2026

What actually caused the Azure West US outage in July 2026?

Microsoft's Preliminary Post Incident Review says a bug in its request conversion system incorrectly included extra devices in a routine maintenance isolation, removing IP routes between the datacentre and the wide-area network for devices not meant to be affected.

How long did the Azure West US outage last?

Connectivity was lost for five hours, from 14:44 UTC to 19:41 UTC on 23 July, affecting 27 services according to The Register's reporting; engineers identified the fault and began reconnecting services within the first hour.

Does multi-region Azure architecture prevent this kind of outage?

Not fully. Because the fault affected traffic entering or leaving the region rather than workloads running inside it, failover traffic could hit the same connectivity problem. UK teams should understand disaster recovery RTO and RPO in terms of shared infrastructure, not just region count.

Was this Microsoft's only major cloud outage in 2026?

No. Microsoft also disclosed a separate 10-hour disruption in February 2026 across US West and US East, caused by a transformer electrical failure and a cascading control-system fault that stopped automatic generator transfer.

Related

Continue reading

More in Backup & DR

Turning this into a buying decision?

One conversation with an engineer who's specced this before. No sales script.

Talk to Servnet →

Talk to a UK specialist

Get expert advice or a no-obligation quote — servers, storage, networking, maintenance, finance and cloud. We reply the same working day.

or call 0800 987 4111