惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

GbyAI
GbyAI
人人都是产品经理
人人都是产品经理
Hugging Face - Blog
Hugging Face - Blog
罗磊的独立博客
博客园 - 【当耐特】
D
Docker
Y
Y Combinator Blog
L
LangChain Blog
博客园 - 三生石上(FineUI控件)
I
InfoQ
阮一峰的网络日志
阮一峰的网络日志
F
Fortinet All Blogs
J
Java Code Geeks
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
V2EX
B
Blog
The GitHub Blog
The GitHub Blog
腾讯CDC
MongoDB | Blog
MongoDB | Blog
博客园 - Franky
爱范儿
爱范儿
A
About on SuperTechFans
量子位
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC

The Register - Off-Prem

Enterprise cloud infrastructure uptake shows no sign of slowing The majority of corporate IT is now off premises for the first time Web app turns your old phone into a new smart display Anyone with a shed, an extension cord, a couple of GPUs and an overdraft is building datacenters. Fujitsu just offloaded five Iran says it Google Cloud outage shows it’s still hard to understand hyperscalers’ real resilience regimes AWS customer learns the hard way how even the smallest oversight can be mission-critical Billing software error sends billion-dollar AWS estimates Top EU court clips YouTube AWS CloudFront outage serves errors instead of websites India’s tech services giant HCL is getting into the AI datacenter business Britain Microsoft shifts to annual exchange rate price revision for cloudy products Amazon’s Mechanical Turk to stop accepting new customers – and not even AI can save it Fire burns Google Cloud India’s network, which remains slow a week later EU sovereignty push gives tech buyers a new alphabet soup to swallow Google, Canonical team up to certify Ubuntu images for TPU VMs Arm moves into the heart of the cloud stack Snowflake to burn $6B on AWS Graviton CPUs and AI accelerators Big Tech extracts retirement-scale wealth from UK internet users, research shows Open Compute urges local government to bask in the warm glow of excess datacenter heat Google Cloud suspended major customer Railway.com without cause, causing outage Broadcom finds a VMware customer willing to stick around: London Stock Exchange Baidu says the quiet part out loud – you can’t build AI infrastructure, so clouds can cash in AWS racks M3 Ultra Macs that boast specs you can’t currently buy Tencent admits GPUs only pay for themselves when powering personalized ads Red Hat blasts RHEL 10.1 into orbit aboard Voyager's micro datacenter Sovereign cloud is only possible if you’re Chinese or American: Gartner Cloudflare to fire 1,100 staff whose jobs just aren’t AI enough AWS warns of EC2 'impairment' as power loss hits notorious US-EAST-1 region
Microsoft fiber foul-up cut off Azure California for almo...
Simon Sharwood · 2026-07-24 · via The Register - Off-Prem

off-prem

Maintenance mistake took out 27 services

Azure users who use resources in Microsoft’s West US region endured an uncomfortable day after the tech giant cut off access to its Californian cloud outpost.

As Microsoft explains in its preliminary post incident review, at 14:44 UTC on July 23rd (07:44 AM Pacific Time), the company started “routine device maintenance.”

That effort immediately produced problems.

“Multiple Azure services began to detect and correlate service degradation,” Microsoft’s review reads.

A minute later, Microsoft says folks from its networking and services teams, plus people it describes as “incident responders” began “reviewing traffic anomalies, routing behavior, packet loss signals, and recent changes.”

Microsoft says this sort of maintenance job requires isolating specific network paths, and that it uses a process that “converts these requests into system-readable requests and verifies that at least one of the two redundant paths remains healthy.” We’re told the company makes “safety checks to confirm the work will be impact-less.”

The mere fact you are reading this story shows that process instead delivered an unintended impact.

“In this case, a bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended,” Microsoft confessed. “The routes were removed between our datacenter and wide-area network, impacting traffic entering or exiting the region.”

Microsoft’s incident report says the issue “initially presented as large-scale route churn in our Wide-Area Network” and that its investigations “later found that route removal was from a datacenter in the West US region.”

Some time between 16:00 UTC and 17:45 UTC Microsoft identified “recent fiber maintenance activity, which we correlated to the identified routing behavior.”

At 17:45 the company started rolling back changes, and by 18:26 Microsoft’s WAN was back to its best.

By 19:41 UTC, all impacted services had fully recovered – and Microsoft joined AWS and Google in offering recent examples of how clouds can be rather more fragile than advertised. ®