惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Apple Machine Learning Research
Apple Machine Learning Research
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
T
Tenable Blog
P
Proofpoint News Feed
爱范儿
爱范儿
The Hacker News
The Hacker News
W
WeLiveSecurity
WordPress大学
WordPress大学
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
The Register - Security
The Register - Security
I
Intezer
S
SegmentFault 最新的问题
P
Privacy & Cybersecurity Law Blog
人人都是产品经理
人人都是产品经理
Security Archives - TechRepublic
Security Archives - TechRepublic
G
Google Developers Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
G
GRAHAM CLULEY
MongoDB | Blog
MongoDB | Blog
Blog — PlanetScale
Blog — PlanetScale
Hugging Face - Blog
Hugging Face - Blog
P
Palo Alto Networks Blog
A
About on SuperTechFans
Attack and Defense Labs
Attack and Defense Labs
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
D
Darknet – Hacking Tools, Hacker News & Cyber Security
博客园_首页
SecWiki News
SecWiki News
博客园 - 聂微东
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Hacker News - Newest:
Hacker News - Newest: "LLM"
Last Week in AI
Last Week in AI
Forbes - Security
Forbes - Security
C
Cyber Attacks, Cyber Crime and Cyber Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
F
Full Disclosure
Help Net Security
Help Net Security
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
P
Proofpoint News Feed
S
Schneier on Security
K
Kaspersky official blog
S
Secure Thoughts
C
Cisco Blogs
T
Tor Project blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
L
LangChain Blog
Hacker News: Ask HN
Hacker News: Ask HN
N
News and Events Feed by Topic
MyScale Blog
MyScale Blog

The Register - Special Features: Datacenter Networking Nexus

How Broadcom is quietly invading AI infrastructure Cisco punts network-security integration as key for agentic A trip through vintage datacenter networking The network is indeed trying to become the computer Cisco fixes two critical make-me-root bugs AI could finally see DPUs take off in enterprise networks An introduction to rack-scale networking HPE Aruba touts new AI agents and network orchestrator Microsoft to retire default outbound access for VMs in Azure Same suspected Chinese spies again attacking Ivanti bugs Hyperconverged infrastructure now needs liquid cooling Asia reaches 50 percent IPv6 capability Rising demand for datacenter capacity sees prefabs sprout The No-Nvidia networking club delivers first spec Nvidia punts silicon photonic switches to keep GPUs fed Chinese snoops spotted on end-of-life Juniper routers
Human error and power glitches to blame for most outages
Dan Robinson Dan Robinson · 2025-05-07 · via The Register - Special Features: Datacenter Networking Nexus

Datacenter Networking Nexus

Blackouts less frequent in 2024, still a PITA when the datacenter downtime demons visit

Datacenter outages are less frequent and severe, but human error remains one of the most persistent challenges, with between two-thirds and four-fifths of major wobbles involving some element of meatbag-related cause.

According to the latest Annual Outage Analysis report from Uptime Institute, the overall picture is one of improving reliability, but with the sting in the tail that when failures do occur they can be significant and costly.

"Outages overall have slowed down," said Andy Lawrence, Uptime executive director of research. "Datacenter operators are facing a growing number of external risks beyond their control, including power grid constraints, extreme weather, network provider failures and third-party software issues. And despite a more volatile risk landscape, improvements are occurring."

Some 53 percent of operators reported an outage in the past three years, but this compares with 60 percent in 2022, 69 percent in 2021, and 78 percent in 2020. Just 9 percent of reported incidents during 2024 were classified as serious or severe, which is the lowest level yet recorded.

But preventing human error remains one of the major stumbling blocks in datacenter operations. Uptime says it views human error as a contributing factor rather than a root cause in outages, though it directly or indirectly plays a part in most of them.

Code changes, for example, played a part in several recent Microsoft incidents, such as problems with Azure cloud services in January and  a Microsoft 365 outage in March.

Nearly 40 percent of organizations have suffered a major outage caused by human error over the past three years, the report says. Staff failing to follow procedures was a feature in 58 percent of those cases, with faulty processes or procedures to blame in 45 percent.

It is also on the increase too, with the proportion of human error-related outages caused by failing to follow procedures up by 10 percentage points from last year. The reason for this may be the rapid growth seen by the datacenter industry recently and the resulting staff shortages in many regions, Uptime suggests.

To combat this, a greater focus on staff training and real-time operational support may reduce risks more effectively than improving documentation and processes, although these are still important.

Backing this up, 80 percent of operators told Uptime they believe that better management and processes might have prevented their organization's most recent downtime disaster.

Power-related issues remain the leading cause of major outages. These account for more than half of all cases, while more than one in four respondents to the 2025 Uptime resiliency survey reported that a serious or severe IT outage was caused by a power glitch within the past three years.

The most frequent factor in these is UPS failure – something that recently led to a six-hour blackout at Google Cloud services in the US east zone in America.

Other elements in the power chain can also cause issues such as intermittent faults in the supply and by mismanaged or misconfigured failover to generators.

Grid instability is also listed as a growing concern by Uptime. Rising demand, aging infrastructure, extreme weather, and the variability of renewable energy sources may increase the frequency of power disruptions - making robust on-site systems even more essential. Datacenters near London's Heathrow airport managed to remain operational despite a power outage that closed the site and caused disruption to a large number of flights in March.

Overall, investments in resiliency and the diligence of operators tell how a success story have led to a reduction in the overall severity and frequency of outages relative to the overall growth in online services.

However, Uptime warns the rising complexity of these environments, driven by AI, automation, and integration between IT and OT systems, is increasing exposure to operational errors and cybersecurity threats. ®