惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
N
Netflix TechBlog - Medium
P
Proofpoint News Feed
D
Docker
J
Java Code Geeks
L
LangChain Blog
Microsoft Security Blog
Microsoft Security Blog
The GitHub Blog
The GitHub Blog
I
InfoQ
Stack Overflow Blog
Stack Overflow Blog
云风的 BLOG
云风的 BLOG
Engineering at Meta
Engineering at Meta
MongoDB | Blog
MongoDB | Blog
月光博客
月光博客
T
Tailwind CSS Blog
M
MIT News - Artificial intelligence
Blog — PlanetScale
Blog — PlanetScale
Google DeepMind News
Google DeepMind News
腾讯CDC
罗磊的独立博客
U
Unit 42
爱范儿
爱范儿
Vercel News
Vercel News
MyScale Blog
MyScale Blog

The New Stack | DevOps, Open Source, and Cloud Native News

Agentic development hinges on verification. For cloud-native software, that is a runtime problem. AI agents need infrastructure: Why Europe’s regional cloud strategy matters Transform your AI coding agent into a deterministic Java Spring expert WeAreDevelopers is coming to the US to give unsung developers a bigger voice Cleaner AI training data, fewer bugs: Sonar’s SonarSweep explained Observability overload is drowning engineers Google’s DiffusionGemma is 4x faster than its other Gemma models Fable 5: Guardrails and burn rate are annoying users, who say it’s still better than Opus 4.8 The Anthropic leader who built Claude Code says he ditched prompting — now he just writes loops. AWS can now mathematically prove your VMs are isolated Microsoft pulled 73 GitHub repos after malware attack — but still won’t say who’s compromised Databricks wants to kill the “email me a file” problem for AI agent skills Ramp bets forward deployed engineers can do what off-the-shelf finance AI can’t Git real: AI agents aren’t just for solo developers anymore Anthropic launches Claude Mythos/Fable 5, but you better try it soon Spring is 23 years old. AI just made it a security emergency. This AI agent startup ditched Anthropic for DeepSeek — and says it’s saving millions When your data model is the bottleneck: lessons from Medium’s feature store How long before we stop reading the code? The tokenmaxxing party is over, and Revenium is mopping up How AI is solving the memory crunch it created Microsoft’s pitch to enterprises: Ditch Azure Repos for GitHub, despite its rocky reliability record Claude Code’s biggest upgrade yet ran 5 agents at once — here’s what happened Why Anthropic just doubled Claude Cowork limits at no charge For years, Apache Cassandra handed this work to your team — 6.0 takes it back “A dangerous combination”: The 2 factors that can “corrupt” AI agent workflows With Foundry, Microsoft bets the enterprise AI battle is about reliability, not capability Microsoft unlocks Visual Studio for developers left behind by its own AI AI teams now deploy 1,000 times a month. Your pipeline wasn’t built for that. Microsoft just made the agent runtime free — and kept everything around it
The reason enterprise outages almost never start where op...
Jennifer Riggins · 2026-05-26 · via The New Stack | DevOps, Open Source, and Cloud Native News

Nothing is Greenfield in an enterprise. Hybrid cloud complexity sits atop siloed teams and systems, making it nearly impossible to observe, understand, remediate, and prevent outages. Understaffed operations and site reliability teams now also have to keep pace with AI. All while the enterprise stack faces more risks of instability and insecurity than ever.

To respond, enterprise operations must fundamentally change how they approach Day 2 and beyond. No longer is it safer to maintain thick boundaries between services or divisions — they block return on AI investment anyway.

To survive, organizations must move from siloed, reactive dashboards to a closed-loop operations model, supported by AI agents, treating orchestration, observability, and remediation as a continuous feedback cycle.

“You’ve got to understand that Day 2 and Day 1 are in a closed loop because what you provision, you need to understand what the current Brownfield is. And you might not want to make a lot of changes when there are a lot of issues going on,” Phanidhar Koganti, senior distinguished technologist in Hewlett Packard Enterprise (HPE) hybrid cloud, tells The New Stack

Under the current restrictions facing ops teams, the only way forward seems to be applying AI techniques to operations, extracting signals from all the noise. This move from DIY to autonomous remediation requires more than just an infusion of AI.

AI-enabled ops demands a platform engineering strategy, predictive analytics, new ops metrics, and more. All is not to replace today’s operators, but to optimize the team’s time, enabling them to be faster, more strategic, and hopefully less stressed in stressful times. 

When every resource is constrained

“Creation is now only limited by the cost of compute, not capacity,” promises the Outcome Engineering Manifesto about the change AI agents are forcing in the tech industry. Yet AI has actually left most operations teams feeling more limited than ever, in both time and available resources. 

“Our customers are under pressure to continue to offer the same SLAs [service level agreements] with a lot less resources,” Sridhar Katere, VP of engineering for HPE’s data center business unit, says. In practice, this means there are fewer ops team members “to troubleshoot an issue when something happens during Day 2.”

Operations agents, managed by ops teams, are an opportunity to scale up troubleshooting and remediation, without expanding teams. 

HPE OpsRamp Software recently went general availability (GA) with its agentic operations copilot, with which, Koganti explains, “You can express what you are trying to achieve in a very high-level intent, and that will be converted to detailed deployment plans, which will include data center, networking-related automation, storage, and various other components for the whole infrastructure to come together.”

AI can also help ops teams move from reactive to proactive, while advocating for right-sized resource budgeting. 

“When we say issues, it doesn’t mean the nature of the failure is service-disruptive only,” Koganti says. “If you’re going to go out of capacity, meaning optics could fall, that’s a binary failure that you’re trying to predict.”

If an ops engineer knows that something, like a switch, is likely to fail in six weeks, Katere explains, then not only do they have time to prevent an outage, but they can also better budget hardware resources, which are extremely constrained by current supply chain uncertainty. 

“This is where we are building predictive analytics to help customers plan and procure the necessary hardware and components ahead of time,” he continues.

Through recent acquisitions and ongoing investment, HPE is building its own operations control center to define this next-gen hybrid, multi-cloud operational model. Its CloudOps Software suite — anchored by HPE OpsRamp Software and HPE Morpheus Software, with HPE Apstra Data Center Director available as an add-on — is designed to help organizations map out what their full cloud operating model actually requires.

Following OpsRamp’s GA release, other HPE products are expected to release agentic systems with conversational interfaces soon.

“Whatever we are talking about in optimizing the full stack will be applicable to capacity issues or even performance issues across the whole stack,” Koganti says, because enterprise systems’ complexity is many-layered.

Ops metrics need redefining 

While the mean time to resolution (MTTR) — total downtime divided by total number of incidents — remains an important metric, it’s insufficient in the face of this enterprise-grade complexity. 

Similarly, the means of identification or detection are still key, but they’re viewed differently when you bring AI into the mix. AI operations must prioritize time to correlate, applying AI ops tooling to analyze logs, metrics, and traces to connect disparate signals into a single, traceable, recommended action

“First instinct is often that network is the culprit, so it’s very important for networking to self-diagnose and come back to say, ‘Hey, I’m not the problem.'”

Like the famous Spider-Man pointing meme, if they were all pointing at networking or storage, AI operations tools must also optimize for mean time to innocence.

“First instinct is often that network is the culprit, so it’s very important for networking to self-diagnose and come back to say: Hey, I’m not the problem,” Katere says.  

In a full-stack environment, symptoms and downtime rarely occur in the same layer as the cause. 

“…the symptom of a failure and the cause of the failure are never in the same layer.”

“What is common that we see among the majority of our customers is that the symptom of a failure and the cause of the failure are never in the same layer. What I mean by that is, for example, the symptom can appear in the application layer as my application transactions are timing out, and so on. While the actual cause of it could be somewhere down in the network layer or the storage layer, most of the time, the network is the culprit,” Koganti says. 

“IT stacks are getting complicated. You need collaboration between various layers and agentic collaborations, as well as regular traditional AIOps collaboration.” 

Enterprise-grade complexity, which can see multiple failures occurring in parallel, requires tighter integrations across the full stack, along with AI to filter the important signals from the noise. Then the next step is AI agents auto-remediating when appropriate — but only when its actions are explainable to the human operator.

HPE also has its own efficacy metrics, including quarterly KPIs that measure how well its platform predicts and troubleshoots issues. Currently, the platform — backed by graph and GraphQL, with a highly contextualized data stack — is about 40% accurate on troubleshooting, Katere says, with a goal of topping 70% by the end of the year. HPE is baking its nearly 100 years of experience in the enterprise space into its new CloudOps Software, along with agentic skills trained at different layers: platform, driver, through, replaceable units, system, and more. 

“We are not living in a static world. It’s a very dynamic world.” Katere admits that it’s all a “moving target in light of how fast things are changing even in the slower enterprise software development lifecycle.” 


This article was originally published on April 29, 2026, and was updated on May 26, 2026.

YOUTUBE.COM/THENEWSTACK

Tech moves fast, don't miss an episode. Subscribe to our YouTube channel to stream all our podcasts, interviews, demos, and more.

Created with Sketch.