惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
DataBreaches.Net
V
Visual Studio Blog
WordPress大学
WordPress大学
Google DeepMind News
Google DeepMind News
K
Kaspersky official blog
博客园 - 叶小钗
月光博客
月光博客
S
Schneier on Security
T
Threat Research - Cisco Blogs
C
CERT Recently Published Vulnerability Notes
量子位
博客园 - 三生石上(FineUI控件)
宝玉的分享
宝玉的分享
P
Privacy & Cybersecurity Law Blog
Cyberwarzone
Cyberwarzone
S
Securelist
Hugging Face - Blog
Hugging Face - Blog
B
Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
V
Vulnerabilities – Threatpost
大猫的无限游戏
大猫的无限游戏
Google DeepMind News
Google DeepMind News
V
V2EX
MongoDB | Blog
MongoDB | Blog
博客园_首页
Recorded Future
Recorded Future
酷 壳 – CoolShell
酷 壳 – CoolShell
F
Fortinet All Blogs
GbyAI
GbyAI
Microsoft Security Blog
Microsoft Security Blog
C
Cybersecurity and Infrastructure Security Agency CISA
T
Troy Hunt's Blog
罗磊的独立博客
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
T
The Blog of Author Tim Ferriss
Application and Cybersecurity Blog
Application and Cybersecurity Blog
P
Proofpoint News Feed
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
Tor Project blog
Microsoft Azure Blog
Microsoft Azure Blog
爱范儿
爱范儿
O
OpenAI News
有赞技术团队
有赞技术团队
Blog — PlanetScale
Blog — PlanetScale
N
News | PayPal Newsroom
G
GRAHAM CLULEY
H
Hacker News: Front Page
Hacker News - Newest:
Hacker News - Newest: "LLM"

The Register - Software: AI + ML

Anthropic, now atop the AI bubble, files for its IPO Sick and wrong: Ontario auditors find doctors' AI note takers routinely blow basic facts OpenAI exec says it will burn $50B on compute this year Astera speaks softly and carries a big switch Anthropic unleashes finance agents for Claude IBM asks DBAs to trust AI to act on their behalf ServiceNow adds agent kill switches to AI control tower British mathematician hands OpenClaw agent a credit card Microsoft fixes VS Code after Copilot credited human code Shadow IT has given way to shadow AI. Enter AI-BOMs AI inference just plays by different rules How TeamViewer ONE transforms IT operations from firefighting to autopilot How TeamViewer ONE transforms IT operations firefighting aut Inference is giving AI chip startups a 2nd chance to shine How to roll your own local AI coding agents CIOs will be the governors for AI agents Govern your bots carefully or chaos could ensue Mozilla pushes back against Google's Prompt API SAP user group slams 'uncertainty' in ERP giant's API policy Microsoft boss tells investors the company is working to 'win back fans' Anthropic tops OpenAI in LLM revenue stakes Amazon's chips become a $20B business Fooling large language models just keeps getting simpler Amazon tells its engineers to review all AI output ZTE powers 2026 Jiangsu Football League with 5G-A & AI robot Future holiday horror: ‘A robot lost my luggage in Tokyo’ The future of software development has less development OpenAI jumps out of Microsoft's bed, into Amazon's Bedrock Vintage chatbot lives in the past like an elderly relative IBM's AI coding 'partner' Bob hits general availability Locked, stocked, and losing budget: AI vendor lock-in bites Ex-AWS legend explains what enterprises need to make AI work DeepSeek's new models offer big inference cost savings Anthropic admits it dumbed down Claude with 'úpgrades' Microsoft gives your Word documents an AI co-author you didn’t ask for Datadog digs down into GPU efficiency as AI costs soar Robotic arm powered by AI bats away ping-pong challenge Partnerships drive ZTE’s strategy to unlock AI potential Gov.uk says AI gaslighting Brits with stale Gov.uk data Google says it has all the answers for AI agent sprawl NeuBird plans a bright future for incident response AI-assisted intruders pwned Vercel via OAuth abuse and a pilfered employee account Vibe coding upstart Lovable denies data leak, cites 'intentional behavior,' then throws HackerOne under the bus Schmoozebots: study finds flattery will get AI everywhere New Android development tool designed for robots, not humans AI is reshaping Britain's datacenter map away from London Just like phishing for gullible humans, prompt injecting AIs is here to stay Anthropic debuts Claude Design, because who needs designers? Mozilla takes on enterprise AI providers with Thunderbolt Anthropic ejects bundled tokens from enterprise seat deal Maine to pause big bit barns as local opposition spreads If you want into Anthropic's Claude club, you may have to show ID Git identity spoof fools Claude into giving bad code the nod Nobody knows how many CVEs Anthropic's Project Glasswing has actually found Allbirds shoe company moving to AI infra is the top Bad teacher bots can leave hidden marks on model students Networks not ready for the challenges of AI traffic US states can't account for datacenter tax breaks. Literally Salesforce debuts Headless 360 agentic platform Waymo's self-driving cars face their toughest test yet: London Commvault has a Ctrl+Z for rogue AI agents Nvidia slaps forehead: AI, that's what quantum needs! OpenAI CEO Sam Altman home attack suspect charged Anthropic: Claude quota drain not caused by cache tweaks AI vs the cold hard reality of the legal profession China wants AI to prepare school lessons and mark homework Linux 7.0 debuts as Linus Torvalds ponders AI's impact Anthropic's Mythos has The Kettle crew curious, skeptical I vibe coded web app: It was enlightening and uncomfortable The AI divide putting open weights models in spotlight Amazon rejects AWS climate disclosure proposal UK to spend £15M on AI mapping in knife crime crackdown UK to spend £15M on AI-powered crime mapping in knife violence crackdown Rebrand automation as 'zero-token architecture' to master AI Call your existing automation ‘zero-token architecture’ to become an instant agentic AI wiz Only 28% of AI infrastructure projects fully pay off UALink delivers 2.0 spec before v. 1.0 silicon ships Only 28% of AI infrastructure projects fully pay off, survey finds No-Nvidia interconnect club delivers 2.0 spec before v1.0 silicon ships Anthropic reveals $30bn run rate and plans to use 3.5GW of new Google AI chips AI slop got better, so now maintainers have more work AMD's AI director slams Claude Code for becoming dumber and lazier since last update Anthropic closes door on subscription use of OpenClaw AI will make anyone a 10x programmer, but with 10x the cleanup PrismML debuts energy-sipping 1-bit LLM in bid to free AI from the cloud Netflix – yes, Netflix – jumps on the AI bandwagon with video editor AI models will deceive you to save their own kind Google battles Chinese open-weights models with Gemma 4 Microsoft shivs OpenAI with three new AI models for speech and images They thought they were downloading Claude Code source. They got a nasty dose of malware instead Even Microsoft knows Copilot shouldn't be trusted with anything important Google's TurboQuant saves memory, but won't save us from DRAM-pricing hell Claude Code bypasses safety rule if given too many commands OpenAI gets $122B to 'just build things' as the world blows them up One in seven Americans are ready for an AI boss, but they might not trust it Claude Code source leak reveals how much info Anthropic can hoover up about you and your system Oracle cuts jobs across sales, engineering, security Anthropic goes nude, exposes Claude Code source by accident GitHub backs down, kills Copilot pull-request ads after backlash Microsoft Fabric Database Hub only a 'partial' solution for admins
NeuBird AI plans a bright future for incident response
Robin Birtstone · 2026-04-22 · via The Register - Software: AI + ML

SPONSORED FEATURE We're at the beginning of the agentic era for operations. Current AIOps tools summarize dashboards and surface correlations, but most don't actually investigate incidents. So engineers still spend hours manually investigating complex incidents.

Vinod Jayaraman, co-founder and head of engineering at NeuBird AI, thinks it's time for that to change. Agentic AI systems can now perform actual SRE investigation work. The companyhas applied agentic AI to production operations. It automatically correlates telemetry data across AWS services without human intervention, surfacing root causes that Jayaraman says save engineers hours.

This marks the first time AI can reason over telemetry like an engineer would. It builds a service map to understand how components connect before starting an investigation, and then explores multiple hypotheses in parallel faster than human experts could.

REG AD

That's already happening, but now he wants to take things further. "One of the topics that is close to our hearts for this year is how we can close the SRE loop and also get closer to code generation," he says. "We also want to have a reasoning graph that explains how we came up with a root cause analysis (RCA)."

REG AD

The future is autonomous investigation, not better dashboards

Alert storms and dashboard sprawl are symptoms of architectural scale, not solutions to it. About 95 percent of alerts generated from low-level metric thresholds don't need to be investigated. High CPU usage might be good (a compilation job running) or bad (a runaway process). Context matters.

Humans cannot keep pace with modern telemetry volume. The next stage of evolution is to reduce human dependency on observability, not to throw more tools at the problem. NeuBird AI streamlines incident response by handling the investigation work autonomously.

AI agents must become the first responder to incidents. When incidents are queued up in the background, the system can refine RCA results using reasoning models without time pressure. Low confidence scores (say, less than 60%) trigger additional investigation passes rather than immediate escalation.

Things need to become more intuitive for SREs, DevOps, and Platform Engineering. "We want them to describe the end outcome that they want to avoid, and do so using natural language, which we call semantic monitoring," he says

That means moving beyond simply setting thresholds on CPU or memory. Instead, an SRE might say 'Monitor for any pod failures in this Kubernetes cluster that last more than two minutes'. The system breaks that down into various things that need to be monitored under the hood.

What the future looks like for SREs &DevOps

Instead of stressing out at the battlefront, site reliability engineers and those tasked with application reliability will soon spend at least some of their time supervising AI agents instead.

REG AD

In fact, while a human stays in control, it's agents all the way down. Supervisor agents can monitor other agents' investigations to keep them on the right track. When an agent gets stuck looking at the same metrics repeatedly, the supervisor redirects its attention elsewhere. It's like having a senior engineer guide a junior through their first incident response.

This agentic support won't replace people, but it will reduce a lot of the correlation work that they currently face when tracking problems across complex distributed systems.

Closing the SRE/DevOps loop

The next step for NeuBird AI is to go beyond solving immediate problems as a discrete practice by integrating with other parts of the incident response chain. He wants to close the problem resolution loop, where engineers diagnose problems, produce resolutions, update software, and prevent recurrence.

"When you come up with a probable RCA, things might get updated, but that's also often where the ball gets dropped," says Jayaraman. "We want to shift left, getting closer to the development cycle where once the RCA is produced, you're able to continue on and surface what you discovered, passing it on to the next step in the pipeline."

This is where his concept of code generation comes in. In the future, an SRE agent might deliver details of a problem to a coding agent that generates a fix and then writes a pull request to fix an issue.

Explainability and reasoning graphs

As NeuBird AI explores these broader automation opportunities, trust will be central to getting engineers on board.

REG AD

Trust remains the limiting factor. Some customers want single, high-confidence root cause analyses rather than multiple hypotheses. Low confidence scores trigger additional investigation passes. For deployment changes versus code changes, the system can apply fixes and verify results.

That's why humans are very much still in this loop. NeuBird AI is developing guardrails for autonomous approval of simple infrastructure changes (adding a single node for resource starvation, for instance). More complex fixes need a person to check the work and pull the lever.

Security permissions determine how far automation goes. Simple problems might get autonomous fixes within strict guardrails. Everything else stays in recommendation mode, with humans making final calls.

That in turn demands transparency, which is something that the company is building more of into its systems. Users can double-click to see why the system arrived at a particular RCA. The system will show its work: which metrics it checked, what logs it parsed, and which correlations it found.

"We present the solution, but also the path by which we arrived at the solution," says Jayaraman. "We present the queries that were executed to get out the information, which led us to believe that this is what the root cause is."

Consider a message queue overflow scenario. The consumer falls behind. NeuBird AI SRE provides context and recommends scaling out by adding another node. It produces specific Terraform script adjustments.

The product generates an internal confidence score on each RCA. High confidence means a clear answer. Low confidence triggers deeper investigation. Either way, the system presents the queries it executed to reach its conclusion, not just the final answer.

Anyone that has seen two AIs argue with each other will appreciate how fascinating this process is. NeuBird AI uses adversarial thinking to sharpen accuracy. Two models analyze the same incident independently. Agreement means high confidence. Disagreement flags uncertainty. The system uses LLMs as judges to evaluate the quality of work done.

Virtual SREs don't need to chat just with each other. Static reports are giving way to conversations, with engineers asking follow-up questions about an RCA, requesting different analyses, or exploring alternative hypotheses. The system learns from each interaction, incorporating feedback to improve future incident analysis.

These reasoning systems are another area that NeuBird AI will take further. The future vision is for these reasoning graphs to evolve into comprehensive automated runbooks. They will go beyond solving specific problems to address similar issues in the future. That way, operational memory persists even as teams change. And it can continuously evolve, updating runbooks as NewBird AI generates more understanding.

Why this shift is happening now

SREs and DevOps have needed capabilities like these for a while. Cloud complexity has exceeded human cognitive limits. The scale isn't a technology problem anymore. "In a complex piece of software, there are many pieces of code. When they all come together with the variability of your environment and the user traffic that comes in, there is no perfect code," Jayaraman points out.

Talent shortages have compounded the problem. It's hard to scale SRE teams when skills are short.

The demand might have been there for some time, but the capability wasn't. That's changing, as LLM maturity enables reasoning across telemetry data. Models can now process information more quickly while maintaining accuracy. NeuBird AI's architecture helps here; the company breaks agent tasks into bite-sized chunks, each using a different model optimized for that task. Smaller reasoning models can often produce superior results to large, cumbersome, foundational models designed to do everything.

The evolution of infrastructure as code also makes autonomous operations both feasible and critical. Agents can now address resource starvation problems through Terraform, while changes tracked in GitHub repositories close the operational loop.

Driving adoption

Aside from trust, removing security friction and deployment constraints is critical for this technology. For AI systems to handle production incidents, they need bulletproof security. NeuBird AI processes telemetry in real-time without storing it persistently. Customer data stays inside their AWS environment. When reasoning happens, only abstracted metadata reaches Amazon Bedrock.

The trust model relies on read-only permissions for AWS services, protected by AWS IAM. Customers control access through their own trust policies and can revoke permissions instantly.

NeuBird AI is also easing adoption through flexible deployment options. Half of the company's customers run purely in the cloud, while the rest use virtual private cloud (VPC) or hybrid setups.

The system connects to Prometheus in a customer's VPC or on-premises just as easily as cloud telemetry. From NeuBird AI's perspective, both look identical.

When those customers get an EC2 instance failure, NeuBird AI gets there first. The AI agent examines CloudWatch metrics, logs, and configuration changes to understand what happened, accessing telemetry sources through read-only connections. By the time the on-call engineer gets to the issue, there's already a root cause analysis waiting.

That's exciting enough, but the company clearly has great things planned for this year. It's working quickly, too, expecting a declarative definition for the problem resolution pipeline ready by end of Q1. All eyes on NeuBird AI to see what it delivers next.

Sponsored by NeuBird AI/em>