





















Sasi Kiran Malladi is a Principal at AWS, focused on Cloud, AI, Observability, Security & Resilience.

getty
Many of the most critical systems in our lives—the power grid, financial networks, hospital systems, transportation platforms—operate quietly in the background, and most people only notice them when something breaks.
With AI now layered into those systems, decisions are being made faster, at greater scale and often without direct human involvement. That raises the question: If AI is helping run critical infrastructure, how do we actually know what it's doing? The answer is that observability has to evolve.
For years, observability has been about understanding system health. Most organizations already collect logs, metrics and traces—the telemetry that tells them whether systems are up and how they are performing.
Traditional monitoring still matters, but it was designed for predictable systems, and it works best when behavior follows known patterns and thresholds. AI changes that equation. These systems adapt, generate new outputs and, in some cases, even make autonomous decisions.
As AI becomes embedded across systems like cloud platforms and customer-facing applications, observability has to move beyond answering what is happening. It has to explain why it is happening—a subtle but important shift.
AI systems introduce a different kind of complexity. Their behavior can vary based on inputs, context and training data. Two similar prompts can produce different results, and outcomes can drift over time. In high-stakes environments, unpredictability creates risk.
Observing AI systems means getting visibility into interactions, not just infrastructure. At the most practical level, it can mean tracking what happens when a user submits a prompt and what response is generated. It also means evaluating whether the output is accurate and whether it contains bias.
This is where familiar concepts like telemetry take on a new role. Logs, metrics and traces are still collected, but they are now combined with additional signals. Prompt-level data, model responses, anomaly detection and output validation all become part of the observability layer.
AI-driven observability works in two directions: Organizations use AI to understand their systems better, and they also use observability to understand their AI better.
On the system side, AI can process massive volumes of telemetry in real time. Instead of relying on static thresholds, it can detect anomalies and identify root causes faster than manual approaches.
On the AI side, observability helps answer harder questions. Is the model producing consistent results? Is it introducing hallucinations or generating information that is not grounded in fact? Is sensitive data being exposed through prompts or responses?
When done correctly, this approach can reduce noise, prioritize meaningful signals and accelerate incident response.
The implications become clearer in critical infrastructure. In cloud and data center environments, AI-driven observability supports autonomous incident response by identifying and resolving issues before they escalate. In power systems, it can help detect early signs of instability and prevent cascading outages. In healthcare, it enables visibility across both clinical workflows and underlying IT systems, where delays or failures can have direct consequences for patient care.
Financial systems, transportation networks, water infrastructure and industrial operations all face similar challenges. As AI becomes more embedded in these systems, observability becomes part of the control layer that keeps them stable.
One of the less obvious roles of observability is trust. Organizations may adopt AI quickly, but they hesitate to rely on it fully without visibility into how it behaves. That hesitation shows up in slower decision-making and increased manual oversight.
Observability helps close that gap. By providing insight into how AI systems operate, it gives leaders the confidence to scale their use. At the same time, it highlights risks that need to be addressed, from biased outputs to gaps in data protection.
There are also pitfalls. Sensitive information can end up in observability pipelines if data is not handled carefully, and some organizations underestimate how much telemetry they need to collect, leaving blind spots in their systems. Minimizing these risks requires being deliberate about what is collected and putting guardrails around sensitive data.
AI is advancing faster than most organizations can fully keep up. Some are early adopters, but many are still working toward integrating these capabilities into their operations.
Traditional monitoring is still critical, but it is no longer enough on its own. Organizations need to upgrade their approach to handle systems that are dynamic and distributed, as well as increasingly autonomous. At the same time, they need to apply observability to AI itself. Being clear on how models behave and how decisions are made is becoming a baseline requirement.
Critical infrastructure has always depended on visibility. What is changing is the level of depth required. As AI becomes part of the systems that keep industries running, observability becomes the way organizations can see, understand and ultimately trust what those systems are doing.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。