惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

K
Kaspersky official blog
V
Visual Studio Blog
宝玉的分享
宝玉的分享
月光博客
月光博客
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
Y
Y Combinator Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
大猫的无限游戏
大猫的无限游戏
H
Help Net Security
博客园_首页
Recent Announcements
Recent Announcements
小众软件
小众软件
MongoDB | Blog
MongoDB | Blog
Attack and Defense Labs
Attack and Defense Labs
The GitHub Blog
The GitHub Blog
Google DeepMind News
Google DeepMind News
Cisco Talos Blog
Cisco Talos Blog
L
LINUX DO - 最新话题
V2EX - 技术
V2EX - 技术
Simon Willison's Weblog
Simon Willison's Weblog
P
Palo Alto Networks Blog
PCI Perspectives
PCI Perspectives
T
Troy Hunt's Blog
Hacker News: Ask HN
Hacker News: Ask HN
S
Security Affairs
量子位
The Register - Security
The Register - Security
腾讯CDC
T
The Exploit Database - CXSecurity.com
P
Privacy & Cybersecurity Law Blog
V
Vulnerabilities – Threatpost
L
LINUX DO - 热门话题
N
News and Events Feed by Topic
Cloudbric
Cloudbric
Cyberwarzone
Cyberwarzone
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Recent Commits to openclaw:main
Recent Commits to openclaw:main
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
D
Darknet – Hacking Tools, Hacker News & Cyber Security
TaoSecurity Blog
TaoSecurity Blog
Scott Helme
Scott Helme
C
Cybersecurity and Infrastructure Security Agency CISA
The Last Watchdog
The Last Watchdog
W
WeLiveSecurity
H
Hacker News: Front Page
T
Tor Project blog
C
Cyber Attacks, Cyber Crime and Cyber Security
NISL@THU
NISL@THU
Know Your Adversary
Know Your Adversary
C
CXSECURITY Database RSS Feed - CXSecurity.com

rss.livelink.threads-in-node

Probably has less bugs than windows 11 | Microsoft Community Hub Quick question about window PC requirements for meta link cable? Is it possible to run ryujinx canary on an administrator account on windows? Why Windows 11 still depends on 1990s code iphone auf pc spiegeln windows 11 – Welche Methode funktioniert zuverlässig? CHERIoT-Ibex: Closing the door on memory safety vulnerabilities with hardware-enforced protection Known issue: Upgrading Microsoft Tunnel version 20260129.1 What's New in Microsoft Entra: May 2026 Carta de validación TSP (aka.ms/TSP_Achievement_Code_Enroll...) restricted across all accounts unable to enroll Class Admin Build observability for scalable AI apps and agents selling through Microsoft Marketplace Inspektor Gadget Completes Its First Independent Security Audit Retirement of Direct Exchange ActiveSync Certificate-Based Authentication by End of 2026 Export mixed text and tabular Excel to PDF Safely Migrating Terraform Managed Disks on Azure Using Stable Keys and Copilot Microsoft 365 & Power Platform Community call Microsoft 365 & Power Platform product updates call Course Retirement Announcement: AI-3022 The End is Nigh for DES and an Update for hunting down RC4 Unable to Access Scheduling Poll Options Title Plan Update - May 8, 2026 Secure Medallion Architecture Pattern on Azure Databricks (Part II) From Observability to Action: Building an AI-Powered AIOps Agent for Customer-Specific Operations General Availability of Mailbox Import and Export Microsoft Graph APIs Why External Participants Can—or Can’t—Join a Microsoft Teams Meeting CRITICAL: Data Loss on Build 26200.8328 - AI Storage Sense deleted 160+ apps with 870GB free space. Why is everyone hating on Windows 11? I was pissed at the Windows 11 context menu so I built this. Windows 11 Shows ASUS LOGO but then goes dark for 5 minutes Windows 11 causes discrete graphics cards to be locked at their base clock speed when idle In Windows 11 Insider Preview 26300.8346, the Start button is offset to the left and misaligned The Windows 10 update failed, so I'm planning to switch to Windows 11 Using the Microsoft Graph PowerShell SDK to Update User Profiles Someone please help me - Windows Insider Program How to Reduce No-Shows for WooCommerce Appointments Windows 11 Insider Program Download Update Beta Channel stuck at 99% How to transfer photos from iphone to pc on Windows 11 without itunes Windows 11 Dell Laptop Latitude 5500 New PC won’t boot up windows 11 Fresh Windows Install – Driver/Service Advice for G14 (2021 R9 5900HS + RTX 3050 Ti) AD Recycle Bin – “The specified value already exists” but Recycle Bin is non‑functional Why deleting files does not free up space on mac How to delete or clear cache on macbook air automatically Announcing Microsoft Host Integration Server 2028: Modern connectivity for IBM Mainframes Midranges From Test Cases to Trust: Elevating Enterprise Quality with GitHub Copilot Bootfähigen usb stick erstellen am mac für Windows 11? My modpack has this error and I can't find the cause Where is the option to attach file for parent Income Statement Form? Watching streamers with 2k resolution enabled makes my PC lag Update Your Exchange SE Hybrid On-premises Rich Coexistence to Graph Partner Blog | Unlock the cloud benefits in your partner benefits package with a streamlined activation experience Turn your Azure commitments into smarter cloud investments with Microsoft Marketplace Drive stronger Microsoft alignment and unlock co-sell growth at Ultimate Partner LIVE RECAP: Microsoft Elevate Partner Community Monthly Call - May 2026 Public Preview: Migrate your regional virtual machines to availability zones Security Dashboard for AI: 3 Ways CISOs Drive Impact Today PLANNER WILL NOT PUBLISH THE SCHEDULES FOR MY PLANS, BY CREATING AN iCALENDAR LINK....HELP | Microsoft Community Hub Firm AI for professional services: governed, agentic workflows built on Microsoft Azure Azure Incident Retrospective - Please register! Session 2 - Tracking ID: 5GP8-W0G Azure Incident Retrospective - Please register! Session 1 - Tracking ID: 5GP8-W0G Autoruns, ProcDump, ZoomIt, DebugView, NotMyFault, ProcExp, Procmon, and Linux tools Edge Canary not auto updating? Edge Dev on Windows - Tab sync not working Available today: GPT-5.5 Instant in Microsoft 365 Copilot Introducing the Azure Resource Manager MCP Server! Defender Threat & Vulnerability Management Reporting A New Chapter for Realtime AI: Reasoning, Translation, and Real-Time Transcription Plan a fun Family Wellness Month with Copilot Released: May 2026 Exchange Server Hotfix Update Service Bus SBMP Retirement: What BizTalk Server 2020 Customers Need to Know Azure Incident Retrospective Please register for one of the 2 sessions below! Planner Agent brings work management directly into Microsoft 365 Copilot Local account and MS account Cannot reset account password Announcing Windows Admin Center: Virtualization Mode Public Preview 2! Azure RBAC Custom Role Best Practices or Common Build Patterns Pivot Table Data Centering How to design production-ready AI architectures for apps and agents on Microsoft Marketplace Troubleshoot with OpenTelemetry in Azure Monitor - Public Preview Secure, Keyless Application Access with Managed Identities - Now GA in Azure Files SMB Help shape the Microsoft 365 Copilot community experience — your input matters Microsoft 365 Security Best Practices for Enterprises Running multimedia AI models on Container Apps with Serverless GPU (A100 & T4) SharePoint List Migration to new Tenant Windows 11: Task Manager: Background Processes: Stop Unnecessary Ones from Starting Automatically Detecting Plain‑Text Password Exposure Using Custom Regex in Microsoft Purview MVP Enthusiast Effective, engaging events in communities and storylines Change Optics Report released into Public Preview to showcase messages impacted by future changes RECAP: Microsoft Elevate Partner Community Monthly Call - April 2026 | Microsoft Community Hub Networking in Windows Server 2025 - Windows Server Summit Want to get more out of Microsoft Copilot Chat without adding another tool—or cost—to your stack? Announcing Public Preview: Security Copilot’s Email Summary in Microsoft Defender Copilot Chat in financial services: Is productivity moving faster than policy? How manufacturers can scale AI from pilot to production with Microsoft Marketplace SSO for a Python Teams Bot (M365 Agents SDK + FastAPI) — Single-Tenant, Multi-Tenant, and UAMI | Microsoft Community Hub Advice required for temp / agency staff | Microsoft Community Hub 'You've tried to sign in too many times with incorrect account/password' | Microsoft Community Hub When will Windows 11 natively support Auracast broadcasting? | Microsoft Community Hub Authenticator | Microsoft Community Hub
Applying Site Reliability Engineering to Autonomous AI Agents
mosiddi · 2026-05-20 · via rss.livelink.threads-in-node

If you practice SRE, you already have a mental model for running reliable production systems. You define SLOs. You track error budgets. You use circuit breakers to stop cascading failures. You run chaos experiments to find weaknesses before customers do. You treat every operational decision as a tradeoff between reliability and velocity.

That mental model transfers directly to AI agents. It just needs four new ideas.

In the Agent Governance Toolkit: Architecture Deep Dive, Policy Engines, Trust, and SRE for AI Agents, we covered Agent SRE briefly as one of AGT's nine packages: SLOs, error budgets, circuit breakers, chaos engineering, and progressive delivery, adapted from the patterns your SRE team already applies to microservices. Several teams asked for the full story. This is it.

Agent SRE is one of the more novel parts of the toolkit. The policy engine, zero-trust identity, and execution sandboxing have clear analogs in existing security practice. Agent SRE explores newer ground. Established patterns for defining SLOs for AI agent behavior, building chaos experiments for LLM provider failures, or applying error budgets to agent autonomy are still emerging across the industry. We built these capabilities because running agents in production without them is the equivalent of running a fleet of microservices without circuit breakers, health checks, or an on-call runbook.

This post is for SRE teams, platform engineers, and anyone responsible for running AI agents in production. You do not need to be an AI specialist. If you know what a burn rate is, you are ready for this.

When a service fails, your observability stack tells you: latency went up, error rate crossed the SLO threshold, the circuit breaker opened. You page the on-call engineer. They look at traces and find the slow database query.

When an AI agent fails, your observability stack is silent. The agent returned HTTP 200. Latency was normal. Error rate was zero. But the agent quietly approved a transaction it was not authorized to approve, hallucinated a database path and wrote to the wrong table, or got stuck in a reasoning loop that consumed $800 of LLM API budget before anyone noticed.

These are not infrastructure failures. They are behavioral failures.

And they are invisible to monitoring tools built for stateless, deterministic services, because those tools only watch for crashes and timeouts. They do not watch for wrong behavior.

This gap is the problem Agent SRE was designed to solve. The solution borrows everything from the SRE playbook and adds one concept that extends it: the Safety SLI.

Traditional SLIs measure system behavior from the user's perspective: latency, availability, error rate, throughput. They answer: did the service respond correctly?

For AI agents, correctness is not enough. An agent that responds correctly but acts outside its authorized scope has not succeeded. It has failed in a way that none of your existing SLIs can detect.

The Safety SLI answers a different question: did the agent act within policy?

from agent_sre import SLO, ErrorBudget
from agent_sre.slo.indicators import PolicyCompliance

# Define a safety SLO: 99% of agent actions must comply with policy
safety_slo = SLO(
    name="safety-compliance",
    indicators=[
        PolicyCompliance(
            target=0.99,
            window="7d",
        ),
    ],
    error_budget=ErrorBudget(
        total=0.01,                      # 1% budget (1 - 0.99 target)
        window_seconds=2592000,          # 30-day window
        burn_rate_alert=2.0,             # warn at 2x sustainable rate
        burn_rate_critical=5.0,          # page at 5x sustainable rate
    ),
)

When an agent's policy compliance rate drops below 99%, the error budget starts burning. The ErrorBudget tracks consumption automatically and exposes burn rate alerts through its firing_alerts() method. When the budget is exhausted, the configured exhaustion_action determines the system response:

from agent_sre.slo.objectives import ExhaustionAction

# Configure what happens when error budget is exhausted
safety_slo = SLO(
    name="safety-compliance",
    indicators=[PolicyCompliance(target=0.99, window="7d")],
    error_budget=ErrorBudget(
        total=0.01,
        window_seconds=2592000,
        burn_rate_alert=2.0,              # fires at 2x sustainable burn rate
        burn_rate_critical=5.0,           # fires at 5x sustainable burn rate
        exhaustion_action=ExhaustionAction.CIRCUIT_BREAK,  # suspend agent when budget is gone
    ),
)

# In your monitoring loop, check for firing alerts
alerts = safety_slo.error_budget.firing_alerts()
for alert in alerts:
    print(f"Alert firing: {alert.name} (severity: {alert.severity})")

# Check budget status
print(f"Budget remaining: {safety_slo.error_budget.remaining_percent:.1f}%")
print(f"Current burn rate: {safety_slo.error_budget.burn_rate():.2f}x")
print(f"Exhausted: {safety_slo.error_budget.is_exhausted}")

This is the governance dial from the other direction. The error budget is not just a metric: it is the mechanism that drives agent autonomy decisions. An agent with a clean 30-day safety record earns autonomy. An agent whose budget is burning at 5x the sustainable rate triggers a critical alert, and when the budget is exhausted, the exhaustion_action fires: ALERT, THROTTLE, FREEZE_DEPLOYMENTS, or CIRCUIT_BREAK. The graduated response mirrors what SRE teams already do with service SLOs, applied to agent behavior.

There are multiple SLI dimensions built into Agent SRE. Safety SLIs and Performance SLIs track different aspects of the same agent:

SLI Type

What It Measures

Target Pattern

When Budget Burns

Safety SLI

PolicyCompliance -- fraction of actions within authorized scope

>= 99%

Restrict capabilities, increase human oversight

Performance SLI

TaskSuccessRate, ResponseLatency, CostPerTask

Configurable per workload

Alert, throttle, or circuit-break LLM provider

Additional built-in indicators include ToolCallAccuracy, DelegationChainDepth, HallucinationRate, and CalibrationDeltaSLI. Both SLOs feed into the same error budget dashboard. An agent can have excellent performance but a degrading safety record, or perfect safety compliance and terrible cost efficiency. You need both dimensions to understand whether an agent is production-ready.

Circuit breakers for services protect against one failure mode: a backend that is slow or unreachable. The pattern is CLOSED -> OPEN -> HALF_OPEN. You know it well.

Agent SRE implements the same state machine for failure modes that are specific to autonomous reasoning systems and do not exist in traditional microservice architectures:

from agent_sre.cascade.circuit_breaker import CircuitBreakerConfig, CircuitBreaker
from agent_sre.chaos.engine import FaultType

config = CircuitBreakerConfig(
    failure_threshold=5,              # Open after 5 failures in the window
    recovery_timeout_seconds=60,      # Stay OPEN for 60s before HALF_OPEN
    half_open_max_calls=3,            # Allow 3 probes in HALF_OPEN
)

breaker = CircuitBreaker(agent_id="analyst-agent-001", config=config)

# Failure modes tracked by the circuit breaker:
tracked_faults = [
    FaultType.POLICY_BYPASS,           # Agent exceeds authorized scope
    FaultType.ERROR_INJECTION,         # Upstream model API fails
    FaultType.TIMEOUT_INJECTION,       # Tool calls exceed time budget
    FaultType.TRUST_PERTURBATION,      # Agent trust score falls below threshold
    FaultType.DEADLOCK_INJECTION,      # Agent stuck in iterative reasoning
]

Each failure mode has different circuit-breaking semantics:

Failure Mode

What Triggers It

Circuit-Break Behavior

Policy bypass

Action denied by policy engine

Count toward threshold; log with full context

LLM provider error

HTTP 5xx from model API

Immediately open; route to fallback model if configured

Tool timeout

Tool call exceeds timeout_ms

Count toward threshold; cancel in-flight call

Trust score degradation

Agent trust score drops below configured floor

Open; escalate to Ring 3 (untrusted) until score recovers

Reasoning loop / deadlock

Token or iteration count exceeds budget

Open; trigger human review before resuming

The reasoning loop breaker deserves attention. A microservice cannot get stuck reasoning. An AI agent absolutely can, and when it does, the failure is not an error code: it is an agent that keeps calling tools, consuming tokens, and generating audit events indefinitely. The circuit breaker detects this pattern from the iteration count and token budget and terminates the loop:

# Reasoning loop detection configuration
loop_detection_config = {
    "max_iterations": 15,             # Hard stop after 15 reasoning steps
    "max_tokens_per_session": 50000,  # Hard stop on token consumption
    "repetition_threshold": 0.85,     # Stop if >85% of recent actions repeat prior ones
    "on_detection": "circuit_break_and_escalate",
}

The state machine behaves identically to what you know from Hystrix or Resilience4j. What changes is the definition of "failure."

CLOSED (serving)
  |
  |  failure_threshold crossed for any tracked fault
  v
OPEN (rejecting -- agent action denied, fallback or human-in-loop fires)
  |
  |  recovery_timeout expires
  v
HALF_OPEN (probe -- limited requests allowed through)
  |
  |-- success_threshold met --> CLOSED
  |-- any failure          --> OPEN (reset timeout)

The only way to know if your agent system is resilient is to break it intentionally. Traditional chaos engineering targets infrastructure: kill a pod, inject network latency, saturate a disk. Agent chaos engineering targets the failure modes specific to autonomous reasoning systems.

Agent SRE ships fault injection templates that cover the failure modes teams consistently underestimate until they hit production:

from agent_sre.chaos.engine import ChaosExperiment, Fault, FaultType

# Experiment 1: LLM provider degrades -- model returns valid responses but with
# increased latency and occasional malformed outputs
experiment = ChaosExperiment(
    name="llm-degradation-resilience",
    target_agent="analyst-agent-001",
    description="Test agent behavior under degraded LLM provider",
    faults=[
        Fault.latency_injection(target="llm-provider", delay_ms=8000),
        Fault.error_injection(target="llm-provider", rate=0.05),
    ],
    duration_seconds=300,
)

# Experiment 2: Trust score manipulation -- simulates an agent receiving
# messages from a peer with a spoofed trust score
trust_experiment = ChaosExperiment(
    name="trust-manipulation-resilience",
    target_agent="orchestrator-001",
    faults=[
        Fault(
            fault_type=FaultType.TRUST_PERTURBATION,
            target="did:mesh:orchestrator-001",
            params={"spoofed_score": 950},
        ),
    ],
    duration_seconds=120,
)

# Experiment 3: Tool timeout cascade -- multiple tools time out simultaneously,
# testing whether the agent abandons gracefully or enters a reasoning loop
cascade_experiment = ChaosExperiment(
    name="tool-timeout-cascade",
    target_agent="analyst-agent-001",
    faults=[
        Fault.timeout_injection(target="database.read", delay_ms=30000),
        Fault.timeout_injection(target="api.call", delay_ms=30000),
    ],
    duration_seconds=180,
)

# Run the experiment
experiment.start()
# ... inject faults during agent execution ...
resilience = experiment.calculate_resilience(
    baseline_success_rate=0.95,
    experiment_success_rate=0.87,
    recovery_time_ms=48000,
)
experiment.complete(resilience=resilience)
print(f"Resilience score: {resilience.overall}/100 -- {'PASSED' if resilience.passed else 'FAILED'}")

Additional fault types built into the chaos engine cover: prompt injection attempts, privilege escalation, data exfiltration attempts, identity spoofing, deadlock injection, and contradictory instruction scenarios. Each maps to a FaultType enum value and can be composed into multi-fault experiments.

Important: The chaos engine records that a fault was injected and triggers the governance response pipeline. Actual infrastructure-level fault injection (network partition, process kill) should be implemented using your existing chaos tooling (Chaos Mesh, Gremlin, Azure Chaos Studio, or similar). Agent SRE governs the agent's behavioral response to faults; it does not own infrastructure manipulation. These two layers are designed to compose.

Each chaos experiment produces a structured resilience score via calculate_resilience(), which compares baseline and experiment success rates. A score of 90+ with passed=True means the agent maintained at least 90% of its baseline performance under fault conditions. Teams use this to set minimum resilience thresholds for production readiness.

Infrastructure incidents are reproducible because infrastructure is deterministic. AI agent incidents are hard to reproduce because agent behavior depends on model state, context window content, and the sequence of tool call results, none of which are preserved by default after a session ends.

Agent SRE's replay engine records every agent session as a replayable artifact: the full trace at each step, every tool call with its inputs and outputs, every policy evaluation with its decision, and every trust score at the time of each inter-agent message.

from agent_sre.replay.capture import TraceStore
from agent_sre.replay.engine import ReplayEngine, ReplayMode

# Traces are captured automatically when SRE tracing is active
store = TraceStore(
    backend="azure_blob",
    retention_days=30,
)

# When an incident occurs, replay the session exactly
engine = ReplayEngine(store=store)

# Full replay: re-run the session against the same recorded inputs
# Uses recorded tool outputs -- no live tool calls -- so replay is deterministic
result = await engine.replay(
    trace_id="trace_2026_05_a7f3b2",
    mode=ReplayMode.FULL,
)

for step in result.steps:
    print(f"Step {step.index}: {step.action} -> {step.decision}")

# Divergence analysis: replay with a policy change applied
# Shows exactly which actions would have been blocked under the new policy
diff_result = await engine.diff(
    trace_id="trace_2026_05_a7f3b2",
    policy_override="policies/stricter-v2.yaml",
)

for diff in diff_result.diffs:
    if diff.description:
        print(f"Step {diff.span_name}: was {diff.original}, "
              f"would be {diff.replayed} under new policy")

The divergence analysis is the feature teams use most. When a policy change is proposed, you replay recent production traces against the new policy to see how many actions would have been blocked, which sessions would have failed, and what the error budget impact would have been. Policy changes stop being guesswork.

When you ship a new service version, you do not send it to all traffic at once. You use canary deployments, feature flags, or traffic splitting. You watch the SLOs. If they degrade, you roll back.

Agent SRE brings the same discipline to agent capability rollout. When you expand an agent's authorized scope, giving it write access it did not have, connecting it to a new tool, or raising its trust floor, you do not expand to the full fleet immediately. You expand progressively, with automated SLO gates controlling each stage.

from agent_sre.delivery.rollout import (
    AnalysisCriterion,
    CanaryRollout,
    RollbackCondition,
    RolloutStep,
)

rollout = CanaryRollout(
    name="database-write-capability",
    steps=[
        RolloutStep(
            name="canary",
            weight=0.05,                   # 5% of agents get the new capability
            duration_seconds=86400,        # 24 hours
            analysis=[
                AnalysisCriterion(metric="safety_sli", threshold=0.995),
                AnalysisCriterion(metric="performance_sli", threshold=0.90),
                AnalysisCriterion(
                    metric="error_budget_consumed",
                    threshold=0.10,
                    comparator="lte",      # canary can burn at most 10%
                ),
            ],
        ),
        RolloutStep(
            name="early-adopters",
            weight=0.25,                   # 25% traffic
            duration_seconds=172800,       # 48 hours
            analysis=[
                AnalysisCriterion(metric="safety_sli", threshold=0.990),
                AnalysisCriterion(metric="performance_sli", threshold=0.88),
            ],
        ),
        RolloutStep(
            name="general-availability",
            weight=1.0,                    # 100% traffic
            duration_seconds=604800,       # 1 week of full observation
            analysis=[
                AnalysisCriterion(metric="safety_sli", threshold=0.990),
                AnalysisCriterion(metric="performance_sli", threshold=0.85),
            ],
        ),
    ],
    rollback_conditions=[
        RollbackCondition(metric="safety_sli", threshold=0.95, comparator="lte"),
    ],
)

# Start the rollout -- SLO gates evaluate at each step
rollout.start()

# Advance to next step when analysis criteria pass
if rollout.advance():
    print(f"Advanced to step: {rollout.current_step.name}")
    print(f"Progress: {rollout.progress_percent:.0f}%")

The SLO gate at each step is the same mechanism as a CI/CD quality gate, but measured on live production behavior rather than test results. An agent capability that degrades the safety SLI during canary does not promote to the next step. If a RollbackCondition fires, the rollout rolls back automatically. This is the mechanism that makes it operationally safe to expand agent autonomy: every expansion is measurable, every measurement gates the next expansion, and rollback is automatic.

Traditional health checks answer: is the service alive? For agents, alive is not enough. A healthy agent is one that is alive, operating within policy, consuming resources within budget, and maintaining a trust score above the Ring threshold it was assigned.

# Agent health check covering multiple dimensions
health = await agent_health_check(
    agent_id="analyst-agent-001",
    dimensions=[
        "liveness",            # Is the agent process running?
        "policy_compliance",   # Is safety SLI above threshold?
        "trust_score",         # Is trust score above Ring floor?
        "resource_budget",     # Is token/API spend within limits?
        "tool_availability",   # Are the tools the agent needs reachable?
    ],
)

# health.status: "healthy" | "degraded" | "unhealthy"
# health.dimensions: per-dimension pass/fail with values
# health.recommended_action: "none" | "restrict" | "suspend" | "terminate"

When health checks report degradation, backpressure controls engage before the circuit breaker opens. Backpressure is the earlier, softer response: accept fewer concurrent tasks, reject low-priority work, drain in-flight tasks gracefully before the situation escalates.

# Backpressure configuration
backpressure_config = {
    "backpressure_threshold": 0.80,    # Engage when resource utilization > 80%
    "max_concurrent": 5,               # Hard cap on simultaneous agent tasks
    "priority_shedding": True,         # Drop low-priority tasks first
    "drain_timeout_seconds": 30,       # Allow in-flight tasks to complete
}

The ordering matters: backpressure first, then circuit breaker, then suspension. Each stage is recoverable. Each stage preserves more agent state than the next. The SRE principle of graduated response applies to agents exactly as it applies to services.

Agent SRE does not ask you to adopt a new observability platform. Governance metrics are exported through the same adapters your infrastructure monitoring already uses, including OpenTelemetry, Prometheus, Datadog, and others.

from agent_sre.tracing.exporters import configure_exporters

configure_exporters(
    backends=[
        {"type": "prometheus", "endpoint": "http://prometheus:9090"},
        {"type": "opentelemetry", "endpoint": "http://otel-collector:4317"},
    ],
    include_metrics=[
        "slo.safety_sli",               # Per-agent safety compliance rate
        "slo.error_budget_remaining",    # Error budget in percentage
        "slo.burn_rate",                 # Current burn rate vs sustainable
        "circuit_breaker.state",         # CLOSED / OPEN / HALF_OPEN
        "circuit_breaker.failure_count",
        "trust_score.current",           # Agent trust score (0-1000)
        "trust_score.ring",              # Current execution ring
        "chaos.experiments_run",         # Chaos experiment telemetry
        "health.status",                 # Aggregate health status
        "backpressure.load",             # Current load vs threshold
    ],
)

Key governance metrics available in your existing dashboards:

Metric

What It Tells You

Alert Condition

slo.safety_sli

Fraction of agent actions within policy

< 0.99

slo.burn_rate

Rate at which error budget is consumed

> 2.0 (warn), > 5.0 (page)

slo.error_budget_remaining

Budget left for the SLO window

< 20%

circuit_breaker.state

Current breaker state per agent

OPEN or HALF_OPEN

trust_score.ring

Execution ring (privilege level)

Ring 3 (untrusted)

health.status

Aggregate health across all dimensions

degraded or unhealthy

If you are already running Grafana dashboards for your services, a governance dashboard for your agent fleet is a new data source and a new set of panels, not a new monitoring stack.

Everything in Agent SRE is built on the SRE mental model you already have, extended with four concepts that adapt traditional reliability thinking for autonomous systems:

Traditional SRE

Agent SRE Equivalent

What Changes

Latency SLI

Safety SLI

Correctness of *action*, not speed of *response*

Error budget

Autonomy budget

Burns on policy violations, not just errors

Circuit breaker

Behavioral circuit breaker

Opens on wrong *behavior*, not just failure codes

Canary deployment

Capability rollout

Rolls out *scope*, not just code

The governance insight is that error budgets work in both directions for agents. A service's error budget only decreases. An agent's autonomy is also a budget: it grows when the safety SLI is strong and shrinks when it degrades. The error budget mechanism becomes the operational mechanism for expanding and contracting agent autonomy in response to evidence, which is exactly what regulated industries and risk-averse enterprise teams need before they will trust an autonomous agent with consequential actions.

pip install agent-sre

A minimal Agent SRE integration requires three things: a safety SLO definition, a circuit breaker, and a health check. The progressive delivery and chaos engineering features layer on top when you are ready for them.

from agent_sre import SLO, ErrorBudget
from agent_sre.slo.indicators import TaskSuccessRate
from agent_sre.cascade.circuit_breaker import CircuitBreakerConfig, CircuitBreaker

# Step 1: Define your safety SLO
slo = SLO(
    name="production-safety",
    indicators=[TaskSuccessRate(target=0.99, window="24h")],
    error_budget=ErrorBudget(total=0.01, burn_rate_alert=2.0, burn_rate_critical=5.0),
)

# Step 2: Configure a circuit breaker
breaker_config = CircuitBreakerConfig(
    failure_threshold=5,
    recovery_timeout_seconds=60,
    half_open_max_calls=3,
)
breaker = CircuitBreaker(agent_id="my-agent", config=breaker_config)

# Step 3: Wire into your existing agent loop
async def governed_agent_loop(agent, task):
    # Check health first
    if not await agent_is_healthy(agent.id):
        return {"error": "agent suspended", "reason": "health check failed"}

    # Run within circuit breaker protection
    async with breaker:
        result = await agent.run(task)
        slo.record_event(good=result.policy_compliant)
        return result

The quickstart in the repository walks through a complete setup with safety SLOs, circuit breakers, and a Prometheus dashboard export in under 50 lines.

Most AI observability tools today focus on what you might call model quality: hallucination rate, latency, token cost, task completion. These are useful metrics. They are not SRE metrics. They do not answer whether the agent acted within its authorized scope, whether its behavioral error budget is burning at a dangerous rate, or whether it would survive the LLM provider going down.

Agent SRE answers those questions using the operational vocabulary that SRE teams already understand: SLOs, error budgets, circuit breakers, chaos experiments, and health checks. The goal is not to replace your observability stack. It is to make agent governance visible inside it.

The reliability of an autonomous agent is not a property of the model. It is a property of the governance infrastructure around it. Agent SRE is that infrastructure.

The Agent Governance Toolkit is an open-source project released under the MIT License. All features described in this post are available in the public repository. The `agent-sre` package is currently in public preview; APIs may change before general availability.

Questions about Agent SRE in your environment? Open an issue at aka.ms/agent-governance-toolkit or start a discussion in the comments below.