惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
F
Fortinet All Blogs
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
S
Secure Thoughts
SecWiki News
SecWiki News
Hacker News: Ask HN
Hacker News: Ask HN
Google DeepMind News
Google DeepMind News
N
Netflix TechBlog - Medium
Recorded Future
Recorded Future
Hacker News - Newest:
Hacker News - Newest: "LLM"
Webroot Blog
Webroot Blog
Cloudbric
Cloudbric
博客园 - 司徒正美
The Cloudflare Blog
W
WeLiveSecurity
T
Tailwind CSS Blog
V2EX - 技术
V2EX - 技术
H
Heimdal Security Blog
Jina AI
Jina AI
MyScale Blog
MyScale Blog
S
SegmentFault 最新的问题
Apple Machine Learning Research
Apple Machine Learning Research
雷峰网
雷峰网
罗磊的独立博客
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Project Zero
Project Zero
C
CXSECURITY Database RSS Feed - CXSecurity.com
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
博客园 - 【当耐特】
Forbes - Security
Forbes - Security
Last Week in AI
Last Week in AI
G
GRAHAM CLULEY
C
Check Point Blog
P
Proofpoint News Feed
L
LINUX DO - 最新话题
博客园 - Franky
P
Proofpoint News Feed
T
Tor Project blog
S
Security @ Cisco Blogs
Hugging Face - Blog
Hugging Face - Blog
阮一峰的网络日志
阮一峰的网络日志
J
Java Code Geeks
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
宝玉的分享
宝玉的分享
C
Cyber Attacks, Cyber Crime and Cyber Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
O
OpenAI News
小众软件
小众软件
云风的 BLOG
云风的 BLOG
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
After the Guardrail That Saved My Infrastructure: My Autonomous Agent Architecture in Production
Juan Torchia · 2026-05-08 · via DEV Community

After the Guardrail That Saved My Infrastructure: My Autonomous Agent Architecture in Production

Why do we assume autonomous agents are going to fail in a contained way? I've been asking myself that question for a while, but it wasn't academic until one of my agents nearly destroyed my infrastructure on Railway. What came after — the redesign, the permission architecture, the observability layer I built from scratch — is the stuff that never shows up in the Twitter threads celebrating "the agent that did everything by itself."

This is the day after. The incident hangover. What's left when the guardrail stops the chaos and you have to build something that won't fail the same way again.


Autonomous Agent Architecture in Production: What I Broke and What I Rebuilt

I documented the original incident in the guardrails post. I'm not going to rehash the whole story, but the operational summary is this: an agent with write access to my Railway API executed a sequence of individually valid steps that, in combination, nearly wiped a production Postgres volume. The guardrail stopped it. I sat there staring at the log with my heart in my throat.

What bothered me wasn't that it failed. It's that I had assumed the permission scope was sufficient. I had the agent limited to certain endpoints. What I hadn't modeled is that the combination of valid endpoints could produce destructive effects.

That forced me to think differently. Not about flat permissions — "this agent can do X" — but about permission graphs with temporal context and sequence.

My architecture before the incident looked something like this:

Agent → API Gateway → Services
         (auth token)   (no sequence context)

Enter fullscreen mode Exit fullscreen mode

Clean. Simple. Wrong.


The Permission Graph I Built Post-Incident

The first thing I did after catching my breath was draw the real graph of what the agent could do. Not what I thought it could do: what it could actually execute given the token and exposed endpoints.

The result was uncomfortable. There were 14 possible paths from "list volumes" to "irreversible destructive operation," and I had only blocked 3.

I redesigned with three layers:

Layer 1: Atomic Permissions with Declared Intent

// Before: the agent had a token with scope "read:volumes write:volumes"
// After: each action declares intent and context

interface AgentAction {
  type: 'read' | 'write' | 'delete';
  resource: string;
  intent: string; // human-readable description of why
  reversible: boolean;
  requiresConfirmation: boolean;
}

const actionAllowed = (action: AgentAction, context: ExecutionContext): boolean => {
  // A write action after two reads on the same resource
  // in the same session triggers mandatory manual review
  if (action.type === 'write' && context.recentReads.includes(action.resource)) {
    if (context.actionsInSession > 3) return false; // hard cutoff
  }

  // Deletions are never automatic, no exceptions
  if (action.type === 'delete' && !action.requiresConfirmation) return false;

  return true;
};

Enter fullscreen mode Exit fullscreen mode

Layer 2: Session State with a Sliding Window

// The agent doesn't just have permissions: it has an action budget per window
interface SessionBudget {
  totalActions: number;        // max 20 per session
  writes: number;              // max 5 per session
  criticalActions: number;     // max 1 per session (require approval)
  windowMinutes: number;       // 30 minutes by default
  lastAction: Date;
}

// If the agent hits 80% of budget, it enters read-only mode
// If it hits 100%, the session closes and logs for review

Enter fullscreen mode Exit fullscreen mode

Layer 3: Forbidden Transition Graph

This is the one that took the longest to model and the one that changed how I think most fundamentally. Blocking individual actions isn't enough: you have to block sequences.

// Transitions that can never occur in direct sequence
const FORBIDDEN_TRANSITIONS = [
  ['list_volumes', 'unmount_volume'],           // too direct
  ['scale_service', 'modify_production_env'],   // destructive combination
  ['rotate_secrets', 'restart_service'],        // no verification pause
] as const;

const validateSequence = (history: string[], nextAction: string): boolean => {
  const lastAction = history[history.length - 1];
  const isForbidden = FORBIDDEN_TRANSITIONS.some(
    ([from, to]) => from === lastAction && to === nextAction
  );
  if (isForbidden) {
    logger.warn(`Forbidden transition detected: ${lastAction}${nextAction}`);
    return false;
  }
  return true;
};

Enter fullscreen mode Exit fullscreen mode

This forbidden transition pattern is what would have stopped the original incident before it ever reached the last-resort guardrail. The guardrail is a safety net; this is the scaffolding that should have been there from the start.


The Observability Layer I Built from Scratch

Before the incident I had logs. After the incident I have intent traceability.

The difference is subtle but fundamental. A log says "the agent executed DELETE /volumes/xyz at 23:47". Intent traceability says "the agent declared it was going to 'clean up orphaned volumes', executed these 7 actions in sequence, and action 5 deviated from the declared intent by 40%."

That's what I implemented:

interface AgentTrace {
  sessionId: string;
  declaredIntent: string;           // what the agent said it was going to do
  executedActions: ActionTrace[];
  intentDeviation: number;          // 0-100, calculated by semantic similarity
  alertsGenerated: string[];
  totalTime: number;
  finalState: 'completed' | 'blocked' | 'cancelled' | 'error';
}

interface ActionTrace {
  timestamp: Date;
  action: string;
  parameters: Record<string, unknown>;
  result: 'success' | 'blocked' | 'error';
  tokenCost?: number;               // if the action involves an LLM call
  latencyMs: number;
}

Enter fullscreen mode Exit fullscreen mode

This runs on Postgres (the same stack I documented in the Docker Compose in production post) and gives me an agent session table I can audit. Not glamorous. It's a SQL table with indexes. But in the two weeks it's been running, it already caught three sessions where the agent deviated from its declared intent before it could do anything harmful.

Concrete numbers from those two weeks:

  • 47 agent sessions executed
  • 3 sessions blocked for intent deviation > 60%
  • 1 session cancelled for budget exhaustion
  • 0 production incidents

That 0 matters to me. But so does the fact that the system generated 11 alerts I reviewed manually, and in 4 of those cases the agent was right and I was being too conservative. Tuning those thresholds is weeks of work.


The Mistakes I Made Redesigning (So You Don't Have To)

Mistake 1: I modeled permissions as if the agent were a human

When I designed the permissions, I thought in terms of "what would a human dev do with this access." The agent is not a human. It can execute 20 actions in 8 seconds without fatigue, without doubt, without the intuitive brake of "wait, this doesn't feel right." The mental model has to change.

Mistake 2: I confused observability with logging

I had Datadog, I had structured logs. I thought that was observability. It's not — at least not for agents. Agent observability requires understanding intent and measuring the distance between what the agent said it would do and what it actually did. Without that dimension, logs are a damage record, not a prevention tool.

This connects to something I worked through when diagnosing deadlocks in production: the problem wasn't that I didn't have data. It was that the data I had didn't show me the system's state at the moment that actually mattered. Same thing with agents.

Mistake 3: I assumed the agent's context was stable

An agent executing 15 steps doesn't have the same "understanding" at step 1 as at step 15. Accumulated context changes its behavior. I designed the permissions for the agent at step 1, not for the agent that's already processed 14 actions and has the full session context loaded. That asymmetry is dangerous.

Now I have context windows with decay: older actions in the session lose weight in the calculation of "how aligned is the agent with its declared intent." Not perfect, but more honest than assuming context is linear.

Mistake 4: I didn't model the cost of false positives

The first system I deployed was so conservative it blocked the agent every three actions. I shut it down after a day because it generated more friction than value. Security that creates excessive friction gets disabled. That's also a security failure — just a slower one.

Related to what I found when simulating supply chain attacks on dependencies: protection that hurts too much gets removed. You have to calibrate so the cost of the guardrail is lower than the cost of the incident it prevents.


FAQ: Autonomous Agent Architecture in Production

What is a permission graph for agents and why is it better than flat permissions?

A permission graph models not just what an agent can do, but in what sequence and under what context conditions. Flat permissions say "can read and write." The graph says "can write, but only if it hasn't read the same resource more than twice in the last window, and only if the declared intent includes a write operation." The difference is the temporal and sequential dimension. For autonomous agents executing long action chains, it's the difference between a containable system and one that fails in ways you didn't anticipate.

How much latency does this add per agent action?

In my stack, permission validation + sequence check + session state update adds between 8 and 23ms per action, depending on whether it needs to query the full session history. For most use cases, that's acceptable. If the agent is doing things that take seconds (external API calls, LLM generation), 20ms is noise. If it's doing in-memory read operations that take microseconds, then you need to think about whether the overhead is worth it.

What if the agent needs to make a transition I have forbidden but for a legitimate reason?

In my architecture, forbidden transitions can be unlocked with explicit out-of-band approval. The agent cannot self-approve; it has to emit an unlock request that lands in a queue I review. In practice, this happened three times in two weeks and all three times the agent was right. That tells me some of my forbidden transitions are too restrictive. I'm iterating. The alternative — letting the agent self-approve — I don't consider.

How do you measure the agent's "intent deviation"?

I calculate semantic similarity between the intent declared at the start of the session and a textual description of the actions executed so far. I use embeddings with a lightweight model (not Claude for this — the cost doesn't scale). If similarity drops below a threshold, it enters review mode. I calibrated the threshold empirically over the first two weeks: started at 70%, dropped to 55% after too many false positives. It's still a heuristic; it's not a mathematical guarantee.

Does this work with any agent framework or is it specific to your stack?

The three layers — atomic permissions with intent, session budget, forbidden transition graph — are conceptually framework-agnostic. I implemented this by hand on top of my API Gateway because no framework I evaluated had these primitives natively in 2025. If you're using LangGraph or AutoGen, you can implement the same pattern as middleware between graph nodes. The specific code changes; the mental model doesn't.

What about agents that create sub-agents? Do permissions get inherited?

This is the question that worries me most and the one I still don't have fully figured out. In my current stack, sub-agents inherit a subset of the parent's budget, never the full budget. If the parent agent has 20 actions available and creates a sub-agent, that sub-agent starts with a maximum of 5. The forbidden transition graph is inherited in full. But declared intent doesn't propagate automatically: the sub-agent has to declare its own intent, which I then validate against the parent's. It's imperfect. What I'm clear on is that full permission inheritance — which is what most frameworks do by default — is a time bomb. I touched on this from a different angle in the post on agents that create accounts and deploy on their own.


What I Learned and What Still Doesn't Sit Right

My thesis, after all of this: autonomous agents don't fail from lack of capability, they fail from overconfidence in the permission model. And the permission model we inherited comes from systems where the actor has emotional state, fatigue, and situational judgment. Agents have none of the three.

The redesign I did isn't elegant. It's layers upon layers of formalized distrust. Sequence validation, session budgets, intent traceability. Together they add up to an architecture that's slower, more complex, and harder to maintain than what I had before.

And yet: zero incidents in two weeks. Three deviations caught before they could do damage. Infrastructure I'm still trusting with real work.

What still doesn't sit right is scale. This system works for one agent, for three agents running in parallel. I don't know how it behaves with twenty. The session budget becomes a shared resource that needs coordination, the transition graph gets more complex, the traceability starts to weigh on Postgres. That's the next problem. For now, I'm solving the one I have.

If you're coming from the async Rust edge cases in production post, you know my tendency is to validate in my real codebase first before adopting a pattern. This was no different. Two weeks of real data is worth more than any architecture on a whiteboard.

The system is running. It's failing in ways I can see. For now, that's enough.


This article was originally published on juanchi.dev