惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

H
Hackread – Cybersecurity News, Data Breaches, AI and More
The Last Watchdog
The Last Watchdog
T
Threatpost
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
S
Security @ Cisco Blogs
S
Securelist
L
LINUX DO - 最新话题
The Hacker News
The Hacker News
S
SegmentFault 最新的问题
C
Cyber Attacks, Cyber Crime and Cyber Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
博客园_首页
博客园 - Franky
H
Heimdal Security Blog
G
Google Developers Blog
N
News and Events Feed by Topic
Cisco Talos Blog
Cisco Talos Blog
Engineering at Meta
Engineering at Meta
B
Blog RSS Feed
S
Schneier on Security
T
Threat Research - Cisco Blogs
D
DataBreaches.Net
Simon Willison's Weblog
Simon Willison's Weblog
Hacker News: Ask HN
Hacker News: Ask HN
WordPress大学
WordPress大学
Latest news
Latest news
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 聂微东
N
Netflix TechBlog - Medium
T
Tor Project blog
月光博客
月光博客
D
Docker
美团技术团队
Recent Announcements
Recent Announcements
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Cyberwarzone
Cyberwarzone
小众软件
小众软件
TaoSecurity Blog
TaoSecurity Blog
Blog — PlanetScale
Blog — PlanetScale
L
LINUX DO - 热门话题
O
OpenAI News
人人都是产品经理
人人都是产品经理
www.infosecurity-magazine.com
www.infosecurity-magazine.com
V
V2EX
C
Cisco Blogs
NISL@THU
NISL@THU
Recent Commits to openclaw:main
Recent Commits to openclaw:main
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 司徒正美

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
SAGA Made Microservices Reliable. Agent Harness Makes AI Agents Reliable.
Seenivasa Ramadurai · 2026-06-14 · via DEV Community

The distributed systems world solved long-running transactions with SAGA. The agentic AI world has a harder version of the same problem. Here's how Agent Harness answers it.

Introduction

I've been deep in agentic AI architecture for a while now & building Digital Workers, designing multi-agent systems, working through the messy production realities of agents that call tools, consult knowledge bases, and loop back on themselves when they're uncertain. And one question keeps coming up when I talk to engineers who come from a microservices background: "Can't we just use SAGA for this?"

It's a fair question. SAGA is one of the more elegant patterns in distributed systems. And on the surface, agentic workflows look similar enough that the analogy is tempting. Both involve coordinating multi-step processes. Both need state management and failure recovery. Both have to deal with partial completions.

But the moment you dig into the details, you realize why SAGA alone isn't enough and why Agent Harness exists.

What SAGA Was Built to Solve

If you've spent time in microservices land, you've lived this problem. Service A completes, Service B completes, Service C fails and now you have a half-committed distributed transaction with no clean rollback and no database level guarantee to save you.

The SAGA pattern was invented for exactly this. The break long-running transactions into a sequence of local steps, and for every step that can succeed, write a compensating action in advance so that if something downstream fails, you can undo the damage cleanly.

It works beautifully because microservices operate in a deterministic world. Every service has a known API contract. Every response has a typed schema. Every failure is a status code or a typed exception. Every retry is predictable. The failure modes are knowable at design time, so you can write compensation logic at design time.

AI agents don't live in that world.

The Problem Is Probabilistic, Not Deterministic

Here's what fundamentally changes when you move from microservices to agentic AI systems, your "services" are now LLM calls, tool invocations, knowledge retrievals, external APIs or MCP Server tool calls **, and **increasingly human approvals. None of these behave like a well defined REST endpoint with a contract you can write compensation logic against.

An LLM call can return an answer that passes every syntax check but is semantically wrong confidently, fluently, plausibly wrong. A tool call might succeed at the HTTP layer but return data that sends the agent down an entirely incorrect reasoning path. A multi-step task might "complete" having taken three hallucinated intermediate steps before landing somewhere that superficially looks like the goal.

And here's the part that should give you pause: a SAGA coordinator would mark all of that as success. No exceptions. No compensation triggered. Workflow complete.

Retrying won't fix it. Compensation logic won't fix it. You need something architecturally different: an Agent Harness.

What SAGA and Agent Harness Actually Share

Before getting into where they diverge, it's worth being honest about the parallel because it isn't just a clever analogy. It's structurally real.

Both patterns exist to solve the same core problem: coordinating multi-step processes where individual steps can fail, state needs to be preserved across the lifecycle, and the overall system needs to recover gracefully when things go sideways.

The SAGA Coordinator manages: state tracking, retries, compensation actions, failure recovery, workflow sequencing, and distributed reliability. The Agent Harness manages all of those same things just mapped to a completely different execution model.

[The architecture maps cleanly. The implementation is night and day.]

What Agent Harness Does That SAGA Cannot

SAGA assumes your workflow steps are atomic and deterministic. Agent Harness has to deal with steps that are neither. That's why it needs an entire category of capabilities that have no real SAGA equivalent:

Memory (Short & Long Term): An agent working a multi-turn task needs to remember what it decided three steps ago, what the user said at the start, and what it already tried that didn't work. That's not transaction state. That's episodic memory and working context interleaved in a way that needs to survive tool calls, retries, and mid-task handoffs.

Reflection & Critique: Before committing to an action or an answer, a well designed harness routes the agent's proposed output through a self critique step. Did the answer actually address the stated goal? Does it contradict something established earlier in the session? Does it fall outside the policy boundaries? SAGA never needs to ask its services whether they feel confident about their output. Agent Harness does.

Guardrails & Policies: In production especially in regulated industries you don't want an agent calling a sensitive external API, accessing PII, or making a consequential decision without policy enforcement at the harness level. This isn't exception handling after the fact. It's proactive constraint evaluation before execution. I've seen this matter enormously in healthcare projects where the consequences of an unguarded tool call are real.

Human-in-the-Loop: SAGA runs unattended by design. Agent Harness needs to know when to stop and ask a human and that decision happens at the semantic level, not the infrastructure level. "I'm not certain this is what the user intended" is a fundamentally different pause condition than "the API returned a 503."

Evaluation & Validation: Did the agent's output actually achieve the goal? Not "did the tool call succeed" did we actually do what we set out to do? This requires goal level evaluation, not just a success/failure bit. It's one of the harder things to operationalize in practice, but skipping it is how you ship agents that complete tasks without accomplishing goals.

Cost & Token Monitoring: LLM calls have variable cost depending on context length, model tier, and how deep the reasoning goes. An agent running a complex multi-step task can burn through budget in ways that are invisible until you get the bill. A production Agent Harness needs token spend guardrails the way a microservices platform needs circuit breakers on latency.

Durable Execution via Checkpointing: If an agent task runs for 40 minutes and the process crashes at minute 39, checkpointing lets you resume from the last stable state rather than starting over. Philosophically similar to SAGA's compensating transactions but the implementation means serializing agent state, tool call history, memory contents, and intermediate reasoning. Substantially more complex, and substantially more necessary for long horizon tasks.

The Concrete Scenario That Makes This Real

Let me give you a specific example, because abstract architecture arguments only go so far.

Imagine an agent tasked with: "Research our top three competitors' pricing pages and prepare a comparison summary for the sales team."

A SAGA style system would model this as: call tool to fetch Page A → call tool to fetch Page B → call tool to fetch Page C → call tool to generate summary → done. If any fetch fails, compensate. If all fetches succeed, the workflow completes.

But here's what can actually happen: Page B returns a cached version from 2 months ago. The agent doesn't know that it just sees valid HTML. It processes the outdated pricing as current. The summary it generates is factually wrong in a way that could embarrass your sales team.

Every step "succeeded." The SAGA coordinator marks it complete. No compensation triggered. And your sales team walks into a meeting with incorrect competitive data.

Agent Harness addresses this at multiple layers. Reflection catches that the retrieved content has anomalous date markers. Evaluation validates whether the output meets the quality criteria defined for the task. Guardrails can flag when retrieved content falls below a freshness threshold. Human-in-the-loop escalation routes the uncertainty to a person rather than silently proceeding.

That's the gap. And it's not a small one.

The Key Difference, Plainly Said

SAGA manages deterministic workflows. Agent Harness manages probabilistic workflows.

In SAGA, failure modes are knowable at design time. You write compensation logic once and trust it to cover the cases. In an Agent Harness, failure can mean: the tool returned a valid response that the agent misread. Or the agent completed every step correctly but arrived at a goal that doesn't satisfy what the user actually wanted. Or the agent is in a soft reasoning loop, re-checking the same condition because it's genuinely uncertain and nobody told it when to escalate.

Handling that requires reflection, self critique, goal validation, and graceful human escalation none of which exist in the SAGA vocabulary, because SAGA was never designed for an execution unit that reasons about the world.

What This Means If You're Building Today

If you're designing an agentic system and you're thinking purely in SAGA terms, you're probably building something that's reliable at the infrastructure layer but brittle at the reasoning layer. Your agents will retry correctly. They'll compensate correctly. But they'll also confidently produce wrong answers, hallucinate tool results, and mark tasks complete that aren't — and your coordinator will have no way to know the difference.

Agent Harness is the layer that closes that gap. It's not a replacement for orchestration. It sits above orchestration and asks: did we actually do the right thing, in the right way, within the right constraints, with the appropriate level of human oversight?

The engineers who built SAGA were solving a genuinely hard distributed systems problem. The people building Agent Harness today are solving a harder version of it because the failure modes are less visible, the state is messier, and "success" is much harder to define when your execution unit is a language model reasoning about an open-ended goal.

But the spirit is exactly the same: build systems that fail gracefully, recover intelligently, and complete what they started.

SAGA made microservices reliable. Agent Harness is what makes AI agents reliable.

One Question Worth Sitting With

Of all the Agent Harness components, I've found that Reflection & Critique and Human-in-the-Loop are the two that teams most consistently underinvest in usually because they're harder to wire up than checkpointing or token monitoring, and the cost of skipping them isn't visible until something goes wrong in production.

Which component do you find hardest to implement in practice and how are you handling it? I'm genuinely curious what patterns the community is landing on. Drop it in the comments.

Thanks
Sreeni Ramadorai