惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
GbyAI
GbyAI
Jina AI
Jina AI
博客园_首页
Y
Y Combinator Blog
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
M
MIT News - Artificial intelligence
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
MyScale Blog
MyScale Blog
MongoDB | Blog
MongoDB | Blog
雷峰网
雷峰网
罗磊的独立博客
博客园 - Franky
Last Week in AI
Last Week in AI
Vercel News
Vercel News
Martin Fowler
Martin Fowler
Stack Overflow Blog
Stack Overflow Blog
Microsoft Security Blog
Microsoft Security Blog
月光博客
月光博客
WordPress大学
WordPress大学
W
WeLiveSecurity
Apple Machine Learning Research
Apple Machine Learning Research
TaoSecurity Blog
TaoSecurity Blog
阮一峰的网络日志
阮一峰的网络日志
博客园 - 聂微东
Schneier on Security
Schneier on Security
Engineering at Meta
Engineering at Meta
Simon Willison's Weblog
Simon Willison's Weblog
博客园 - 【当耐特】
宝玉的分享
宝玉的分享
大猫的无限游戏
大猫的无限游戏
S
SegmentFault 最新的问题
P
Proofpoint News Feed
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - 叶小钗
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
L
LINUX DO - 最新话题
S
Security @ Cisco Blogs
B
Blog RSS Feed
B
Blog
爱范儿
爱范儿
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Tailwind CSS Blog
Attack and Defense Labs
Attack and Defense Labs
V
Visual Studio Blog
NISL@THU
NISL@THU
U
Unit 42
Hugging Face - Blog
Hugging Face - Blog
A
Arctic Wolf

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Debugging Multi-Agent Systems in TypeScript: From Flat Logs to Execution Trees
chintanonweb · 2026-05-18 · via DEV Community

AI agents are easy to demo when they follow a clean path: receive a task, call a tool, produce an answer, and finish successfully.

They become much harder to reason about when multiple agents run together.

In a real system, agents may plan, call tools, retry failures, make decisions from stale state, run in parallel, or touch the same resource from different paths. When something breaks, flat logs usually tell us what happened, but they rarely show why it happened.

That is the debugging gap I wanted to explore.

So I built a small TypeScript-based multi-agent incident-response simulator. The goal was simple: simulate a production incident where multiple agents diagnose and remediate infrastructure problems. The system had a diagnostic agent, database agent, network agent, scaling agent, and coordinator agent.

On paper, the design looked reasonable.

The DiagnosticAgent analyzed the incoming incident. The DatabaseAgent handled database-related issues. The NetworkAgent managed load balancer or routing problems. The ScalingAgent handled capacity decisions. The CoordinatorAgent orchestrated everything and was responsible for avoiding conflicting actions.

The architecture looked clean until the agents started working at the same time.

The Problem With Flat Logs

In the first version, the simulator emitted logs like this:

\[2:47:23\] DiagnosticAgent: High DB latency detected  
\[2:47:24\] DatabaseAgent: Initiating replica scale-up  
\[2:47:25\] DiagnosticAgent: Connection pool exhaustion detected  
\[2:47:26\] DatabaseAgent: Taking node-3 offline for maintenance  
\[2:47:27\] ScalingAgent: Database performance degraded, scaling up  
\[2:47:28\] NetworkAgent: Detected backend failures, restarting load balancer  
\[2:47:29\] CoordinatorAgent: Conflict detected  
\[2:47:32\] ERROR: Cluster quorum lost

Enter fullscreen mode Exit fullscreen mode

These logs were useful, but only up to a point.

They showed that the database agent scaled replicas. They showed that another agent also tried to scale. They showed that a node was taken offline. They showed that the coordinator noticed a conflict.

But they did not clearly answer the important questions:

Which agent made a decision from stale state?

Did the coordinator run before or after the conflicting tool calls?

Were the database and scaling agents truly running in parallel?

Which exact tool call caused the final failure?

Was the problem an LLM decision, a tool execution issue, or a coordination issue?

This is where normal logging started to feel too flat. The system behavior was no longer a simple list of events. It was a tree of decisions, tool calls, retries, and parallel branches.

That is when I tried agent-inspect.

Adding Local Execution Tracing

agent-inspect is a local-first execution tree debugger for TypeScript and Node.js AI agents. Instead of sending traces to a hosted dashboard, it writes local traces that can be inspected from the terminal.

That local-first model is important during development. I did not want to set up a full observability platform just to understand one local agent run. I wanted something closer to a structured debugging layer between console.log and production-grade observability.

The first step was to wrap the coordinator flow.

import { inspectRun, step } from "agent-inspect";

async function handleIncident(incident: Incident) {  
 return inspectRun(  
   "incident-response-coordinator",  
   async () \=\> {  
     const diagnosis \= await step("diagnose-incident", async () \=\> {  
       return diagnosticAgent.analyze(incident);  
     });

     const actions \= await step("execute-remediation", async () \=\> {  
       return Promise.all(\[  
         step.tool("database-remediation", () \=\>  
           databaseAgent.handleIssue(diagnosis.dbIssues)  
         ),  
         step.tool("network-remediation", () \=\>  
           networkAgent.handleIssue(diagnosis.networkIssues)  
         ),  
         step.tool("scaling-remediation", () \=\>  
           scalingAgent.handleIssue(diagnosis.scalingIssues)  
         ),  
       \]);  
     });

     return step("resolve-conflicts", async () \=\> {  
       return resolveConflicts(actions);  
     });  
   },  
   {  
     traceDir: "./.agent-inspect",  
   }  
 );  
}

Enter fullscreen mode Exit fullscreen mode

The code did not need a full rewrite. The main change was adding meaningful boundaries around the work.

The outer inspectRun represented one agent run. The normal step calls represented logical phases. The step.tool calls marked operations that touched external systems or simulated infrastructure.

Then I instrumented the database agent.

class DatabaseAgent {  
 async handleIssue(issues: DbIssue\[\]) {  
   return step("database-agent-execution", async () \=\> {  
     const dbState \= await step.tool("check-db-state", async () \=\> {  
       return this.getClusterState();  
     });

     const decision \= await step.llm("decide-db-action", async () \=\> {  
       return this.llm.chat({  
         messages: \[  
           {  
             role: "user",  
             content: JSON.stringify({  
               task: "Decide the safest database remediation action",  
               issues,  
               dbState,  
             }),  
           },  
         \],  
       });  
     });

     if (decision.action \=== "scale-up") {  
       return step.tool("scale-database", async () \=\> {  
         return this.scaleUpReplicas(decision.targetCount);  
       });  
     }

     if (decision.action \=== "restart-node") {  
       return step.tool("restart-node", async () \=\> {  
         return this.restartNode(decision.nodeId);  
       });  
     }

     return {  
       action: "no-op",  
       reason: "No safe database action selected",  
     };  
   });  
 }  
}

Enter fullscreen mode Exit fullscreen mode

The important part is not just the tracing. It is the naming.

A trace is only useful if the steps describe the system in the same language engineers use during debugging. check-db-state, decide-db-action, scale-database, and restart-node are much more useful than generic messages like running task or tool call started.

Inspecting the Failed Run

After running the simulator, I listed the local traces:

npx agent-inspect list --dir ./.agent-inspect

Then I inspected the failed run:

npx agent-inspect view <run-id> --dir ./.agent-inspect

The execution tree made the issue much easier to understand:

incident-response-coordinator                              \[47.2s\] ✗  
├─ diagnose-incident                                       \[3.1s\] ✓  
├─ execute-remediation                                     \[41.8s\] ✗  
│  ├─ database-remediation                                 \[23.2s\] ✓  
│  │  └─ database-agent-execution                          \[23.1s\] ✓  
│  │     ├─ check-db-state                                 \[0.4s\] ✓  
│  │     ├─ decide-db-action                               \[2.1s\] ✓  
│  │     ├─ scale-database                                 \[18.3s\] ✓  
│  │     ├─ check-db-state                                 \[0.3s\] ✓  
│  │     ├─ decide-db-action                               \[1.9s\] ✓  
│  │     └─ restart-node                                   \[0.3s\] ✓  
│  ├─ network-remediation                                  \[5.2s\] ✓  
│  └─ scaling-remediation                                  \[41.7s\] ✗  
│     └─ scaling-agent-execution                           \[41.6s\] ✗  
│        ├─ check-scaling-state                            \[0.3s\] ✓  
│        ├─ decide-scaling-action                          \[2.2s\] ✓  
│        └─ scale-database                                 \[39.1s\] ✗  
│           └─ Error: Operation timeout \- cluster in inconsistent state  
└─ resolve-conflicts                                       \[not reached\]

Enter fullscreen mode Exit fullscreen mode

This view showed the problem more clearly than the logs.

The database agent checked the state, decided to scale up, and started a database scaling operation. Then it checked state again and decided to restart a node. At the same time, the scaling agent also detected database pressure and started another scaling operation.

Both agents were acting on the same resource. Both believed their action was valid. The coordinator was supposed to resolve conflicts, but the trace showed that resolve-conflicts was never reached because the failure happened inside the parallel remediation step.

That was the real bug.

It was not simply a bad prompt. It was not only a database operation failure. It was a coordination bug caused by parallel agents acting on the same resource without a proper resource-level guard.

Fixing the Coordination Model

Once the execution tree made the failure visible, the fix became much more direct.

The first change was to add a state refresh guard. If the database cluster already had an operation in progress, the agent should wait for stable state before making another decision.

async function handleIssue(issues: DbIssue\[\]) {  
 return step("database-agent-execution", async () \=\> {  
   const dbState \= await step.tool("check-db-state", async () \=\> {  
     return this.getClusterState();  
   });

   if (dbState.hasInProgressOperations) {  
     return step("wait-for-stability", async () \=\> {  
       await this.waitForStableState();  
       return this.handleIssue(issues);  
     });  
   }

   return this.decideAndExecute(issues, dbState);  
 });  
}

Enter fullscreen mode Exit fullscreen mode

The second change was to protect critical operations with a lock.

async function scaleUpReplicas(targetCount: number) {  
 return step.tool("scale-database", async () \=\> {  
   const lock \= await this.acquireLock("database-scaling", 60\_000);

   try {  
     return this.performScaleUp(targetCount);  
   } finally {  
     await lock.release();  
   }  
 });  
}

Enter fullscreen mode Exit fullscreen mode

The third change was at the coordinator level. If multiple agents wanted to touch the same resource, the coordinator should not blindly run them in parallel.

const actions \= await step("execute-remediation-sequenced", async () \=\> {  
 const targets \= identifyResourceTargets(diagnosis);

 if (targets.database.length \> 0\) {  
   const dbActions \= await step.tool("database-remediation", () \=\>  
     databaseAgent.handleIssue(diagnosis.dbIssues)  
   );

   const networkActions \= await step.tool("network-remediation", () \=\>  
     networkAgent.handleIssue(diagnosis.networkIssues)  
   );

   return {  
     dbActions,  
     networkActions,  
   };  
 }

 return Promise.all(\[  
   step.tool("network-remediation", () \=\>  
     networkAgent.handleIssue(diagnosis.networkIssues)  
   ),  
   step.tool("scaling-remediation", () \=\>  
     scalingAgent.handleIssue(diagnosis.scalingIssues)  
   ),  
 \]);  
});

Enter fullscreen mode Exit fullscreen mode

After the fix, the trace looked different:

incident-response-coordinator                              \[15.3s\] ✓  
├─ diagnose-incident                                       \[2.8s\] ✓  
├─ execute-remediation-sequenced                           \[11.2s\] ✓  
│  └─ database-remediation                                 \[8.4s\] ✓  
│     └─ database-agent-execution                          \[8.3s\] ✓  
│        ├─ check-db-state                                 \[0.3s\] ✓  
│        ├─ acquire-lock                                   \[0.1s\] ✓  
│        ├─ decide-db-action                               \[1.9s\] ✓  
│        ├─ scale-database                                 \[5.8s\] ✓  
│        └─ release-lock                                   \[0.1s\] ✓  
└─ resolve-conflicts                                       \[1.3s\] ✓

Enter fullscreen mode Exit fullscreen mode

This is the kind of output I want during agent development.

Not just “something failed,” but where it failed. Not just “the tool timed out,” but what sequence caused the timeout. Not just “agents ran in parallel,” but which branches actually overlapped.

Why This Matters for AI Agent Engineering

As agent systems become more common, debugging needs to move beyond raw logs.

A single-agent workflow can often be debugged with a few log statements. But multi-agent systems introduce coordination problems. A bug may not live inside one function. It may live between two valid decisions that become unsafe when executed together.

That is why execution trees are useful.

They show the structure of the run. They show parent-child relationships. They separate normal logic from tool calls and LLM calls. They make retries, skipped steps, failed branches, and slow operations easier to reason about.

This also changes how we think about observability.

Production observability platforms are still important. Tools like LangSmith, Langfuse, OpenTelemetry-based pipelines, and APM platforms solve important team and production problems. But during local development, I often want something lighter. I want to run the agent, inspect the trace, make a change, and compare the result.

That is the space where a local-first tool like agent-inspect fits naturally.

It is not trying to replace production monitoring. It is closer to a developer workflow tool for understanding agent behavior before it reaches production.

Practical Lessons From the Project

The first lesson is that flat logs hide structure. In a multi-agent workflow, order alone is not enough. You need to know which step belonged to which agent, which steps were siblings, and which operation blocked or failed.

The second lesson is that not every agent bug is an LLM bug. In this simulator, the expensive failure came from tool coordination and stale state, not from a slow model call. Without tracing, it would have been easy to spend time tuning prompts while ignoring the actual failure path.

The third lesson is that instrumentation can become living documentation. A well-named step() call describes the architecture. When a new engineer reads the trace, they can understand the runtime behavior faster than reading scattered logs.

The fourth lesson is that local-first debugging is still valuable. Not every debugging session needs a dashboard, collector, account, or cloud upload. Sometimes the fastest path is a local trace file and a terminal command.

Final Thoughts

The more I build with AI agents, the more I feel that debugging is becoming an architecture problem.

It is not enough to know that an agent produced the wrong answer. We need to know what it planned, which tools it called, which state it observed, which branches ran in parallel, where retries happened, and what changed between two runs.

For TypeScript and Node.js teams building agentic systems, agent-inspect is a useful tool to explore that workflow. It gives you a lightweight way to turn agent runs into readable execution trees without committing to a hosted observability setup on day one.

For my multi-agent incident-response simulator, the biggest value was simple: it turned a confusing wall of logs into a system I could reason about.

And that is usually the first step toward making agent systems reliable.