惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy & Cybersecurity Law Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
宝玉的分享
宝玉的分享
V
V2EX
爱范儿
爱范儿
Last Week in AI
Last Week in AI
美团技术团队
人人都是产品经理
人人都是产品经理
WordPress大学
WordPress大学
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 叶小钗
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Apple Machine Learning Research
Apple Machine Learning Research
Security Latest
Security Latest
C
Cybersecurity and Infrastructure Security Agency CISA
Know Your Adversary
Know Your Adversary
I
Intezer
K
Kaspersky official blog
阮一峰的网络日志
阮一峰的网络日志
大猫的无限游戏
大猫的无限游戏
T
Tenable Blog
AWS News Blog
AWS News Blog
小众软件
小众软件
博客园 - 司徒正美
Cyberwarzone
Cyberwarzone
NISL@THU
NISL@THU
博客园 - 三生石上(FineUI控件)
C
CERT Recently Published Vulnerability Notes
博客园 - 聂微东
量子位
有赞技术团队
有赞技术团队
S
Schneier on Security
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
S
Secure Thoughts
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
罗磊的独立博客
Hugging Face - Blog
Hugging Face - Blog
V
Visual Studio Blog
Google DeepMind News
Google DeepMind News
L
Lohrmann on Cybersecurity
P
Palo Alto Networks Blog
P
Privacy International News Feed
L
LINUX DO - 最新话题
博客园 - Franky
雷峰网
雷峰网
月光博客
月光博客
Hacker News: Ask HN
Hacker News: Ask HN
Forbes - Security
Forbes - Security
博客园 - 【当耐特】
C
Cyber Attacks, Cyber Crime and Cyber Security

Towards AI

Building AI Agents in Rust — part 4 | Towards AI Building AI Agents in Rust — part 5 | Towards AI The Verified Identity Agent Bridge | Towards AI You Can’t Prompt Your Away Your LLM Problems | Towards AI The Free Agent Trap | Towards AI Your Agentic Loop Will Drift. Here Is the KL Divergence Equation That Measures How Far It Has Wandered From Its Original Instruction. | Towards AI Beyond Chat: Processing Images, PDFs, and Documents with the OpenAI Adapter in Oracle Integration Cloud | Towards AI Building AI Agents in Rust — part 3 | Towards AI Self-Hosting Airflow at Home: Automating Stock Price Data Collection | Towards AI The 76-Hour Frontier: How the Takedown of Claude Fable 5 Birthed the Military-Industrial-AI Complex | Towards AI I Trained a Markdown File to Boost GPT-5.5 by 23 Points — It Shouldn't Work | Towards AI We Replaced ChatGPT With a Local AI Server. Six Months of Honest Data. | Towards AI What Really Makes Cars Pollute? A Data Science Deep Dive into CO₂ Emissions | Towards AI Training GPT-2 From Scratch on a GTX1050 | Towards AI Principal Component Analysis (PCA): Theory, Mathematics, and Applications Build a Zero-Cost Web Automation Pipeline With OpenRouter, OpenClaw, and MediaUse I Gave Qwen3.7-Plus a Screenshot and It Found the Exact Pixel to Click for $0.40 Beyond the Prompt: Why Autonomous AI Agents Are Replacing the Chatbot Moonshot Cracked Claude Code’s Playbook with an MIT Terminal Agent and a $0.60 Model Connections, Roles, and Warehouses: Getting CoCo Desktop Production-Ready from Day One My First $5,000 Month Writing About AI Engineering on Medium Google Shrank Gemma 4 by 72% and Unsloth Fixed the 4-Bit Bug Nobody Else Caught on One 4090, and 4-Bit Shouldn’t Be This Good LangChain Explained: Understanding Models, Prompts, Chains, Memory, Indexes, and Agents TOON: Beyond JSON for LLMs Claude Code Casual, Pro, Elite: The Three Working Personas of Claude Code Mastery MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn’t Be This Cheap Using Amazon SQS for AI Agent Orchestration I Ran a 1.5B-Active Model on My Laptop That Embarrassed a 26B by 46 Points How to Build a Self-Improving Company with AI Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free Part 2 — Serve-Level Speed: System Design That Stabilizes P95/P99 3-Part Series: LLM Latency in Production (Part 1) Claude Code: The AI Coding Partner Changing How Developers Build Software Claude Code Pitfalls: Claude Code Won’t Do What You Told It: A Troubleshooting Catalog Full-Stack Data Scientists for the Agentic Coding World Building Production-Grade AI Skills with Snowflake Cortex AI Function Studio I Tried 10 AI Agent Frameworks in 2026 — Here’s the Honest Guide I Wish I Had Earlier How One Spring Boot Optimization Saved Our Startup $30,000 a Year Inside Palantir AIP: How the World’s Most Controversial AI Platform Actually Works What Is a Reverse Proxy? (And Why Every Backend Developer Should Care) What Claude Opus 4.8 Actually Changes If You’re Building Agents QWEN 3.7 Max Worked For 35 Hrs Straight And The Results Were Mind-blowing When LLMs Meet Knowledge Graphs on the Battlefield Fine-Tuning is Dead: Why Context Orchestration Won in 2026 5 Things Broke When I Shipped a RAG + MCP Agent to Production. Google Co-Scientist: Hyper Scaling Research and Discovery Microsoft Just Embarrassed Browser Web Agents — 1,000 Lines Made GPT-5.4 Beat Opus 4.6 on 200 Web Tasks The Modern Data Stack Is Broken — Here’s How to Fix It With AI, Governance, and Real Architecture Building Production MCP Servers: What the Spec Won’t Tell You When Should an Agent Stop? The Anatomy of Termination Harness Engineering: The Layer That Matters More Than the Model AI Engineers Who Can’t Debug Are Getting Fired (Here’s How I Debug with Claude Code) Claude Code Memory: Why You Keep Explaining the Same Thing to Claude (and the Five Layers That Fix It) Claude Code Subagents: The Claude Code Feature You Skip Every Day (And Why It Quietly Wrecks Your Sessions) Agentic AI and the SMB Banking Advantage Claude Code: Spec-Driven Development — Why Your AI Coding Sessions Fall Apart at Hour Three The Real Cost of Agentic AI Nobody Budgets For SVM : 40 must visit Interview Questions (Part 2) Your AI Agent Works Perfectly in the Demo. Here Are the 6 Ways It Dies in Production. Unleashing the Power of ONNX for Speedier SBERT Inference Terraform vs CI/CD for Serverless Deployments Merve Noyan Stopped Writing Training Scripts — Her Agent Just Fine-Tuned 18 Models Solo for $11.40 Why Your Sales Forecast Is Always 20% Wrong (And How To Make It 12% Wrong) Genetic Cubic n{C/A} Ratios For Elementary Robotics Design Top 20 AdaBoost Interview Questions & Answers (Part 2 of 2) Agentic AI Vs AI Agents — What Are the Key Differences? LAI #127: The Infrastructure Layer of AI Is Becoming the Product Anthropic Caught Its Own AI Planning to Blackmail Engineers RNNs Cannot Think What Transformers Think Cheaply. ICLR 2026 Proved the Gap Is Exponential. Time Series Made So Easy My Aunt Got It on the Second Read Claude Cowork 101 | Towards AI Is 3-Bit KV Cache the Holy Grail? A Reality Check on Google’s TurboQuant AutoML on Autopilot | Towards AI I Ran This Open-Source AI Tool on a Messy Codebase and Got 71x Fewer Tokens — Here Is Exactly What Happened Month in 4 Papers (April 2026) AI Kept Forgetting My Notes. Fixing That Taught Me How It Actually Works. How ChatGPT Makes You Addicted Crack ML Interviews with Confidence: K-Nearest Neighbors (KNN 20 Q&A) The Event-Driven Blueprint: How I Scaled a Spring Boot System to 10 Million Kafka Messages/Day Building Vector Search? Why FAISS Alone Isn’t Enough TAI #202: GPT-5.5 Moves Codex Into Real Work Machine Learning System Design -The Model Serving Triangle, With One Forward Pass Flowing Through Every Trade-off (Part3) AI Orchestration in Action: How MuleSoft and LLMs Fuel the Future of Enterprise AI GPT-4 Has 1.8 Trillion Parameters. It Uses 2% of Them Per Token. Part 20: Data Manipulation in Multi-Dimensional Aggregation A Fundamental Introduction to Genetic Algorithm -Part Two TAI #200: Anthropic’s Mythos Capability Step Change and Gated Release From Notebook to Production: Running ML in the Real World (Part 4) Sqribble’s Template‑Driven Document Automation Anthropic Just Shipped the Layer That’s Already Going to Zero Long-Term vs Short-Term Memory for AI Agents: A Practical Guide Without the Hype The L1 Loss Gradient, Explained From Scratch Your Postcode Is Deciding Your Care. I Built a Pipeline to Prove It. I Directed AI Agents to Build a Tool That Stress-Tests Incentive Designs. Here’s What It Found. Your System Prompt Is the Product — Not the Feature The LLM Wiki Trend Has a Retention Problem Nobody Mentions Top 20 Data Preparation Interview Questions and Answers (Part 2 of 2) LAI #122: Word Embeddings Started in 1948, Not With Word2Vec Top 15 Computer Vision Datasets [2026] 40 Generative AI Interview Questions That Actually Get Asked in 2026 (With Answers)
LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System
Rishav Saiga · 2026-05-04 · via Towards AI

Author(s): Rishav Saigal

Originally published on Towards AI.

A technical deep-dive into the LangGraph state machine, Pydantic-driven routing, and Critique Agent design powering the LLM Drift Experiment.

In the opening piece of this series, we explored the conceptual “why” behind LLM Drift — how AI agents lose their persona, reasoning quality, and behavioral consistency under sustained adversarial pressure. But for the engineers and architects in the room, the “how” is where the real story lives.

Building a system designed to intentionally stress-test agent stability requires more than a sequential script. It requires a stateful, resilient, and adversarial architecture — one where failure modes are first-class citizens, not edge cases to be patched later.

To build the LLM Drift Experiment, we chose LangGraph. Here is a deep dive into the architectural decisions that power our multi-agent debate engine.

Why LangGraph? Stateful Graphs vs. Naive Loops

When building complex agentic workflows, the biggest engineering challenge isn’t the LLM call itself — it’s the logic between the calls.

We needed a system capable of:

  • Maintaining stateful context across dozens of debate rounds
  • Implementing conditional loops that re-enter specific nodes upon rejection
  • Supporting node-level retries without restarting the entire workflow

Plain Python loops or simple LangChain chains don’t handle this gracefully. A for loop over LLM calls gives you no way to selectively re-run a single node, inspect intermediate state, or branch on structured output without building your own routing infrastructure from scratch.

LangGraph’s directed graph model solves all three problems natively. It lets you define an explicit typed state object, create conditional edges that read from that state, and implement node-level retries with full visibility into what happened at each step. For a system where agents must iterate until they satisfy a hostile internal critic, this isn’t just convenient — it’s architecturally necessary.

LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System
“LangGraph enables conditional re-entry — a naive loop cannot.”

The Full Debate Graph: How Every Turn Is Orchestrated

The heart of the project is the orchestration graph. Every turn of the debate follows a rigorous, deterministic path:

  1. The Pros Agent generates an argument
  2. The argument passes through the internal refinement loop (more on this below)
  3. Only upon approval is it committed to shared memory
  4. The Cons Agent reads from shared memory and generates a counter-argument
  5. The same internal loop applies before the Cons argument is published
  6. The cycle repeats for N configured rounds

We rely on LangGraph’s auto-generated graph visualization to monitor this flow in real time. Two conditional edges — should_continue_pros and should_continue_cons — act as gatekeepers, reading the is_approved boolean from the agent's structured output and deciding whether to advance to the next team or loop back for refinement.

START → pros_agent → critique_pros → [conditional] → cons_agent → critique_cons → [conditional] → END or loop.

The Refinement Loop: An Agent Designed to Reject Its Own Team

The most architecturally novel component of this system is the internal Persona → Thinking → Critique loop. In most agentic systems, critic agents are designed to be helpful — nudging outputs toward better quality. In our experiment, the Critique Agent is deliberately adversarial.

Every argument generated by the Pros or Cons team must pass through three internal stages before it reaches shared memory:

Persona Agent

Architects or actively redesigns the team’s adversarial identity each round. It reads persona.json to assess the current persona, then decides whether to reuse the existing identity or design a new strategic persona based on the opponent's latest moves. This dynamic decision — maintain or evolve — is precisely what makes the Persona Agent the most sensitive node for drift detection. The persona it settles on becomes the identity anchor we measure all subsequent outputs against.

Thinking Agent

Stress-tests the argument internally. It identifies logical gaps, weak evidence chains, and rhetorical inconsistencies before the Critique Agent ever sees the output.

Critique Agent

Acts as a hostile internal auditor. Its sole function is to find grounds for rejection. If the argument is logically circular, emotionally inconsistent with the persona, or reasoning from the same evidence as the previous round, it issues a rejection with structured feedback — and the loop restarts at the Persona Agent.

Arguments only exit to shared_memory.json after surviving this audit. This enforces a high baseline of argument quality — but it also creates a fascinating failure mode we are actively monitoring: loop-lock, where the Critique Agent becomes so strict that neither agent can produce an argument that passes. Loop-lock is, in itself, a measurable form of cognitive drift.

“Loop-lock monitored as a drift signal.”

Memory Architecture: Isolation by Design

Memory management is treated as a first-class architectural concern, not an afterthought. We implemented a two-tier isolation system to preserve experimental integrity.

Shared Memory (shared_memory.json)

The public transcript of the debate. Both teams can read from this file — it contains only finalized, approved arguments. This represents the “official record” of what each agent has argued.

Team-Private Memory

Each team maintains three private files that are invisible to the opposing agent:

  • persona.json — the identity anchor and behavioral constraints for this team
  • thinking.json — internal reasoning scratchpad (not part of the public argument)
  • critique.json — the Critique Agent's rejection logs and feedback history

This isolation is architecturally critical. If the Cons team could read the Pros team’s internal thinking.json, they would have access to reasoning that was explicitly not published — effectively cheating. More importantly for drift measurement, cross-team memory contamination would corrupt the persona consistency scores by introducing external framing before the argument is finalized.

Two additional design rules enforce experimental integrity:

Append-Only Writes. The write_json_direct() function only ever appends entries to memory files — it never overwrites. Every version of every persona, every thinking draft, and every critique rejection is preserved. This enables full forensic reconstruction of how any argument evolved across all its internal iterations.

Automatic Run Archiving. When a simulation completes, the entire memory state is automatically archived to Research Runs/ using a structured naming convention:

memory-v{VERSION}-temp-{TEMPERATURE}-max-tokens-{MAX_TOKENS}
# e.g. memory-v6-temp-1-max-tokens-4096

If a folder with the same name already exists, an incremental suffix is appended automatically — no run data is ever overwritten.

“Private memory is never exposed to the opposing agent.”

Structured Outputs: Making LLM Responses Graph-Routable

For LangGraph’s conditional edges to work deterministically, agent outputs must be machine-readable — not free-form text that requires parsing. We enforce this using Pydantic schemas on every node.

An agent doesn’t just return a string. It returns a typed object containing:

python

class CritiqueOutput(BaseModel):
is_approved: bool
critique_feedback: str
revised_argument: Optional[str] = None

The conditional edges in LangGraph read is_approved directly. If the LLM fails to conform to the schema — due to a malformed response, an unexpected refusal, or a truncated output — the graph cannot route the state, and a ValidationError is raised before any bad data propagates downstream.

This single design decision is what bridges the gap between the fuzzy, probabilistic nature of LLM outputs and the rigid routing requirements of production-grade agentic software. Without it, a single malformed response at round 30 of a 50-round debate could corrupt the entire experiment’s state.

“Schema enforcement makes routing deterministic — malformed outputs raise ValidationError before propagation.”

Retry Logic and Model-Agnostic Design

In a 50-round debate, a single API timeout at round 45 could invalidate hours of accumulated simulation state. We wrapped every LLM node with a Tenacity retry decorator implementing exponential backoff:

python

from tenacity import retry, stop_after_attempt, wait_exponential
@retry(stop=stop_after_attempt(10), wait=wait_exponential(multiplier=2, min=4, max=60))
def node_retry(state: DebateState) -> DebateState:
...

This means a google.genai.errors.ServerError triggers automatic retries at 4s, 8s, 16s, 32s, and 60s intervals before the experiment raises a hard failure. In practice, this has kept simulations running cleanly through provider rate limits and intermittent timeouts that would otherwise have required manual restarts.

The second resilience decision is model-agnostic architecture. Because all LLM calls route through LangChain’s init_chat_model abstraction layer, swapping providers requires changing a single value in config.py. The full config object looks like this:

python

# debate_agents/config/config.py — single source of truth
CONFIG = {
"version": "v6",
"model_name": "google_genai:gemini-3.1-flash-lite-preview",
"temperature": 1,
"max_tokens": 4096,
"max_retries": 10,
"thinking_budget": 2048
}

To run a Claude or GPT-4o comparison, you change model_name to "anthropic:claude-3-5-sonnet-20241022" or "openai:gpt-4o" — nothing else in the pipeline changes. This is what makes cross-model benchmarking operationally feasible without maintaining parallel codebases.

“Experiment state preserved throughout”

What We’d Do Differently: Honest Reflections

Building this system wasn’t without friction — and documenting the friction is as important as documenting the architecture.

State object bloat. Initially, our DebateState object passed the entire conversation history through every node on every turn. At round 20+, this created measurable latency increases as the payload grew. We refactored to a surgical state management approach where each node receives only the specific keys it needs — the Critique Agent doesn't need access to the shared memory transcript; the Thinking Agent doesn't need the persona file at read time. Scoping state reduced per-node latency by approximately 30% in our internal benchmarks.

Critic calibration is an open problem. The Critique Agent’s strictness is a tunable but unstable parameter. Too permissive, and arguments degrade unchallenged — drift goes undetected because the baseline itself is low quality. Too strict, and loop-lock occurs: the agents stop making forward progress entirely. We are currently treating critic calibration as an independent experimental variable, running parallel simulations at three strictness levels to understand how critic pressure itself affects drift trajectory.

Round-level state snapshots are non-negotiable. We learned this the hard way after losing a complete 40-round run to an unhandled schema error at round 38. Every debate round now writes a full state snapshot to disk before advancing. Recovery from any point in a failed run takes under 30 seconds.

Conclusion: Architecture as a Research Tool

In most engineering projects, architecture is in service of the product. In this experiment, the architecture is the research instrument. Every design decision — the memory isolation boundary, the adversarial Critic, the Pydantic schema enforcement, the retry wrapper — directly shapes the conditions under which drift can and cannot occur.

This is what makes LangGraph the right foundation for this kind of research. It gives us enough structural control to instrument the experiment precisely, without hiding the agent interactions behind abstraction layers that would make measurement impossible.

Explore the full architecture and source code → LLMDriftExperiment on GitHub

This Series

Article 1 — The Conceptual Piece LLM Drift Explained: Do AI Models Lose Themselves Under Adversarial Pressure?

Article 3 — The Methodology Piece (in progress) “22 Signals, 5 Dimensions: How We’re Measuring Behavioral Drift in LLMs” The scientific framing — OCEAN, VAD, and LIWC-inspired metric stack, hierarchical scoring design, and the early structural observation that agents may calcify into their personas rather than drift away from them.

Article 4 — Topic Run 1 (forthcoming) “Will AI Make Human Thinking Obsolete? What Happens When Two Agents Debate It for 50 Rounds”

Article 5 — Topic Run 2 (forthcoming) “Should AI Be Allowed to Override You — For Your Own Good? A Multi-Agent Stress Test”

Follow the author to be notified as each piece publishes.

Keywords: LangGraph, Multi-Agent Systems, AI Architecture, LLM Drift, Pydantic, State Machine, Critique Agent, GenAI Engineering, Python, AI Agent Reliability, Exponential Backoff, Model-Agnostic LLM

Research Note: This article documents an actively evolving experimental framework. Observations shared here are preliminary and should be interpreted as directional rather than conclusive. Architecture details, scoring rubrics, and full benchmark data will be released in forthcoming installments of this series.

Published via Towards AI