惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

B
Blog RSS Feed
Martin Fowler
Martin Fowler
爱范儿
爱范儿
IT之家
IT之家
Last Week in AI
Last Week in AI
A
About on SuperTechFans
Google DeepMind News
Google DeepMind News
阮一峰的网络日志
阮一峰的网络日志
V
V2EX
aimingoo的专栏
aimingoo的专栏
G
Google Developers Blog
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
美团技术团队
The Cloudflare Blog
MyScale Blog
MyScale Blog
T
The Blog of Author Tim Ferriss
Hugging Face - Blog
Hugging Face - Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
云风的 BLOG
云风的 BLOG
Y
Y Combinator Blog
The GitHub Blog
The GitHub Blog
腾讯CDC
Microsoft Security Blog
Microsoft Security Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
I Built a Debugger for LLM Agents — Here's Why "Observabi...
Raju Shaniga · 2026-05-19 · via DEV Community

Raju Shanigarapu

Every time I changed a prompt, I was running a hypothesis test.

But I had no debugger. No way to pause execution. No structural comparison between "before" and "after." Just two terminal windows and a vague feeling that maybe it was better now.

I built agent-lens to fix this.


The Problem with "Observability"

Langfuse, LangSmith, Phoenix — these are great tools. They show you what happened. Traces, spans, token counts.

But none of them answer the question I actually had: did this change make it better?

That requires something different:

  • A way to compare two runs structurally
  • A record of why you made the change (the hypothesis)
  • A verdict — not just "here are the numbers," but "this was an improvement"

What agent-lens Does Differently

1. Pause a live agent mid-run

import agent_lens
from openai import OpenAI

agent_lens.install()          # auto-patches OpenAI + Anthropic
agent_lens.dashboard.start()  # localhost:7878

client = OpenAI()

@agent_lens.trace
def my_agent(query: str) -> str:
    return client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": query}]
    ).choices[0].message.content

Enter fullscreen mode Exit fullscreen mode

Open the dashboard, click Pause. The agent blocks at the next LLM call.

2. State a hypothesis before you change anything

POST /runs/{run_id}/fork
{
  "span_id": "abc123",
  "edited_messages": [{"role": "system", "content": "Be concise."}],
  "notes": "Hypothesis: shorter system prompt reduces hallucination",
  "expected_output": "concise"
}

Enter fullscreen mode Exit fullscreen mode

The note travels with the run forever. Future you can read your reasoning.

3. GET /diff — one call, one verdict

GET /runs/{run_a}/diff/{run_b}

Enter fullscreen mode Exit fullscreen mode

{
  "metrics_delta": {
    "latency_ms":   {"a": 1847, "b": 820,  "pct_change": -55.6},
    "total_tokens": {"a": 453,  "b": 87,   "pct_change": -80.8},
    "cost_usd":     {"a": 0.0045, "b": 0.00087, "pct_change": -80.7}
  },
  "assertion_result": {
    "expected_output": "concise",
    "passed_in_a": false,
    "passed_in_b": true,
    "verdict": "improved"
  }
}

Enter fullscreen mode Exit fullscreen mode

Hypothesis confirmed. With numbers.


The Full Flow

[Agent running] → Pause → agent blocks at next LLM call
                              ↓
                    [Edit messages in dashboard]
                              ↓
                    Fork → new run diverges
                              ↓
                    Resume → original continues
                              ↓
              [Two runs. GET /diff. Get verdict.]

Enter fullscreen mode Exit fullscreen mode

No restarts. No re-running preceding steps.


Zero Infrastructure

Everything runs locally. SQLite at ~/.agent-lens/runs.db. No Docker. No cloud. No API keys needed to start exploring:

pip install agentlens-tracer
python examples/07_demo_mock.py  # runs a full demo with no API key

Enter fullscreen mode Exit fullscreen mode


Works with LangChain and LlamaIndex Too

from agent_lens.integrations.langchain import AgentLensCallbackHandler
from agent_lens.integrations.llamaindex import AgentLensLlamaIndexHandler

Enter fullscreen mode Exit fullscreen mode

Pass as a callback — every LLM call is traced automatically.


Why This Matters

You're not debugging a function. You're debugging a probabilistic system. Every prompt change is a hypothesis test.

Today you run that test by eyeballing outputs. agent-lens makes it structural, repeatable, and recorded.

Vibes-based prompt engineering is debugging without a debugger.
agent-lens is the debugger.


GitHub: https://github.com/RAJUSHANIGARAPU/agent-lens
Install: pip install agentlens-tracer

Would love to hear how you're currently debugging LLM agents — drop a comment below.