惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园_首页
PCI Perspectives
PCI Perspectives
T
Tailwind CSS Blog
月光博客
月光博客
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
V
V2EX
D
Docker
P
Proofpoint News Feed
阮一峰的网络日志
阮一峰的网络日志
博客园 - 司徒正美
酷 壳 – CoolShell
酷 壳 – CoolShell
云风的 BLOG
云风的 BLOG
H
Help Net Security
The Register - Security
The Register - Security
宝玉的分享
宝玉的分享
C
Check Point Blog
T
Threatpost
The GitHub Blog
The GitHub Blog
P
Privacy International News Feed
G
Google Developers Blog
博客园 - Franky
爱范儿
爱范儿
T
Tor Project blog
博客园 - 聂微东
Google DeepMind News
Google DeepMind News
G
GRAHAM CLULEY
雷峰网
雷峰网
Cyberwarzone
Cyberwarzone
人人都是产品经理
人人都是产品经理
C
Cybersecurity and Infrastructure Security Agency CISA
Vercel News
Vercel News
Scott Helme
Scott Helme
aimingoo的专栏
aimingoo的专栏
Martin Fowler
Martin Fowler
MyScale Blog
MyScale Blog
Last Week in AI
Last Week in AI
GbyAI
GbyAI
Microsoft Azure Blog
Microsoft Azure Blog
腾讯CDC
K
Kaspersky official blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Project Zero
Project Zero
F
Fortinet All Blogs
AWS News Blog
AWS News Blog
The Cloudflare Blog
C
CERT Recently Published Vulnerability Notes
I
InfoQ
Spread Privacy
Spread Privacy
T
Tenable Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
AI Agent 出问题时,不要只看最终回答:一次请求级调试的思路
侯垒 · 2026-06-22 · via DEV Community

这两年 AI 编程工具变化很快。

从 Copilot 式的代码补全,到 Claude Code、Codex、OpenCode、Cursor、Cline 这类 Agent 工具,AI 已经不只是“帮你补几行代码”了。它可以读项目、改文件、跑命令、调用工具、分析报错,然后继续下一轮。

但 Agent 越像一个“会干活的人”,问题也越明显:

它出错的时候,你很难判断它到底错在哪里。

很多人遇到 AI 编程工具跑偏时,第一反应是改 prompt。

比如:

你要更认真一点。
不要偷懒。
必须先读代码再修改。
不要乱改无关文件。

这些提示有时有用,但更多时候只是把问题往后推。

因为真正的问题可能根本不在你写给它的那句话里,而在它实际发给模型的完整请求里。

为什么只看最终回答不够?

一个 AI Agent 完成任务,通常不是一次请求结束。

它的流程更像这样:

用户提出任务
-> 模型收到 system prompt、上下文、工具列表
-> 模型决定调用某个工具
-> 本地执行工具,比如读文件、搜索代码、运行命令
-> 工具结果再进入下一轮模型请求
-> 模型继续判断
-> 直到给出最终回答或修改代码

我们在终端或编辑器里看到的,通常只是这个过程的一小部分。

最终回答只能告诉你“它做了什么”,但很难告诉你“它为什么这么做”。

比如下面这些问题,只看最终回答基本看不出来:

  • 它到底有没有看到关键文件?
  • 它看到的是完整上下文,还是被截断后的上下文?
  • system prompt 里有没有某条规则影响了它?
  • 工具 schema 是不是描述不清,导致它选错工具?
  • tool result 有没有进入下一轮?
  • 哪一轮开始 token 暴涨?
  • provider 返回的 400 到底是模型问题,还是请求格式问题?
  • 它是不是每一轮都重复带入了大量无效上下文?

这些都属于“请求级问题”。

AI Agent 的常见问题,其实可以分层看

我现在更倾向于把 AI Agent 的问题分成几层,而不是笼统地说“模型不行”。

第一层:模型没有看到该看的东西

这是最常见的问题之一。

你以为 Agent 已经读了某个文件,但实际发给模型的请求里没有那段内容。

或者它读过,但后续请求里没有保留。

于是模型只能基于不完整上下文做判断,结果自然不稳定。

这种时候继续强调“请仔细阅读代码”没有太大意义。你真正要确认的是:

关键上下文有没有进入下一轮请求?

第二层:工具列表或工具描述有问题

Agent 能调用工具,并不代表它会正确调用工具。

模型看到的是一组工具 schema,包括工具名、描述、参数结构。

如果工具描述太模糊,模型可能不知道什么时候该用。

如果参数结构太复杂,模型可能生成错误参数。

如果工具太多,模型也可能选错。

这类问题在 MCP、插件、子任务工具越来越多以后会更明显。

第三层:tool call 和 tool result 没有正确闭环

一个正常的工具调用闭环应该是:

assistant: tool_use
tool: tool_result
assistant: 根据 tool_result 继续

如果中间某一环丢了,Agent 就会开始出现奇怪行为。

比如:

  • 工具没有真正执行,但模型以为执行了
  • tool result 太长,进入下一轮后污染上下文
  • 多个 tool_result 顺序错了
  • tool_use 和 tool_result 没有正确配对
  • provider 要求的消息格式和客户端保存的格式不一致

这类 bug 最难靠肉眼看最终回答判断。

你必须看到真实的请求和响应。

第四层:token 和成本问题

很多人用 AI 编程工具时,一开始只关心效果。

但一旦使用频率上来,token 成本就会变成实际问题。

尤其是 Agent 场景里,token 并不只花在最终回答上。

大量 token 可能花在:

  • system prompt
  • 工具 schema
  • 历史消息
  • 文件内容
  • 搜索结果
  • 命令输出
  • tool result
  • 子 Agent 的上下文

有时一次任务很贵,不是因为最终回答很长,而是因为某一轮请求带了巨大的上下文。

所以看 session 总成本还不够,最好能看到每一次请求的 token 使用。

请求级调试应该看什么?

如果一个 AI Agent 跑偏,我通常会先看这几类信息。

1. system prompt

system prompt 是 Agent 行为的底层约束。

很多看似“模型自己决定”的行为,其实是 system prompt 影响的。

比如:

  • 它为什么总是先写计划?
  • 它为什么不直接执行命令?
  • 它为什么遇到某类文件不修改?
  • 它为什么要反复确认?

这些可能都和 system prompt 有关。

2. messages

messages 决定模型当前能看到什么。

重点不是“这个 Agent 曾经读过什么”,而是:

当前这一轮请求里到底包含了什么?

这两个问题不一样。

Agent 本地读过一个文件,不代表后续每一轮模型请求都带着这个文件。

3. tools schema

如果 Agent 没有调用你期望的工具,先不要急着骂模型。

可以先看:

  • 工具是否真的出现在 tools 里?
  • 工具描述是否清楚?
  • 参数 schema 是否合理?
  • 是否有多个类似工具让模型混淆?
  • provider 是否支持这种 schema?

4. tool call / tool result

这一层最适合排查 Agent 为什么做错。

你可以看:

  • 模型选择了哪个工具
  • 传入了什么参数
  • 工具返回了什么
  • 返回结果有没有进入下一轮
  • 下一轮模型有没有基于这个结果继续

如果这里断了,后面再怎么调 prompt 都不稳定。

5. token / cache / cost / latency

这些指标能帮你判断问题发生在哪一轮。

比如:

  • 哪一轮 input token 突然升高?
  • 哪一轮 tool result 特别大?
  • cache 有没有命中?
  • 哪个模型最贵?
  • 哪个请求延迟最高?
  • 失败请求有没有 usage 信息?

这比只看“今天花了多少钱”有用得多。

ccglass 解决的就是这一层

最近我发现了一个开源工具:ccglass。

项目地址:

https://github.com/jianshuo/ccglass

它不是另一个 AI 编程助手,而是一个本地观测工具。

简单说,它会在本地启动一个代理和 Dashboard,让你看到 Claude Code、Codex、OpenCode、CodeBuddy、Qoder 等工具实际发给模型的请求。

你可以看到:

  • system prompt
  • messages
  • tools schema
  • tool calls
  • tool results
  • request / response body
  • token / cache / cost
  • latency
  • turn-to-turn diff

也就是说,它关心的不是“再帮你写一段代码”,而是“让你看清 Agent 到底是怎么工作的”。

一个典型使用场景

假设你让 Claude Code 修一个 bug。

它最后确实改了代码,但你觉得改得很奇怪。

这时你可以不急着重跑,而是打开 ccglass 看这几件事:

  1. 第一轮请求里有没有带上相关文件?
  2. system prompt 有没有影响它的修改策略?
  3. 它调用了哪些工具?
  4. 每个工具返回了什么?
  5. 哪个 tool result 进入了下一轮?
  6. 修改代码前,它是否真的看到了测试失败信息?
  7. token 是不是在某一轮突然暴涨?

如果你能回答这些问题,调试 Agent 就从“感觉不对”变成了“证据链不对”。

它和普通抓包工具有什么区别?

很多人会问:这是不是 Charles、mitmproxy、Proxyman 也能做?

某些情况下可以,但 AI 编程 Agent 有几个麻烦点:

  • 不一定遵守系统代理
  • 有些客户端走自定义 base URL
  • 有些使用 OpenAI / Anthropic 兼容接口
  • 有些 provider 的 streaming 格式不一样
  • 有些工具调用信息需要按 Agent 语义展示,而不是只看 HTTP 包

ccglass 的目标不是替代通用抓包工具,而是专门围绕 AI Agent 请求结构做展示。

它会把 system、messages、tools、tool call、tool result、usage、cost、diff 这些东西整理成适合调试 Agent 的视图。

为什么我觉得这件事会越来越重要?

AI 编程工具的发展方向很明显:Agent 会越来越自动。

它们会读更多文件,调用更多工具,跑更长任务,甚至调度子 Agent。

这当然会提高效率。

但如果没有观测能力,复杂度也会一起上升。

以后我们遇到的问题可能不再是:

这段代码补全得对不对?

而是:

为什么这个 Agent 在第 7 轮调用了这个工具?
为什么它把这个 tool result 带到了后面 12 轮?
为什么这次任务花了 30 万 token?
为什么同样的任务换 provider 就失败?
为什么上下文压缩后它丢了关键约束?

这些问题都需要请求级视角。

结语

我觉得 AI 编程工具接下来的一个重要方向,不只是让 Agent 更强,而是让 Agent 更可观察。

只看最终回答,就像只看程序崩溃后的最后一行日志。

真正要调试复杂系统,还是要看输入、输出、中间状态和调用链。

AI Agent 也是一样。

如果你也在用 Claude Code、Codex、OpenCode、CodeBuddy、Qoder 这类工具,可以试试 ccglass:

https://github.com/jianshuo/ccglass

它解决的不是“让 AI 更聪明”,而是让使用者更清楚地知道:

AI 到底看到了什么,又基于什么做出了下一步。