惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cyberwarzone
Cyberwarzone
Vercel News
Vercel News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
aimingoo的专栏
aimingoo的专栏
B
Blog RSS Feed
A
About on SuperTechFans
T
The Blog of Author Tim Ferriss
爱范儿
爱范儿
腾讯CDC
S
SegmentFault 最新的问题
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
The Hacker News
The Hacker News
J
Java Code Geeks
大猫的无限游戏
大猫的无限游戏
B
Blog
IT之家
IT之家
Spread Privacy
Spread Privacy
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
C
Cisco Blogs
Recent Announcements
Recent Announcements
H
Hacker News: Front Page
AI
AI
I
InfoQ
H
Heimdal Security Blog
T
Threatpost
Cisco Talos Blog
Cisco Talos Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
I
Intezer
W
WeLiveSecurity
SecWiki News
SecWiki News
MongoDB | Blog
MongoDB | Blog
宝玉的分享
宝玉的分享
博客园 - 【当耐特】
云风的 BLOG
云风的 BLOG
T
Threat Research - Cisco Blogs
V2EX - 技术
V2EX - 技术
N
News and Events Feed by Topic
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
O
OpenAI News
阮一峰的网络日志
阮一峰的网络日志
T
Troy Hunt's Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
博客园 - 司徒正美
Apple Machine Learning Research
Apple Machine Learning Research
雷峰网
雷峰网
T
Tor Project blog
有赞技术团队
有赞技术团队
Schneier on Security
Schneier on Security
Last Week in AI
Last Week in AI

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
A voice agent is not a chatbot with a phone number
Arthur · 2026-06-19 · via DEV Community

The cleanest illustration of why this matters comes from a small, ordinary failure on a small, ordinary outbound campaign that I've been reading about: roughly one day, a few hundred cold-call attempts, and about $100 of telephony plus STT plus TTS plus model spend, evaporated by a voice agent that occasionally found itself dialing into someone else's voicemail or IVR or, the most expensive case, another voice agent. The exchange the operator screenshotted is the kind of thing that reads as a comedy bit until you remember it's billing the whole time:

— Hello.
— Hello, how can I help you?
— I'm calling because…
— Hello, how can I help you?
— Sure, could you tell me…

In a chat window this would be a funny screenshot. On a phone it's billing the whole time — telephony plus STT plus TTS plus model tokens, on two endpoints, both confidently polite, neither programmed to recognise the other side as a peer and hang up. The lesson the operator pulled from this, and the one I want to walk through, is the larger one: a voice agent is not a chatbot with a phone number. It's a realtime system, and almost every "voice agent failure in production" I've now read about reduces to chat-architecture assumptions being applied to a medium that doesn't tolerate them.

Let me unpack what specifically doesn't translate.

Latency in chat and latency on a call are different objects

In a chat the unit of cost is "time until the model starts streaming a reply." A two-second pause is fine. The user is reading the previous turn or sipping coffee or alt-tabbed away. In a phone call the unit of cost is time of silence on an open audio channel, and that has a perceptual budget set by human conversational physiology, not by your latency dashboards.

The specific budget is well-studied. Levinson and Torreira's 2015 paper Timing in turn-taking and its implications for processing models of language, drawing on a corpus across ten languages, reports that the typical gap between turns in natural conversation is around 200 milliseconds, with modal values clustering in the 100–300ms range — and overlap is more common than long pauses. The authors note the cognitive trick that makes this possible: speakers begin planning their response before the previous turn ends. Two hundred milliseconds is an interaction signature, not a latency target you choose.

Once you exceed that, perceptual breakdowns happen on a sliding scale. The voice-AI industry — see, e.g., AssemblyAI's "300ms rule" writeup — converges on a perceptual gradient: by 300–400ms the listener is starting to notice the silence; by 500ms they're starting to assume something is wrong with their own line; sub-500ms is the working threshold below which an agent feels live. Retell AI, one of the larger commercial platforms, claims about 600ms end-to-end and frames that as competitive. It is competitive, and that also tells you the ceiling: even the leading systems are sitting just above the perceptual breakdown line, not below it.

Now look at what the chat-style architecture has to fit inside that budget on every turn:

  • streaming STT to recognise the user's speech;
  • LLM call (with potentially several tool calls — CRM lookup, calendar check, database query);
  • response generation;
  • streaming TTS, with first-byte audio out the door before the rest is ready.

In a chat you can spend several seconds on this and the user waits. On a call you cannot, because the other party is not waiting; they are filling the silence with "hello?" and starting to repeat themselves and asking if you're still there. The streaming transcript captures all of it. The model now has to respond to a turn that is partly the original question and partly the interruption-and-repeat, and the conversation begins to liquefy.

The big-prompt problem doesn't translate

The chat reflex when an agent isn't reasoning well is to make the prompt longer. Add more rules. Add more examples. Add more tools. A long context-rich system prompt is the standard chat-deployment pattern.

In voice, this fails for a separate reason, distinct from the context-rot problem Anthropic and the Lost-in-the-Middle line of research have written about elsewhere (although that problem is also present). The voice-specific failure is goal drift mid-call. The operator I'm retelling here used a low-latency Gemini Flash–class model for one project (Google documents a separate Live Preview line for native realtime audio; it's not clear from the source whether the operator was on Live Preview or on the standard Flash variant adapted to a voice pipeline). What the operator observed was that the model could keep up with the latency budget but, given a long playbook stuffed into one prompt, would lose track within a few turns of which stage of the call it was in: had it asked about budget yet, was it still confirming identity, was it allowed to close. The model wasn't slow; it was disoriented. A fast model with a long prompt is not the same as a fast, focused model.

The substitution that works isn't a smarter model. It's an explicit graph.

Calls are graphs, not soup

A voice agent that holds up in production does not look like a single "be helpful and talk to the customer" prompt. It looks like a set of named stages with explicit transitions, each stage carrying a short instruction, restricted tool access, and explicit fallbacks. The platforms that ship voice agents (Retell's flow editor, ElevenLabs's Conversational AI workflow editor) make this graph structure visible, because that's what works:

[Greeting]
    │
    ▼
[Identity check] ── wrong person ──▶ [Apologise] ─▶ [End call]
    │
    ▼ identity confirmed
[Consent] ── not given ──▶ [Apologise] ─▶ [End call]
    │
    ▼ consent given
[Question 1] ─▶ [Question 2] ─▶ [Question 3]
    │
    ▼
[Closing]
    │
    ▼
[End call]

Fallbacks (any state):
  voicemail detected   ─▶ [Leave message] ─▶ [End call]
  human IVR             ─▶ [Press digit / wait for transfer]
  technical issue       ─▶ [Apologise + "we'll call back"] ─▶ [End call]
  another bot detected  ─▶ [End call]
  budget cap reached    ─▶ [End call]   (hard limit; not a prompt instruction)

This is dull engineering. It is also the engineering that turns "the agent sometimes gets confused" into "the agent's behaviour is auditable and the failure modes are named." Each stage has a budget — both in tokens and in real seconds — and each transition is explicit. No stage's instruction is "use your judgement"; if a stage needs judgement, that's a sign it should be split into two stages.

What works in voice (and what doesn't)

The same operator's piece is candid about which categories of voice deployment they made work and which they couldn't. The patterns are clean enough to tabulate; what's interesting is why the column splits look the way they do.

Category What changes vs. chat Result
Inbound lead qualification (small fixed questionnaire) Closed-world flow; user has consented to the call by submitting the form; small graph with a clear success criterion Worked. ~40 hours/week saved on a four-rep team.
Webinar attendance reminders (call N minutes before start) Single objective, single FAQ branch ("who are you / what's the webinar about"), short call Worked. Attendance lifted from ~10% to ~30%.
Cold outbound (open-world dial) Voicemail, IVR, gatekeepers, other bots, "send us an email instead," "I don't make those decisions," "who gave you my number" — each needs explicit behaviour Did not work. $100/day burning on indeterminate paths.

The pattern is structural, not coincidental. Inbound and reminders have a closed world: you control the flow because you also control the entry point. The user dialled or opted in; they're inside your graph from second one. Cold outbound has the opposite property: the world dials back, and the world contains things your graph does not. The right default for cold outbound is therefore not a smarter agent or a better prompt; it's a more aggressive exit policy — every recognisable open-world input maps to a transition that ends the call without burning cycles.

The hidden cost is that every one of those open-world inputs has to be recognised before you can transition on it. Recognising "this is voicemail and not a person" is itself a hard signal-processing task, and getting it wrong on either side is expensive: false positives end calls with real prospects, false negatives leave the agent monologuing to a beep for the maximum call duration the platform allows. (And if the platform has no maximum, which is somebody's first oversight, the bill is the limit.)

Why managed voice platforms are not just "Twilio with a wrapper"

You can build all of this directly on Twilio's media streams and your own STT/TTS/LLM pipes. The case for using a managed voice-agent platform (Retell, ElevenLabs, or one of several others that have appeared in the last 18 months) isn't that they're hard to imitate. It's that the things they ship under the hood are exactly the things that make the difference between a demo and a production deployment, and you only realise this after you've discovered them yourself:

  • Interruption handling. When the user talks over the agent, the TTS has to actually stop, the STT has to absorb the new turn, and the agent state has to update. "The TTS stops mid-syllable" is not a free behaviour; it's the result of a tightly integrated audio pipeline.
  • Streaming STT/TTS coordination with first-byte targets. Generating a full response and then sending it to TTS is fatal for latency. Streaming the text as it's generated, and beginning TTS on the first sentence, is fatal *un*tested. There is no architecture-on-paper that gets this right; it has to be tuned.
  • Regression tests for prompts and tool calls. When you change the wording in the consent stage, you want to know that the budget-question stage didn't silently start failing. The platforms ship saved-conversation regression tests precisely because hand-written voice tests are unreasonably hard to maintain.
  • Hard limits on call duration and spend. Not a prompt — a limit. If the agent enters an infinite politeness loop with another bot, the call has to end because the limit said so, not because the agent reasoned its way out.
  • Post-call extraction. A consistent set of fields pulled from the transcript at end-of-call, rather than asked of the model live.

What the platforms are actually selling is the boring stuff that turns out to be load-bearing. It is much cheaper to buy this than to discover what each piece is for and rebuild it badly.

The pre-launch checklist

If I were starting on a voice agent today, this is the order I'd want answered before I picked a model:

  1. What are the named stages of a typical successful call?
  2. What's the one objective of each stage?
  3. What inputs does each stage expect, and what data does it have access to?
  4. Which tools are valid in which stages? (Most stages should have zero.)
  5. What are the legal transitions? Which transitions are explicitly forbidden?
  6. What counts as success? What counts as a dead end?
  7. When is the agent required to end the call?
  8. How is voicemail recognised? IVR? Another bot?
  9. What's the latency budget for each stage, and how do we know we hit it?
  10. Which conversations do we save as regression tests?
  11. What's the per-call spend cap that auto-terminates? (This is not optional.)
  12. What's the per-day spend cap on the campaign?

The last two are not jokes. The classic voice-agent incident is the cost-disaster one — not because the agent did something dramatic, but because nobody set the limit.

What I'm taking from this

The framing I keep coming back to is that voice agents are not the natural successor to chatbots. They're a different class of system that happens to share an LLM. The chat lineage tells you to hand the model a long prompt, give it broad tool access, and let it figure out the conversation; the voice medium punishes every one of those choices. The systems that work in voice tend to be small, explicit graphs with named stages, narrow tool grants per stage, hard time-and-money limits, and an aggressive exit policy when the world doesn't behave like the graph expected.

The one-line summary the operator I read closes on is the one I'd keep: making the agent call is not the hard part; making it stop calling, in the right way, at the right time, when it's clearly off the rails, is the hard part. That's the engineering. Everything before it is plumbing.