惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
The Exploit Database - CXSecurity.com
S
Secure Thoughts
A
Arctic Wolf
V
Vulnerabilities – Threatpost
S
Schneier on Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threat Research - Cisco Blogs
AWS News Blog
AWS News Blog
NISL@THU
NISL@THU
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Palo Alto Networks Blog
L
Lohrmann on Cybersecurity
Schneier on Security
Schneier on Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Scott Helme
Scott Helme
L
LINUX DO - 最新话题
L
LangChain Blog
量子位
T
Threatpost
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
腾讯CDC
W
WeLiveSecurity
Last Week in AI
Last Week in AI
美团技术团队
The GitHub Blog
The GitHub Blog
The Last Watchdog
The Last Watchdog
C
CERT Recently Published Vulnerability Notes
月光博客
月光博客
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
L
LINUX DO - 热门话题
Microsoft Security Blog
Microsoft Security Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
T
Troy Hunt's Blog
Webroot Blog
Webroot Blog
云风的 BLOG
云风的 BLOG
博客园 - 叶小钗
V2EX - 技术
V2EX - 技术
雷峰网
雷峰网
Security Latest
Security Latest
小众软件
小众软件
J
Java Code Geeks
博客园 - Franky
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Spread Privacy
Spread Privacy
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Recent Commits to openclaw:main
Recent Commits to openclaw:main
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
The Blog of Author Tim Ferriss
M
MIT News - Artificial intelligence

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Why AI Agents Cost More Than LLMs (And How to Stop Bleeding Tokens)
Abhishek · 2026-05-11 · via DEV Community

AI Agents vs LLM pricing

I was building a small bookmark app last weekend. You send it a URL, Gemini
summarizes and tags the page, the result goes into Postgres. A few hundred lines
of TypeScript.

The first version cost almost nothing. One LLM call per URL, that's it. Then I
added "tools" so the model could fetch pages, look up similar bookmarks, or
check things against Google Search.

My token bill quadrupled.

That's where most people building agents land for the first time. Going from a
plain chat call to an agent loop is way more expensive than docs make it sound,
and the reason isn't obvious until you watch the round trips happen one by one.
Let's do that.

What a plain LLM call costs

Here's the simplest LLM call in TypeScript with @google/genai:

import { GoogleGenAI } from '@google/genai';

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY! });

const res = await ai.models.generateContent({
  model: 'gemini-2.5-flash',
  contents: 'Summarize this article: ...',
});

console.log(res.text);

Enter fullscreen mode Exit fullscreen mode

One request out, one response back. You pay for two things:

  • Input tokens for your prompt
  • Output tokens for the model's reply

That's it. Two numbers on your bill. If your prompt is 500 tokens and the answer
is 200, you pay for 700 tokens. Done.

Now add a single tool

Tools are how the model talks to the outside world. Calling an API, querying a
database, fetching a URL, anything. You describe each tool with a small JSON
schema, and the model can ask to "call" one mid-conversation. You actually run
the function, send the result back, and the model writes its final answer using
that result.

The basic version:

import { GoogleGenAI, Type } from '@google/genai';

const tools = [{
  functionDeclarations: [{
    name: 'getWeather',
    description: 'Get the weather of any city',
    parameters: {
      type: Type.OBJECT,
      properties: {
        location: { type: Type.STRING },
      },
      required: ['location'],
    },
  }],
}];

const first = await ai.models.generateContent({
  model: 'gemini-2.5-flash',
  contents: 'What is the weather in Tokyo?',
  config: { tools },
});

console.log(first.text);
// → undefined

Enter fullscreen mode Exit fullscreen mode

undefined?

The model didn't answer. It returned a structured request:

first.functionCalls
// → [{ name: 'getWeather', args: { location: 'Tokyo' } }]

Enter fullscreen mode Exit fullscreen mode

This is the part that surprises people. The model got asked a question, and
instead of answering, it asked you to run a function. So you do that and
ship the result back:

const result = getWeather('Tokyo');  // { temperature: 23, condition: 'sunny' }

const second = await ai.models.generateContent({
  model: 'gemini-2.5-flash',
  contents: [
    { role: 'user',  parts: [{ text: 'What is the weather in Tokyo?' }] },
    { role: 'model', parts: [{ functionCall: { name: 'getWeather', args: { location: 'Tokyo' } } }] },
    { role: 'user',  parts: [{ functionResponse: { name: 'getWeather', response: result } }] },
  ],
  config: { tools },
});

console.log(second.text);
// → "It's 23°C and sunny in Tokyo."

Enter fullscreen mode Exit fullscreen mode

Two LLM calls. One question. That's the agent tax.

Why we can't just do it in one call

The first reaction (mine too): why can't the model just answer in one shot?

The reason is simple. The model can't predict what the tool will return. The
temperature in Tokyo isn't in its training data, the API hasn't been hit yet,
the result doesn't exist. You can't write "It's 23°C in Tokyo" before you know
it's 23°C.

So turn 1 is "decide what to do." Turn 2 is "use what you learned." They can't
be merged. The model has no memory between calls.

One exception is worth knowing about: server-side tools. Things like
googleSearch or urlContext in Gemini run inside Google's own servers, and
the API returns one merged response. From your side it looks like a single call.
You lose some control (you can't see exactly what got searched), but you save
a round trip.

Counting the actual tokens

Here's where the cost lives. Look at what turn 2 has to send compared to turn 1:

Turn 1 in Turn 1 out Turn 2 in Turn 2 out
System prompt yes yes, billed again
Tool schemas yes yes, billed again
User question yes yes, billed again
Model's tool call yes yes, as input
Your tool result yes
Final answer yes

Your system prompt and tool definitions get sent to the API twice. Turn 1
doesn't free you from re-sending everything in turn 2, because the model is
stateless. It forgets the whole conversation between calls.

Real numbers from my bookmark agent:

  • System prompt: ~200 tokens
  • 4 tool declarations: ~400 tokens
  • User question: ~50 tokens
  • Tool result (a few rows from Postgres): ~300 tokens
Plain LLM call:    ~650 in  +  ~200 out  =  ~850 tokens
One-tool agent:   ~1300 in  +  ~230 out  = ~1530 tokens (about 1.8x)

Enter fullscreen mode Exit fullscreen mode

And that's the best case. Exactly one tool call, no follow-ups. Real agents are
worse. A lot worse.

Real agents grow quadratically

The bookmark agent does three things on a new URL:

  1. Fetch the page (fetchUrl tool)
  2. Look for similar existing bookmarks in the DB (searchSimilar tool)
  3. Pick a category from the user's existing taxonomy (getTaxonomy tool)

That's 4 LLM turns total. Ask, get tool calls, send back results, ask again,
get more calls, send results, finally write the summary.

What the cumulative input size looks like each turn:

Turn What gets sent Input tokens
1 system + schemas + URL 700
2 + previous calls + fetchUrl result (~1500 of page) 2200
3 + searchSimilar result 2400
4 + getTaxonomy result 2600

Total input across all turns: about 7900 tokens to summarize one webpage.

For comparison, a plain generateContent({ contents: "summarize this:\n" + pageText })
costs ~1500 input + 200 output. About 1700 tokens.

Same task. Almost 5x the bill.

It gets worse. Cost grows quadratically with the number of turns, because
each turn replays everything that came before. A 10-turn agent isn't 10x the
cost. It's closer to 30x.

Three ways to stop the bleeding

You're not stuck. Here's what actually works.

1. Prompt caching

The biggest lever by far. Every major provider supports it now: OpenAI,
Anthropic, Google. The system prompt and tool schemas don't change between
turns, so cache them once and pay about 25% of the input cost on every reuse.

With @google/genai:

const cache = await ai.caches.create({
  model: 'gemini-2.5-flash',
  config: {
    systemInstruction: 'You are a bookmark organizer...',
    tools,  // these never change across turns
  },
});

const res = await ai.models.generateContent({
  model: 'gemini-2.5-flash',
  contents: history,
  config: { cachedContent: cache.name },
});

Enter fullscreen mode Exit fullscreen mode

For my 4-turn flow this cuts input costs by roughly half. Anthropic and OpenAI
do the same thing with different syntax.

Gemini also has implicit caching. It auto-caches recent prefixes for you with
zero code changes. You just see cheaper retries. Check if your provider has it
on before reinventing the wheel.

2. Different model per turn

The "decide which tool to call" turn is dumb work. It barely needs reasoning.
It's pattern matching on a question. The final synthesis turn is where you
actually want a smart model.

// Cheap, fast: decides what to do
const decision = await ai.models.generateContent({
  model: 'gemini-2.5-flash-lite',
  // ...
});

// Smarter: writes the actual answer
const finalAnswer = await ai.models.generateContent({
  model: 'gemini-2.5-pro',
  // ...
});

Enter fullscreen mode Exit fullscreen mode

In a 4-turn flow, three of the turns can run on the cheap model. Only the last
one, the user-facing answer, needs the expensive one. For high-volume agents
this saves more than caching does.

3. Parallel tool calls

The model can ask for multiple tools in a single response. Code I see in
tutorials usually does functionCalls[0] and silently drops the rest, turning
what could be one round trip into many.

The fix is one line of Promise.all:

const results = await Promise.all(
  resp.functionCalls.map(async (c) => ({
    name: c.name!,
    response: await dispatchers[c.name!](c.args),
  }))
);

Enter fullscreen mode Exit fullscreen mode

For "summarize all my React bookmarks from last month," the model might call
searchBookmarks and getDateRange in parallel. Handle both, and you save a
whole round trip.

What you can't optimize away

Tools have a real cost, and they buy you real value. The reason you reach for
them is the same reason they're expensive. You're forcing the model to use
facts that exist outside its head instead of making them up.

A plain LLM call will happily tell you the weather in Tokyo. It'll just be
wrong.

Quick way to think about it when picking an architecture:

  • Plain LLM is a guess from training data. Cheap, fast, hallucinates.
  • Tools / agent is real data. Expensive, slower, honest.

Most apps shouldn't be agents. If your task is "summarize this text I'm pasting
in" or "rewrite this email," you don't need tools. You need one call. A lot of
agent frameworks make it really easy to add tools by default, which makes it
really easy to spend 5x what you should.

Tools earn their cost when you have side effects (writing to a DB, sending a
message), grounded data (today's weather, this user's bookmarks, current docs),
or chained reasoning where intermediate steps actually need verification.

They don't earn it on anything you could solve with one good prompt.

The receipt

Last week I added one tool to a Gemini call and watched the cost go from 850
tokens to 1530 for the same question. Once I started parallelizing calls and
caching the system prompt, I got the bookmark agent down to about 4500 tokens
across all four turns. Still 2.5x a plain call, but way better than the 7900
the naive version was burning.

Your agent isn't a smarter LLM. It's the same LLM with a longer receipt. Once
you can read the receipt, every optimization becomes obvious.

If you like my content support by like and share 💟 also dont forget to follow me on Twitter/X and LinkedIn. If you want me to connect, checkout my site. See you in next one.