惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Google DeepMind News
Google DeepMind News
S
Secure Thoughts
Application and Cybersecurity Blog
Application and Cybersecurity Blog
AWS News Blog
AWS News Blog
GbyAI
GbyAI
腾讯CDC
WordPress大学
WordPress大学
V
V2EX
小众软件
小众软件
C
CXSECURITY Database RSS Feed - CXSecurity.com
M
MIT News - Artificial intelligence
T
Troy Hunt's Blog
H
Hacker News: Front Page
Scott Helme
Scott Helme
The Hacker News
The Hacker News
Schneier on Security
Schneier on Security
K
Kaspersky official blog
S
Security @ Cisco Blogs
N
News | PayPal Newsroom
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
T
Threatpost
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Martin Fowler
Martin Fowler
P
Privacy International News Feed
Blog — PlanetScale
Blog — PlanetScale
D
DataBreaches.Net
O
OpenAI News
P
Proofpoint News Feed
J
Java Code Geeks
B
Blog RSS Feed
Attack and Defense Labs
Attack and Defense Labs
L
Lohrmann on Cybersecurity
T
Threat Research - Cisco Blogs
C
Cybersecurity and Infrastructure Security Agency CISA
阮一峰的网络日志
阮一峰的网络日志
P
Palo Alto Networks Blog
Engineering at Meta
Engineering at Meta
AI
AI
Google DeepMind News
Google DeepMind News
雷峰网
雷峰网
Microsoft Azure Blog
Microsoft Azure Blog
Simon Willison's Weblog
Simon Willison's Weblog
人人都是产品经理
人人都是产品经理
T
Tailwind CSS Blog
G
GRAHAM CLULEY
A
Arctic Wolf
N
Netflix TechBlog - Medium
L
LINUX DO - 最新话题
爱范儿
爱范儿
博客园 - 聂微东

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I Accidentally Spent $400 on GPT-4o in One Month. Here's How to Never Do That.
Muhammad Awais · 2026-06-04 · via DEV Community

Last year I deployed a GPT-4o powered support chatbot for a small SaaS. Traffic was modest — maybe 500 active users. I checked the OpenAI bill at the end of the month.

$400.

The app was supposed to cost $40/month. I'd done the mental math: "1,000 API calls a day, maybe 500 tokens each, GPT-4o is $10/M output... that's like $50/month, fine."

What I forgot: my system prompt was 600 tokens. It was sent on every single call. At 8,000 daily calls (users were chatty), that's 4.8M extra input tokens per day just from the system prompt. At $2.50/M, that's $12/day — $360/month — from a prompt I copy-pasted and never measured.

That's when I built the LLM API Cost Calculator — a free tool that covers 18 models, 7 currencies, a live token counter, and a Compare tab that ranks every model by cost for your exact workload. Let me show you how to use it properly, and walk through the math so you never have a surprise bill month.


The Mental Model Most Developers Get Wrong

Before touching any calculator, you need to understand one thing: output tokens cost 3–5× more than input tokens, and your architecture determines which one dominates your bill.

Take Claude Sonnet 4:

  • Input: $3.00 per million tokens
  • Output: $15.00 per million tokens

That's a 5× ratio. Now look at two very different workloads:

Sentiment classification:
- Input: 300 tokens (review + system prompt)
- Output: 10 tokens ("positive" / "negative")
- 5,000 calls/day

Input cost:  300 × 5,000 / 1,000,000 × $3.00  = $4.50/day
Output cost: 10  × 5,000 / 1,000,000 × $15.00 = $0.75/day
→ Input dominates. Optimize your system prompt.

Enter fullscreen mode Exit fullscreen mode

Support chatbot:
- Input: 800 tokens (history + system prompt + user message)
- Output: 400 tokens (detailed response)
- 1,000 calls/day

Input cost:  800 × 1,000 / 1,000,000 × $3.00  = $2.40/day
Output cost: 400 × 1,000 / 1,000,000 × $15.00 = $6.00/day
→ Output dominates. Cap response length with max_tokens.

Enter fullscreen mode Exit fullscreen mode

Same model, radically different cost structure. The cost calculator's breakdown bar shows you exactly this split so you know which side to optimize.


LLM Pricing Snapshot — June 2026

Here's where the major models sit right now. All prices are per million tokens (input / output):

Model Input Output Best For
Mistral Small 3 $0.10 $0.30 Cheapest input overall
Llama 4 Scout $0.11 $0.34 High-volume classification
GPT-4o mini $0.15 $0.60 Budget general-purpose
Gemini 2.5 Flash $0.15 $0.60 Cheap + thinking mode
DeepSeek V3 $0.27 $1.10 Coding, open-source value
DeepSeek R1 $0.55 $2.19 Reasoning tasks
Llama 3.3 70B $0.59 $0.79 Balanced open-source
o4-mini $1.10 $4.40 OpenAI reasoning, budget
Gemini 2.5 Pro $1.25 $10.00 1M context window
o3 $2.00 $8.00 Advanced reasoning
Mistral Large 2 $2.00 $6.00 EU data residency
GPT-4o $2.50 $10.00 Flagship OpenAI
Claude Sonnet 4 $3.00 $15.00 Coding, instruction-following
Grok 3 Mini $0.30 $0.50 Small gap in/out pricing
Grok 3 $3.00 $15.00 xAI flagship
Claude Haiku 3.5 $0.80 $4.00 Budget Anthropic
Claude Opus 4 $15.00 $75.00 Most expensive overall

The price gap is brutal: Claude Opus 4 output costs 250× more than Llama 4 Scout output. For most real workloads, that premium is unjustifiable.


4 Real Scenarios With Real Numbers

Rather than abstract pricing, let me run 4 actual workloads through the calculator. You can replicate all of these yourself — just open the Compare Models tab and enter these numbers.

Scenario 1 — Startup Chatbot (500 users/day)

800 input / 400 output / 1,000 calls/day

Model Monthly Cost
Llama 4 Scout ~$7
GPT-4o mini ~$7
Claude Haiku 3.5 ~$40
GPT-4o ~$120
Claude Sonnet 4 ~$144
Claude Opus 4 ~$720

If your chatbot handles general questions, GPT-4o mini vs GPT-4o is literally $113/month saved — $1,356/year — for quality most users won't notice.

Scenario 2 — RAG Document Search (Enterprise)

3,000 input (chunked docs) / 500 output / 500 calls/day

Model Monthly Cost
DeepSeek V3 ~$25
Gemini 2.5 Flash ~$36
Gemini 2.5 Pro ~$112
Claude Sonnet 4 ~$189

DeepSeek V3 at $25/month vs Gemini 2.5 Pro at $112/month for RAG. Unless you need Gemini's 1M context window for very large documents, that's a 78% cost reduction for identical architecture.

Scenario 3 — Code Review in CI/CD

2,000 input (diff + context) / 800 output / 200 calls/day

Model Monthly Cost
DeepSeek V3 ~$14
GPT-4o ~$54
Claude Sonnet 4 ~$72

This is a case where quality might justify cost. Claude Sonnet 4 genuinely outperforms on nuanced code review. But $72 vs $14/month is a real conversation — benchmark 100 real diffs before committing.

Scenario 4 — High-Volume Classification (5,000 calls/day)

300 input / 50 output / 5,000 calls/day

Model Monthly Cost
Mistral Small 3 ~$4.50
Llama 4 Scout ~$5.40
GPT-4o mini ~$6.75
Claude Haiku 3.5 ~$12

For simple classification, you're choosing between $4.50 and $12/month. Mistral Small 3 wins unless you specifically need Anthropic's API capabilities.


The System Prompt Tax (What Killed My Budget)

Here's the thing I got wrong — and it's the most common mistake I see in production AI apps.

A 400-token system prompt at 10,000 calls/day:

400 tokens × 10,000 calls = 4,000,000 input tokens/day
At GPT-4o ($2.50/M):        $10/day = $300/month

Enter fullscreen mode Exit fullscreen mode

$300/month just from your system prompt. Before a single user message.

What to do about it:

  1. Measure first. Paste your system prompt into the Token Counter tab and see the exact count.
  2. Compress it. You can often cut a 400-token system prompt to 200 tokens without losing behavior. Use the AI Prompt Optimizer to reduce token footprint without losing quality — that alone saves $150/month in this example.
  3. Cache it. OpenAI's Prompt Caching gives you 50% off cached input tokens{:target="_blank"}{:rel="noopener"} for prompts over 1,024 tokens. Anthropic has similar caching. If your system prompt is fixed across calls — and it usually is — you could cut input costs by 40–50% overnight.

Conversation History: The Silent Cost Multiplier

Here's the second thing developers consistently underestimate:

Turn 1:  800 input tokens sent
Turn 2:  800 + 300 (turn 1 response) = 1,100 tokens sent
Turn 3:  1,100 + 300 = 1,400 tokens sent
Turn 4:  1,400 + 300 = 1,700 tokens sent
Turn 5:  1,700 + 300 = 2,000 tokens sent

Enter fullscreen mode Exit fullscreen mode

A 5-turn conversation that starts at 800 tokens averages 1,400 tokens per call — not 800. Your cost estimate needs to reflect the average turn depth, not just the first message.

Strategies:

  • Sliding window: Only keep the last N turns in context
  • Summarization: After turn 5, summarize history into 200 tokens and continue
  • Topic detection: Reset context when the topic changes

The agentic workflows guide covers context management for multi-step agents in detail — the same principles apply to chatbot history.


Agentic Loops: Where Estimates Fall Apart Completely

If you're building AI agents — tools that make multiple API calls in a loop — standard cost estimation breaks down fast.

An agent that makes 5 tool calls per user request, each with a growing context:

Call 1: 2,000 tokens in, 500 out (tool invocation)
Call 2: 2,500 tokens in, 500 out (tool result added)
Call 3: 3,000 tokens in, 500 out
Call 4: 3,500 tokens in, 500 out
Call 5: 4,000 tokens in, 800 out (final response)

Total: 15,000 input + 2,800 output per "1 user request"

Enter fullscreen mode Exit fullscreen mode

At Claude Sonnet 4: $0.045 input + $0.042 output = $0.087 per user request.

Looks small. At 1,000 agent tasks/day: $87/day = $2,610/month.

Use the AI Agent preset in the calculator (8,000 in / 2,000 out / 100 calls) as a starting point, but measure your actual loop depth before estimating production costs.


Sharing Estimates With Your Team

One feature I find genuinely useful: the Share Estimate button encodes everything into the URL:

?m=claude-sonnet-4&in=800&out=400&d=1000&c=PKR

Enter fullscreen mode Exit fullscreen mode

Open that URL and all settings restore automatically. Useful for:

  • Sending a cost comparison to a teammate: "Here's Gemini vs DeepSeek for our RAG workload — click the link and switch models"
  • Client proposals: pre-fill their expected volume so they can explore numbers themselves
  • Budget reviews: bookmark the URL with your current production numbers

Nothing sensitive in the URL — just model name, token counts, call volume, and currency.


Quick Checklist Before You Pick a Model

Before committing to any LLM in production:

  • [ ] Measure your actual system prompt token count (Token Counter tab)
  • [ ] Estimate real output length from 5+ sample responses — don't guess
  • [ ] Account for conversation history growth (average turn depth, not turn 1)
  • [ ] Run the Compare Models tab with your actual numbers
  • [ ] Check if cheaper models pass a 100-sample quality benchmark for your task
  • [ ] Verify if your provider offers caching discounts for your prompt pattern
  • [ ] Model the worst-case volume (peak traffic, not average)
  • [ ] Switch to the cheapest model during development

Try It

The LLM API Cost Calculator is free, no signup, runs entirely in your browser. 18 models, 7 currencies, token counter, shareable URLs, downloadable CSV reports.

If you're building something where costs matter — and they always do eventually — spending 5 minutes with the Compare tab before picking a model is the highest ROI activity in your planning process.

What's your current monthly API bill? And which model surprised you most with its cost in production? Drop it in the comments — I'm genuinely curious what architectures people are running.