惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Hugging Face - Blog
Hugging Face - Blog
B
Blog
博客园_首页
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
G
GRAHAM CLULEY
Microsoft Azure Blog
Microsoft Azure Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
WordPress大学
WordPress大学
The GitHub Blog
The GitHub Blog
Security Latest
Security Latest
F
Full Disclosure
云风的 BLOG
云风的 BLOG
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
C
Cyber Attacks, Cyber Crime and Cyber Security
L
LINUX DO - 热门话题
V
Visual Studio Blog
有赞技术团队
有赞技术团队
腾讯CDC
V
V2EX
Vercel News
Vercel News
C
Cisco Blogs
V2EX - 技术
V2EX - 技术
Scott Helme
Scott Helme
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
AWS News Blog
AWS News Blog
S
Schneier on Security
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
小众软件
小众软件
G
Google Developers Blog
C
Check Point Blog
C
CERT Recently Published Vulnerability Notes
博客园 - 叶小钗
S
SegmentFault 最新的问题
T
Tor Project blog
J
Java Code Geeks
L
Lohrmann on Cybersecurity
Application and Cybersecurity Blog
Application and Cybersecurity Blog
T
The Exploit Database - CXSecurity.com
Apple Machine Learning Research
Apple Machine Learning Research
T
Tailwind CSS Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园 - 司徒正美
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
Secure Thoughts
量子位
N
News and Events Feed by Topic
MyScale Blog
MyScale Blog
TaoSecurity Blog
TaoSecurity Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Engineering at Meta
Engineering at Meta

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
The AI Cost Paradox: 280x Cheaper, Bills Still Rising
Matthias | StudioMeyer · 2026-06-20 · via DEV Community

The cost of running a capable AI model fell by roughly 280 times in two years. Over the same stretch, the average company's AI bill went up, not down. Both numbers are real, both come from credible research, and the space between them is the single most useful thing an operator can understand about AI economics in 2026. It explains why "the models keep getting cheaper" and "our AI spend is out of control" are being said in the same meeting, by the same people, about the same systems.

I watch this play out in client projects every month. Someone reads that token prices collapsed, assumes their costs are about to fall off a cliff, and then opens an invoice that did the opposite. The confusion is not a billing error. It is a structural feature of how AI is now built, and once you see the mechanism you can plan around it instead of being surprised by it.

The Number That Should Have Lowered Your Bill

Start with the collapse, because it is genuinely staggering. Stanford's 2026 AI Index pegs the price of GPT-3.5-level performance at about 280 times cheaper between November 2022 and October 2024, falling from roughly 20 dollars per million tokens to about 7 cents. That is not a typo and it is not a one-off. Epoch AI measures a median decline near 50 times per year for equal capability. The venture firm a16z frames the same trend more conservatively at around 10 times per year, which they point out is still faster than compute fell in the PC era or bandwidth fell during the dotcom build-out.

The frontier did the same thing in public. When Anthropic shipped Claude Opus 4.5 in November 2025, it cut the flagship price from 15 and 75 dollars per million input and output tokens to 5 and 25, a 67 percent reduction in a single release. What happened next is the part people miss. Anthropic then held that 5-and-25 price across Opus 4.6, 4.7, and 4.8 while the model kept getting better. The per-token price stopped falling and capability kept climbing, which is its own kind of price cut.

The trigger for most of this was competition from below. DeepSeek R1 landed in January 2025 at 55 cents per million tokens while scoring around 95 percent of OpenAI's o1, and the major labs responded with emergency price moves. By mid-2026 the floor is remarkable. OpenAI's GPT-5.4-nano runs at 20 cents input and 1.25 dollars output per million. DeepSeek V4 Pro, an open-weights model you can host yourself, sits near 44 cents input. Google's Gemini 3.5 Flash beats the previous generation's Pro tier on agent benchmarks at 1.50 and 9 dollars. On paper, intelligence has never been this cheap to rent.

Why the Bill Went Up Instead

Here is the paradox stated plainly. Per-token prices fell by a factor of hundreds, and by one estimate the average enterprise AI bill still rose more than 300 percent over the same window. I treat the exact magnitude of that spend figure as indicative rather than gospel, because it comes from a secondary analysis, but the direction is confirmed everywhere and the reason is structural, not accidental.

Cheaper tokens get spent, not saved. The thing you are buying changed shape. In 2023 a typical interaction was one prompt and one answer, a few thousand tokens, one model call. In 2026 the same business outcome runs through an agent that fires somewhere between 10 and 20 model calls for a single user task. It plans, it calls a tool, it reads the result, it re-plans, it checks its own work, it writes a commit message. Retrieval-augmented generation inflates the context of each of those calls by stuffing in three to five times more reference text. And the agent does not go home at night. Monitoring agents and always-on assistants bill around the clock whether anyone is watching or not.

So the unit got 280 times cheaper and the number of units per job went up by more than that. This is the same pattern every efficiency gain in computing has followed. Cheaper storage did not shrink data centers, it gave us video everywhere. Cheaper bandwidth did not lower the average person's internet bill, it gave us streaming. Cheaper intelligence is not lowering AI spend, it is making agents economically possible, and agents are hungry. For anyone running a product on top of an API, that is the line that matters: a workload that cost a cent yesterday is a loop that costs fifteen cents today, and the loop is what makes the product good.

The Unlimited Era Just Ended

If you want a single event that marks the turn, it is GitHub Copilot. On the first of June 2026, GitHub moved every Copilot plan to usage-based billing. Premium request units were replaced by AI Credits priced at one cent each, metered against input, output, and cached tokens at each model's published rate. The cheaper fallback model that used to absorb overflow is gone. When your credits run out you either set a budget or you stop.

The reason GitHub gave is the clearest sentence anyone has written about this whole shift. With agents and subagents in the picture, the company said, "it is now common for a handful of requests to incur costs that exceed the plan price." Read that again with your own product in mind. A flat monthly subscription assumes a roughly predictable amount of work per user. Agentic software breaks that assumption, because one motivated user pointing an agent at a hard problem can burn a month of margin in an afternoon.

Everyone building on these APIs is now living in the world GitHub just formalized. Providers split pricing into short-context and long-context tiers. They charge per tool call for search and computer use. They sell priority lanes at 2.5 times the base rate and offer cached-input discounts up to 90 percent to reward architectures that reuse prompts. The flat-rate, all-you-can-eat plan was a product of an era when a call was a call. That era is closing, and pricing your own AI product as if it were still open is how you wake up subsidizing your heaviest users.

Open Weights Caught Up, and That Changes the Math

The second force reshaping the economics is that the cheap option got genuinely good. For most of the last three years, "open-weights" meant "almost as good, if you squint." That is no longer true at the top. On Artificial Analysis's intelligence benchmark in April 2026, the best open models scored around 54 against 60 for the strongest closed flagship, a gap of a few points rather than a generation. Nine of the thirteen models on the intelligence-versus-price frontier are open weight. Stanford's same index puts the gap between the top US and top Chinese model at 2.7 percent as of March 2026, down from 17 to 31 points in 2023.

What this means in practice is that you are no longer choosing between an expensive model that works and a free one that does not. You are choosing along a curve, and most of that curve is now usable. A model like DeepSeek V4 ships with a million-token context, runs at a fraction of frontier pricing, and can be self-hosted inside your own infrastructure. The strategic question stopped being "can we afford a good model" and became "which good model fits this specific job, at this volume, under these privacy rules."

That last clause matters more here than in most places. For a business in the EU handling client data, the ability to run a competent model on your own server or inside a private cloud is not just a cost decision, it is a compliance one. The cost math on a self-hosted AI server looks very different when the alternative is shipping regulated data to a third-party API, and the models that make it viable are now good enough that the tradeoff is real rather than theoretical.

The Move Is the Right Model for Each Job

Put the two forces together, cheaper-but-hungrier tokens and a deep bench of usable models, and the winning strategy stops being a single choice and becomes an architecture. The pattern practitioners keep converging on is the cascade, and it is simple to state. Send the high-volume, predictable 80 to 90 percent of work to a small or open or on-device model. Reserve the expensive frontier model for the hard tail that actually needs it. Done well, this captures most of the cost savings while keeping frontier reasoning available for the cases that justify it.

The dividing line is not glamour, it is task shape. Classification, extraction, routing, and short summaries are exactly what small models do well now. Microsoft's Phi-4-mini matches the quality of a far larger model on structured extraction while running in 8 gigabytes of memory. Google's Gemma 4 edge variants are multimodal and run on a phone. These are not toys, they are the right tool for the 80 percent. The frontier model earns its price on multi-step reasoning, long-document synthesis, and open-ended agent work where the inputs are wide and unpredictable and 80 percent accuracy is not good enough.

This is also why I am wary of two common reactions to the cost news. The first is "wait for prices to drop more," which misreads the paradox entirely, because your bill is driven by how many calls your design makes, not by the price of one call. The second is "just use the most expensive model for everything to be safe," which is how you turn a 2-cent task into a 20-cent one at scale for no quality gain. The discipline is matching model to job, and it is the same instinct behind treating model choice as a resilience decision rather than a brand loyalty. The agency that picks the right model for each step, and builds metering and routing in from the start, ends up with both lower costs and a system that does not fall over when one provider changes its terms.

What This Actually Means

The cost of intelligence will keep falling, and your AI bill will keep being a real line item, and both of those will stay true at the same time. That is not a contradiction to resolve, it is the operating condition to design for. The teams that internalize it will build agentic products with budget caps, cascade routing, and a clear-eyed view of which model belongs on which step. The teams that wait for the technology to get cheap enough to stop thinking about cost will keep being surprised by their invoices, because the technology already got cheap and the surprise is structural.

My prediction for the back half of 2026 is that "model strategy" becomes a normal part of any serious AI build, the way "database choice" is now, and that the wrapper-tax conversation gets loud. When a customer can see that their seat of tokens costs you 2 dollars, a flat 24-dollar plan starts to look like markup, and the products that survive will be the ones that separate the value they add from the inference they pass through. The cheap-model era did not make cost irrelevant. It moved cost from a price you look up to a decision you architect, and that is a better problem to have, as long as you actually treat it as one.


Written by Matthias Meyer of StudioMeyer, a web and AI agency on Mallorca building MCP servers, agent fleets and AI products for small and mid-size businesses. This article was originally published on the StudioMeyer blog.