惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

U
Unit 42
博客园 - Franky
T
Tailwind CSS Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
月光博客
月光博客
人人都是产品经理
人人都是产品经理
雷峰网
雷峰网
Hugging Face - Blog
Hugging Face - Blog
有赞技术团队
有赞技术团队
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
阮一峰的网络日志
阮一峰的网络日志
C
Check Point Blog
爱范儿
爱范儿
T
The Blog of Author Tim Ferriss
aimingoo的专栏
aimingoo的专栏
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
L
LangChain Blog
云风的 BLOG
云风的 BLOG
MyScale Blog
MyScale Blog
Microsoft Security Blog
Microsoft Security Blog
The Cloudflare Blog
博客园 - 三生石上(FineUI控件)

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers
AI Pricing: Why Cost Optimization Is The Wrong Battle
Scott Breitenother · 2026-05-08 · via Forbes - Innovation

Scott Breitenother, Co-founder and CEO of Kilo Code.

getty

​While the conversation around AI costs swings between two extremes (are we spending too much, or not enough?), leaders are forgetting the only metric that matters: ROI.

Venture capitalist Chamath Palihapitiya vented frustration over his software company’s AI costs, which had tripled since November, trending toward $10M annually. At the other end of the spectrum, "tokenmaxxers" are competing on the number of agents they have running, and some companies are celebrating their big spenders.

AI cost optimizers want to know: Which pricing model is more predictable, and which tools or models are burning the most tokens? Is consumption-based billing finally going to sink the AI coding tool market?

These are questions worth asking, but neither penny-pinching nor throwing money at AI will help you answer the most important one: What are you actually getting back from your AI investment?

The pricing debate is a distraction.

Seat-based and consumption-based pricing models each have their merits.

Fixed subscriptions offer predictability (for now) as they shift cost risk onto providers—and most are absorbing losses on their heaviest users, having bet that inference costs will fall fast enough to catch up.

Most platforms are likely losing money on that wager, because the engineers who use AI most intensively keep upgrading to newer, more expensive frontier models rather than staying on the cheaper ones the pricing assumed.

Consumption-based billing is more transparent about this reality, but introduces its own anxiety: The bill shock problem, where costs become unpredictable as teams scale and experiment.

Neither model has fully solved for the fact that we are still in the early innings of understanding what AI productivity actually looks like at scale.

Just a few years ago, a human might have one chatbot or coding assistant asking for feedback every two minutes. This meant that ~95% of the time it was inactive. Today, with agents, task duration may be two hours. With less human input, one engineer can keep an agent busy 50% of the time.

Even as newer models become more token-efficient—Gartner predicts a 90% reduction in inference costs with some models by 2030—agents demand more tokens per task, driving up inference costs per worker.

Companies are optimizing for the wrong metric.

The economics are shifting in ways that make cost optimization a moving target.

Focusing on inputs has never been as meaningful as measuring output, and the same is true for AI: The engineers who use AI most heavily could be the ones driving the most value. Capping their access to protect a budget line is optimizing for the wrong variable.

You’re already paying about $25K each month for an engineer—worrying about $2K to $3K in token spend is coming at it from the wrong angle. AI seems like a cost center, but deployed effectively, it’s a productivity multiplier. ​

What matters is return on investment: factors like output per engineer, speed of shipping and their ability to take on more ambitious work. Restricting AI use just hamstrings your team’s potential.

Token inefficiency is a phase, not a policy failure.

There is a learning curve to working effectively with AI tools. In the early stages, engineers are inevitably inefficient. Without experience, they may run longer sessions than necessary, miss opportunities to manage context or default to frontier models for simpler tasks when older, cheaper ones would do.

This is normal—and often temporary. In my experience, provided teams are given room to experiment, users of agents often consume fewer tokens after they understand better how the tools work.

As the old joke goes, “The CFO asks the CEO, ‘What happens if we invest in developing our people and they leave us?’ The CEO responds, ‘What happens if we don’t, and they stay?’”

The engineering industry is undergoing a fundamental transformation, and everyone is learning on the job. Organizations that restrict AI investment during this period don’t avoid the cost—they defer the capability. Their teams would likely never develop the context management skills, the model routing judgment or the workflow discipline that make AI use efficient. The bill stays high because the skills never develop.

Measure what actually matters.

The question engineering leaders should be asking is not “How much are we spending on tokens?” but “What are we shipping, and is it adding value?”

Traditional DevOps Research and Assessment (DORA) metrics—pull requests merged, deployment frequency, change failure rate—are still useful anchors.

Combine them with uptime and performance data, and you start to get a real picture of AI’s impact. One founder estimated on LinkedIn that $3K in token spend per engineer has allowed to improve their output by five times. ​

The conversation about AI economics will look different in a year or even six months. As inference commoditizes, the pricing model debate will fade in relevance, much as the question of which cloud server your application runs on became largely invisible to developers.

Competition will shift toward systems that route tasks to the right model, balance cost against quality automatically and abstract the infrastructure decisions that currently eat engineering attention.

However, it's important to remember that cost visibility still matters. Rising token spend without a corresponding increase in output, long sessions with no commits or consistent use of frontier models for simple tasks are all signals worth investigating. The goal is to understand why rather than restricting access.

In practice, the highest-leverage skill engineers can develop is context management: knowing how much of the context window they’re using, when to start a fresh session and which model is appropriate for a given task. Teams that develop those habits spend less and ship faster. The CFO conversation should evolve beyond setting spending limits to aligning on what the spend is producing.​

The goal posts will move, but software engineering is still about delivering customer value. There will be ways to measure that are unique to your business, and that’s what you should focus on tracking. Token spend alone is no more insightful here than lines of code generated. ​​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?