惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

阮一峰的网络日志
阮一峰的网络日志
博客园 - 司徒正美
D
DataBreaches.Net
宝玉的分享
宝玉的分享
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 【当耐特】
人人都是产品经理
人人都是产品经理
博客园 - Franky
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
IT之家
IT之家
博客园 - 三生石上(FineUI控件)
J
Java Code Geeks
腾讯CDC
博客园_首页
The Cloudflare Blog
S
SegmentFault 最新的问题
C
Check Point Blog
美团技术团队
爱范儿
爱范儿
大猫的无限游戏
大猫的无限游戏
Hugging Face - Blog
Hugging Face - Blog
T
The Blog of Author Tim Ferriss
A
About on SuperTechFans
Blog — PlanetScale
Blog — PlanetScale

The New Stack | DevOps, Open Source, and Cloud Native News

Agentic development hinges on verification. For cloud-native software, that is a runtime problem. AI agents need infrastructure: Why Europe’s regional cloud strategy matters Transform your AI coding agent into a deterministic Java Spring expert WeAreDevelopers is coming to the US to give unsung developers a bigger voice Cleaner AI training data, fewer bugs: Sonar’s SonarSweep explained Observability overload is drowning engineers Google’s DiffusionGemma is 4x faster than its other Gemma models Fable 5: Guardrails and burn rate are annoying users, who say it’s still better than Opus 4.8 The Anthropic leader who built Claude Code says he ditched prompting — now he just writes loops. AWS can now mathematically prove your VMs are isolated Microsoft pulled 73 GitHub repos after malware attack — but still won’t say who’s compromised Databricks wants to kill the “email me a file” problem for AI agent skills Ramp bets forward deployed engineers can do what off-the-shelf finance AI can’t Git real: AI agents aren’t just for solo developers anymore Anthropic launches Claude Mythos/Fable 5, but you better try it soon Spring is 23 years old. AI just made it a security emergency. This AI agent startup ditched Anthropic for DeepSeek — and says it’s saving millions When your data model is the bottleneck: lessons from Medium’s feature store How long before we stop reading the code? The tokenmaxxing party is over, and Revenium is mopping up How AI is solving the memory crunch it created Microsoft’s pitch to enterprises: Ditch Azure Repos for GitHub, despite its rocky reliability record Claude Code’s biggest upgrade yet ran 5 agents at once — here’s what happened Why Anthropic just doubled Claude Cowork limits at no charge For years, Apache Cassandra handed this work to your team — 6.0 takes it back “A dangerous combination”: The 2 factors that can “corrupt” AI agent workflows With Foundry, Microsoft bets the enterprise AI battle is about reliability, not capability Microsoft unlocks Visual Studio for developers left behind by its own AI AI teams now deploy 1,000 times a month. Your pipeline wasn’t built for that. Microsoft just made the agent runtime free — and kept everything around it
“Tokenmaxxing is real, expensive & it’s spreading”: AI bu...
Adrian Bridgwater · 2026-05-28 · via The New Stack | DevOps, Open Source, and Cloud Native News

There’s a new weapon in the fight against tokenmaxxing.

Tokenmaxxing, of course, occurs when an enterprise decides that AI token usage equates to productivity. But token usage can quickly become a vanity metric, and a business that treats token gluttony as a direct measure of productivity will likely fail to map token usage to desired outcomes.

As a fad, Tokenmaxxing was wildly popular for a while, but it seems cooler heads are prevailing as the focus shifts to outcomes rather than just using AI for its own sake. Take the recent case of tokenmaxxing at Uber:

Uber CTO Neppalli Naga told The Information last month that he’s “back to the drawing board because the budget [he] thought [he] would need is blown away already.” That budget was earmarked for Uber’s use of Anthropic Claude Code.

For his part, Uber COO Andrew Macdonald responded a few weeks later, saying in a Rapid Response interview first reported on by Business Insider that Naga’s comments about blowing through the Claude budget created a “head-exploding moment” for the operations team.

“Everyone was like, ‘Oh, head-exploding moment,'” Macdonald said. “We’re going to have to start talking about token consumption and the associated costs vs. headcount, and making trades on that as an engineering organization.

“If you’re not able to draw a direct line to how many useful features and functionalities you’re shipping to your users, that trade can feel harder to justify.”

Lexi Reese, co-founder and CEO of Lanai, underscores that the problem is occurring everywhere. Uber is just the latest high-profile company to experience it.

“Tokenmaxxing is real, it’s expensive, and it’s spreading beyond just a few engineers or companies,” Reese tells The New Stack.

Likely to create zones of bloated code, agentic sprawl, and areas where software applications might eventually become brittle or even vulnerable, tokenmaxxing is expensive and reduces visibility into the total system state.

Lanai, an AI accountability company, aims to help enterprises understand where AI spend occurs, which workflows AI is applied to, and at what cost.

The company recently debuted Token Tuner to identify where lower-cost models can reduce unnecessary token costs. It’s the latest tool that developers and leaders can use to control token usage by engineers and end users. The internet is full of top-ten lists for how to reduce token usage. Companies and organizations like Kong, Braintrust, LiteLLM, and Dynatrace, among others, offer tools to ensure token usage is being budgeted.

“Tokenmaxxing is real, it’s expensive and it’s spreading beyond just a few engineers or companies.”

Reese and team have positioned Token Tuner as a service that fills the missing context gap for enterprises by mapping token spend to workflows, model choices, efficiency, and value created. The software ties each AI interaction to a measurable outcome and generates a productivity score based on how well each user matched token usage and model choice to the task at hand. 

For example, an employee using Opus 4.7 for email responses is likely to receive a lower efficiency score than if they used a smaller model for the task. 

From tokenmaxxing to outcomemaxxing

Instead of tokenmaxxing, Reese would like to see companies focus on outcomemaxxing to analyze which workflows are actually improving productivity.

Currently in beta, one Lanai Token Tuner user delegated 4.2% of all AI leverage hours across the organization while using only 0.7% of tokens. Their efficiency score was 6.0, indicating they were matching tasks to the right models, while others were burning 10x as many tokens for half the efficiency.

Lanai Chief Product Officer Mohit Mehta tells The New Stack that Token Tuner is an all-terrain vehicle, i.e., its scoring engine can calculate productivity scores when a single workflow spans multiple models simultaneously.

“Productivity is estimated by the complexity of work delegated to AI as observed through prompt and tool activity by Lanai’s proprietary models,” says Mehta. “The model operates at the level of prompts and tool invocations independent of models and applications.”

Tracking AI usage for business tasks

As we start to place greater emphasis on business results from applied technology deployments (even politicians have begun using the term “measurable outcomes” in recent times), we need to question which instrumentation is required at the API layer for Token Tuner to attribute tokens to specific business outcomes.

“Lanai aggregates prompt interactions and associated tool activity for a given session and then runs proprietary models to calculate the task type and associated productivity gain and complexity,” explains Mehta. “This enables customers to go from contextless vendor invoice to connecting intent to value to cost at the interaction level.  No custom instrumentation is required for this functionality.”

“Rather than relying on synthetic evaluations, we utilize observed outcome data, Our recommendations are grounded in how actual users within an organization achieve comparable results across different models.”

In terms of how this technology drives business efficiency, business users may ask – when Token Tuner recommends a lower-cost model, is there a benchmark in place to assess output quality equivalence before surfacing the recommendation?

“Rather than relying on synthetic evaluations, we utilize observed outcome data,” clarifies Mehta. “Our recommendations are grounded in how actual users within an organization achieve comparable results across different models.

“Rather than a recommendation like ‘this will work for you,’ we provide empirical evidence that ‘teams in your company performed this exact workflow on Haiku with equal success,’ for example. This represents real-world preference at scale over synthetic benchmarks.”

Key features include workflow-level value visibility, a service that shows which teams, workflows, and use cases are driving AI spend and whether that usage is tied to measurable business value. 

Productivity and efficiency measurement compares token spend with the leverage gained by users, teams, and workflows to show where AI creates the most value per dollar. A spend optimization recommendation function identifies runaway workflows, mismatched tasks, and premium model usage for work that lower-cost models could handle.

AI’s next killer service: efficiency?

First, the Earth cooled, and we just wanted AI… and the plain old predictive version was fine. Then, the dinosaurs died off, and we wanted domain-specific RAG-based intelligence with what then became agentic AI services that could work for us with human-in-the-loop oversight to ward off the rise of the robots. Now, perhaps, we want AI that is fit for purpose in the most applied sense of the term, so that we don’t use it where we don’t need to, and we use high-octane services only when we can really justify the turbocharge.

In truth, AI’s next killer app factor will come down to a whole lot more than just business efficiency, but this could become a more prevalent part of the mix.

YOUTUBE.COM/THENEWSTACK

Tech moves fast, don't miss an episode. Subscribe to our YouTube channel to stream all our podcasts, interviews, demos, and more.

Created with Sketch.