




















For weeks now, a new term has been stirring debate across the tech industry: tokenmaxxing. It refers to the practice of maximizing AI token consumption — whether to hit internal productivity metrics or to climb to the top of in-house leaderboards. Following the recent Google I/O keynote, in which CEO Sundar Pichai explicitly picked up the term, warning voices from across the industry are growing louder.
Most recently, Peter Steinberger, founder of OpenClaw and now at OpenAI, drew attention with a screenshot he shared on X showing that he had burned through the equivalent of $1.3 million in tokens for OpenAI’s coding agent Codex over the past 30 days.
The term — riffing on Gen Z slang like “looksmaxxing” or “sleepmaxxing” — broke into the mainstream in April 2026, after industry outlet The Information reported on an internal dashboard at Meta Platforms. An employee had set up a leaderboard there on his own initiative, named “Claudeonomics,” ranking colleagues by individual token consumption and awarding titles such as “Token Legend,” “Model Connoisseur,” or “Cache Wizard.” According to The Information, Meta employees consumed around 60 trillion tokens in 30 days; the top-ranked user alone accounted for roughly 281 billion tokens — a volume that, at standard API prices, can translate into costs ranging from several hundred thousand to several million US dollars. The dashboard was taken offline a few days later.
Similar internal competitions have since been documented at Microsoft and Amazon. At Google itself, Pichai acknowledged on the I/O stage: “Some out there might call this tokenmaxxing, and there’s probably some truth to it.” According to its own figures, Google now processes 3.2 quadrillion tokens per month — two years ago, the number stood at 9.7 trillion.
How quickly the game turns into a business problem is illustrated by the most prominent example of recent weeks: Uber. CTO Praveen Neppalli Naga had disclosed in an April interview with The Information that the mobility group had already used up its entire 2026 annual budget for tools such as Claude Code and Cursor within just four months. In the first quarter of 2026, the share of engineers using Claude Code rose from 32 to 84 percent. With roughly 5,000 engineers, each individual currently spends between $500 and $2,000 per month on AI tools alone — adding up to millions of dollars per month.
More striking than the numbers, however, is the retrospective assessment: Uber President and COO Andrew Macdonald spoke on the Rapid Response podcast of a “head-exploding moment” and publicly questioned whether higher token spending actually translates into a proportional productivity gain. His conclusion after conversations with the CTO’s team: implicitly, more features were being shipped, but a direct line between token consumption and “25 percent more useful consumer features” simply could not be drawn. Macdonald’s pointed remark: “AI seems free when you’re just sitting there coming up with interesting scenarios. But ultimately the company pays for it.”
In response to similar cost blow-ups, Microsoft has revoked Claude Code access for thousands of internal engineers, shifting them to GitHub Copilot CLI in order to save money ahead of the new fiscal year.
It is into precisely this debate that Eugene Cheah, CEO and co-founder of Featherless.ai, now steps with a clear warning to the industry: using token consumption as a yardstick for success, he argues, misleads companies about the actual economic value of their AI deployments.
“Token use is one metric, but extreme usage in the guise of tokenmaxxing is, in most cases, not a sustainable business model and an imprecise way to understand real value,” says Cheah. “It is a clumsy way of gauging success. Not all tokens are created equal; different actions create different returns for companies. Chasing these numbers shows that some still do not understand the actual mechanics of AI return on investment.”
Cheah argues that the next phase of enterprise AI will not be defined by maximization but by token minimization: “While engineering teams often treat massive context windows and high throughput as vanity metrics, the next phase will actually be about the opposite. Every unnecessary token generated is a direct tax on corporate productivity, slowing down latency and draining unit economics.”
And further: “The approach of relying on one massive model to handle every task simply encourages wasteful generation. Instead, smarter architectures use smaller, specialised models designed to achieve pinpoint precision with a fraction of the compute. In the near future, the most sophisticated AI frameworks will be judged by how little they actually need to generate to get the job done.”
Cheah also points to an effect that is becoming particularly visible in the industry right now: “A surge in token usage volume is completely normal in the early days of a high profile new AI product launch, particularly when introductory costs are minimal. However, the real demand and long-term viability of any AI platform only becomes clear when pricing normalises and the true costs kick in for businesses.”
Observers are increasingly framing the tokenmaxxing phenomenon as a textbook case of Goodhart’s Law — the observation that a measure ceases to be a good measure as soon as it becomes a target. Linear COO Cristina Cordova summed it up on X: ranking engineers by token spend, she wrote, is like ranking a marketing team by who spent the most money.
At the same time, the movement is not without its defenders: Y Combinator CEO Garry Tan, for one, has embraced the term, and Meta CTO Andrew Bosworth told Forbes that his best engineer was spending the equivalent of his salary in tokens — but was, in return, “five to ten times more productive.”
That the hyperscalers are taking the headwinds seriously became apparent on the I/O stage: Pichai positioned Gemini 3.5 Flash explicitly as a way out of the tokenmaxxing hangover. A customer running a trillion tokens a day, he said, could save more than a billion US dollars annually by shifting 80 percent of its workloads onto Flash.
The message emerging from the Meta, Uber, and Microsoft cases, and from Cheah’s warning, is hard to miss: anyone still convinced in 2026 that more tokens automatically mean more productivity may be in for a surprise when the next bill arrives.
Aus Datenschutz-Gründen ist dieser Inhalt ausgeblendet. Die Einbettung von externen Inhalten kann in den Datenschutz-Einstellungen aktiviert werden:
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。