惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Check Point Blog
美团技术团队
Jina AI
Jina AI
人人都是产品经理
人人都是产品经理
The Cloudflare Blog
V
Visual Studio Blog
Google DeepMind News
Google DeepMind News
Hugging Face - Blog
Hugging Face - Blog
云风的 BLOG
云风的 BLOG
有赞技术团队
有赞技术团队
T
The Blog of Author Tim Ferriss
WordPress大学
WordPress大学
月光博客
月光博客
宝玉的分享
宝玉的分享
小众软件
小众软件
MongoDB | Blog
MongoDB | Blog
Apple Machine Learning Research
Apple Machine Learning Research
A
About on SuperTechFans
J
Java Code Geeks
博客园_首页
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
N
Netflix TechBlog - Medium
Vercel News
Vercel News
博客园 - 聂微东

Computerworld

Microsoft 365: A guide to the updates Windows 11 Insider Previews: What’s in the latest build? Windows 11: A guide to the updates Apple can't make chips fast enough, but that's only part of the story AI-led job cuts don’t always mean stronger ROI — Gartner Microsoft, Google push AI agent governance into enterprise IT mainstream Microsoft now has more than 20M paying Copilot users AI is more accurate than doctors in emergency diagnoses — study Start small, but start now: How to bring AI into your small business Apple is preparing to spend, but not necessarily on AI 10 quick productivity tips for Microsoft 365 mobile apps Relying on LLMs is nearly impossible when AI vendors keep changing things Apple breaks records, admits it can’t make Macs fast enough Spotlight report: Transforming software development with AI - Whitepaper Repository - 25 great uses for an old Android device AI chatbots need ‘deception mode’ Friendlier chatbots can be less reliable, study says Gartner sees untamed growth in agentic AI Apple reportedly abandons Vision Pro AI venture funding to shoot up this year as bubble looms Scaling up a tech startup in Europe is hard — 'EU Inc.' aims to help Apple will be behind on AI — until it isn’t EU lawmakers fail to agree on watered-down AI Act, talks pushed to May Android reminders, reinvented Who’s the better CEO, Apple’s Tim Cook or Microsoft’s Satya Nadella? AWS unveils trio of key AI strategy announcements SAS makes AI governance the centerpiece of its agent strategy Can Apple’s new CEO turn things around? Enterprises need to think beyond GPUs for agentic AI, analysts say Fleet hopes to be the MDM provider for the AI Era
How companies are racing to solve the AI token problem
Agam Shah · 2026-06-18 · via Computerworld

Generative AI services and tools that use tokens to produce results can get expensive quickly. That’s spurring IT leaders to look for new ways to reduce token use — and save money.

Because generative AI (genAI) tools and services have become so ubiquitous (and popular), the costs of using them are going through the roof — leading to an insatiable appetite for tokens.

Tokens represent a common way to measure and price AI use. Much like letters and words in English, large language models (LLMs) grasp a sentence or query by breaking words into tokens.

With the AI explosion well under way, tokens are now “the fundamental units of data our models process, many representing a problem being solved,” according to Google CEO Sundar Pichai. (Google, by the way, processes about 3.2 quadrillion tokens a month.)

But as the price of all those tokens adds up, business and IT execs are looking for ways to cut costs while keeping corporate productivity up. Uncontrolled token use has already landed one company with an unexpected $500 million AI bill

There are a number of ways companies can rein in the price of AI at the model, infrastructure, silicon, and business levels. Here’s a look at how some of those savings might actually be achieved.

Switch to lower-cost models

One way of potentially saving money is by re-routing AI work to a cheaper model, Pichai said. At Google that would Gemini 3.5 Flash. It delivers “frontier-level capabilities at less than half the price of comparable frontier models.

“If companies use a mix of [Gemini 3.5] Flash and other frontier models, they could save a lot of money,” Pichai said.

Those kinds of models provide cheaper tokens, with reasoning that’s good enough for many users — if not as strong as mainstream Gemini 3.5 — to deliver useful results.

“There is sometimes overkill with the [LLMs],” said Deepak Seth, senior director analyst at Gartner. “I don’t always need a large language model which has been trained on the works of Charles Dickens and Shakespeare and Harry Potter.”

Hyperframe Research principal analyst Steven Dickens can’t stop using Amazon’s Quick, which costs $20 a month, for personal tasks. “It is great personal ROI as it has not only made tasks faster, but unlocked tasks I would never have even attempted previously,” Dickens said.

Don’t forget the hardware and software part of the equation

The token crisis isn’t new, said Dheeraj Pandey, CEO of DevRev, who likens what’s going on now in the AI market to the disruptions that emerged with the arrival of cloud computing and virtualization years ago. 

“We let chaos reign and then we had to rein in the chaos,” Pandey said. “The word that people started using was server consolidation and virtualization.”

The answer to the token problem, he said, is the same: “Anything in systems can be solved with caching and indirection.”

DevRev, for example, is building a memory layer between AI agents and primary data sources, such as Salesforce or ERP records; that can cut token load and make data movement more efficient. The layer holds a knowledge graph with answers to common agent questions and runs on cheaper CPUs, avoiding more costly GPU cycles.

Sending agents straight at systems like ServiceNow and Salesforce “will burn a lot more tokens. It’s also not precise. And finally, it’s not safe enough where I can roll it back in case an agent has committed a mistake,” Pandey said.

Network automation firm NetBrains uses a different method: It uses conventional computing to map a network’s layout then feeds only key information to models for planning and reasoning, where AI excels. “So you don’t have to spend all the tokens,” said Netbrains CTO Sang Peng.

Focus on prompt efficiency

Staffing firm ManpowerGroup has found that prompt efficiency can be an effective tool for improving token use, both internally and externally for clients.

For example, users accessing its internal labor-market tool initially needed 10 follow-up questions to drill into a query. A year later, more efficient use of prompts has brought that number down to an average of four, said Max Leaming, head of data science and AI solutions at ManpowerGroup. 

“They’re using fewer tokens and they’re simply more efficient,” he said. “And that in large part has to do with your ability to prompt efficiently.”

Go local

New AI hardware that generates free tokens at home could ease some of the cost crisis.

At GTC Taipei earlier this month, Nvidia and Microsoft unveiled RTX Spark, an agentic AI desktop PC that runs agents and 120-billion-parameter models locally on Windows. The goal is “to deliver unmetered intelligence to every home and every desk with Windows,” Microsoft CEO Satya Nadella said in a statement.

Some companies are looking to reduce cloud AI costs by putting their own hardware in data centers, with vendors such as HPE and Dell providing servers installed in independent facilities. (On-premise AI is gaining ground amid sovereign AI and geopolitical concerns, including the recent conflict in the Middle East, where large data centers were struck with missiles.)

“There are local, region-specific and multiple vendor AI solutions. All of those things can help mitigate the risk. But they’re not going to eliminate it,” said Max Goss, senior director analyst at Gartner.

Use forward-deployed engineers

Reducing token costs is something that may fall to forward-deployed engineers (FDEs) in customer environments, said Taimur Rashid, managing director of AWS’s Generative AI Innovation Center.

“I expect these teams to be able to architect systems that have those cost requirements in mind, whether it’s use a different model or a different use case that doesn’t increase the per-token cost,” Rashid said.

Companies may spend heavily on token consumption, “but if you’re generating revenue, as long as the economics work out, then you’re at peace,” Rashid said.

The use of FDEs is gaining ground as IT decision-makers look to both rollout successful AI deployments while also keeping an eye on costs.

Change the measure of success from tokens to outcomes

Even with the current emphasis on reducing token use to save money, the metrics used to measure AI success are likely to shift, Gartner’s Seth said. At some point, token-based pricing will move more toward an outcome-based model, where the unit of value is outcomes, not fragments of words.

“Some companies are moving towards outcome-based pricing,” Seth said. “When people start realizing the real cost of tokens, then companies will start looking at token efficiency.”

SUBSCRIBE TO OUR NEWSLETTER

From our editors straight to your inbox

Get started by entering your email address below.