惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The Register - Security
The Register - Security
Cisco Talos Blog
Cisco Talos Blog
P
Proofpoint News Feed
Vercel News
Vercel News
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
C
Check Point Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Y
Y Combinator Blog
V
Visual Studio Blog
H
Help Net Security
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Stack Overflow Blog
Stack Overflow Blog
The Cloudflare Blog
The Last Watchdog
The Last Watchdog
博客园 - 司徒正美
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Apple Machine Learning Research
Apple Machine Learning Research
SecWiki News
SecWiki News
博客园 - 叶小钗
V
Vulnerabilities – Threatpost
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
IT之家
IT之家
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
罗磊的独立博客
C
CXSECURITY Database RSS Feed - CXSecurity.com
V2EX - 技术
V2EX - 技术
T
The Blog of Author Tim Ferriss
小众软件
小众软件
The GitHub Blog
The GitHub Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
D
DataBreaches.Net
L
LINUX DO - 热门话题
大猫的无限游戏
大猫的无限游戏
V
V2EX
Latest news
Latest news
NISL@THU
NISL@THU
Last Week in AI
Last Week in AI
Spread Privacy
Spread Privacy
云风的 BLOG
云风的 BLOG
S
Secure Thoughts
W
WeLiveSecurity
S
Security @ Cisco Blogs
C
CERT Recently Published Vulnerability Notes
AWS News Blog
AWS News Blog
I
InfoQ
A
About on SuperTechFans
K
Kaspersky official blog
Security Latest
Security Latest
P
Proofpoint News Feed

Computerworld

Microsoft 365: A guide to the updates Windows 11 Insider Previews: What’s in the latest build? Windows 11: A guide to the updates Apple can't make chips fast enough, but that's only part of the story AI-led job cuts don’t always mean stronger ROI — Gartner Microsoft, Google push AI agent governance into enterprise IT mainstream Microsoft now has more than 20M paying Copilot users AI is more accurate than doctors in emergency diagnoses — study Start small, but start now: How to bring AI into your small business Apple is preparing to spend, but not necessarily on AI 10 quick productivity tips for Microsoft 365 mobile apps Relying on LLMs is nearly impossible when AI vendors keep changing things Apple breaks records, admits it can’t make Macs fast enough Spotlight report: Transforming software development with AI - Whitepaper Repository - 25 great uses for an old Android device AI chatbots need ‘deception mode’ Friendlier chatbots can be less reliable, study says Gartner sees untamed growth in agentic AI Apple reportedly abandons Vision Pro AI venture funding to shoot up this year as bubble looms Scaling up a tech startup in Europe is hard — 'EU Inc.' aims to help Apple will be behind on AI — until it isn’t EU lawmakers fail to agree on watered-down AI Act, talks pushed to May Android reminders, reinvented Who’s the better CEO, Apple’s Tim Cook or Microsoft’s Satya Nadella? AWS unveils trio of key AI strategy announcements SAS makes AI governance the centerpiece of its agent strategy Can Apple’s new CEO turn things around? Enterprises need to think beyond GPUs for agentic AI, analysts say Fleet hopes to be the MDM provider for the AI Era Why simplicity is the silent driver of hybrid workplace success Why security matters in the meeting room Can everyday IT decisions turn sustainability from intent into impact? Why the meeting room has become the true test of hybrid work Why smart meeting rooms are becoming strategic IT assets How collaboration technology defines the next phase of hybrid work Microsoft, OpenAI change contract terms–again OpenAI plans its own ‘iPhone killer’ Your AI strategy is all wrong Agent Mode is now available in Microsoft Word, Excel, and PowerPoint Adobe bets on AI agents to stay at the center of marketing workflows Microsoft to offer voluntary retirement buyouts to about 7% of the US workforce Google Keep cheat sheet: How to get started The AI workplace paradox: Higher productivity, higher anxiety Gartner: Global IT spending to grow by 13.5% this year Apple may be the only laptop vendor to grow in 2026 Tim Cook’s legacy: a successful CEO who stumbled over AI Google Chat becomes an agent interface for Workspace Gemini Enterprise update brings AI agents into collaborative workflows Meta to track employee keystrokes, screen activity to train AI agents The smartest ways to sync your Android and computer clipboards Microsoft trims cloud desktop pricing, even as it boosts AI costs Adobe builds an ‘agentic content supply chain’ for the AI era You can now test and compare AI models on LinkedIn With John Ternus as CEO, expect Apple’s platforms to proliferate Apple CEO Tim Cook stepping down, to be replaced by John Ternus Global RAM shortage appears set to continue through 2027 Is this where Apple Silicon will be in 5 years? AI-ready skills are not what you think World ID expands its ‘proof of human’ vision for the AI era Microsoft's Patch Tuesday updates: Keeping up with the latest fixes Microsoft’s Patch Tuesday release for April is a whopper Robot Zuckerberg shows how IT can free up CEOs’ time UK wants to build sovereign AI — with just 0.08% of OpenAI’s market cap 20 tricks for more efficient Android messaging AI is finally delivering productivity — for remote employees Google should share search data to break its monopoly, European Commission suggests How to think about Apple Business Microsoft Teams cheat sheet: How to get started Reporter’s notebook: In Nepal and Sri Lanka, AI boom brings hope How to create your own custom Android air gesture Can Microsoft really meet its carbon-negative goal by 2030? About the Best Places to Work in IT Microsoft to cut Windows 365 price for SMBs Blancco confirms Mac adoption is accelerating Apple devices’ satellite link is under new ownership IBM’s government DEI settlement could increase pressure to avoid tech hiring diversity Microsoft is developing Copilot features inspired by Openclaw Global RAM shortage prompts Microsoft to hike Surface prices Apple Business rolls out to 200+ countries today Windows 10: A guide to the updates Nvidia’s Stephen Jones on the toolkit powering GPUs: ‘A wild ride’ The French government eyes alternatives to Windows Apple preps for the face race How to build your own AI agents with Google Workspace Studio Adobe Summit 2026: How Adobe hopes to redesign marketing and creativity with AI DARPA wants to help AI agents to talk to one another Apple unveiled a new high-end market opportunity this week Microsoft adds hidden feature flags to Windows Insider builds Meta moves fast toward a world where AI builds the software PC sales rise in Q1 despite memory shortage — IDC Agentic AI – Ongoing coverage of its impact on the enterprise Google’s new AI app is a glimpse of the future This problem might not need a solution: customer-service bots that code for free Chrome, Vivaldi, and the challenge of changing browsers The new M5-based MacBook Air is built to last — and perform Apple worst, Asus best for laptop repairability US court refuses to stay Pentagon’s ‘supply-chain risk’ blacklisting of Anthropic The top priority for Adobe’s next CEO? Prepping for the ‘age of agents’ It's iPhone speculation time: flips, flaps — and Fold
How companies are racing to solve the AI token problem
Agam Shah · 2026-06-18 · via Computerworld

Generative AI services and tools that use tokens to produce results can get expensive quickly. That’s spurring IT leaders to look for new ways to reduce token use — and save money.

Because generative AI (genAI) tools and services have become so ubiquitous (and popular), the costs of using them are going through the roof — leading to an insatiable appetite for tokens.

Tokens represent a common way to measure and price AI use. Much like letters and words in English, large language models (LLMs) grasp a sentence or query by breaking words into tokens.

With the AI explosion well under way, tokens are now “the fundamental units of data our models process, many representing a problem being solved,” according to Google CEO Sundar Pichai. (Google, by the way, processes about 3.2 quadrillion tokens a month.)

But as the price of all those tokens adds up, business and IT execs are looking for ways to cut costs while keeping corporate productivity up. Uncontrolled token use has already landed one company with an unexpected $500 million AI bill

There are a number of ways companies can rein in the price of AI at the model, infrastructure, silicon, and business levels. Here’s a look at how some of those savings might actually be achieved.

Switch to lower-cost models

One way of potentially saving money is by re-routing AI work to a cheaper model, Pichai said. At Google that would Gemini 3.5 Flash. It delivers “frontier-level capabilities at less than half the price of comparable frontier models.

“If companies use a mix of [Gemini 3.5] Flash and other frontier models, they could save a lot of money,” Pichai said.

Those kinds of models provide cheaper tokens, with reasoning that’s good enough for many users — if not as strong as mainstream Gemini 3.5 — to deliver useful results.

“There is sometimes overkill with the [LLMs],” said Deepak Seth, senior director analyst at Gartner. “I don’t always need a large language model which has been trained on the works of Charles Dickens and Shakespeare and Harry Potter.”

Hyperframe Research principal analyst Steven Dickens can’t stop using Amazon’s Quick, which costs $20 a month, for personal tasks. “It is great personal ROI as it has not only made tasks faster, but unlocked tasks I would never have even attempted previously,” Dickens said.

Don’t forget the hardware and software part of the equation

The token crisis isn’t new, said Dheeraj Pandey, CEO of DevRev, who likens what’s going on now in the AI market to the disruptions that emerged with the arrival of cloud computing and virtualization years ago. 

“We let chaos reign and then we had to rein in the chaos,” Pandey said. “The word that people started using was server consolidation and virtualization.”

The answer to the token problem, he said, is the same: “Anything in systems can be solved with caching and indirection.”

DevRev, for example, is building a memory layer between AI agents and primary data sources, such as Salesforce or ERP records; that can cut token load and make data movement more efficient. The layer holds a knowledge graph with answers to common agent questions and runs on cheaper CPUs, avoiding more costly GPU cycles.

Sending agents straight at systems like ServiceNow and Salesforce “will burn a lot more tokens. It’s also not precise. And finally, it’s not safe enough where I can roll it back in case an agent has committed a mistake,” Pandey said.

Network automation firm NetBrains uses a different method: It uses conventional computing to map a network’s layout then feeds only key information to models for planning and reasoning, where AI excels. “So you don’t have to spend all the tokens,” said Netbrains CTO Sang Peng.

Focus on prompt efficiency

Staffing firm ManpowerGroup has found that prompt efficiency can be an effective tool for improving token use, both internally and externally for clients.

For example, users accessing its internal labor-market tool initially needed 10 follow-up questions to drill into a query. A year later, more efficient use of prompts has brought that number down to an average of four, said Max Leaming, head of data science and AI solutions at ManpowerGroup. 

“They’re using fewer tokens and they’re simply more efficient,” he said. “And that in large part has to do with your ability to prompt efficiently.”

Go local

New AI hardware that generates free tokens at home could ease some of the cost crisis.

At GTC Taipei earlier this month, Nvidia and Microsoft unveiled RTX Spark, an agentic AI desktop PC that runs agents and 120-billion-parameter models locally on Windows. The goal is “to deliver unmetered intelligence to every home and every desk with Windows,” Microsoft CEO Satya Nadella said in a statement.

Some companies are looking to reduce cloud AI costs by putting their own hardware in data centers, with vendors such as HPE and Dell providing servers installed in independent facilities. (On-premise AI is gaining ground amid sovereign AI and geopolitical concerns, including the recent conflict in the Middle East, where large data centers were struck with missiles.)

“There are local, region-specific and multiple vendor AI solutions. All of those things can help mitigate the risk. But they’re not going to eliminate it,” said Max Goss, senior director analyst at Gartner.

Use forward-deployed engineers

Reducing token costs is something that may fall to forward-deployed engineers (FDEs) in customer environments, said Taimur Rashid, managing director of AWS’s Generative AI Innovation Center.

“I expect these teams to be able to architect systems that have those cost requirements in mind, whether it’s use a different model or a different use case that doesn’t increase the per-token cost,” Rashid said.

Companies may spend heavily on token consumption, “but if you’re generating revenue, as long as the economics work out, then you’re at peace,” Rashid said.

The use of FDEs is gaining ground as IT decision-makers look to both rollout successful AI deployments while also keeping an eye on costs.

Change the measure of success from tokens to outcomes

Even with the current emphasis on reducing token use to save money, the metrics used to measure AI success are likely to shift, Gartner’s Seth said. At some point, token-based pricing will move more toward an outcome-based model, where the unit of value is outcomes, not fragments of words.

“Some companies are moving towards outcome-based pricing,” Seth said. “When people start realizing the real cost of tokens, then companies will start looking at token efficiency.”

SUBSCRIBE TO OUR NEWSLETTER

From our editors straight to your inbox

Get started by entering your email address below.