惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
腾讯CDC
博客园 - 聂微东
爱范儿
爱范儿
罗磊的独立博客
P
Proofpoint News Feed
博客园 - Franky
博客园 - 三生石上(FineUI控件)
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
酷 壳 – CoolShell
酷 壳 – CoolShell
Jina AI
Jina AI
Blog — PlanetScale
Blog — PlanetScale
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 司徒正美
美团技术团队
MongoDB | Blog
MongoDB | Blog
WordPress大学
WordPress大学
A
About on SuperTechFans
I
InfoQ
博客园_首页
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
H
Help Net Security
Microsoft Azure Blog
Microsoft Azure Blog
G
Google Developers Blog

... eeNews Europe

Molex Teramount deal targets co-packaged optics NVIDIA and ServiceNow extend AI governance from desktops to data centres Faraday Future launches Physical AI robotics institute with BIBS System Check: Should engineers learn analog? Decoupled by Design: How Gateworks and NXP are rethinking edge AI architecture NVIDIA and Corning partner on AI photonics expansion GCT taps satellite partner to speed 5G rollout Sodankylä supersite to support ESA Earth observation SiTime posts 88% revenue growth on AI infrastructure demand Infrared LEDs support in-cabin sensing for vehicle safety Anthropic compute deal taps SpaceXAI Colossus 1 Quantum Brilliance CEO Mark Luo on deployable quantum systems and the future of diamond-based computing SEMI Summit spotlights Europe’s chiplet and packaging push AI data center infrastructure drives Pennsylvania energy expansion NXP CoreRide gains Vector software support for SDV platforms Elektor Lab Talk covers Red Pitaya and reconfigurable test gear SEMI names Julie Rogers to lead ESD Alliance ASML CEO backs joint call for Europe tech competitiveness push FlexIC RFID inlays bring NFC to paper packaging ROHM targets smart rings with ultra-compact NFC wireless power chipset China silicon wafers push boosts Eswin capacity ESD Alliance outlook spotlights agentic AI in chip design AI robotics sales growth rises as Faraday Future expands into education Microchip expands dsPIC33A controllers for AI data center power sensiBel MEMS microphone heads to Silex production SEMI: Global silicon wafer shipments jump 13% on AI demand AI drives photonics innovation Advantech adds Intel Core Series 3 to edge AI systems NI CHESS enables software-driven RF channel emulation into aerospace testing Forsee Power battery system powers new electric fire pump
AI token costs force rethink at Uber and Microsoft
Brian Tristam Williams · 2026-05-29 · via ... eeNews Europe

AI token costs force rethink at Uber and Microsoft

Business news |

By Brian Tristam Williams



AI token costs are becoming harder to treat as a rounding error, as agentic coding tools and enterprise AI workflows push usage from simple prompts into long, multi-step inference jobs.

Goldman Sachs Research says agentic AI could drive a 24-fold increase in token consumption by 2030, reaching 120 quadrillion tokens per month as consumer and enterprise adoption grows. The bank’s analysis, published earlier this month, argues that the same trend could improve the economics of hyperscalers and model providers if inference costs keep falling faster than demand rises.

AI token costs move from hype to budget line

The problem for customers is that lower unit costs do not automatically mean lower bills. Agentic tools can call models repeatedly, review context, generate code, run checks, and revise their own output. That turns a single developer request into a chain of token-consuming actions. This is why token-based billing is becoming a practical issue for engineering organisations rather than a narrow cloud-infrastructure concern.

Uber has become one of the more visible examples. The company is reassessing parts of its AI spending after reports that its 2026 AI budget had been exhausted within the first few months of the year. Uber president and COO Andrew Macdonald has said the company does not yet see a clear link between higher token consumption and more useful consumer-facing features. That does not mean the tools are useless, but it does make the cost-benefit argument less automatic.

Microsoft is facing a related issue inside its own engineering operations. The company is reportedly winding down most internal Claude Code licences for parts of its Experiences + Devices group and steering developers towards GitHub Copilot CLI by the end of June. Separately, GitHub has announced that Copilot plans will move to usage-based billing from 1 June 2026, with GitHub AI Credits consumed according to token usage across input, output, and cached tokens.

AI token costs expose the hardware gap

The Goldman Sachs view is not simply bearish. It expects semiconductor providers to cut inference cost per token by 60% to 70% per year through chip and architecture improvements. It also expects chip supply to remain constrained for the next 12 to 18 months as production capacity catches up with the pace of new AI use cases.

That makes the story relevant well beyond software procurement. If agents become a default interface for coding, customer service, search, and enterprise workflow automation, the load shifts back into datacentre silicon, networking, memory, storage, and power infrastructure. As previously reported by eeNews Europe when ARM set out its datacentre CPU plans, agentic AI is already being used to justify new processor strategies for AI datacentres.

The immediate lesson is more prosaic. Businesses are being pushed to measure AI against shipped features, resolved support cases, reduced engineering time, or revenue impact, not against token volume. AI token costs may fall at the hardware level, but agentic workflows can easily spend the savings before finance teams see them.

If you enjoyed this article, you will like the following ones: don't miss them by subscribing to :

   eeNews on Google News