惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

雷峰网
雷峰网
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 三生石上(FineUI控件)
博客园 - 聂微东
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hugging Face - Blog
Hugging Face - Blog
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - Franky
MyScale Blog
MyScale Blog
A
About on SuperTechFans
博客园_首页
B
Blog RSS Feed
Martin Fowler
Martin Fowler
大猫的无限游戏
大猫的无限游戏
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Vercel News
Vercel News
C
Check Point Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
博客园 - 【当耐特】
M
MIT News - Artificial intelligence
宝玉的分享
宝玉的分享
T
Tailwind CSS Blog
I
InfoQ
罗磊的独立博客

Latest from TechRadar

Quordle hints and answers for Monday, April 13 (game #1540) NYT Strands hints and answers for Monday, April 13 (game #771) NYT Connections hints and answers for Monday, April 13 (game #1037) Morbid Metal developer explains why he ditched an origami art direction in favor of gritty sci-fi — 'It worked, but it didn't really feel like me' '71% of US households get routers from ISPs': Why new FCC rules could leave millions stuck with outdated,… 'The CPU is the system’s executive layer': Intel joins SambaNova as both face existential threat from… ‘More bang for your buck’: 7 easy ways to boost your MacBook Neo’s performance for free DJI Romo P vs Roborock Saros 10R — which robot vacuum comes out on top when it comes to dodging obstacles? I put… I spent 6 hours with Genshin Impact on the Galaxy S26 Ultra, and I can't believe how far mobile gaming has come What is the release date for The Testaments episode 4 on Hulu and Disney+? I reviewed the LG G6 for 3 weeks, and it's a fantastic OLED TV that's the new best option for brighter rooms Is your bird feeder camera doing more harm than good? 3 tips for using it safely as RSPB issues urgent disease warning Chelsea vs Man City Live Streams: How to watch Premier League 2025/26 from anywhere in the world, team news How to watch Alcaraz vs Sinner for FREE: TV Channels for Monte-Carlo Masters Final Sunderland vs Tottenham Live Streams: How to watch Premier League 2025/26 from anywhere in the world, team news Are these the best-designed workout headphones ever? I used them for a month to find out How to watch Snooker 900 John Virgo online (it's free) – stream O'Sullivan vs Higgins anywhere I've only just discovered the Walk With Frodo app on Garmin's Connect IQ store — and as as a huge LOTR nerd, it's going to make the next 1,800 miles fly by 'Just not sustainable': Why your monthly £25 broadband internet bill could soon hit £45 How to watch Paris-Roubaix 2026: Free Streams & TV Info as Tadej Pogacar chases third Monument How to watch Euphoria season 3 online – stream Zendaya & Sydney Sweeney drama from anywhere today '$15K bill destroyed a solo developer’s startup': How hackers are using leaked Google API keys to… There's a sneaky way to watch UFC 327 really cheap... NYT Connections hints and answers for Sunday, April 12 (game #1036) NYT Strands hints and answers for Sunday, April 12 (game #770) Quordle hints and answers for Sunday, April 12 (game #1539) Amazon's Ring cameras are the perfect solution to secure your home on a budget — shop today's best deals… I've tested every iPhone since the iPhone 12, and Ceramic Shield 2 is the first iPhone glass I fully trust UFC 327 live stream: how to watch Procházka vs Ulberg, start time, preview, full card We're officially getting the DJI Pocket 4 on April 16, but here's how Insta360 could beat it
'5% utilization is a math fail': Millions of GPUs worth b...
Efosa Udinmw · 2026-04-22 · via Latest from TechRadar
Nvidia GPUs
(Image credit: Rude Baguette)

  • Most AI GPUs run at shockingly low utilization across production systems
  • Companies are paying for twenty times more GPU capacity than needed
  • Overprovisioning is rising sharply instead of improving year after year

Companies across the tech industry are racing to buy massive amounts of AI infrastructure, but most of it does barely any useful work at all.

A report from Cast AI, based on tens of thousands of Kubernetes clusters across AWS, Azure, and GCP, found that average GPU utilization sits at just 5%.

Many teams deploy sophisticated AI tools to manage their applications, yet those same tools are not used to optimize the underlying infrastructure.

Article continues below

The numbers are getting worse, not better

Organizations pay for roughly 20x more GPU capacity than their workloads actually use at any given moment.

The numbers come from direct measurements of production clusters and millions of compute resources before any optimization was applied.

"This is the third year we've published this report. The numbers are worse," said Laurent Gil, co-founder and President of Cast AI. "CPU utilization fell to 8%, down from 10%. Memory dropped from 23% to 20%."

The report also measured something called overprovisioning, which is the gap between what workloads actually need and what teams allocate to them.

Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!

CPU overprovisioning rose from 40% to 69% year over year, while memory overprovisioning now stands at 79%.

This means organizations reserve nearly twice as many CPU resources and four times as much memory as their workloads actually consume.

In short, organizations pay for infrastructure that their workloads do not even request, and the trend is accelerating instead of improving.

The situation gets even more expensive when comparing CPU and GPU costs directly. A CPU core sitting idle costs only cents per hour, but a GPU sitting idle costs dollars per hour.

For the first time since EC2 launched in 2006, GPU prices are rising instead of falling.

In January 2026, AWS raised H200 Capacity Block prices by 15%, citing supply and demand, which broke a two-decade precedent.

"At 5% utilization, the math doesn't work," the report states. The hoarding instinct makes sense because lead times are long, yet that same hoarding feeds the scarcity loop that drives prices even higher.

Not every cluster performs this badly, and one organization hit 49% utilization on H200s and 30% on H100s, well above the 5% average.

The difference comes down to automation rather than luck or better hardware. The tools to fix this already exist, including automated rightsizing, GPU sharing or time slicing, and Spot management.

However, most teams never get there because overprovisioning feels safer than running out of capacity, but that safety comes at a steep price.

The teams that closed the gap stopped treating resource efficiency as a manual, one-time task and started treating it as an automated, continuous process.

But Cast AI data reveals that most companies seem willing to keep paying large fees rather than change their habits.


Google logo on a black background next to text reading 'Click to follow TechRadar'

Follow TechRadar on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.


Efosa has been writing about technology for over 7 years, initially driven by curiosity but now fueled by a strong passion for the field. He holds both a Master's and a PhD in sciences, which provided him with a solid foundation in analytical thinking.

You must confirm your public display name before commenting

Please logout and then login again, you will then be prompted to enter your display name.