惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

MyScale Blog
MyScale Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
人人都是产品经理
人人都是产品经理
V
Visual Studio Blog
博客园 - 叶小钗
A
About on SuperTechFans
Last Week in AI
Last Week in AI
量子位
博客园 - 三生石上(FineUI控件)
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
MongoDB | Blog
MongoDB | Blog
T
The Blog of Author Tim Ferriss
Vercel News
Vercel News
博客园 - 司徒正美
博客园 - Franky
博客园 - 【当耐特】
月光博客
月光博客
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Apple Machine Learning Research
Apple Machine Learning Research
Hugging Face - Blog
Hugging Face - Blog
S
SegmentFault 最新的问题
大猫的无限游戏
大猫的无限游戏
博客园 - 聂微东
J
Java Code Geeks

The Register

Grafana offers AI assistant for free, warns users not to go mad Right to repair champ Framework punts modular 13in laptop with Core Ultra Series 3 Scotland Yard can keep using live facial recognition on Londoners, say judges UK tribunal sends £2B claim accusing Microsoft of overcharging for licensing to trial Nation-states want to cause harm, not just steal cash - stop handing your cyber defenses to the cheapest contractor Murder, she wrote: Ex-FBI chief wants some ransomware crims charged with homicide Phone-to-satellite use goes into orbit, growing 25% in 8 months macOS ClickFix attacks deliver AppleScript stealers to snarf credentials, wallets Anthropic bakes memory fixes into Bun 1.1.13 as developers complain of leaks The spaghettified DBMS chart that shows Oracle's crown is slowly slipping Yet another ex-ransomware negotiator admits turning rogue after payoff from crimelords FAA grounds Blue Origin's New Glenn as it probes missed satellite delivery 'mishap' AMD's Ryzen 9 9950X3D2 Dual Edition tested: Gratuitous overkill with a price to match AI-assisted intruders pwned Vercel via OAuth abuse and a pilfered employee account Crook claims to leak 'video surveillance footage' of companies Met police trials snoop tech platform in push to cuff more London shoplifters England's school phone ban gets teeth, just in time to bite no one Adaptavist Group breach spawns imposter emails as ransomware crew claims mega-haul Panasonic creates device-locked QR codes to speed facial biometric capture Iran claims US used backdoors to knock out networking equipment during war NASA Inspector fears new spacesuits won’t be ready for Moon landing Vibe coding upstart Lovable denies data leak, cites 'intentional behavior,' then throws HackerOne under the bus Trump-branded datacenter project fails to make itself great, again World's blandest man steps down from CEO job to spend more time in tastefully appointed home Chase got a spiff of $77 million to create one job with New York datacenter Scot becomes second Scattered Spider-linked crook to plead guilty in US You too can build a nuclear battery from junk you have lying around the house Schmoozebots: study finds flattery will get AI everywhere One of Europe's sovereign cloud picks may not be so-sovereign after all New Android development tool designed for robots, not humans
Datadog digs down into GPU efficiency as AI costs soar
Joe Fay · 2026-04-23 · via The Register

Datadog has added GPU monitoring to its observability stack, giving AI-hungry organizations more insight into exactly what's happening on their most expensive silicon.

AMD Ryzen 9950X3D2-DE

AMD's Ryzen 9 9950X3D2 Dual Edition tested: Gratuitous overkill with a price to match

READ MORE

The observability vendor says GPU instances now make up 14 percent of cloud compute costs as companies clamber on the AI bandwagon, and that GPU spend will take up an even bigger proportion of cloud compute spend in the future.

Earlier this month, IDC said: "Worldwide spending on artificial intelligence (AI) infrastructure reached $89.9 billion in Q4 2025" up 62 percent on the year. And accelerated compute – mainly GPUs – is the "structural backbone" of this.

But there's plenty of debate over what value – if any – companies are deriving from their massive AI investments.

Datadog is not getting into that bearpit. But as chief product officer Yanbing Li puts it, "While these companies can see their costs climbing, they can't chargeback GPU spend across business units, see workload context or identify clear next steps for improvement."

To address that, Datadog claims, its latest tool offers unified visibility across the AI stack, "giving customers a single view linking GPU fleet health, cost, and performance directly to the teams relying on them for faster troubleshooting of slow workloads and cost savings."

A longer explainer says the tooling works across both cloud and neocloud instances as well as on-prem GPU fleets – handy if sovereignty concerns are making you wary of AI in the cloud.

"It's easy to see how much of your fleet is sitting completely idle or being ineffectively consumed by a workload that doesn't require GPUs at all," it says. "You can drill into the Fleet Explorer to hold each team accountable for their GPU utilization and spend."

As well as identifying stalled or zombie processes soaking on GPU time, it will spot workloads that were never configured for GPUs in the first place, effectively burning cash.

"Internally at Datadog, GPU Monitoring helped us save tens of thousands in monthly expenses by identifying and removing a serving pod that had been stuck in the initialization phase," the explainer said.

"Rising costs are often driven by operational inefficiency rather than hardware alone. By linking cost to utilization and workload behavior, teams can reduce waste while maintaining performance."

Datadog is certainly not alone in extending observability further down the AI stack. This week also saw Grafana launch observability tools for AI, provide insights into agent behavior, while its Grafana Cloud platform offers GPU observability tools covering hardware utilization and resource allocation, as well as cost optimization.

Earlier this month, Nutanix unveiled a multi-tenancy framework to allow organizations to run more workloads on their previous GPUs, and provide more insight into how AI systems are chewing through tokens.

So, it's getting easier to work out how much individual AI workloads are costing you, and what processes and software misconfigurations could be making bills higher than necessary.

This means enterprises can ensure their AI infrastructure and their associated apps and agent are running as efficiently as possible. Whether this means enterprises can actually start working out whether they're getting value from AI investments may be quite another question. ®