惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
SegmentFault 最新的问题
Jina AI
Jina AI
罗磊的独立博客
V
Visual Studio Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
J
Java Code Geeks
U
Unit 42
Microsoft Azure Blog
Microsoft Azure Blog
B
Blog RSS Feed
爱范儿
爱范儿
酷 壳 – CoolShell
酷 壳 – CoolShell
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
腾讯CDC
Hugging Face - Blog
Hugging Face - Blog
T
Tailwind CSS Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
I
InfoQ
月光博客
月光博客
博客园_首页
Vercel News
Vercel News
P
Proofpoint News Feed
GbyAI
GbyAI
Y
Y Combinator Blog

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers
AI’s Hidden Tax: Why Your Observability Stack Can’t See Y...
Satyabrat Chowdhury · 2026-06-01 · via Forbes - Innovation

Satyabrat Chowdhury, Field CTO at CoreStack | Mentor, Mentorship EDGE Program—University of Washington Bothell School of Business.

getty

​A CTO at a mid-market financial services firm showed me his cloud bill last fall. His engineering team had shipped a customer-facing AI feature three months earlier—a smart document summarization tool that was getting strong adoption. The billing surprise wasn’t that the feature was expensive. Nobody knew why the costs kept climbing. The observability tool showed green. Latency was fine. Error rates were low. His observability stack told him the system was healthy. His finance team told him he was $400K over budget.

That gap—between “operationally healthy” and “financially visible”—is where I spend most of my time now.

The Numbers Are Catching Up

For a long time, AI cloud costs were a rounding error. Training runs were expensive but infrequent. Inference was cheap. That equation has been inverted.

IDC projects that G1000 organizations will face up to a 30% rise in underestimated AI infrastructure costs by 2027—not because they’re overspending, but because they’re under-forecasting expenses that don’t fit their existing cost models. The FinOps Foundation’s most recent State of FinOps report makes the scale of the shift concrete: 98% of organizations now manage AI spend as a formal practice, up from 31% just two years ago. AI cost management isn’t an emerging discipline anymore. It’s table stakes—and most organizations are still scrambling to build the muscle.​

Why Your Observability Stack Is Flying Blind

Traditional observability was designed to answer three questions: Is the system up? Is it slow? Is it throwing errors? For a decade, that was enough. APM platforms, distributed tracing and infrastructure monitors—all built around CPU cycles, request latency and error logs.

AI inference workloads don’t follow that logic. An agentic workflow that completes successfully might make six redundant API calls to get there, each costing fractions of a cent that compound to thousands of dollars at scale.

The financial signal lives in a layer those tools simply weren’t instrumented to read: token consumption patterns, model routing decisions, cache hit rates, batch-versus-real-time inference splits. These aren’t operational metrics. They’re economic ones.

​FinOps And Observability Have To Merge

​FinOps teams have cost data. They can tell you GPU spend climbed 40% last month. They can’t tell you which model, which feature or which engineering team drove it. Observability teams have signal data. They know a model call completed in 280 milliseconds. They have no idea what it cost or why that cost mattered.

The organizations getting this right have stopped treating FinOps and observability as parallel tracks. They’re building what I’d call cost observability—an instrumentation layer that ties every model invocation to a cost center, a business function and a unit economic metric. Not a dashboard that displays spend. A system that explains spend causality.

This isn’t just a rebranding of existing practices. It requires tagging model invocations with the same rigor you tag cloud resources, building token-level cost attribution into your telemetry pipeline and treating cost anomalies with the same urgency as latency anomalies.

What This Looks Like In Practice

In conversations with engineering leaders across financial services, healthcare and retail, a consistent set of patterns is emerging among the teams making real progress.

They’ve moved from dollar alerts to rate alerts. They’re not watching total monthly spend—they’re watching token consumption per user session, per workflow, per API endpoint. They treat cache hit rate as a first-class financial metric, not just an engineering efficiency signal. They’ve built cost attribution into their deployment pipelines so every new model version ships with a spend profile alongside its performance profile.

Agentic workflows deserve particular attention. A single orchestrated agent task can involve a dozen model calls, tool executions and retrieval operations. At test scale this looks manageable. At production scale, with thousands of concurrent users, the cost multiplier can reach 100x compared to a single API call—a reality the FinOps Foundation now flags as one of the fastest-growing sources of AI budget overrun. If you don’t have token-level attribution across the full chain, you’re not managing cost. You’re discovering it after the fact.

The Ask For Executives

If you’re a CTO or vice president of engineering reading this, I’d push you to ask one question in your next architecture review: Can we trace a dollar of AI spend back to a specific product decision?

For most organizations, the answer is still no. That’s the gap to close—not by buying another monitoring tool, but by changing how your engineering and finance teams instrument and share signal together. The goal isn’t a prettier cost dashboard. It’s accountability at the model invocation level, so that when AI costs climb, you know exactly what drove them and can make a deliberate choice about whether that spend is justified.

The organizations that build this capability in the next 18 months will carry a structural cost advantage into every AI investment that follows. AI economics are only going to get more complex as agentic systems mature and token volumes scale. Building the observability foundation now—before the next surprise bill lands—is the only way to stay ahead of it.​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?