惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
V
Vulnerabilities – Threatpost
C
Cisco Blogs
A
Arctic Wolf
L
LINUX DO - 热门话题
P
Proofpoint News Feed
Security Latest
Security Latest
AWS News Blog
AWS News Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Cisco Talos Blog
Cisco Talos Blog
L
Lohrmann on Cybersecurity
W
WeLiveSecurity
爱范儿
爱范儿
Last Week in AI
Last Week in AI
Hacker News - Newest:
Hacker News - Newest: "LLM"
S
Security Affairs
PCI Perspectives
PCI Perspectives
C
Cybersecurity and Infrastructure Security Agency CISA
Spread Privacy
Spread Privacy
IT之家
IT之家
月光博客
月光博客
云风的 BLOG
云风的 BLOG
宝玉的分享
宝玉的分享
J
Java Code Geeks
美团技术团队
酷 壳 – CoolShell
酷 壳 – CoolShell
I
Intezer
博客园_首页
博客园 - 司徒正美
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Palo Alto Networks Blog
NISL@THU
NISL@THU
Recent Commits to openclaw:main
Recent Commits to openclaw:main
有赞技术团队
有赞技术团队
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
量子位
The Last Watchdog
The Last Watchdog
Google Online Security Blog
Google Online Security Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 聂微东
N
News and Events Feed by Topic
Webroot Blog
Webroot Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Security @ Cisco Blogs
罗磊的独立博客
大猫的无限游戏
大猫的无限游戏
The Cloudflare Blog
V
V2EX
Jina AI
Jina AI

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
AI Costs Are Cloud Costs Now: Why FinOps Is the New Playbook for AI Spend
Khushi Dubey · 2026-06-24 · via DEV Community

Something quietly changed inside finance dashboards over the last eighteen months. The line item for AI tools used to be small and predictable. Now it sits right next to the cloud bill, growing at a pace nobody fully forecasted, and looking suspiciously similar to how AWS looked back in 2015.

This is not a coincidence. AI coding assistants, model APIs, and agent platforms all bill on usage. They are variable. They are skewed by power users. And most teams have almost no visibility into who is spending what, on which models, for which projects.

In this guide, you will learn why AI spend behaves exactly like cloud infrastructure spend, what FinOps lessons apply directly, and a practical framework you can use this quarter to bring AI costs under control without slowing your engineering teams down.

Why Today's AI Spend Looks Identical to Early Cloud Bills
Ten years ago, most finance teams treated cloud as a single line item. Engineering ran the show. Spend grew quietly until it didn't, and then everyone scrambled.

AI is repeating this exact pattern, just on a faster clock.

A few drivers explain why:

AI tools moved from fixed-seat pricing to usage-based pricing in under two years
Power-user skew is severe; a small percentage of developers often drive most of the consumption
Multiple models with very different price points create silent cost differences
Agentic workflows accumulate token cost in ways that linger long after a session ends
No engineer is incentivized to think about cost while they code
According to research from McKinsey's State of AI series, generative AI adoption inside companies more than doubled in a single year. That kind of growth curve mirrors the early AWS era, when teams discovered that elastic also meant expensive at scale.

The takeaway is simple. AI spend is not a new beast. It is the next chapter of cloud cost management, and the playbook that worked for EC2 and S3 already works for tokens and prompts.

The Visibility Gap That Is Quietly Costing Companies Millions
Walk into most engineering organizations and ask a simple question.

"Which team spent the most on AI last week?"

Silence usually follows. Or someone pulls up a vendor dashboard that shows total seats and a token total, but nothing useful below the surface.

This is the same gap cloud teams had a decade ago. The bill arrives. The total goes up. Nobody knows exactly why.

Common blind spots include:

Which developers are responsible for the largest spend
How spend splits across input tokens, output tokens, and cached tokens
Which models drive the cost across Claude, GPT, Gemini, and open-source families
Whether spend is mostly autocomplete or mostly long-running agent sessions
How spend correlates with actual engineering output
Gartner has been warning for years about shadow IT growing inside organizations. The new version is shadow AI. Developers find a tool, use it, expense it, and finance discovers it three quarters later when the consolidated invoice arrives.

The fix is not new technology. The fix is visibility, the same kind cloud cost programs built years ago

Five FinOps Principles That Apply Directly to AI
The FinOps Foundation spent years codifying what good cloud cost management looks like. Most of it transfers cleanly to AI.

Here are five principles worth lifting straight off the shelf:

Visibility comes before control. You cannot manage what you cannot see. Get the data first.
Allocate spend to teams, projects, and outcomes. Top-line totals are useless; team-level breakdowns are actionable.
Measure unit economics, not raw spend. Dollars per PR, dollars per ticket, dollars per deploy.
Detect anomalies early. Use a daily or weekly cadence, not monthly.
Use informed guardrails, not hard caps. Educate engineers; do not lock them out.
The pattern that emerges here is not technological. It is cultural. Finance and engineering have to share the same numbers.

Tagging and Allocation for AI: Treat Tokens Like EC2 Hours
In cloud cost work, tagging is the foundation. Without it, allocation is impossible.

AI spend actually has better attribution data than most cloud services. Every API request typically includes:

The model used
Input and output token counts
Latency and request metadata
Optional custom metadata fields
Caller identity, when API keys are scoped correctly
The raw signal is rich. The challenge is converting it into something a non-technical stakeholder can actually use.

A simple mapping looks like this:

Raw AI Data Business Translation
2.3M Opus input tokens, dev_id 472 Payments squad refactor, week 14
800K cached tokens on agent runs Docs team migration, ongoing
1.1M output tokens, GPT family Support ticket triage automation
50K tokens, Haiku model Inline autocomplete, all engineering
hat kind of breakdown turns a single invoice into a story finance can understand, and a budget engineering can own.

If your organization already has cost allocation workflows for cloud, you do not need to start from zero. Add AI as another provider with another set of dimensions, and feed it into the same reports.

If you are still building cloud allocation muscle, the opslyft blog covers tagging strategies that translate naturally to AI spend management.

Unit Economics: The Metric That Actually Matters
Raw spend numbers do not tell you whether your AI investment is working. Unit economics do.

Consider two teams.

Team A spends $4,000 per month on AI tools and ships 80 PRs.
Team B spends $4,000 per month on AI tools and ships 35 PRs.
Same spend. Very different efficiency. Without unit economics, the dashboards look identical.

The metrics that matter most include:

Cost per PR merged. How much does it cost in AI tokens to ship a unit of code?
Cost per ticket closed. How much does it cost to resolve a unit of planned work?
Cost per deploy. Measured across the full pipeline from prompt to production.
AI cost per developer per sprint. Is utilization rising as the team learns?
Cost per AI-assisted feature. End-to-end, including review and rework.
Computing these requires connecting two data sources. The cost side comes from your AI providers (Anthropic, OpenAI, Cursor, GitHub Copilot, and so on). The output side comes from GitHub, GitLab, Linear, Jira, or your CI pipeline.

When you put them together, conversations change. Instead of asking why AI costs are going up, the question becomes whether each dollar is producing more output than it did last quarter.

That is a question finance and engineering can actually answer together.

Detecting Anomalies Before They Become Invoices
Usage-based spend produces surprises. Cloud taught us this. AI is no different.

Common AI cost spikes include:

A developer leaves an agentic session running overnight with a runaway retry loop
A team switches from a lower-tier to a higher-tier model and the cost jumps 10x without anyone noticing
A long-running agent accumulates context until each turn costs five times the first
An automated workflow hits an edge case and retries hundreds of times
A new feature ships with verbose prompts and silently triples cost per request
Most of these are invisible until the monthly invoice arrives. By then the damage is done.

Anomaly detection works the same way it does in cloud. Set baselines, monitor daily or weekly, flag deviations, and surface them to the right team owner. The detection logic is identical. Only the patterns differ.

A few quick wins to set up immediately:

Daily per-developer spend baseline with a 2x threshold
Per-team weekly trend with month-over-month comparison
Model mix alert that notifies when premium model usage exceeds a percentage
Session-length alert that flags when a single agentic session exceeds a token threshold
None of this requires fancy machine learning. Simple thresholds catch the vast majority of cost surprises

Why Hard Caps Fail, and What to Use Instead
One of the harder lessons in cloud cost work was that blunt controls backfire.

Restrict instance types and engineers spin up larger instances less often, often using more compute than the cap was meant to save. Cap spend at a hard limit and entire projects stall on the last week of the month.

The same applies to AI.

If you cut off a developer's access to a high-quality model, they will fall back to a cheaper one, take longer to ship, and burn more total tokens in the process. The productivity gain that justified the tool evaporates.

Better alternatives include:

Soft budgets with alerts. "You are at 80% of your typical monthly spend with two weeks left" is useful. A shutoff is not.
Task-aware model guidance. Heavy reasoning warrants a premium model. Inline autocomplete does not. Make this explicit.
Real-time session cost visibility. Show developers what a session is costing as it runs.
Default to cheaper models with easy escalation. Use the cheapest model that meets the task, with a clear path to upgrade when needed.
Education over restriction. A short internal guide on model selection beats any cap.
The pattern here is the same one that worked in cloud. Trust engineers, give them the data, and let them make informed decisions.

Three Real Scenarios Where Companies Burn Money on AI
A few patterns come up over and over in conversations with engineering and finance leaders.

  1. The Forgotten Agent A developer kicks off an agent on Friday afternoon to refactor a service. They go home. The agent hits a flaky test, retries, escalates context, retries again, and runs all weekend. Monday morning brings a single-developer spend equal to the rest of the team for the month.

The fix: a session-length alert and a per-session budget cap, not a per-developer cap.

  1. The Silent Model Upgrade A team's tooling defaults change after a vendor update. What used to call the cheaper model now calls the premium model. Output quality goes up. Nobody notices the cost has gone up 8x until the invoice arrives.

The fix: model mix monitoring with a week-over-week trend alert.

  1. The Context-Bloat Session An agent works on a large codebase. Each turn appends more context. By turn 40, a single message costs more than the entire first hour of the session. Productivity feels normal. Cost is exponential.

The fix: real-time per-session cost surfacing, plus guidance on when to reset context.

These are not edge cases. They are the new normal. Every team running AI tools at scale will hit some version of each within their first year.

How opslyft Helps Businesses Manage AI and Cloud Costs Together
Most companies trying to manage AI spend today face a familiar problem. The data sits in many places. Cursor has one dashboard. Anthropic has another. OpenAI has another. AWS has fifty. None of them talk to each other.

opslyft brings these data sources into a single view, applies cost allocation, and connects spend to engineering output. The platform was built for cloud cost management and extends naturally to AI tools, treating AI as another provider in a unified FinOps program.

Specific capabilities include:

Multi-source integration across cloud providers, AI tools, and developer platforms
Cost allocation by team, project, environment, and developer
Unit economics dashboards linking spend to PRs, tickets, and deploys
Anomaly detection with daily and weekly cadence
Soft budgets and informed guardrails that protect productivity
Optimization recommendations with measurable savings impact
Security-first deployment with read-only access patterns and SOC 2 controls
The principle is the same one that worked for cloud. Visibility first, then allocation, then unit economics, then targeted action. AI is just the next provider on the list.

Conclusion
AI spending is not a new problem. It is the next chapter of the same cloud cost story finance and engineering teams have been working through for a decade.

The companies that treat AI as just another provider inside their FinOps program will move faster, spend smarter, and avoid the budget shocks that catch everyone else by surprise.