惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

N
News and Events Feed by Topic
WordPress大学
WordPress大学
Vercel News
Vercel News
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
小众软件
小众软件
L
LangChain Blog
雷峰网
雷峰网
D
DataBreaches.Net
博客园 - 三生石上(FineUI控件)
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
Tor Project blog
NISL@THU
NISL@THU
Scott Helme
Scott Helme
量子位
S
Security Affairs
T
Threat Research - Cisco Blogs
博客园_首页
云风的 BLOG
云风的 BLOG
D
Docker
AWS News Blog
AWS News Blog
腾讯CDC
博客园 - 聂微东
The GitHub Blog
The GitHub Blog
U
Unit 42
Recent Announcements
Recent Announcements
Apple Machine Learning Research
Apple Machine Learning Research
G
Google Developers Blog
T
The Exploit Database - CXSecurity.com
MongoDB | Blog
MongoDB | Blog
Stack Overflow Blog
Stack Overflow Blog
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
L
LINUX DO - 热门话题
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
The Last Watchdog
The Last Watchdog
C
Cybersecurity and Infrastructure Security Agency CISA
IT之家
IT之家
W
WeLiveSecurity
P
Privacy & Cybersecurity Law Blog
F
Full Disclosure
L
Lohrmann on Cybersecurity
The Hacker News
The Hacker News
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Y
Y Combinator Blog
S
Security @ Cisco Blogs
C
Cyber Attacks, Cyber Crime and Cyber Security
C
Check Point Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
N
News and Events Feed by Topic
PCI Perspectives
PCI Perspectives
I
InfoQ

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
LLM Cost Optimization: Cut AI Inference Costs 47–80% Without Sacrificing Quality
Dishant Sethi · 2026-06-01 · via DEV Community

Key Takeaways

  • LLM API spending doubled from $3.5B to $8.4B in 2025 — most of the growth is from production deployments, not experiments
  • Semantic caching + model routing alone cut spend 47–80% without any change to model quality or user experience
  • Eight techniques ranked by cost impact and implementation complexity — sequence them starting with the fastest wins
  • Prompt caching, batch inference, and output length control are each deployable in under a week with minimal architectural change

LLM cost optimization in production reduces per-inference spend without degrading output quality. The three highest-impact techniques — model routing, semantic caching, and prompt prefix caching — deliver 40–90% savings on the token categories they address. Applied together on a typical enterprise workload, they produce the 47–80% total cost reduction most teams are targeting.


Why LLM API Bills Are Out of Control

Global LLM API spending doubled from $3.5B to $8.4B in 2025, driven by enterprises moving from proof-of-concept to production at scale. The cost growth is not from model improvements — it is from production architectures designed for experimentation: every request routed to the most expensive model, identical prompts recomputed on every call, and no caching layer in place.

The typical production LLM system makes several expensive mistakes simultaneously: it routes every request to the most capable (and most expensive) model regardless of task complexity, it recomputes identical prompt prefixes on every call, and it generates responses from scratch even when a semantically equivalent query was answered thirty seconds ago.

The result is a bill that scales roughly quadratically with request volume. Every engineering team eventually hits an inflection point where the cost per user is incompatible with unit economics, and they have to go back and redesign the inference layer they should have built correctly the first time.

This guide covers eight techniques that production teams use to reduce LLM costs without degrading output quality. The techniques compound: applying all eight to a typical enterprise workload produces the 47–80% savings figure, and the first two techniques — semantic caching and model routing — often deliver half that savings on their own.


Technique 1: Model Routing

Model routing directs each incoming request to the cheapest model capable of handling it reliably, reducing per-request spend by 40–70%. GPT-4o costs $5–15 per million input tokens; Claude 3 Haiku costs $0.25 per million. For 60–80% of production requests — classification, extraction, short-form generation — the cheap model produces indistinguishable output.

The implementation pattern is a router that wraps your existing LLM client. Each request is scored before dispatch. The score is cached so repeat queries skip the classification overhead entirely. When a small-model response fails a downstream quality check, the router escalates and logs the feature vector so the classifier can learn from the failure.

Production teams running this pattern report 40–70% reductions in spend with no measurable degradation in user-facing quality metrics, because the tasks that land on the cheap model were never hard enough to require the expensive one.


Technique 2: Prompt Caching (Anthropic Prefix Caching)

Prompt caching eliminates the cost of recomputing stable prompt prefixes on every request. On Anthropic's API, cached token reads cost 90% less than uncached — $0.03 per million for Claude 3 Haiku versus $0.30 for a cache miss. Cache writes cost 25% more than standard input tokens, so the breakeven is any prefix used twice within a 5-minute TTL window.

The prompt caching benefits are largest in workloads with long, stable system prompts — RAG systems with large retrieved context blocks, coding assistants with large repo context windows, customer support agents with extensive policy documentation. Any prefix that appears in more than two requests per cache TTL window (5 minutes on Anthropic's current implementation) is a candidate for caching.

Enabling prefix caching on Anthropic's API requires placing a cache_control breakpoint at the end of the prefix you want cached. The breakpoint tells the API where the stable prefix ends and the dynamic user content begins. You can place up to four breakpoints per request, allowing fine-grained control over what gets cached.

Teams with long system prompts (4,000+ tokens) that reuse across sessions report 60–90% reductions in input token costs after enabling caching, with no change to output quality because the model never sees the cache boundary — only the billing layer does.


Technique 3: Semantic Caching

Semantic caching intercepts LLM calls by embedding each query, searching a vector index for near-identical past queries, and returning a cached response when cosine similarity exceeds a threshold. Research on enterprise LLM workloads finds roughly 31% of queries are semantically equivalent to one answered in the past 24 hours — exact-string caching misses almost all of them.

The implementation requires three components: an embedding model (OpenAI text-embedding-3-small at $0.02 per million tokens, or a self-hosted model at near-zero marginal cost), a vector store (Redis with vector search, Pinecone, or Qdrant), and a similarity threshold tuned to your tolerance for stale or slightly mismatched responses.

The threshold is the key operational decision. A threshold of 0.95 is conservative — it only serves cached responses for near-identical queries and misses many reusable answers. A threshold of 0.85 captures more cache hits but occasionally serves a response that is subtly wrong for the reworded query. Most production teams run 0.90 with a human feedback loop that flags responses where the user immediately refines their question — a signal that the cache hit was low quality.

Semantic caching compounds well with model routing: a routing decision that should go to a cheap model often hits the semantic cache first and costs nothing at all.


Technique 4: Quantization

Quantization reduces model weight precision from FP16 to INT8 or INT4, cutting VRAM requirements and increasing GPU throughput on self-hosted inference. A 70B parameter model in FP16 requires roughly 140GB of VRAM; the same model quantized to INT4 fits in 35GB on a single A100 80GB, enabling more concurrent requests per GPU at lower cost per token.

The tradeoff is accuracy degradation. A February 2025 Amazon study found INT4 quantization caused a 39.46% accuracy drop on Llama-3.3 70B on certain benchmarks. INT8 quantization is safer — typical accuracy degradation is 0.5–2% on general benchmarks, which is often acceptable for production tasks. GPTQ and AWQ are the two dominant quantization schemes for LLMs; both are well-supported by vLLM and HuggingFace Transformers.

The cost reduction applies only to self-hosted inference. If you are calling a managed API (OpenAI, Anthropic), quantization is already applied by the provider and you cannot further tune it. Quantization is the right technique for teams that have moved workloads on-premises or to a dedicated GPU cluster.


Technique 5: Batch Inference

Batch inference processes asynchronous LLM requests at 50% of real-time API pricing via OpenAI's Batch API or Anthropic's Message Batches API. The tradeoff is a 24-hour completion window, making it suitable for offline workloads only: nightly document classification, bulk content generation, dataset enrichment, evaluation suite runs, and any pipeline where the requester does not need a response before the next step.

Identify your offline LLM workloads and route them to the batch endpoint. Many engineering teams are running expensive real-time API calls for workloads that could be batched — simply because the batch endpoint was added after the original integration was built.


Technique 6: Context Compression

Context compression reduces input token count by removing redundant or low-relevance content from the context window before it reaches the main model. In RAG systems, retrieved context is typically the largest cost driver — and reranker models like Cohere Rerank or BGE-Reranker reduce context size by 50–70%, producing net savings whenever their per-request cost is less than the tokens they eliminate.

The primary approaches are: (1) extractive compression — running a smaller model or BM25 reranker to select only the most relevant passages from retrieved documents; (2) abstractive compression — using a small model to summarize retrieved passages before passing them to the main model; and (3) conversation summarization — replacing long multi-turn conversation histories with a running summary.

The math works in your favor whenever the reranker costs less than the tokens it eliminates.


Technique 7: Output Length Control

Output tokens cost 3–4× more than input tokens on most API pricing schedules — on Claude 3.5 Sonnet, output tokens cost $15 per million versus $3 for inputs. Explicit length instructions, structured output formats like JSON mode, and stop sequences reduce output token counts by 15–30% on tasks where verbosity adds no information value.

First, explicit length instructions in the system prompt ("respond in 2–3 sentences", "use bullet points, not paragraphs") reliably reduce output tokens for tasks where brevity is acceptable. Second, structured output formats (JSON mode, function calling) eliminate the model's tendency to wrap answers in prose scaffolding. A response that would have been 400 tokens in natural language is often 150 tokens as a JSON object. Third, stop sequences terminate generation early once the required information has been produced.

Output length control pairs well with model routing: the large model that produces unnecessarily verbose output for a simple task is both more expensive per token and produces more tokens than needed.


Technique 8: OSS Models for Narrow Tasks

Open-source models fine-tuned for narrow, well-defined tasks cost 70–95% less than frontier API calls and typically outperform them on those specific workloads. A fine-tuned Llama-3 8B on a single A10G GPU handles approximately 500 requests per minute at roughly $0.0002 per request — versus $0.005–0.015 for GPT-4o on equivalent input, a 25–75× cost difference.

The economics work for tasks like document classification, sentiment analysis, named entity recognition, and translation into common language pairs. The investment is fine-tuning effort and inference infrastructure management. The payoff is only positive when the task is sufficiently narrow and high-volume. Teams that deploy OSS models for broad, general-purpose tasks — where the frontier model's generalization is actually needed — typically see quality degradation that erodes the cost savings through rework and escalation.


LLM Cost Optimization: Technique Summary and Sequencing

Eight techniques ranked by estimated savings and implementation complexity. Start with the low-complexity wins — prompt caching, batch inference, and output length control deliver significant savings with minimal architectural change. Model routing and semantic caching require more infrastructure but produce the largest absolute cost reductions for most production workloads. Prodinit applies this sequence across every AI infrastructure engagement: instrumentation first, then caching, then routing, then model substitution.

Technique Estimated Savings Implementation Complexity
Model Routing 40–70% on routed requests Medium
Prompt Caching 60–90% on cached tokens Low
Semantic Caching 31–47% reduction in LLM calls Medium
Quantization 30–60% on self-hosted compute High
Batch Inference 50% on batchable workloads Low
Context Compression 20–40% on input tokens Medium
Output Length Control 15–30% on output tokens Low
OSS Models for Narrow Tasks 70–95% on targeted workloads High

Apply low-complexity techniques first: prompt caching, batch inference, and output length control can be enabled in a week with minimal architectural change. Quantization and OSS model deployment make sense after the low-hanging fruit is captured and you have strong monitoring in place to catch quality degradation.