惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
C
CXSECURITY Database RSS Feed - CXSecurity.com
S
Schneier on Security
博客园 - 叶小钗
博客园 - 三生石上(FineUI控件)
C
Cybersecurity and Infrastructure Security Agency CISA
Engineering at Meta
Engineering at Meta
Google DeepMind News
Google DeepMind News
酷 壳 – CoolShell
酷 壳 – CoolShell
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园_首页
T
Threat Research - Cisco Blogs
C
Cisco Blogs
Recent Announcements
Recent Announcements
S
Securelist
N
Netflix TechBlog - Medium
The Register - Security
The Register - Security
P
Privacy & Cybersecurity Law Blog
宝玉的分享
宝玉的分享
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
LINUX DO - 热门话题
T
Tor Project blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
月光博客
月光博客
AWS News Blog
AWS News Blog
P
Proofpoint News Feed
博客园 - 司徒正美
L
LINUX DO - 最新话题
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
H
Help Net Security
Spread Privacy
Spread Privacy
PCI Perspectives
PCI Perspectives
Project Zero
Project Zero
I
Intezer
T
The Blog of Author Tim Ferriss
有赞技术团队
有赞技术团队
The Last Watchdog
The Last Watchdog
C
Check Point Blog
Blog — PlanetScale
Blog — PlanetScale
B
Blog RSS Feed
MyScale Blog
MyScale Blog
V
Vulnerabilities – Threatpost
Recorded Future
Recorded Future
T
Tenable Blog
Jina AI
Jina AI
D
DataBreaches.Net
阮一峰的网络日志
阮一峰的网络日志

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
10 Ways To Reduce Your LLM API Costs
Bruno Pérez · 2026-05-21 · via DEV Community

So you made it, you have your AI app on prod, you are onboarding your users and they like what you've done, cheers! Now comes the hard-to-swallow part, the AI bill.

Serving those users consumes AI inference and it's literally eating all your margins. Let's see 10 ways to reduce that AI bill of yours.

Table of contents

  1. Choose a well-fitted AI model
  2. Use your Pro subscriptions
  3. Reduce output tokens to cut your LLM bill
  4. Use prompt caching when you can
  5. Use Batch API for nightly workflows
  6. Be FLEX-ible and accept slow tiers
  7. Don't use AI
  8. Use free models and free tiers
  9. Don't miss those Cloud provider credits
  10. Observe your AI costs and take back control

1. Choose a well-fitted AI model

Choosing the perfectly fitted model is not that easy. While realizing that your model is not "smart enough" for your requirements is the easy part, the other side is trickier: you may use an overkill model and you are overspending without noticing it.

Potential reductions are huge here, for example switching from GPT-5.5 to GPT-5.4 reduces costs by 50%. Use GPT-5.4 Mini instead? That's 85% cheaper.

However nothing is magical here: the cheaper models produce lower-quality outputs. Spending time to benchmark models on your particular use cases is the key to evaluate the impact of the quality loss. The closer you are to real production data, the better. Here are some ideas to put you on the right track:

  • Downshift within your model's family: flagship -> mini -> nano
  • Look for other providers' models: For example Chinese providers tend to be cheaper for an equivalent quality
  • Go for the previous generation: Opus 4.7 -> Opus 4.6 -> Opus 4.5

In some cases, the query is variable and you can't anticipate it. Using LLM routers that evaluate your query on-the-fly can be a good choice.

2. Use your Pro subscriptions

Are you paying for a ChatGPT, MiniMax or other pro subscription? Make the most of it and plug it into your app. With Manifest, you can plug several pro subscriptions into your apps, while keeping your code and favorite SDK as it is. As of today Manifest allows you to plug Anthropic, GitHub Copilot, MiniMax, Ollama Cloud, OpenAI, OpenCode Go and Z.ai pro subscription directly into your AI app. Make sure it's ok with the terms of service of your providers.

The Connect providers modal in Manifest — toggle Anthropic, GitHub Copilot, MiniMax, Ollama Cloud, OpenAI and other subscriptions on for routing

The thing to watch here is the rate limit. Those subscriptions are often way cheaper than their API key alternative, but they are limited in usage. You need to set up a fallback model in case you hit the API rate limit. See how MyTrainer did exactly this in production.

3. Reduce output tokens to cut your LLM bill

Did you notice that you don't actually pay directly for the inference of the model you call? You pay for what goes in and for what comes out of it, not the "amount of thinking" itself. What if we can tweak the system to produce the same value with less?

The output tokens cost on average 5 times more than the input tokens. It's totally worth it to spend more on the input tokens to reduce the output ones.

The easiest way is to simply tell the LLM explicitly to "be concise" and even specify the format that you are expecting. Ask for structured JSON or CSV instead of prose. Prose tends to be long and adds verbosity to the outputs. If you need prose, Caveman (60k+ GH stars) reduces output tokens by 75% by… speaking like a caveman. Caveman tool real. Make agent talk short. 🪨

By the way most providers allow a "max_tokens" that truncate outputs. While this doesn't make the model concise — it just truncates — it can still prevent long outputs.

4. Use prompt caching when you can

Caching has been used for decades to reduce compute and thus latency. For example when a website shows you a list of items, like the latest news for example, it will load from the database only for the first user, and then store it somewhere. Next users will receive that list from the cache directly, and have it served almost instantaneously.

It works the same with LLM prompts: Models do heavy work on every token you send to them. If part of it doesn't change between requests, they do less work and you pay less. Generally cached input costs between 50% and 90% less.

The #1 rule here is that your static content (the one that doesn't change) should come before the dynamic content (the volatile one). The first changed character and you're paying full price for following tokens. Structure your call correctly so that system prompts and knowledge bases come first:

messages = [
    {"role": "system", "content": SYSTEM_PROMPT + KNOWLEDGE_BASE},  # Static
    {"role": "user", "content": f"{user_question}"},  # Dynamic
]

resp = client.chat.completions.create(
    model="gpt-5.4-mini",
    messages=messages,
)

Enter fullscreen mode Exit fullscreen mode

5. Use Batch API for nightly workflows

Batch API is a straight 50% discount on inference if you accept to receive the response within 24 hours.

You're not going to implement that on a chat or on any actions that need real-time answers, however any background process is a good candidate. It requires a bit of change in your code but it's totally worth it if you have a significant amount of inference in scheduled tasks, nightly workflows or routines.

Here is the link to the providers' docs that support batch API:

6. Be FLEX-ible and accept slow tiers

Some providers even have a "flex" tier: a synchronous tier explicitly slower than the standard. Unlike Batch API, here you do get the real-time response, but it is significantly slower. However the discount is usually the same: 50% off.

Now it's up to you to decide if you can afford that extra latency on some calls. Here is a list of providers that offer that trade-off:

7. Don't use AI

This one can sound a bit provocative but the cheapest inference will always be the one you don't use. Are those AI features genuinely needed? Or are they just decorative nice-to-haves?

Limiting LLM calls and using algorithms (aka good old software) when possible is the biggest win on this list. Of course it's easier said than done, but many operations can be done programmatically: validation, regex, rules, heuristics…

Pro tip: go hybrid. Sometimes algorithms cannot handle all cases, but work well on some of them. Use conditions in your code to use them when possible, and fall back on LLMs.

8. Use free models and free tiers

If you are just hacking around, maybe you don't need to pay at all! Some providers offer inference for free on some models. Of course the rate limit is quite low but if your app has a low AI inference consumption, it may work. If it doesn't, you can still use model fallback.

Awesome Free LLM APIs — a curated list of LLM providers with a permanent free tier

And you know what? We prepared a very cool list of Awesome Free LLM APIs just for you 😎

9. Don't miss those Cloud provider credits

Building a startup? All major cloud providers can give you up to $300k in cloud credits, most of them can be used in inference. The terms vary from one provider to another, and you probably need to fill applications and even meet them.

The concept here is that the faster you can burn those tokens, the more they'll offer you. It makes sense as they are looking for new potential big customers that will stick with their models.

You'll probably never get the $300k on the first shot, it goes incrementally. One important thing: incubators, accelerators and startup programs tend to have agreements on credit packages with providers. Reach out to your program manager if you're into one of those!

10. Observe your AI costs and take back control

Last but not least, you can't optimize what you can't see. Maybe one single recurrent LLM call is burning all your budget and you haven't identified it yet.

30-day messages and usage chart from the Manifest dashboard — track your AI inference volume and cost over time

There are many tools (open source or proprietary) like Manifest that let you visualize your costs, analyse them, and even set some budget limits. Some of them are simple and easy to understand, others more complicated and let you go into details. Up to you to find your fit!

Conclusion

Reducing cloud AI inference costs is key for an app's profitability. Consider the few lines of each point in this post as a starting point, an idea to dig into. Not all of them are necessarily applicable in your case, but they are worth knowing and understanding.

At Manifest, we think that AI is an incredible technology, and that it deserves to be affordable. Our platform is open source and gives total control to our users. Check out our website and give us a star on GitHub to support us! Or simply share this post, as it helps others pay less for AI.

Happy hacking!