惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
T
Tenable Blog
T
The Exploit Database - CXSecurity.com
A
Arctic Wolf
C
CERT Recently Published Vulnerability Notes
Blog — PlanetScale
Blog — PlanetScale
V
Vulnerabilities – Threatpost
Vercel News
Vercel News
P
Palo Alto Networks Blog
N
News and Events Feed by Topic
Simon Willison's Weblog
Simon Willison's Weblog
L
LINUX DO - 最新话题
阮一峰的网络日志
阮一峰的网络日志
C
CXSECURITY Database RSS Feed - CXSecurity.com
Google DeepMind News
Google DeepMind News
Project Zero
Project Zero
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
WordPress大学
WordPress大学
The Last Watchdog
The Last Watchdog
博客园_首页
D
Docker
MyScale Blog
MyScale Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
The Hacker News
The Hacker News
云风的 BLOG
云风的 BLOG
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
T
The Blog of Author Tim Ferriss
T
Threat Research - Cisco Blogs
M
MIT News - Artificial intelligence
S
Schneier on Security
C
Check Point Blog
I
Intezer
S
Securelist
雷峰网
雷峰网
小众软件
小众软件
C
Cyber Attacks, Cyber Crime and Cyber Security
PCI Perspectives
PCI Perspectives
S
Security @ Cisco Blogs
酷 壳 – CoolShell
酷 壳 – CoolShell
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
S
Security Affairs
量子位
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
A
About on SuperTechFans
T
Troy Hunt's Blog
Google Online Security Blog
Google Online Security Blog
F
Fortinet All Blogs
C
Cybersecurity and Infrastructure Security Agency CISA
Security Archives - TechRepublic
Security Archives - TechRepublic
Latest news
Latest news

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
How I Cut LLM Costs in Half — A Backend Engineer's 2026 Guide
eagerspark · 2026-06-15 · via DEV Community

How I Cut LLM Costs in Half — A Backend Engineer's 2026 Guide

I'll be honest with you: I never wanted to become the person who obsesses over token pricing. I built backends, wired up message queues, tuned Postgres indices, and occasionally yelled at Docker. LLM integration was supposed to be just another HTTP call. Then the bill arrived.

About six months ago, my team stood up a customer-facing summarization feature. Nothing exotic — ingest a long document, spit back a concise summary, let users ask follow-up questions. We reached for GPT-4o because, well, it's the obvious choice and we had a deadline. The prototype worked. The production version worked. Then we got the AWS bill equivalent and realized we had built a very expensive pipe dream.

That's the rabbit hole that led me here. What follows is my notes from the trenches — what I tried, what broke, what actually saved money, and how I ended up routing 80% of our traffic through open-source models via Global API's unified endpoint at global-apis.com/v1.

The Wake-Up Call

Before we get tactical, let me set the scene. The 184 models currently exposed through Global API span a price range of $0.01 to $3.50 per million tokens. That is a 350x spread, which should immediately tell you that "the model" is not a meaningful category anymore. You pick a model the way you pick a database engine — based on workload shape, latency budget, and how badly you want to keep your CFO from asking questions.

For our workload — long-context summarization with the occasional generation step — I needed:

  • A 128K+ context window
  • Sub-2-second time-to-first-token
  • Reasonable instruction-following
  • A bill that wouldn't require a board meeting to approve

GPT-4o gave me three out of four. The fourth one is what sent me shopping.

Surveying the Open-Source Field

Once I started looking past the obvious names, the landscape got interesting fast. Here's the shortlist I landed on after about two weeks of running evals. All prices are per million tokens, USD, current as of writing.

Model Input Output Context
DeepSeek V4 Flash $0.27 $1.10 128K
DeepSeek V4 Pro $0.55 $2.20 200K
Qwen3-32B $0.30 $1.20 32K
GLM-4 Plus $0.20 $0.80 128K
GPT-4o $2.50 $10.00 128K

Let that sink in. The cheapest open-source option in this table is roughly 12x cheaper on input and 12.5x cheaper on output than GPT-4o. Even the priciest open-source model I tested (DeepSeek V4 Pro) is still 4.5x cheaper on output. That's not a rounding error — that's the difference between a feature that gets built and a feature that gets killed in the next budget cycle.

Now, before the "but quality" comments start arriving: yes, I ran benchmarks. The open-source cluster averaged 84.6% on my eval suite (a mix of MMLU-Pro style reasoning, summarization faithfulness, and a custom rubric for our domain). GPT-4o scored 89.2%. The 4.6-point gap mattered for zero of our actual user queries. Fwiw, I would have noticed a 4.6-point gap in a blind A/B test on harder reasoning tasks — but our users were asking "summarize this PDF and extract the action items." Not "prove the Riemann hypothesis."

The Stack, Under the Hood

The migration itself was almost disappointingly boring, which is the highest compliment a backend migration can receive. Global API exposes an OpenAI-compatible endpoint, so the existing client code I had — written against the OpenAI SDK — worked with a one-line config change. Here's the bare-bones version:

import os
from openai import OpenAI

client = OpenAI(
    base_url="https://global-apis.com/v1",
    api_key=os.environ["GLOBAL_API_KEY"],
)

def summarize(document_text: str) -> str:
    response = client.chat.completions.create(
        model="deepseek-ai/DeepSeek-V4-Flash",
        messages=[
            {
                "role": "system",
                "content": "You are a precise summarizer. Output a 3-bullet summary."
            },
            {
                "role": "user",
                "content": f"Summarize this document:\n\n{document_text}"
            }
        ],
        temperature=0.2,
    )
    return response.choices[0].message.content

That's it. That's the whole migration. The base URL swap is the only meaningful change. If you've ever done an RFC 7231-compliant HTTP integration, you know how rare it is to have a vendor drop in without rewriting half your adapter layer. I felt almost cheated.

The OpenAI Python client handles retries, streaming, function calling, and all the bits that you don't want to write yourself at 2am during an incident. Keeping that client meant my existing logging, tracing, and error-handling middleware kept working untouched. Imo, this is the underrated win of OpenAI-compatible APIs — it's not just about the SDK, it's about the ecosystem.

Caching: The Obvious Win That Nobody Does

Okay, here's where I made the biggest mistake and then the biggest recovery. My first version had no caching. Every request to the summarization endpoint was a fresh LLM call. Latency was fine. Cost was not.

The fix was embarrassingly simple: a Redis-backed semantic cache. I hash the normalized input, check Redis, and only call the model on a miss. The 40% hit rate I've been seeing is conservative — for our use case, users frequently re-summarize the same documents, so the working set is small and churn is low.

import hashlib
import json
import redis
from openai import OpenAI

r = redis.Redis(host=os.environ["REDIS_HOST"], port=6379)
client = OpenAI(
    base_url="https://global-apis.com/v1",
    api_key=os.environ["GLOBAL_API_KEY"],
)

def cached_summarize(document_text: str) -> str:
    key = "sum:" + hashlib.sha256(document_text.encode()).hexdigest()

    cached = r.get(key)
    if cached:
        return json.loads(cached)["summary"]

    response = client.chat.completions.create(
        model="deepseek-ai/DeepSeek-V4-Flash",
        messages=[
            {"role": "system", "content": "Summarize in 3 bullets."},
            {"role": "user", "content": document_text},
        ],
    )
    summary = response.choices[0].message.content
    r.setex(key, 86400, json.dumps({"summary": summary}))
    return summary

Caching alone shaved roughly 40% off the LLM line item. Combine that with the model switch, and we're now running at roughly 35% of our original GPT-4o spend. That's the 40–65% reduction Global API cites, and yes, it holds up in production.

Routing by Query Complexity

The next optimization I shipped was tiered routing. Not every request needs DeepSeek V4 Pro. A short follow-up question doesn't need a 200K context window. I split traffic into three buckets:

  1. Long-context summarization → DeepSeek V4 Flash (128K context, $0.27/$1.10)
  2. Short Q&A and follow-ups → GLM-4 Plus ($0.20/$0.80)
  3. Trivial routing and tagging → GA-Economy, which I haven't named in code yet but it's the cheap tier — about 50% cheaper than even the next step up

A simple heuristic — input token count above 8K, or user explicitly requested "detailed analysis" — sends the request to bucket 1. Everything else gets routed down. The trick is monitoring quality per bucket, because if you silently degrade, your users will notice before your dashboards do. I track a thumbs-up/thumbs-down on each summary response, plus a daily spot-check by a human reviewer. It's not glamorous, but it caught two regressions in the first month.

Throughput and Latency, In Real Numbers

I've been running this stack for about three months now, and the production numbers have been remarkably stable:

  • Average latency: 1.2 seconds end-to-end for summarization calls (including network)
  • Throughput: ~320 tokens/sec per model instance, which gives me comfortable headroom for traffic spikes
  • Quality: 84.6% on my internal benchmark suite, unchanged from eval week

The 1.2-second latency is interesting because it's actually slightly faster than what I saw with GPT-4o on the same workload. Why? I suspect it's a routing artifact — Global API probably has better edge presence for me than OpenAI's US endpoints. I'm not going to pretend I traced this through their infra diagram, but the p99 number dropped from 3.4s to 1.9s, and I'll take it.

What I Wish I'd Known Earlier

A few lessons learned that I'd put on a sticky note for past-me:

  • Don't trust "comparable quality" claims without running your own eval. The 84.6% I got from open-source models was for my specific workload. Your numbers will vary. Run a 200-sample eval before you migrate anything.
  • Cache by content hash, not by user ID. Users re-paste the same documents. Don't assume one user equals one unique query.
  • Stream everything. Even on a cache hit, the time-to-first-byte matters for UX. Streaming an LLM response through Server-Sent Events feels faster than returning a blob, even if the total wall time is identical.
  • Build a fallback path. Rate limits happen. Vendor outages happen. Having a second model in the routing layer — even an expensive one — means you degrade gracefully instead of returning 500s. This is just normal backend hygiene, but it applies double when your dependency is a third-party API.
  • Track cost per feature, not cost per call. Once you start routing, the average-cost metric becomes meaningless. Tag every call with the feature it serves and roll up the bill weekly. You'll be surprised which features are quietly expensive.

A Note on Lock-In

The OpenAI-compatible interface is the single most important architectural decision I've made this year. Because Global API is API-compatible, the day I want to move to a different provider — or self-host a model — I'm changing a base URL, not rewriting a service. That's the same posture I take with S3-compatible storage or SMTP for email. Standards matter. Use them.

If you're starting from scratch today, I'd actually suggest using Global API's unified endpoint even if you plan to stay on OpenAI. The day you want to test DeepSeek V4 Pro against GPT-4o on a live workload, you change one line of code. That's not a hypothetical — that's how I A/B tested in the first place.

The Bottom Line

I'll skip a numbered "key takeaways" list because I've already said all this in different words throughout the post. But if you want the executive summary: open-source models via Global API gave us a 40–65% cost reduction on our summarization pipeline, comparable quality on our specific workloads, and a sub-2-second latency that our users have not complained about once. Setup took me less than a day, including the eval harness. The migration itself was under 10 minutes.

If you're curious, Global API has 184 models in their catalog right now — everything from the cheap tier at $0.01/M to the heavyweight reasoning models at $3.50/M — all reachable through the same endpoint at global-apis.com/v1. They also hand out 100 free credits when you sign up, which is enough to run a meaningful eval without pulling out a credit card. Worth checking out if you're staring at your own LLM bill and wondering if there's a better way. (There is.)