惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Schneier on Security
Security Archives - TechRepublic
Security Archives - TechRepublic
T
Threat Research - Cisco Blogs
G
GRAHAM CLULEY
P
Privacy & Cybersecurity Law Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
Cisco Talos Blog
Cisco Talos Blog
The Hacker News
The Hacker News
L
Lohrmann on Cybersecurity
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
Security Latest
Security Latest
Know Your Adversary
Know Your Adversary
P
Palo Alto Networks Blog
C
Cisco Blogs
AWS News Blog
AWS News Blog
T
Threatpost
L
LINUX DO - 热门话题
Simon Willison's Weblog
Simon Willison's Weblog
Scott Helme
Scott Helme
C
Cybersecurity and Infrastructure Security Agency CISA
T
Tor Project blog
Cyberwarzone
Cyberwarzone
P
Proofpoint News Feed
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
T
The Exploit Database - CXSecurity.com
The Register - Security
The Register - Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
罗磊的独立博客
云风的 BLOG
云风的 BLOG
V
Vulnerabilities – Threatpost
N
News | PayPal Newsroom
Project Zero
Project Zero
NISL@THU
NISL@THU
博客园_首页
MyScale Blog
MyScale Blog
V2EX - 技术
V2EX - 技术
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
F
Full Disclosure
T
Troy Hunt's Blog
Recorded Future
Recorded Future
N
Netflix TechBlog - Medium
P
Privacy International News Feed
H
Hackread – Cybersecurity News, Data Breaches, AI and More
A
Arctic Wolf
C
Check Point Blog
W
WeLiveSecurity
Apple Machine Learning Research
Apple Machine Learning Research
C
CERT Recently Published Vulnerability Notes

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
How I Cut Our Recommendation Engine Bill 60% Without Losing Quality
fiercedash · 2026-06-14 · via DEV Community

How I Cut Our Recommendation Engine Bill 60% Without Losing Quality

I still remember the Slack thread where our finance lead pinged me at 11pm on a Thursday. Our monthly AI spend had crossed six figures, and the recommendation engine alone was responsible for nearly 40% of that. I'm a cloud architect, not a magician, but the next morning I started digging into whether we really needed what we were paying for. What I found over the following weeks changed how I approach AI infrastructure entirely, and I want to walk you through the lessons because if you're running recommendation workloads at scale, you're probably leaving a lot of money on the table.

The uncomfortable truth about recommendation systems in 2026 is that the generic solutions everyone reaches for first are wildly overpriced for what they actually do. When I audited our stack, I realized we were using a top-tier model to do classification, basic similarity scoring, and content matching — tasks that don't require the cognitive horsepower of something like GPT-4o. We were paying $10.00 per million output tokens for work that a $0.80 model could handle with comparable quality. That's a 12.5x cost multiplier on workloads that process millions of requests daily. No wonder the bill was scary.

The shift I made was moving from a "one model for everything" architecture to a tiered routing system. I split our recommendation pipeline into three lanes: cheap-and-fast for the bulk of straightforward requests, mid-tier for the nuanced cases, and premium only when we genuinely needed it. This is the same pattern I use for any high-throughput system — you don't run a Ferrari to deliver pizza. You run a Honda Civic. The interesting part was discovering that the Honda Civic in this metaphor, running on models like DeepSeek V4 Flash or GLM-4 Plus, could genuinely keep up with the Ferrari on most tasks.

Let me give you the raw numbers because I know that's what you're here for. Through Global API's unified interface, we now have access to 184 AI models, with prices ranging from $0.01 to $3.50 per million tokens. That ceiling-to-floor spread is what made the tiered approach possible. Here's the pricing table that lives in our team's Notion and gets referenced every time someone proposes a new model integration:

Model Input ($/M) Output ($/M) Context Window
DeepSeek V4 Flash 0.27 1.10 128K
DeepSeek V4 Pro 0.55 2.20 200K
Qwen3-32B 0.30 1.20 32K
GLM-4 Plus 0.20 0.80 128K
GPT-4o 2.50 10.00 128K

Look at the spread between GLM-4 Plus and GPT-4o on output tokens: $0.80 versus $10.00. That's not a marginal difference, that's an order of magnitude. When you're processing five million recommendation requests a day and each one generates 500 output tokens, the math becomes obvious very quickly. We went from spending roughly $25,000 a day on output tokens to around $9,000, and the recommendation quality actually went up in some segments because we could afford to run the models at higher temperature settings and explore more candidate items per request.

Now, before you think I'm just chasing the cheapest option, let me talk about the quality story. Across our benchmark suite, the tiered approach delivered a 40-65% cost reduction compared to our previous single-model setup, with quality scores that were either equivalent or slightly better. The blended benchmark average for our new pipeline hit 84.6%, and our internal A/B tests showed no statistically significant degradation in user engagement metrics. If anything, click-through rates on recommendations ticked up by 1.3% because we could generate more diverse suggestions per session.

The throughput numbers were equally important for me. Our recommendation engine now sustains 320 tokens per second on the cheap lane, with an average response time of 1.2 seconds end-to-end. For a system that serves real-time suggestions to millions of users, that's the kind of latency that keeps your p99 dashboards green. And this is where my cloud architect brain really kicked in — once I had the cost and quality story sorted, I needed to make sure the deployment story was bulletproof.

Let me show you the actual client setup I use. It's embarrassingly simple, which is the point:

import openai
import os
from typing import Optional

class TieredRecommendationClient:
    def __init__(self):
        self.client = openai.OpenAI(
            base_url="https://global-apis.com/v1",
            api_key=os.environ["GLOBAL_API_KEY"],
        )
        self.routes = {
            "fast": "deepseek-ai/DeepSeek-V4-Flash",
            "balanced": "Qwen3-32B",
            "premium": "openai/gpt-4o",
        }

    def recommend(self, prompt: str, tier: str = "fast") -> str:
        model = self.routes.get(tier, self.routes["fast"])
        response = self.client.chat.completions.create(
            model=model,
            messages=[{"role": "user", "content": prompt}],
            max_tokens=512,
        )
        return response.choices[0].message.content

That client object is the foundation, but in production I wrap it with retry logic, circuit breakers, and — this is the part I'm most proud of — automatic fallback chains. The whole point of going multi-tier isn't just cost optimization, it's resilience. If our premium lane gets rate-limited or starts returning degraded results, the system automatically shifts to the balanced lane. If that fails, it drops to the fast lane. Our uptime has been 99.97% over the last quarter, and I'd argue a chunk of that is directly attributable to having multiple model backends behind a single abstraction layer.

The SLA conversation is where this gets really interesting from an enterprise perspective. When I talk to other architects about AI infrastructure, the number one concern I hear is vendor lock-in. Everyone's terrified of waking up one morning to find that their provider has jacked prices, deprecated a model, or — worst case — gone out of business. By routing everything through a unified API endpoint, we sidestep that problem almost entirely. If DeepSeek V4 Flash disappears tomorrow, I change one line in my routing config and we're on Qwen3-32B or whatever the next best thing is. The lock-in risk goes from existential to trivial.

Multi-region deployment was another huge win. I run our recommendation service across three AWS regions (us-east-1, eu-west-1, ap-southeast-1) with active-active traffic shaping. The unified API endpoint means I'm not maintaining three separate client libraries, three different auth schemes, or three different rate limit tracking systems. Everything flows through the same client, and I can shift regional load based on latency, cost, or capacity in seconds. Our p99 latency for users in Asia dropped from 3.8 seconds to under 1.5 seconds once we started routing through the closest region.

Here's a more complete picture of how the fallback chain looks in our production code:

import openai
import os
import time

client = openai.OpenAI(
    base_url="https://global-apis.com/v1",
    api_key=os.environ["GLOBAL_API_KEY"],
)

FALLBACK_CHAIN = [
    "deepseek-ai/DeepSeek-V4-Flash",
    "Qwen3-32B",
    "glm-4-plus",
    "openai/gpt-4o",
]

def robust_recommend(prompt: str, max_retries: int = 2) -> str:
    last_error = None
    for model in FALLBACK_CHAIN:
        for attempt in range(max_retries):
            try:
                response = client.chat.completions.create(
                    model=model,
                    messages=[{"role": "user", "content": prompt}],
                    timeout=10,
                )
                return response.choices[0].message.content
            except Exception as e:
                last_error = e
                time.sleep(0.5 * (attempt + 1))
                continue
    raise RuntimeError(f"All fallback models exhausted: {last_error}")

That robust_recommend function is the workhorse of our recommendation service. It tries the cheapest viable option first, escalates only when necessary, and gives up gracefully with a clear error if every tier fails. In practice, the fast lane handles about 72% of requests, the balanced lane picks up another 21%, and the premium lane only fires for the remaining 7% — typically the complex personalization cases where context really matters. That 7% is the only segment where I let the $10.00/M output cost stand, and even there I cap token usage aggressively.

Caching is the other piece of the puzzle I want to talk about, because it's where the compounding savings come from. We cache recommendation responses at multiple layers — Redis for hot data, S3 for warm data, and a CDN edge layer for our most popular content. Across the whole stack, we're hitting a 40% cache hit rate, which means four out of every ten requests never even touch the model. That single optimization saves us roughly $2,400 a day. I know that number because I track it obsessively in Grafana.

The other architectural decision that paid off was streaming. For real-time recommendation interfaces, perceived latency matters more than actual latency, and streaming responses cut our perceived latency by about 60%. Users see the first suggestion in under 200ms even when the full response takes 1.2 seconds. The OpenAI-compatible client makes this trivial — you just set stream=True and iterate over the chunks. I won't bore you with another code block, but the implementation is maybe four lines.

Let me talk briefly about monitoring because no recommendation system survives contact with production without it. We track five core signals: model latency at p50, p95, and p99, cost per thousand requests, quality scores from our offline evaluation suite, user satisfaction signals from thumbs-up/thumbs-down buttons, and fallback rate per tier. That last metric is my favorite because it tells me when the cheap lane is struggling and I need to investigate. Last month it spiked to 8% for about an hour, and I caught it because the dashboard was red. Turns out there was a subtle prompt injection attack on one of our endpoints, and the fast lane was correctly refusing to process the malicious payloads while the balanced lane was getting confused. The fallback chain worked exactly as designed.

The thing I want you to take away from all this is that AI recommendation systems in 2026 don't require you to pick one model and accept whatever it costs. The economics have changed dramatically, and the tooling has caught up. Through Global API, you have 184 models at your fingertips, with prices that range from pocket change to premium, all behind a single OpenAI-compatible endpoint. You can build a tiered, multi-region, auto-scaling recommendation engine in under 10 minutes — I timed my last greenfield deployment at 7 minutes and 42 seconds, including the Terraform apply.

If you're building or maintaining a recommendation system, I'd genuinely encourage you to look at what Global API has put together. The unified SDK, the model selection, the pricing transparency — it checked every box on my architectural requirements list. I don't say this often about infrastructure providers, but they made my job easier, and the bill my CFO sees every month is finally something I don't dread opening. Check it out at global-apis.com if you want to see for yourself. The 100 free credits they offer are more than enough to run a meaningful benchmark against your current setup, and the worst case is you learn something interesting about your existing pipeline. Best case, you cut your costs in half and sleep better at night. I know I do.