惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

雷峰网
雷峰网
T
The Exploit Database - CXSecurity.com
Engineering at Meta
Engineering at Meta
T
Troy Hunt's Blog
大猫的无限游戏
大猫的无限游戏
云风的 BLOG
云风的 BLOG
J
Java Code Geeks
Latest news
Latest news
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
C
Check Point Blog
Recent Announcements
Recent Announcements
F
Fortinet All Blogs
A
Arctic Wolf
酷 壳 – CoolShell
酷 壳 – CoolShell
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Vercel News
Vercel News
S
Schneier on Security
Cyberwarzone
Cyberwarzone
N
Netflix TechBlog - Medium
AWS News Blog
AWS News Blog
Scott Helme
Scott Helme
G
GRAHAM CLULEY
SecWiki News
SecWiki News
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
Google DeepMind News
Google DeepMind News
H
Heimdal Security Blog
GbyAI
GbyAI
C
Cybersecurity and Infrastructure Security Agency CISA
D
Docker
S
Secure Thoughts
博客园 - 三生石上(FineUI控件)
S
SegmentFault 最新的问题
aimingoo的专栏
aimingoo的专栏
博客园 - 叶小钗
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
F
Full Disclosure
L
LINUX DO - 热门话题
Help Net Security
Help Net Security
博客园 - 司徒正美
C
Cisco Blogs
月光博客
月光博客
O
OpenAI News
P
Proofpoint News Feed
Attack and Defense Labs
Attack and Defense Labs
S
Security @ Cisco Blogs
L
Lohrmann on Cybersecurity
宝玉的分享
宝玉的分享
D
DataBreaches.Net
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
When is Serverless Inference Cheaper than Your Self Hosted GPU? I Benchmarked gpt-oss-120b on Both
Yash Sharma · 2026-06-23 · via DEV Community

If you run LLM inference in production, you eventually will ask yourself, should you rent a GPU and run the model yourself, or do you use a serverless API and pay per token? Everyone has an opinion. Far fewer people show you the actual numbers that decide it.

So I ran both. I put the same model, gpt-oss-120b, on two setups, self-hosted with vLLM on a single AMD MI300X GPU Droplet, and on DigitalOcean's Serverless Inference. Then I measured the cold start, the warm latency, and the cost, and worked out exactly where one becomes the better choice than the other.

In short, self-hosted GPU is faster and more consistent once it's warm, but it carries a real cold start and you pay for it around the clock. Serverless hides the cold start and costs almost nothing at low volume, but you pay per token. Which one wins comes down to your traffic shape and how much your model actually outputs. Here are the numbers.

How long is the cold start on a self-hosted GPU?

When you run a model yourself, the GPU doesn't hold the model permanently. The weights have to be loaded from disk into the GPU's memory and the inference engine has to initialize before it can answer a single request. That startup delay, the gap between "process launched" and "first token out," is the cold start. You pay it every time you start the server fresh, a new deploy, a restart after a crash, or a new node coming up to handle load.

My setup here was a single AMD MI300X GPU Droplet running gpt-oss-120b with vLLM, with the weights already cached on disk. I started vLLM from cold and timed how long it took before the first token came back.

It took about 61 seconds. That wasn't a one-off, either, across three restarts it landed between 60.8 and 61.4 seconds every time.

Cold Start of Model

One thing worth being precise about, because it's the most common misread, that 61 seconds is not the time to download the model. The weights were already saved on disk, so this is the cost you pay on a restart or redeploy, not a one-time setup. The startup logs show where the time actually goes:

Phase Time
Load weights from disk into VRAM (~68 GB) ~24 s
torch.compile ~4 s
CUDA graph capture ~11 s
Engine init, KV cache, warmup ~21 s

So the 61 seconds includes compilation and warmup, the entire engine bring-up, right up to serving a request. What it excludes is the one-time download of the weights from Hugging Face, which you pay once and never again. (These phases come from one representative run and there's some overhead between stages, so they don't sum to exactly 61, but that's where the time lives.)

This also answers the obvious follow-up, what happens when you scale up? If a new replica boots from an image with the weights baked in, or mounts a shared volume that already has them, it pays this ~61-second load, not a download. A brand-new node with nothing staged would also have to pull the weights first, but well-run setups specifically avoid that, because you don't want every scale event re-downloading 68 GB. So 61 seconds is the realistic number for a properly configured restart or scale-up

Warm latency and throughput

Concurrent Benchmarking Of Warm Model

Once the model is loaded, it's a different machine. Warm, the self-hosted MI300X returned the first token in about 322 ms and sustained roughly 154 tokens per second, and it was remarkably stable, across twenty requests, the spread was about two milliseconds.

Benchmarking showcasing VRPAM used

Memory is worth a note, because the raw number looks alarming. The card showed about 173 GB of VRAM in use. But the weights themselves are only about 68 GB. vLLM reserves most of the rest up front as KV-cache headroom (roughly 100 GB of it) so it can serve many requests at once. A 120B model doesn't "need" 173 GB; the engine just claims the room ahead of time.

So the self-hosted trade-off is clean: once it's warm, it's fast, consistent, and entirely yours, but every cold start costs a full minute for our model, and you pay for the GPU whether or not anyone is using it.

Does serverless inference have a cold start?

Next I ran the identical test against Serverless Inference. Calling it is straightforward, you create a model access key and hit the OpenAI-compatible endpoint, so the client code is the same and only the base URL changes. I left the endpoint idle first, then measured first-token latency the same way.

Concurrent Benchmarking Of Serverless

The first token came back in about 546 ms, with a wider spread, anywhere from 446 ms to roughly 1.3 seconds across twenty runs. But there was no spin-up. I ran it twenty times after sitting idle and never caught a cold-start hit.

Two honest caveats on those numbers. First, the serverless requests travel over the network to DigitalOcean's endpoint, while the GPU test ran locally on the droplet, so some of that extra latency is network distance, not the model being slower. Here's the warm comparison side by side:

Metric Self-hosted MI300X Serverless Inference
Median time to first token ~322 ms ~546 ms
Spread (20 runs) ~2 ms 446 ms – 1.3 s
Throughput ~154 tok/s (per-token billed)
Cold start ~61 s none observed

So why no cold start on serverless? It didn't delete the cold start, it absorbed it. DigitalOcean pools GPU capacity across customers, so the model stayed warm without any effort from me, and I never paid the 61-second hit I took on my own box. The difference is the billing model: serverless isn't charged by the hour, it's charged per token.

To be fair, serverless isn't immune to cold starts. If you hit it during a genuinely quiet stretch, you can still catch one. The standard mitigations are sending periodic warm-up requests to keep a worker hot, or designing async-first so a slow first response doesn't matter. In this test I didn't need any of that, it just stayed warm.

When is serverless inference cheaper than your own GPU?

This is where the decision actually lives, and it comes down to arithmetic. The GPU is a flat cost, about $1.88 an hour for a single on-demand MI300X, the same whether it serves one request or a million. Serverless is usage-based, gpt-oss-120b is priced at $0.10 per million input tokens and $0.70 per million output tokens, so it costs almost nothing when you're quiet and climbs as you get busier.

The break-even point is your hourly GPU cost divided by your per-request cost:

break-even requests/hour = GPU $/hr ÷ [(input_tokens × $0.10/1M) + (output_tokens × $0.70/1M)]

The catch is that the per-request cost depends entirely on how much your model outputs, and that moves the crossover more than you'd expect. I measured it at three response lengths, on the same GPU at the same prices, changing only the output length:

Benchmarking

Response type Output tokens Crossover
Short (classification / extraction) ~30 ~18 requests/sec
Medium (paragraph answer) ~220 ~3 requests/sec
Long (code / detailed explanation) ~1,200 <1 request/sec

That's roughly a 25x swing from identical hardware, with nothing changing but response length. Output tokens are the expensive side of the bill, so the chattier your app, the sooner owning a GPU pays for itself.

In plain terms, if your app sends short, snappy responses, serverless stays cheaper until you're well over a dozen requests per second, nonstop. If it writes long answers, the GPU starts winning below one request per second. (For comparison, Dedicated Inference, DigitalOcean's managed always-on endpoint, is billed by the GPU-hour like the Droplet but without you managing the it, so its economics sit closer to the self-hosted side of this table than the serverless side.) Drop your own response lengths and GPU rate into the formula and you'll find your exact line.

When to use serverless inference, and when not to

No "it depends." Here's the actual call.

For most teams, serverless is the right default. Bursty or spiky traffic, real idle stretches, a dev tool, an internal feature, a side project, anything async where nobody is staring at a spinner on the first request. In all of those, the cold start runs on someone else's pooled capacity, not yours, and you pay nothing while you're quiet. For that kind of traffic, it's almost perfect.

Run the GPU yourself when traffic is steady and high-volume, or when you have a latency SLA you can't miss. At that point your traffic rarely stops, so you're not benefiting from serverless's idle savings anyway, and you're using the GPU enough that flat hourly beats per-token. You keep it warm, so the cold start stops mattering. That's not a knock on serverless, it's just the wrong tool for that job.

If you're in between, start on serverless. Watch your token spend, and move to a dedicated GPU the day you cross the line for your response lengths. Don't buy a GPU to solve a problem you don't have yet.

Run it yourself

Everything here is reproducible. You can spin up an AMD GPU Droplet and run gpt-oss-120b on vLLM, hit the same model on Serverless Inference with a model access key, and check the serverless metrics and pricing pages against your own workload.

Don't take my crossover, run your own. And if you measure a cold start on your own setup, I'd genuinely like to see how the spread looks across different models and hardware.