惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
SegmentFault 最新的问题
B
Blog
P
Proofpoint News Feed
美团技术团队
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
A
About on SuperTechFans
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Vercel News
Vercel News
有赞技术团队
有赞技术团队
小众软件
小众软件
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Google DeepMind News
Google DeepMind News
Martin Fowler
Martin Fowler
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
aimingoo的专栏
aimingoo的专栏
H
Help Net Security
罗磊的独立博客
L
LangChain Blog
GbyAI
GbyAI
腾讯CDC
T
The Blog of Author Tim Ferriss
Microsoft Security Blog
Microsoft Security Blog

Runpod Blog.

DeepSeek V4 in the wild, and how to run it on Runpod New Runpod datacenter now live: AP-IN-1 Track GPU spend across your team with Cost Centers The GPU supply supercycle is here. Here’s what AI builders need to know. Community Spotlight: One-click AI image and video generation on Runpod with SwarmUI | Runpod Blog Community Spotlight: LoRA Pilot Data Prep to Inference Introducing the Runpod Assistant: Manage Your Cloud GPU Resources with Natural Language OpenAI's Parameter Golf: Train the Best Language Model That Fits in 16MB on Runpod LLM inference optimization: techniques that actually reduce latency and cost Pruna P-Video and Vidu Q3 public endpoints now available on Runpod Runpod brand spelling guide Quickstart - Runpod Documentation The AI market looks nothing like the narrative Training StyleGAN3 with Vision-Aided GAN on Runpod KoboldAI – The Other Roleplay Front End, And Why You May Want to Use It How to Connect Cursor to LLM Pods on Runpod for Seamless AI Dev Community Spotlight: How AnonAI Scaled Its Private Chatbot Platform with Runpod Prompt Scheduling with Disco Diffusion on Runpod Runpod's Latest Innovation: Dockerless CLI for Streamlined AI Development Run Your Own AI from Your iPhone Using Runpod Introducing Flash: Run GPU workloads on Runpod Serverless: No Docker required Use Claude Code with your own model on Runpod: No Anthropic account required Avoid Errors by Selecting the Proper Resources for Your Pod What hackers built on Runpod at TreeHacks 2026 Easily Back Up and Restore Your Pod with Cloud Sync + Backblaze B2 The Complete Guide to GPU Requirements for LLM Fine-Tuning AI Guides, Tutorials & GPU Infrastructure Insights | Runpod Your first Claude Code project within Runpod: a complete setup guide Building for resilience: Runpod’s response to the AWS us-east-1 outage How to Connect Google Colab to Runpod
10 billion Serverless requests and counting
Brendan McKeag · 2026-02-18 · via Runpod Blog.

We just served our 10 billionth serverless request.

That's 10 billion images generated.

10 billion videos created.

10 billion training steps.

10 billion moments where someone had an idea and our infrastructure helped make it real.

But we didn't build this.

You did.

Built by Builders

Every one of those requests represents a developer who trusted us with their workload. A startup that bet on us to scale with them. A creator who chose Runpod when they could have gone anywhere else.

Three years ago, serverless was an experiment. Today, it's powering production workloads for teams building the future of AI. From solo developers training their first model to infrastructure teams at companies processing millions of requests per day, serverless has become the way modern AI gets built.

We've watched this evolution happen in real-time. The first serverless requests were tentative—developers testing the waters, seeing if this whole "pay per second" thing actually worked. Then came the hockey stick. Suddenly we were seeing endpoints that processed thousands of images per hour, video generation pipelines handling viral traffic spikes, and code generation tools serving entire development teams.

Why Serverless Matters

And why is this important? Serverless represents the perfect bite-size segmentation of workloads, letting you put that GPU to work directly on what matters rather than scaffolding around what you're actually after.

Traditional GPU infrastructure makes you think about the wrong things. How many instances do I need? What if traffic spikes? What about idle time? You end up spending more time being a cloud architect than building your actual product.

Serverless flips that model. No idle costs. No infrastructure headaches. No guessing at capacity. Just your code, running exactly when it needs to, scaling from zero to hundreds of workers in seconds.

The math is simple: if your workload is bursty, unpredictable, or event-driven—which most AI workloads are—you shouldn't be paying for GPUs sitting idle. You should be paying for compute only when you're actually computing.

That's what 10 billion requests looks like when infrastructure gets out of your way.

What's Next

Thank you. For building with us, for pushing us to be better, and for showing us what's possible when great tools meet great builders.

We're not stopping here. We're working on faster cold starts, more flexible scaling policies, and deeper integrations with the tools you're already using. Because every one of those 10 billion requests taught us something about what you need.

Here's to the next 10 billion.

Want to learn more about serverless? Check out our docs or our YouTube channel.

Author profile: Brendan McKeag