惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
The Cloudflare Blog
G
Google Developers Blog
博客园_首页
Martin Fowler
Martin Fowler
Apple Machine Learning Research
Apple Machine Learning Research
L
LangChain Blog
D
Docker
C
Check Point Blog
T
Tailwind CSS Blog
博客园 - 司徒正美
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Hugging Face - Blog
Hugging Face - Blog
Microsoft Security Blog
Microsoft Security Blog
V
V2EX
博客园 - 叶小钗
T
The Blog of Author Tim Ferriss
酷 壳 – CoolShell
酷 壳 – CoolShell
IT之家
IT之家
M
MIT News - Artificial intelligence
Microsoft Azure Blog
Microsoft Azure Blog
博客园 - 【当耐特】
GbyAI
GbyAI

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Async Batching Is the Real Latency Win Nobody's Talking A...
Aamer Mihays · 2026-05-15 · via DEV Community

Synchronous batching is a throughput hack that became a design constraint. Hugging Face's latest work on asynchronous continuous batching shows why the distinction matters more than the batch size.

Most inference servers treat batching as a queuing problem. Requests pile up, you wait for N items or a timeout, then you process them together. This works until it doesn't—when your tail latency spikes because one long request blocks the entire batch, or when your GPU sits idle waiting for that last straggler to arrive.

The move to continuous batching helped. Instead of fixed windows, you could add and evict requests dynamically. But it was still fundamentally synchronous: every forward pass had to wait for the slowest sequence in the batch to complete its decode step. The GPU utilization looked good on dashboards, but the latency distribution told a different story.

The Async Shift

Asynchronous continuous batching decouples the scheduling loop from the forward pass. Requests enter a pool, the scheduler makes decisions about what to run, and the GPU executes independently. This sounds subtle but changes everything about how you think about inference throughput.

First, you can pipeline. While the GPU is working on step T, the scheduler is already preparing the batch for step T+1. The overhead doesn't disappear, but it overlaps with useful work. On modern GPUs with async copy engines, this matters more than most benchmarks capture.

Second, you can preempt. Not in the OS sense, but in the ability to yank a completed sequence from the batch mid-flight and replace it with a fresh one. The synchronous model forced you to wait for the entire batch to finish before anyone could leave. Async lets you maintain a full batch even when individual sequences have wildly different lengths.

Why This Matters for Agents

Agent workloads break traditional batching assumptions. Tool calls introduce non-deterministic latency. A request might pause for 500ms waiting for a search result, then resume with a burst of generation. Synchronous batching either holds the slot (wasting GPU memory) or evicts the request (paying recompute costs). Neither is acceptable at scale.

Async batching treats these pauses as first-class citizens. The request steps aside, the GPU keeps working on other sequences, and the scheduler brings it back when the tool responds. The memory stays allocated, but the compute doesn't stall.

This is particularly relevant for the emerging class of "always-on" agents that maintain long-running sessions. You can't batch these traditionally—they're perpetual. But you can interleave them with short-turnaround requests if your scheduler understands async completion.

The Implementation Reality

Hugging Face's TGI and vLLM have both moved toward async scheduling, though the implementations differ. TGI uses a dedicated scheduling thread that runs ahead of the GPU, while vLLM's recent iterations push more of the async logic into the CUDA graph itself. The tradeoffs are familiar: thread overhead versus kernel launch latency, complexity versus control.

What both approaches acknowledge is that the synchronous abstraction was a convenience, not a requirement. The hardware has been capable of async execution for years. The software is catching up.

The Takeaway

If you're running inference at scale, look at your tail latency percentiles, not your average throughput. If p99 is more than 3x your median, you're probably suffering from synchronous batching artifacts. Async continuous batching won't fix everything—memory bandwidth is still a bottleneck, and attention costs don't disappear—but it removes a class of scheduling-induced latency that has no business existing in 2026.

The best part: for many workloads, this is a software upgrade, not a hardware purchase. Your A100s or H100s get immediately more useful when the scheduler stops waiting for permission to work.