惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

TaoSecurity Blog
TaoSecurity Blog
L
LINUX DO - 热门话题
Spread Privacy
Spread Privacy
C
Cybersecurity and Infrastructure Security Agency CISA
B
Blog RSS Feed
P
Proofpoint News Feed
AWS News Blog
AWS News Blog
GbyAI
GbyAI
D
DataBreaches.Net
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
aimingoo的专栏
aimingoo的专栏
C
CERT Recently Published Vulnerability Notes
A
About on SuperTechFans
NISL@THU
NISL@THU
Google DeepMind News
Google DeepMind News
P
Privacy International News Feed
Martin Fowler
Martin Fowler
Hacker News - Newest:
Hacker News - Newest: "LLM"
H
Help Net Security
Cisco Talos Blog
Cisco Talos Blog
T
Troy Hunt's Blog
博客园 - 三生石上(FineUI控件)
Help Net Security
Help Net Security
V2EX - 技术
V2EX - 技术
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
云风的 BLOG
云风的 BLOG
N
News and Events Feed by Topic
C
Cyber Attacks, Cyber Crime and Cyber Security
Cloudbric
Cloudbric
H
Hacker News: Front Page
T
The Blog of Author Tim Ferriss
罗磊的独立博客
MongoDB | Blog
MongoDB | Blog
P
Proofpoint News Feed
博客园_首页
C
CXSECURITY Database RSS Feed - CXSecurity.com
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Application and Cybersecurity Blog
Application and Cybersecurity Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
L
LangChain Blog
MyScale Blog
MyScale Blog
S
Security Affairs
L
Lohrmann on Cybersecurity
Recorded Future
Recorded Future
Webroot Blog
Webroot Blog
L
LINUX DO - 最新话题
腾讯CDC
Google Online Security Blog
Google Online Security Blog
Google DeepMind News
Google DeepMind News
T
Tor Project blog

Towards AI

Building AI Agents in Rust — part 4 | Towards AI Building AI Agents in Rust — part 5 | Towards AI The Verified Identity Agent Bridge | Towards AI You Can’t Prompt Your Away Your LLM Problems | Towards AI The Free Agent Trap | Towards AI Your Agentic Loop Will Drift. Here Is the KL Divergence Equation That Measures How Far It Has Wandered From Its Original Instruction. | Towards AI Beyond Chat: Processing Images, PDFs, and Documents with the OpenAI Adapter in Oracle Integration Cloud | Towards AI Building AI Agents in Rust — part 3 | Towards AI Self-Hosting Airflow at Home: Automating Stock Price Data Collection | Towards AI The 76-Hour Frontier: How the Takedown of Claude Fable 5 Birthed the Military-Industrial-AI Complex | Towards AI I Trained a Markdown File to Boost GPT-5.5 by 23 Points — It Shouldn't Work | Towards AI We Replaced ChatGPT With a Local AI Server. Six Months of Honest Data. | Towards AI What Really Makes Cars Pollute? A Data Science Deep Dive into CO₂ Emissions | Towards AI Training GPT-2 From Scratch on a GTX1050 | Towards AI Principal Component Analysis (PCA): Theory, Mathematics, and Applications Build a Zero-Cost Web Automation Pipeline With OpenRouter, OpenClaw, and MediaUse I Gave Qwen3.7-Plus a Screenshot and It Found the Exact Pixel to Click for $0.40 Beyond the Prompt: Why Autonomous AI Agents Are Replacing the Chatbot Moonshot Cracked Claude Code’s Playbook with an MIT Terminal Agent and a $0.60 Model Connections, Roles, and Warehouses: Getting CoCo Desktop Production-Ready from Day One My First $5,000 Month Writing About AI Engineering on Medium Google Shrank Gemma 4 by 72% and Unsloth Fixed the 4-Bit Bug Nobody Else Caught on One 4090, and 4-Bit Shouldn’t Be This Good LangChain Explained: Understanding Models, Prompts, Chains, Memory, Indexes, and Agents TOON: Beyond JSON for LLMs Claude Code Casual, Pro, Elite: The Three Working Personas of Claude Code Mastery MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn’t Be This Cheap Using Amazon SQS for AI Agent Orchestration I Ran a 1.5B-Active Model on My Laptop That Embarrassed a 26B by 46 Points How to Build a Self-Improving Company with AI Part 2 — Serve-Level Speed: System Design That Stabilizes P95/P99 3-Part Series: LLM Latency in Production (Part 1) Claude Code: The AI Coding Partner Changing How Developers Build Software Claude Code Pitfalls: Claude Code Won’t Do What You Told It: A Troubleshooting Catalog Full-Stack Data Scientists for the Agentic Coding World Building Production-Grade AI Skills with Snowflake Cortex AI Function Studio I Tried 10 AI Agent Frameworks in 2026 — Here’s the Honest Guide I Wish I Had Earlier How One Spring Boot Optimization Saved Our Startup $30,000 a Year Inside Palantir AIP: How the World’s Most Controversial AI Platform Actually Works What Is a Reverse Proxy? (And Why Every Backend Developer Should Care) What Claude Opus 4.8 Actually Changes If You’re Building Agents QWEN 3.7 Max Worked For 35 Hrs Straight And The Results Were Mind-blowing When LLMs Meet Knowledge Graphs on the Battlefield Fine-Tuning is Dead: Why Context Orchestration Won in 2026 5 Things Broke When I Shipped a RAG + MCP Agent to Production. Google Co-Scientist: Hyper Scaling Research and Discovery Microsoft Just Embarrassed Browser Web Agents — 1,000 Lines Made GPT-5.4 Beat Opus 4.6 on 200 Web Tasks The Modern Data Stack Is Broken — Here’s How to Fix It With AI, Governance, and Real Architecture Building Production MCP Servers: What the Spec Won’t Tell You When Should an Agent Stop? The Anatomy of Termination Harness Engineering: The Layer That Matters More Than the Model AI Engineers Who Can’t Debug Are Getting Fired (Here’s How I Debug with Claude Code) Claude Code Memory: Why You Keep Explaining the Same Thing to Claude (and the Five Layers That Fix It) Claude Code Subagents: The Claude Code Feature You Skip Every Day (And Why It Quietly Wrecks Your Sessions) Agentic AI and the SMB Banking Advantage Claude Code: Spec-Driven Development — Why Your AI Coding Sessions Fall Apart at Hour Three The Real Cost of Agentic AI Nobody Budgets For SVM : 40 must visit Interview Questions (Part 2) Your AI Agent Works Perfectly in the Demo. Here Are the 6 Ways It Dies in Production. Unleashing the Power of ONNX for Speedier SBERT Inference Terraform vs CI/CD for Serverless Deployments Merve Noyan Stopped Writing Training Scripts — Her Agent Just Fine-Tuned 18 Models Solo for $11.40 Why Your Sales Forecast Is Always 20% Wrong (And How To Make It 12% Wrong) Genetic Cubic n{C/A} Ratios For Elementary Robotics Design Top 20 AdaBoost Interview Questions & Answers (Part 2 of 2) Agentic AI Vs AI Agents — What Are the Key Differences? LAI #127: The Infrastructure Layer of AI Is Becoming the Product Anthropic Caught Its Own AI Planning to Blackmail Engineers RNNs Cannot Think What Transformers Think Cheaply. ICLR 2026 Proved the Gap Is Exponential. Time Series Made So Easy My Aunt Got It on the Second Read Claude Cowork 101 | Towards AI Is 3-Bit KV Cache the Holy Grail? A Reality Check on Google’s TurboQuant LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System AutoML on Autopilot | Towards AI I Ran This Open-Source AI Tool on a Messy Codebase and Got 71x Fewer Tokens — Here Is Exactly What Happened Month in 4 Papers (April 2026) AI Kept Forgetting My Notes. Fixing That Taught Me How It Actually Works. How ChatGPT Makes You Addicted Crack ML Interviews with Confidence: K-Nearest Neighbors (KNN 20 Q&A) The Event-Driven Blueprint: How I Scaled a Spring Boot System to 10 Million Kafka Messages/Day Building Vector Search? Why FAISS Alone Isn’t Enough TAI #202: GPT-5.5 Moves Codex Into Real Work Machine Learning System Design -The Model Serving Triangle, With One Forward Pass Flowing Through Every Trade-off (Part3) AI Orchestration in Action: How MuleSoft and LLMs Fuel the Future of Enterprise AI GPT-4 Has 1.8 Trillion Parameters. It Uses 2% of Them Per Token. Part 20: Data Manipulation in Multi-Dimensional Aggregation A Fundamental Introduction to Genetic Algorithm -Part Two TAI #200: Anthropic’s Mythos Capability Step Change and Gated Release From Notebook to Production: Running ML in the Real World (Part 4) Sqribble’s Template‑Driven Document Automation Anthropic Just Shipped the Layer That’s Already Going to Zero Long-Term vs Short-Term Memory for AI Agents: A Practical Guide Without the Hype The L1 Loss Gradient, Explained From Scratch Your Postcode Is Deciding Your Care. I Built a Pipeline to Prove It. I Directed AI Agents to Build a Tool That Stress-Tests Incentive Designs. Here’s What It Found. Your System Prompt Is the Product — Not the Feature The LLM Wiki Trend Has a Retention Problem Nobody Mentions Top 20 Data Preparation Interview Questions and Answers (Part 2 of 2) LAI #122: Word Embeddings Started in 1948, Not With Word2Vec Top 15 Computer Vision Datasets [2026] 40 Generative AI Interview Questions That Actually Get Asked in 2026 (With Answers)
Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free
Editorial Team · 2026-06-03 · via Towards AI

Author(s): Mehedi Hasan

Originally published on Towards AI.

You now know how to make the model fast (Part 1) and how to build a stable serving layer around it (Part 2). The final question is: which engine actually implements all of this without forcing you to write a custom scheduler from scratch?

The theme of this part: inference engines are not neutral wrappers. They bake in specific opinions about batching, KV cache memory layout, prefix caching, and kernel selection. Pick the engine that aligns with your pain points, and you get chunked prefill, continuous batching, and paged KV cache for free. Pick the wrong one, and you’ll spend sprints reimplementing features the right engine already has.

Here is how the four major runtimes compare in 2026, with exact configs and the tradeoffs that matter for production.

vLLM: The Production Default

vLLM is the safest starting point for most teams. Its core innovation — PagedAttention — treats the KV cache like virtual memory with fixed-size blocks, reducing fragmentation from 60–80% in naive systems to under 4%. This directly translates to 2–4x higher concurrency on the same GPU.

What you get out of the box:

  • Continuous batching (iteration-level scheduling): requests enter and leave the GPU every token step, not every batch
  • Chunked prefill (v0.4+): long prompts are broken into chunks and interleaved with decode steps, so a 3K-token prefill doesn’t starve short chat requests
  • Automatic prefix caching (APC): the engine detects shared prompt prefixes and reuses KV cache automatically
  • Speculative decoding (EAGLE, Medusa, n-gram): 2–3x latency reduction for memory-bound decode
  • Multi-LoRA serving: serve hundreds of fine-tuned adapters on one base model
  • Broad quantization support: GPTQ, AWQ, FP8, INT8, INT4, AutoRound
  • 200+ model architectures: Llama, Qwen, DeepSeek, Mixtral, MoE, VLMs, embedding models

The config that matters:

python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.3-70B-Instruct \
--tensor-parallel-size 2 \
--gpu-memory-utilization 0.85 \
--max-model-len 8192 \
--max-num-seqs 64 \
--enable-prefix-caching \
--enable-chunked-prefill \
--quantization fp8

Key flags explained:

  • --enable-prefix-caching: turns on automatic prefix caching for shared system prompts (massive for RAG)
  • --enable-chunked-prefill: prevents long prefill monopolization; interleaves prefill chunks with decode
  • --gpu-memory-utilization 0.85: leaves 15% headroom for CUDA graph capture and KV cache growth; going to 0.95 often causes OOM during graph compilation
  • --max-num-seqs 64: caps concurrent sequences. Higher isn’t always better—if you hit memory limits, the engine will evict blocks and thrash.

MRV2 (Model Runner V2): In v0.17.0+, enable VLLM_USE_V2_MODEL_RUNNER=1 for a rewritten backend that delivers significant throughput gains, especially on newer architectures like GB200.

When to choose vLLM:

  • You support many different models and need one engine to handle them all
  • You run on heterogeneous hardware (NVIDIA, AMD, Intel Gaudi, AWS Trainium)
  • Your team wants the largest community, best documentation, and fastest debugging
  • You need to be online in under 90 seconds from a cold start

Limitation: Peak throughput on dedicated H100 clusters is ~29% lower than SGLang or LMDeploy in some benchmarks, primarily due to Python orchestration overhead. If you have a fixed model and a specialized team, you can squeeze more out of other engines. But for most teams, vLLM’s breadth outweighs that gap.

SGLang: The Throughput Challenger with Automatic Prefix Caching

SGLang, developed by LMSYS (the team behind Chatbot Arena), is no longer a niche alternative. It powers xAI’s Grok 3 and Microsoft Azure’s DeepSeek R1 deployments, running on over 400,000 GPUs worldwide.

What differentiates it: RadixAttention. Instead of manually configuring prefix caches, SGLang builds a radix tree from request prefixes and automatically reuses KV cache across any requests that share token sequences. This is transformative for multi-turn chat, agent loops, and RAG pipelines where system prompts and retrieved contexts repeat.

What you get out of the box:

  • RadixAttention: automatic, dynamic prefix caching without manual key management
  • Chunked prefill: same interleaving benefit as vLLM
  • EAGLE/EAGLE3 speculative decoding: state-of-the-art draft-model speculation
  • Prefill-decode disaggregation: separate prefill and decode across different GPU pools for independent scaling
  • MLA-optimized kernels: specifically tuned for DeepSeek models
  • Zero-overhead CPU scheduler: moves scheduling logic off the GPU thread

The config that matters:

python -m sglang.launch_server \
--model-path meta-llama/Llama-3.3-70B-Instruct \
--tp 2 \
--quantization fp8 \
--context-length 8192 \
--mem-fraction-static 0.92 \
--enable-flashinfer-mla \
--host 0.0.0.0 \
--port 8000

Key flags explained:

  • --tp 2: tensor parallelism across 2 GPUs
  • --mem-fraction-static 0.92: SGLang’s memory allocator is more aggressive than vLLM’s; 0.92 is typically stable on H100
  • --enable-flashinfer-mla: enables optimized Multi-Head Latent Attention kernels for DeepSeek-class models

Performance reality check: In H100 benchmarks with unique prompts (no prefix sharing), SGLang achieves roughly 29% higher throughput than vLLM. However, the gap narrows or reverses on workloads with high memory pressure where vLLM’s PagedAttention is more mature. The real win is in shared-prefix workloads — multi-turn conversations, agent loops, and RAG with fixed retrievers — where RadixAttention provides gains no other engine matches automatically.

When to choose SGLang:

  • Your workload is dominated by multi-turn conversations or shared system prompts
  • You are serving DeepSeek models (MLA kernels are best-in-class)
  • You have a dedicated inference team that can manage dependencies (FlashInfer can be finicky to install)
  • You need prefill-decode disaggregation at scale

Limitation: Model coverage is narrower than vLLM. If you serve exotic architectures or need to swap models frequently, vLLM is safer.

TensorRT-LLM: The NVIDIA Optimizer (With a Catch)

TensorRT-LLM is NVIDIA’s official inference SDK. It delivers the highest raw throughput and lowest TTFT on NVIDIA hardware when fully tuned. But it makes very specific tradeoffs.

The compiled engine tradeoff: Traditionally, TensorRT-LLM required compiling a model into a serialized engine — a process that takes ~28 minutes for a 70B model. This is a one-time cost per model version, but it breaks auto-scaling and blue-green deploys unless you precompile and cache engines.

The PyTorch backend (v1.0+): This changed the game. TensorRT-LLM now defaults to a PyTorch backend that loads HuggingFace weights directly, cutting cold start to ~60–90 seconds (comparable to vLLM). You lose some peak throughput compared to the compiled engine, but you gain deployment flexibility.

What you get out of the box:

  • Fused kernels: aggressive kernel fusion for attention and MLP layers
  • FP8 quantization: native, optimized FP8 on Hopper (H100/H200/GB200)
  • Tensor parallelism + pipeline parallelism: mature multi-GPU orchestration
  • Speculative decoding: supported via draft models
  • CUDA graph capture: minimal CPU launch overhead

The config that matters (compiled engine):

# Step 1: Quantize and compile (one-time, ~28 min for 70B)
python quantize.py --model_dir ./llama-3.3-70b \
--output_dir ./quantized \
--qformat fp8
trtllm-build --checkpoint_dir ./quantized \
--output_dir ./engine \
--gemm_plugin fp8
# Step 2: Serve
trtllm-serve --engine_dir ./engine \
--max_batch_size 32 \
--max_input_len 4096 \
--max_output_len 1024

The config that matters (PyTorch backend, no compile):

trtllm-serve --model ./llama-3.3-70b \
--quantization fp8 \
--tp 2 \
--max_batch_size 32

When to choose TensorRT-LLM:

  • You have a single model that won’t change for months
  • You are on NVIDIA-only infrastructure (Hopper or newer)
  • Your team can invest 1–2 weeks in tuning and compilation pipelines
  • You need the absolute highest throughput at 100+ concurrent requests

When to avoid it:

  • You auto-scale from zero (unless you use the PyTorch backend)
  • You serve multiple models and need to swap them daily
  • You are not on NVIDIA hardware

NVIDIA NIM: If you want TensorRT-LLM performance without the compilation headache, NVIDIA NIM bundles precompiled engines, weights, and an API server into a single container. It is essentially TensorRT-LLM with DevOps handled for you.

TGI (Text Generation Inference): The Maintenance Mode Legacy

TGI was HuggingFace’s production serving engine, powering Hugging Chat and the Inference API. It introduced continuous batching and Flash Attention to a wide audience. But as of 2026, TGI is officially in maintenance mode.

HuggingFace’s own guidance: accept pull requests for minor bug fixes only, and recommend migrating to vLLM or SGLang for new deployments.

What this means for you:

  • If you are already running TGI in production, plan a migration path
  • If you are starting a new project, do not choose TGI
  • TGI’s ecosystem contributions (quantization support, model architectures) have been upstreamed into vLLM and SGLang

TGI remains a respectable piece of engineering, but it is no longer the future.

llama.cpp: The Edge and Local Workhorse

llama.cpp is not a datacenter serving engine. It is optimized for running quantized models (GGUF format) on consumer hardware, CPUs, and edge devices.

What it does well:

  • GGUF quantization: runs 70B models on 24GB consumer GPUs via aggressive quantization
  • CPU inference: AVX/AVX2 optimized paths for machines without GPUs
  • Metal backend: runs on Apple Silicon (M3/M4 Ultra)
  • Local server mode: exposes an HTTP API for local development

When to choose it:

  • You need inference on a laptop, edge device, or embedded system
  • You are building a local AI assistant (e.g., Ollama, which wraps llama.cpp)
  • You want to avoid cloud costs entirely for personal use

When to avoid it:

  • Multi-tenant GPU serving
  • High-throughput API backends
  • Workloads requiring continuous batching across hundreds of concurrent users

LMDeploy: The Dark Horse (C++ Native, Minimal Friction)

LMDeploy is a pure C++ inference engine that achieves near-SGLang throughput with trivial installation (pip install lmdeploy). It is the practical choice if you want maximum performance without dependency hell.

What you get:

  • Native C++ backend: zero Python orchestration overhead
  • First-class quantization: AWQ, GPTQ, FP8, INT4
  • Turbomind engine: optimized CUDA kernels for decode
  • One-line deployment: simpler setup than SGLang or TensorRT-LLM

The config:

lmdeploy serve api_server \
meta-llama/Llama-3.3-70B-Instruct \
--model-format hf \
--quant-config dict(type='fp8') \
--tp 2

When to choose LMDeploy:

  • You want 99% of SGLang’s throughput with 10% of the setup complexity
  • You are on NVIDIA hardware and don’t need vLLM’s broad hardware portability
  • Your team values installation simplicity over ecosystem size

Decision Guide: Map Your Pain Point to the Engine

Here is how to choose based on the problems you identified in Parts 1 and 2.

Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free
Comprehensive comparison of major LLM serving engines

The Boring Choice Is Usually the Right Choice

If you are a tech lead making this decision for a team, here is the empirical advice:

Start with vLLM. It is not the fastest engine on any single benchmark, but it is the fastest to deploy, the easiest to debug, and the most forgiving when your requirements change. You can switch to SGLang later if you measure that RadixAttention would materially improve your workload. You can switch to TensorRT-LLM later if you have a fixed model and a dedicated team to manage compilation.

Do not build your own engine. The gap between a naive FastAPI wrapper around model.generate() and vLLM is 10-24x in throughput. The gap between vLLM and a custom C++ scheduler you wrote in a month is that your custom scheduler has bugs vLLM already fixed.

Measure before you optimize. Run your actual workload — your actual prompts, your actual concurrency patterns — through vLLM first. Log TTFT, TPOT, and queue depth. If your P99 is dominated by KV cache exhaustion, tune --max-model-len and --gpu-memory-utilization. If your P99 is dominated by long prompts blocking short ones, enable chunked prefill. Only after you have exhausted the engine’s built-in optimizations should you consider switching engines.

The engine is the last 10% of the optimization stack. Parts 1 and 2 gave you the 90%: quantization, kernel selection, traffic lanes, batching discipline, and backpressure. Get those right with any modern engine, and your users will see sub-second TTFT and stable streaming. Get those wrong, and the fastest engine in the world will still feel broken.

Series Summary:

  • Part 1 made the model fast: quantization, Flash Attention, paged KV cache, and GPU kernel tuning.
  • Part 2 made the serving stable: traffic lanes, continuous batching, backpressure, cold-start avoidance, and output control.
  • Part 3 gave you the engine map: vLLM for breadth, SGLang for shared-prefix workloads, TensorRT-LLM for peak NVIDIA throughput, and llama.cpp for the edge.

Pick the engine, deploy the configs, and ship.

Published via Towards AI