惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Martin Fowler
Martin Fowler
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
A
About on SuperTechFans
Apple Machine Learning Research
Apple Machine Learning Research
The Register - Security
The Register - Security
Vercel News
Vercel News
H
Hackread – Cybersecurity News, Data Breaches, AI and More
人人都是产品经理
人人都是产品经理
MyScale Blog
MyScale Blog
云风的 BLOG
云风的 BLOG
博客园_首页
U
Unit 42
T
Tailwind CSS Blog
G
GRAHAM CLULEY
F
Full Disclosure
V
Vulnerabilities – Threatpost
T
Tenable Blog
月光博客
月光博客
P
Privacy & Cybersecurity Law Blog
P
Privacy International News Feed
K
Kaspersky official blog
Scott Helme
Scott Helme
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
N
News and Events Feed by Topic
T
The Exploit Database - CXSecurity.com
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
Recent Commits to openclaw:main
Recent Commits to openclaw:main
L
LINUX DO - 最新话题
Recorded Future
Recorded Future
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Help Net Security
Help Net Security
The GitHub Blog
The GitHub Blog
Cisco Talos Blog
Cisco Talos Blog
SecWiki News
SecWiki News
P
Proofpoint News Feed
Security Latest
Security Latest
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
罗磊的独立博客
S
Security Affairs
M
MIT News - Artificial intelligence
L
LINUX DO - 热门话题
美团技术团队
Simon Willison's Weblog
Simon Willison's Weblog
T
Threat Research - Cisco Blogs
Stack Overflow Blog
Stack Overflow Blog
Forbes - Security
Forbes - Security
Hugging Face - Blog
Hugging Face - Blog
博客园 - Franky
V
Visual Studio Blog

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Why LLM Inference Needs a New Kind of Router - Part 3
No items found. · 2026-06-05 · via Modular Blog

Qualcomm to Acquire Modular. Read More →

June 5, 2026

Aayush Deshpande

Deep Dhillon

Alexandr Nikitin

Michael Dunn-OConnor

In Part 2 of this series, we built a data structure that can query which pods have the user’s KVCache blocks in microseconds for every request. This post goes beyond the cluster and cache state to explain how Modular Cloud generates routing decisions and dispatches them across pods.

Most routing stacks ship with a fixed set of algorithms: round-robin, least-requests, consistent hashing, etc. These are generally independent implementations rather than composable components. As a result, when a customer asks for "consistent hashing with a concurrency cap" or "cache-aware with session stickiness," it requires adding a new algorithm from scratch. Disaggregated prefill/decode increases this proliferation. Every variant traditionally has its own HTTP handler, discovery logic, proxy code, and session management. That requires hundreds of lines of additional plumbing per variant.

Instead, Modular Cloud’s routing layer is built from a small number of stages where behaviors are expressed as composable plugins. New requirements can be satisfied with plugins, and new execution patterns are created by composing those plugins.

The five stages

Every routing decision in Modular Cloud goes through the same five stages, in the same order. PrepareFilterScorePickExecute.

Prepare enriches the routing context with whatever the downstream stages need: tokenizing the prompt, hashing it into blocks, extracting a session key from a header, and computing a hash key for consistent hashing. It runs once per request, before any candidate evaluation.

Filter removes candidates that can't serve the request due to health checks, hardware-role matching (prefill requests must not go to decode-only pods), and concurrency. It then outputs a smaller candidate list.

Score assigns a quality score to each remaining candidate based on factors like cache affinity, load, and node locality. A routing profile can include multiple scorers, each evaluating candidates on different factors.

Pick selects one or more candidates from the scored list using one of several strategies: MaxScorePicker (deterministic, highest score wins), RoundRobinPicker (stateful cycle), or SessionPicker (sticky lookup). Multiple scores can be composed with explicit weights to produce the final per-candidate score. Pick then outputs a RoutingPlan, which is a list of tuples representing the role, candidate, failure policy, and more.

Execute dispatches the RoutingPlan. For single-dispatch routing, this is one HTTP proxy call. For multi-step plans like disaggregated prefill/decode, it runs a sequenced flow: call the prefill pod, wait, call the decode pod, then stream the response.

Much of our inspiration comes from Endpoint Picker (EPP), the routing component of the Gateway API Inference Extension (GAIE). EPP effectively returns a single endpoint, whereas we wanted the routing layer to support execution of multiple endpoints. To enable that flexibility, we expanded our profiles with additional stages like Prepare and Execute.

Unlike EPP, Prepare is a distinct stage from scoring. Tokenization and block hashing are expensive, while extracting a session key from a header is cheap. Mixing these into scoring forces you to redo expensive work for each scorer and complicates the dependency graph for scorers that share derived inputs. Separating Prepare gives the framework a single place to run expensive transforms and cache their results for the rest of the pipeline.

Additionally, Execute is a first-class stage, not an implicit step. EPP's pipeline ends at Pick, and whatever executes the result lives outside the pipeline. That works for single-dispatch routing, where "execute" is a single HTTP proxy. It breaks down when execution has its own structure, as it does for disaggregated prefill/decode, where the plan is a sequence rather than a single step. By putting Execute in the pipeline, the same framework can handle both simple and complex dispatch with the same abstraction.

Plugins can compose, but algorithms can't

These five stages are composable because of their defined interfaces and clear separation of concerns. A scorer evaluates each candidate on one criterion and produces a per-candidate score. The framework combines multiple scorers with explicit weights, and a picker selects the winner. Different routing strategies are created by combining these reusable plugins, making it easier to implement new patterns.

RoundRobinScorer prioritizes the next endpoint in a rotation. When combined with LeastLoadScorer, that priority is weighted toward emptier queues.

Consistent hashing has a preparer to derive a hash key from the request and a picker to select a candidate matching that key. Potential candidates can be scored by load, cache affinity, locality, or any weighted combination.

Cache-aware routing utilizes two preparers and two scorers. The preparers tokenize and hash the prompt up front. CacheAffinityScorer rewards pods that already hold the blocks, and LeastLoadScorer prevents hot-spotting. The weights control that tradeoff between reuse and capacity.

Similarly, combining cache-aware routing with sticky sessions can be achieved by combining existing preparers and scorers. Once you have five stages with clean interfaces, familiar routing behaviors decompose into stage-level plugins. And the plugins compose into new execution patterns.

Typed state between plugins

Plugins need to talk to each other. TokenizePreparer produces a token array that BlockHashPreparer consumes, which is later consumed by CacheAffinityScorer.

How do they communicate without being coupled? The framework provides typed slots on the RoutingContext. A slot is a typed key with a name and a compile-time type. TokenizePreparer writes into the Tokens slot. BlockHashPreparer reads Tokens and writes BlockHashes. CacheAffinityScorer reads BlockHashes. None of the three plugins needs to reference each other to complete these read/writes.

There are two benefits of this indirection:

  • Decoupling: You can swap TokenizePreparer for a different implementation (a tokenizer tuned for a specific model family) without any downstream plugin knowing. All that matters is that something writes to the Tokens slot.
  • Compile-time typing: The slot is typed, and the reader gets the right type out. Mismatches are build-time errors, and therefore easier to trace and resolve.

Compare this to EPP's approach. EPP shares per-request plugin state through a structure called CycleState, which is backed by a key-value store keyed by opaque strings with interface{}-typed values. Type-safe generic accessors were added on top, which catch type mismatches at the call site. However, the storage is still string-keyed at heart. Two plugins can independently choose the same key and conflict, and nothing in the system notices until something goes wrong in production. The Gateway API Inference Extension (GAIE) community is aware of this and is actively discussing improvements.

We can close that gap with typed slots resolved at build time.

Validation at build time

When the framework composes a routing profile, it can check a set of invariants statically: every slot reader has a writer, every declared dependency appears in the composition, the dependency graph has no cycles, and no two plugins are mis-ordered or declare conflicting dependencies.

If any check fails, the service doesn't start. The operator sees a structured error at deployment time, not during a failed read in production. You can build new profiles with confidence that errors surface immediately.

The Selector / Workflow / Executor split

So far, we've described a single-dispatch pipeline: one pass through PrepareFilterScorePickExecute, one pod selected, one HTTP call. Most routing looks like this, but disaggregated prefill/decode doesn't. Disaggregation means one client request is served by two pods. Prefill builds the KV cache, and decode generates tokens. Routing has to pick both, and the second pick depends on the first.

The framework handles this with a three-layer abstraction on top of the five-stage pipeline.

Selector: A Selector is one pass through Prepare → Filter → Score → Pick. It takes a set of candidates and returns a chosen candidate (or a ranked list).

Workflow: A Workflow composes one or more Selectors and produces the full RoutingPlan. A single-dispatch Workflow uses one Selector. A prefill/decode Workflow uses two. One for the prefill pod, one for the decode pod. The decode Selector's decisions can depend on what the prefill Selector chose, because they share a RoutingContext.

Executor: An Executor takes the RoutingPlan and dispatches it. A single-dispatch Executor does one HTTP proxy. A sequential Executor does the prefill-wait-then-decode pattern, with hooks for request mutation, body injection, and response streaming at each step.

The composition scales linearly. Single-dispatch uses one Selector, and disaggregated prefill/decode uses two. A hypothetical three-step workflow (route through a preprocessor, then prefill, then decode) would use three. The Workflow abstraction handles the composition of any number of required Selectors. Adding a new execution pattern requires a different Workflow and a different Executor, not a new HTTP handler.

The execution layer: multi-step coordination

The Selector, Workflow, and Executor split is most valuable in disaggregated prefill/decode, where a single request is served by two pods, requiring additional coordination. Decode can't start until prefill has built the KV cache. Selections for both pods can happen concurrently (prefill and decode selection are independent decisions), but execution is sequential:

  1. Send the request to the prefill pod with max_tokens=1.
  2. Wait for prefill to finish.
  3. Send the request to the decode pod with a cache hint pointing to the prefill pod's cache.
  4. Stream the decode pod's tokens to the client.

A disaggregated Workflow uses two Selectors sharing one RoutingContext. Sharing one context lets the second decision build on the first. The prefill Selector records the pod it picks, and the decode Selector reads that value so it can prefer a decode pod close to the prefill pod, minimizing how far the cache must travel. The sequential Executor then carries out the steps above and streams the response back to the client.

Each Selector runs the same Prepare, Filter, Score, and Pick stages as single-dispatch routing, using the same plugins and scoring.

Whether disaggregation helps depends on the workload. The Hao AI Lab at UC San Diego published an 18-month retrospective on where disaggregation helps and where it doesn't. The short version: disaggregation wins on long-context, cache-cold workloads and loses on short-context, high-hit-rate ones.

When disaggregation does make sense, the framework uses the same plugins, the same RoutingContext, and the same scoring logic as single-dispatch routing. No parallel system is required.

Adding a new behavior

Adding a new routing behavior to this framework takes four steps.

Say we want a geographic proximity scorer that prefers pods in the same region as the client. We would have to implement the Scorer interface with a Score(ctx, candidates) []float64 method, register the plugin type with the framework's plugin registry, add the scorer to whichever profiles want it, and weight it against other scorers. Profile authors decide the weight, and the framework validates the composition at startup.

Three layers, one framework

Over these three posts, we demonstrated how Modular Cloud’s routing layer allows us to rapidly implement new routing optimizations. The data layer tracks which KV cache blocks live on which pods at microsecond query latency. The decision layer expresses routing behaviors as composable plugins rather than hard-coded algorithms. The execution layer coordinates multi-step flows when a single request touches multiple pods, as with disaggregated prefill/decode.

We plan to follow this series with a deep dive into inference scheduling. Routing decides which pod handles a request. Scheduling decides which requests run next, in what batch, against what cache state. This blog series will explore how we approach scheduling as a single system, allowing us to perform holistic optimizations for large-scale inference.

The routing layer described in this series is running in production on Modular Cloud today. If you're serving text, image, or video models and want to test our vertically integrated stack against your workloads, request access to Modular Cloud.

  • Person with blonde hair using a laptop with an Apple logo.

    Sign up today

    Signup to our Cloud Platform today to get started easily.

    Sign Up

  • Magnifying glass emoji with black handle and round clear lens.

    Browse open models

    Browse our model catalog, or deploy your own custom model

    Browse models

Sign up for our newsletter

Get all our latest news, announcements and updates delivered directly to your inbox. Unsubscribe at anytime.

Thanks for signing up to our newsletter! 🚀

Thank you,

Modular Sales Team

Oops! Something went wrong while submitting the form.