惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

A
Arctic Wolf
博客园 - 聂微东
F
Fortinet All Blogs
云风的 BLOG
云风的 BLOG
小众软件
小众软件
V
Visual Studio Blog
博客园 - 三生石上(FineUI控件)
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Apple Machine Learning Research
Apple Machine Learning Research
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Cloudflare Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园_首页
L
LangChain Blog
A
About on SuperTechFans
阮一峰的网络日志
阮一峰的网络日志
I
Intezer
T
The Blog of Author Tim Ferriss
Security Latest
Security Latest
C
CXSECURITY Database RSS Feed - CXSecurity.com
Know Your Adversary
Know Your Adversary
Simon Willison's Weblog
Simon Willison's Weblog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
P
Palo Alto Networks Blog
Scott Helme
Scott Helme
S
Secure Thoughts
Spread Privacy
Spread Privacy
T
Threat Research - Cisco Blogs
Attack and Defense Labs
Attack and Defense Labs
P
Privacy & Cybersecurity Law Blog
O
OpenAI News
H
Heimdal Security Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
Help Net Security
Help Net Security
C
Cyber Attacks, Cyber Crime and Cyber Security
Blog — PlanetScale
Blog — PlanetScale
GbyAI
GbyAI
G
Google Developers Blog
博客园 - Franky
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
K
Kaspersky official blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
T
Tor Project blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Tenable Blog
Google Online Security Blog
Google Online Security Blog
PCI Perspectives
PCI Perspectives

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Three trends from MLSys 2026
No items found. · 2026-05-29 · via Modular Blog

May 29, 2026

Michael Dunn-OConnor

Brian Zhang

Shouzheng Liu

MLSys 2026 provided an excellent overview of the current state of inference across research and industry. With six sessions on LLM serving this year (twice as many as last year) the program covered opportunities and challenges at the core of Modular’s mission. Modular was glad to sponsor the conference and our Chief Scientist Abdul Dakkak presented a lightning talk on the state-of-the-art performance of the Modular platform. Our team who attended MLSys noted three trends that stood out across the talks, posters, and keynotes. These are all topics that Modular has been addressing from first principles, with the advantage of our unique stack.

Trend 1: Agents are writing everything from kernels to systems

Monday’s keynote set the tone. Mark Saroufim's When AI Starts Writing Systems Code showed examples of novice kernel developers using AI agents to write kernels that could place them near the top of competitive hackathons. He then comically undercut some of these agentic achievements by demonstrating how agents would cheat the benchmarks and optimize results that would never generalize outside of the test cases provided. Rather than waiting for a generation of agents that always play by the rules, Saroufim outlined the need for “zero trust” verification by creating comprehensive enough benchmarks to not rely on good faith submissions.

Lidong Zhou's Tuesday keynote The Next Horizon of Systems: From MLSys to System Intelligence argued that the systems community needs to plan for AI agents as the primary authors of low-level code. Zhou presented a Rust microkernel (Nanvix) where AI-generated specifications and proofs are verified module by module, with a pass rate on a 150-task proof generation benchmark that climbed from 2 percent (GPT-4o, prompt-based) to 91.3 percent (fine-tuned LLaMA-3.1 8B with self-debugging). The talk also documented the shortcuts the model takes when it cannot complete a proof: wrapping code in external_body to bypass the verifier, planting false postconditions, and shifting proof burden to callers.

The shared conclusion of these talks was that agentic engineering requires substantially greater rigor in specification, design, and validation.

Subsequent talks showcased specific applications of agentic engineering across kernels and systems. AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization describes a closed-loop system where an LLM proposes accelerator kernel variants, profiles them, and feeds the results back to itself. FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems in the ML for Systems session frames this as a feedback loop: the benchmark exists to give the agents something to optimize against.

The kernel-author pain that motivates all of this also showed up directly. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling moved its implementation off CUDA C++ templates and into CuTe-DSL embedded in Python with the explicit goal of letting downstream developers extend the kernel without modifying the core framework. HipKittens: Fast and Furious AMD Kernels and ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels both argued for a simpler set of abstractions for kernel engineering. The shared assumption is that human kernel authors are not going to write a new template forest for every accelerator generation, and that abstraction benefits agents at least as much as human developers.

Modular’s Solution

Modular engineers and the community are already writing Mojo code with agents and leveraging language features that make it optimal for agentic engineering. Mojo’s robust type system, efficient compilation, and clear error messages support the tight feedback loops of agentic development with human verification. Our blog post on Translating to Mojo via AI Agents provides a practical guide to this workflow. The official modular/skills package plugs into Claude Code, Cursor, and other coding agents and corrects any misconceptions and out-of-date patterns that models may produce. In the post, Brad Larson walks an agent from a CUDA softmax kernel (Szymon Ożóg's FastSoftmax) to a portable Mojo version that runs on NVIDIA, AMD, and Apple silicon in a single session. Automatika Robotics did the same to autonomous-navigation kernels for their EMOS / kompass-core workload and reported 15.973 ms versus a 16.358 ms SYCL/CUDA baseline on the agent's first pass, with no Mojo-side optimization. Another blog post from Ehsan Kermani demonstrates effective agentic engineering to create Mojo libraries that are thoroughly tested and meet a real community need.

The second connection is on the kernel side. The composable abstractions the Modular kernel team created are documented in our Structured Mojo Kernels blog series. This pattern breaks production kernels into three components (TileIO, TilePipeline, TileOp) with context managers that make incorrect synchronization unrepresentable. Rewriting the B200 matmul cut 14,683 lines down to 7,634 with consistent performance at 1770 TFLOPS. The Structured Mojo Kernels patterns apply to both NVIDIA and AMD and they are open source in the Modular repository. This is the direct answer to the FlashAttention-4 maintenance argument: a kernel layer that is concise enough for a human (or an agent) to reason about, with abstractions that hold up under hardware changes, without a tradeoff in performance.

Trend 2: KV cache became the dominant subsystem

KV cache was the single most discussed topic at the conference. The LLM Serving 3 session on Wednesday afternoon was effectively a KV cache session, with six papers in a row presented on the topic.

Yuhan Liu's talk on LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference made the architectural argument directly. The paper treats the KV cache as a first-class data structure rather than an internal byproduct of inference, and reports that real-world deployments have made KV cache usage outgrow GPU memory: over five weeks of telemetry, the dominant majority of stored KV cache no longer fit on the GPUs, while reuses per token grew by more than 19 percent across users. The system now supports eight storage backends (NFS, WEKA, GPU-Direct Storage, Mooncake Store, NIXL, S3, InfiniStore, Valkey) across four processor types (NVIDIA, AMD, Ascend, TPU) and two inference engines (vLLM, SGLang). That breadth of coverage is essential. The cache layer has to be portable because the storage and compute hardware underneath it is not homogenous.

Read together, these papers describe the same shift. The KV cache is increasingly distributed, heterogeneous, and complicated, with placement, transfer, eviction, and reuse policies that span GPU memory, host DRAM, local disk, and distributed storage. The trend has gone far enough that there is now a market for custom silicon dedicated to the cache: Netpreme is building a Memory Processing Unit with networked memory tiering whose pitch is extending XPU memory by 100x. When a subsystem starts pulling in its own hardware vendors, it is no longer an implementation detail.

The other techniques for optimizing the cache are quantization and sparsity, used together to tame the memory demand directly. Kitty: Accurate and Efficient 2-bit KV Cache Quantization with Dynamic Channel-wise Precision Boost improves throughput by enabling 8x larger batches under the same memory budget. The SGLang team's HiSparse: Turbocharging Sparse Attention with Hierarchical Memory (which Christos Kozyrakis flagged in his Friday keynote on inference efficiency) pairs sparse attention with a hierarchical memory tier: the GPU keeps a hot device buffer of frequently accessed KV regions while inactive entries offload to host memory. The result is up to 5x throughput on long-context workloads with GLM-5.1-FP8. Quantization shrinks what the cache costs to hold, sparsity shrinks what the engine has to attend over, and pairing them is the practical answer for inference at long context.

The common thread is that the cache is being pulled out of the engine and turned into a distributed system in its own right, with its own scheduling, its own storage tiering, and its own reuse semantics.

Modular’s Solution

This is the trend Brian Zhang described in his blog post The Five Eras of KVCache. The post traces the cache's evolution from a local in-engine optimization in 2023, through PagedAttention, prefix caching, and offloading, ending in an era where the cache is unified, distributed, and composable across heterogeneous infrastructure. The MLSys 2026 papers document the urgency of these optimizations, which were built into Modular Cloud from the start.

For context on how distributed KV cache works at the cluster level, check out the inference routing series on the Modular blog. Why LLM Inference Needs a New Kind of Router makes the case that routing decisions and KV cache state are inseparable, that the router has to be cache-aware, and that cumulative chaining and bitmap indexing are what make that practical at scale. Zooming out, Kyle Caverly’s walkthrough of MAX Serve from prompt to response is a great primer to where the KV cache fits into the overall inference platform. The MLSys papers describe a multitude of approaches to efficient KV cache management. These papers validate the urgency of what Modular has been building for years. Modular Cloud implements a composable approach to large-scale distributed inference that works across hardware and is flexible to ever-changing optimizations.

Trend 3: Inference workloads are leveraging heterogeneous hardware

Heterogeneity at MLSys 2026 ran through all six LLM Serving sessions and both Industry Track sessions on serving. Esha Choukse's Day 1 invited talk Beyond Model Serving: Cross-Stack Co-Design for Agentic Systems set the frame: hardware diversity is a prerequisite for efficiently serving interactive, multimodal, and agentic systems.

Meta's Industry Track paper Optimizing Deployment Configurations for LLM Inference showed the benefits of heterogeneity even in single-model deployments. The paper documents 15 to 25 percent TCO (total cost of ownership) improvements from running prefill on one accelerator type and decode on another, because prefill is compute-bound (favoring high FLOP/s) and decode is memory-bandwidth-bound (favoring high HBM bandwidth). NVIDIA's Beyond the Buzz: A Pragmatic Take on Inference Disaggregation in the same session conceded that prefill-decode disaggregation is real and valuable, but only if rate matching, KV transfer, cache routing, and elastic scaling are solved simultaneously. BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization treated the choice of which model runs on which accelerator as a search problem. TriInfer: Hybrid EPD Disaggregation for Efficient Multimodal Large Language Model Inference in the Multimodal and Generative Models session extended disaggregation past two phases into encode-prefill-decode for multimodal workloads.

The hardware-specific kernel work showed why this matters. SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips found that existing offloading frameworks utilize the NVLink-C2C interconnect on GH200 at less than 5 percent of its 900 GB/s capacity because they treat it as if it were PCIe. The bottleneck, the paper concluded, is in the software stack. SHIP: SRAM-Based Huge Inference Pipelines for Fast LLM Serving described Groq's LPU-based serving stack, where the entire model fits in SRAM and the compiler statically schedules collective communication at cycle granularity. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference depends on PyTorch's SymmetricMemory API and NVLink4's NVSHARP engines, both of which exist precisely because vendor-specific collective communication paths no longer fit how kernel authors want to work.

Heterogeneous hardware provides an edge in multimodal inference and cost optimization, but it can be constrained by vendor-specific software stacks.

Modular’s Solution

MAX is built from the ground up for hardware plurality. One container runs on NVIDIA GPUs, AMD GPUs, and CPUs. The kernel layer is written in Mojo, which compiles to each hardware target rather than wrapping vendor-specific libraries. The runtime is hardware-aware but not hardware-specific.

When run on real workloads (e.g. Hippocratic AI), MAX delivers sub-500ms mean TTFT, roughly 30 percent faster P99 end-to-end, and 22 percent faster mean end-to-end against SGLang on NVIDIA B300 GPUs for 400B-plus parameter models. When compared against vLLM on B200, the Modular stack is 5.5x faster P50 TTFT on Kimi-K2.5 and 2.5x faster P99 TTFT on Gemma-4-31B-it, with 1.5x throughput on both. The Flux.2-dev image generation workload runs 6.9x faster than PyTorch Diffusers with torch.compile on B200 and 3.8x faster on AMD MI355x, in the same software stack. You can read more about MAX’s state-of-the-art performance in Modular’s MLSys lightning talk.

Efficient inference serving requires getting the full potential of all available hardware to meet ever-changing industry needs. The MLSys 2026 sessions look at components of the inference stack and optimize them in isolation: kernels, KVCache, serving, etc. Modular rewrote the entire stack as a single system, allowing us to perform holistic optimizations from kernel to cloud. This approach allows us to be at the leading edge of research and achieve optimizations that cut across the entire stack.

Thanks to all those who stopped at the Modular booth to discuss our stack and open roles at Modular. Please reach out on our forums if you’re interested in contributing to our open-source projects or collaborating on research. We look forward to MLSys 2027, both to see how research has developed and to share our own ongoing innovations. See you next year!


  • Person with blonde hair using a laptop with an Apple logo.

    Sign up today

    Signup to our Cloud Platform today to get started easily.

    Sign Up

  • Magnifying glass emoji with black handle and round clear lens.

    Browse open models

    Browse our model catalog, or deploy your own custom model

    Browse models

Sign up for our newsletter

Get all our latest news, announcements and updates delivered directly to your inbox. Unsubscribe at anytime.

Thanks for signing up to our newsletter! 🚀

Thank you,

Modular Sales Team

Oops! Something went wrong while submitting the form.