惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

AI
AI
博客园 - 叶小钗
Blog — PlanetScale
Blog — PlanetScale
Microsoft Azure Blog
Microsoft Azure Blog
Vercel News
Vercel News
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
MyScale Blog
MyScale Blog
大猫的无限游戏
大猫的无限游戏
A
About on SuperTechFans
量子位
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园 - 【当耐特】
Martin Fowler
Martin Fowler
阮一峰的网络日志
阮一峰的网络日志
D
Docker
Jina AI
Jina AI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
The Register - Security
The Register - Security
J
Java Code Geeks
S
SegmentFault 最新的问题
月光博客
月光博客
G
Google Developers Blog
美团技术团队
Last Week in AI
Last Week in AI
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
T
The Blog of Author Tim Ferriss
腾讯CDC
Recent Announcements
Recent Announcements
Recorded Future
Recorded Future
The Cloudflare Blog
有赞技术团队
有赞技术团队
博客园_首页
博客园 - 聂微东
人人都是产品经理
人人都是产品经理
B
Blog
I
InfoQ
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
F
Fortinet All Blogs
B
Blog RSS Feed
Engineering at Meta
Engineering at Meta
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Microsoft Security Blog
Microsoft Security Blog
MongoDB | Blog
MongoDB | Blog
爱范儿
爱范儿
D
DataBreaches.Net
F
Full Disclosure
M
MIT News - Artificial intelligence
博客园 - 司徒正美
H
Help Net Security

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure
No items found. · 2026-01-29 · via Modular Blog

January 29, 2026

Modular Team

Today we’re releasing Modular 26.1, a major step toward making high-performance AI computing easier to build, debug, and deploy across heterogeneous hardware. This release is focused squarely on developer velocity and programmability—helping advanced AI teams reduce time to market for their most important innovations.

Modular 26.1 centers on a new MAX Python API that simplifies building and deploying high-performance GenAI models across heterogeneous hardware. This release also improves the strength and ergonomics of our APIs within Mojo and includes new DevEx improvements to error reporting, language features to catch more bugs at compile time, and expanded Apple silicon GPU support.

26.1 release highlights include:

  • MAX Python API out of experimental – PyTorch-like modeling with eager mode for debugging, model.compile() for production
  • MAX LLM Book now stable – build transformers from scratch at llm.modular.com
  • Growing community contributions – Qwen3 embeddings, BERT, Mamba, visual generation pipelines, and more from external contributors
  • Mojo API and ergonomics improvements – including compile-time reflection, linear types, typed errors, better error messages

From Prototype to Production, Faster with MAX

We generally describe MAX as an AI modeling and serving framework that delivers state-of-the-art performance and cost efficiency for production inference. But one of MAX’s most distinctive strengths is its programmability and extensibility, and Modular 26.1 brings that into sharper focus.

The MAX Python API makes it easy to port custom GenAI models — often trained in PyTorch — into a high-performance format that runs across diverse hardware. With 26.1, the eager execution APIs graduate out of experimental, offering a PyTorch-like modeling interface backed by more robust documentation and a growing developer community.

The result is that MAX is no longer just a faster way to serve models — it’s a platform for building and serving GenAI models end to end, without sacrificing performance, portability, or control.

A More Natural Python Modeling Experience

In 26.1, MAX takes a major step toward feeling intuitive to users coming from PyTorch:

  • PyTorch-like modeling APIs are no longer experimental while retaining MAX’s portability and performance characteristics. See the model developer guide for more details.
  • Eager mode support reduces friction during experimentation and interactive development. This significantly improves the developer experience and makes MAX feel far more natural during debugging, experimentation, and model bring-up. Although we’re still improving compile time, his release meaningfully closes the gap between PyTorch’s eager UX and MAX’s performance-first execution model. Check out the tensor fundamental guide and get the latest performance improvements in the nightly builds.
  • Compile your MAX model for production by simply calling model.compile(). You’ll get all the benefits of ahead-of-time graph compilation with the full speed and memory efficiency benefits you can expect from the MAX graph compiler and its high-performance GPU kernels.

Together, these improvements bring MAX much closer to the “just works” experience users expect, while preserving the ability to scale seamlessly into fully compiled, production-grade execution when performance and efficiency matter most.

A Hands-On Guide to Building an LLM with MAX

Our comprehensive guide to building an LLM from scratch (the MAX LLM book) is now stable and maintained alongside API changes in the nightly builds. If you saw the experimental version, take another look—we completely updated the sequence so the code you write in each step builds upon the last one, until you've built a fully executable OpenAI LLM in MAX.

The MAX LLM Book walks through every component of a transformer—from tokenization and embeddings to attention and decoding—while teaching you how to express these ideas using the MAX Python APIs. It’s designed for two audiences:

  1. Developers who want a deep, concrete understanding of how transformers actually work.
  2. Practitioners who need to customize or extend models for real production use cases.

Each chapter includes executable code and detailed explanations, making it both a learning resource and a practical reference. Start building now.

Expanded Apple Silicon GPU Support

Building on our initial Apple silicon GPU support in 25.7, we've significantly expanded coverage in 26.1. Simple MAX graphs can compile and run on Apple silicon GPUs, and all the Mojo GPU puzzles now run on Apple GPUs (excluding NVIDIA-specific puzzles). Future updates will expand our support all the way to LLM inference on Apple Silicon GPUs. Please join our community if you'd like to help build into this support.

The latest features are landing regularly in our nightly releases, so keep up with those to get the best support. You can also help by contributing new Apple silicon GPU support to many of our existing open source Mojo kernels.

A Growing Community of MAX Models

The past few months have marked a turning point in MAX’s evolution — from an internal framework to a community-grown modeling platform. Contributors from across the ecosystem are extending MAX with new model architectures, performance improvements, and infrastructure enhancements that benefit everyone. Here are a few highlights:

  • Community member Sören Brunk contributed production-ready Qwen3 embedding support, including performance optimizations that match vLLM throughput, and infrastructure enhancements that improve MAX's architecture registry for all future multi-task models.
  • Ryan Wayne added BERT embedding model support specifically to replace his CUDA-dependent text embeddings infrastructure, demonstrating MAX's appeal for simplifying production ML stacks.
  • Tolga Cangoz is developing an ambitious Z-Image visual generation pipeline that will bring state-of-the-art diffusion-based image generation to MAX
  • The Qwerky AI team is exploring state-space models in MAX through the Mamba architecture. State-space models provide an intriguing alternative to traditional Transformer-based architectures, promising potentially faster inference and support for much longer context windows. Qwerky AI has been working on custom Mojo kernels, a new model architecture, and serving pipeline enhancements to support these novel models in MAX.
  • We’ve already open-sourced the MAX Python API, our entire GPU kernel library (for NVIDIA, AMD, and Apple silicon), all our model architectures, serving pipelines, and more. If you're interested in contributing model architectures, performance optimizations, or expanding MAX's capabilities, check out our contribution guidelines and join the conversation in our MAX forum.

    Mojo: More powerful and more ergonomic APIs

    Mojo 26.1 delivers a set of foundational language features that move the language meaningfully closer to Mojo 1.0. Our objective with 1.0 is to provide a language that’s capable of high-performance computing for diverse hardware but is still a joy to use.

    This release drives us toward that by incorporating research directions in language design for compile-time safety, improving the experience when you encounter errors, and enhancing the overall ergonomics of Mojo code. Specifically, this release adds:

    • Compile-time reflection: A new system for compile-time reflection on types further enhances Mojo’s already versatile metaprogramming system. This powerful new capability allows, among other things, automatic conformance to traits that are defined in libraries. For example: automatic equatability, JSON serialization, or CLI argument parsing. Members of the Mojo community are already doing exciting things with this using nightly builds.
    • Explicitly destroyed types: Mojo now supports explicitly destroyed types (aka "Linear Types" in programming language jargon), enabling compile-time guarantees that certain values cannot be forgotten. This is an example of Mojo drawing from leading research languages, exceeding the safety and power of most other widely used languages.
    • Typed errors: Functions can now raise types other than Error, enabling error-handling on GPUs and embedded systems without overhead. Mojo also supports “parametric raise-ability”, enabling powerful and concise expression of generic algorithms.
    • Improved error messages: We’ve fixed the biggest user-painpoint in Mojo error messages, where it would complain that it couldn’t infer a parameter instead of saying types don’t match. Mojo now also will diff two similar types and tell you which sub-parameter disagree when the types are almost the same!
    • Mojo LSP server improvements: LSP users will see dramatically reduced CPU usage when typing, the compiler has experimental support for applying fix-it suggestions.

    As always, there are far more interesting and useful additions to Mojo and its tooling than those listed above, so check out the full Mojo changelog.

    Try 26.1 Today

    Get everything you need to build LLMs with MAX and write high-performance GPU kernels with Mojo by installing the modular package with pip, uv, pixi, or conda. For more details, see our quickstart guide.

    shell

    uv pip install modular

    Once you’re set up, you can explore everything in Modular Platform 26.1:

    For a complete breakdown, see the MAX and Mojo 26.1 changelogs.

    Modular 26.1 is another step toward making high-performance AI development accessible to everyone. Your questions, feedback, and contributions directly shape the platform — join the discussion on our forum and report any issues or feature requests.

    We can’t wait to see what you build!

    Discover what Modular can do for you

    Request a demo

    • Person with blonde hair using a laptop with an Apple logo.

      Sign up today

      Signup to our Cloud Platform today to get started easily.

      Sign Up

    • Magnifying glass emoji with black handle and round clear lens.

      Browse open models

      Browse our model catalog, or deploy your own custom model

      Browse models

    Sign up for our newsletter

    Get all our latest news, announcements and updates delivered directly to your inbox. Unsubscribe at anytime.

    Thanks for signing up to our newsletter! 🚀

    Thank you,

    Modular Sales Team

    Oops! Something went wrong while submitting the form.