惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Recent Announcements
Recent Announcements
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Last Week in AI
Last Week in AI
Scott Helme
Scott Helme
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
L
LINUX DO - 最新话题
S
Security @ Cisco Blogs
Webroot Blog
Webroot Blog
S
Security Affairs
H
Hacker News: Front Page
TaoSecurity Blog
TaoSecurity Blog
W
WeLiveSecurity
G
GRAHAM CLULEY
T
Tenable Blog
Schneier on Security
Schneier on Security
S
Securelist
Cyberwarzone
Cyberwarzone
P
Privacy International News Feed
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
S
Schneier on Security
Hacker News - Newest:
Hacker News - Newest: "LLM"
Recent Commits to openclaw:main
Recent Commits to openclaw:main
O
OpenAI News
N
News and Events Feed by Topic
AWS News Blog
AWS News Blog
C
Cisco Blogs
T
Threat Research - Cisco Blogs
S
Secure Thoughts
大猫的无限游戏
大猫的无限游戏
C
Check Point Blog
The GitHub Blog
The GitHub Blog
G
Google Developers Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
美团技术团队
Martin Fowler
Martin Fowler
Microsoft Security Blog
Microsoft Security Blog
L
LangChain Blog
Apple Machine Learning Research
Apple Machine Learning Research
爱范儿
爱范儿
D
DataBreaches.Net
博客园_首页
MyScale Blog
MyScale Blog
博客园 - 叶小钗
博客园 - 三生石上(FineUI控件)
P
Proofpoint News Feed
J
Java Code Geeks
SecWiki News
SecWiki News
P
Palo Alto Networks Blog
Know Your Adversary
Know Your Adversary
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience
No items found. · 2025-11-20 · via Modular Blog

Today, we’re excited to release Modular Platform 25.7, an update that deepens our vision of a unified, high-performance compute layer for AI. With a fully open MAX Python API, an experimental next-generation modeling API, expanded hardware support for NVIDIA Grace superchips, and a safer, more capable Mojo GPU programming experience, this release moves us closer to an ecosystem where developers spend less time fighting infrastructure and more time advancing what AI can do.

MAX: Faster, Easier, and More Open

MAX continues to evolve into the fastest cross-vendor inference framework available. In 25.7, we’ve deepened MAX across three fronts: openness, modeling flexibility, and production-grade control.

MAX: Great way to build high performance inference models

With this release, the entire Python interface to MAX is now fully open-source. This gives developers:

  • Visibility into how MAX models are built, executed, and served across hardware vendors.
  • Real examples of how features like MAX ↔ PyTorch interoperability were implemented.
  • Greater clarity on internal abstractions through newly-visible unit tests, making it easier to reason about performance and correctness.
  • Better bug reporting and contribution paths, enabling faster iteration across the ecosystem.

We've been working hard to refine our APIs, enable a first class "eager" programming model, and really lean into the power of our next generation technology stack.  It is now time to open up to more developers, so we are open sourcing the full API and starting to teach how to build MAX-native models.

The New (Experimental) Model API and Development Workflow

25.7 introduces a redesigned Model API consisting of the new max.nn.module_v3 and max.experimental.tensor – together, they deliver our most significant modeling upgrade since MAX launched.

Our new API gives developers an experience more aligned with the higher-level model abstractions they are familiar with, bringing better modeling ergonomics to MAX. You write models the way you’re used to, and MAX handles the heavy lifting underneath:

  • PyTorch-like syntax with minimal cognitive overhead for defining and composing models
  • Lazy evaluation with eager semantics that matches PyTorch’s eager mode for enhanced debugging
  • No more manual graph construction or session-load-run ceremonies
  • Intuitive weight loading with load_static_dict() similar to PyTorch
  • Pure MAX/Mojo stack: no dependencies on PyTorch, NumPy or other frameworks

When you’re ready for production, just add model.compile(input_type) to unlock the performance benefits of full model compilation. This enables full ahead-of-time graph compilation with the full speed and memory efficiency benefits you come to expect from MAX.

python

import max.nn.module_v3 as nn
from max.experimental import functional as F, tensor, random

class MyModel(nn.Module):
    def __init__(self):
        self.fc = nn.Linear(10, 10)
        self.proj = nn.Linear(10, 1)

    def __call__(self, x: tensor.Tensor) -> tensor.Tensor:
        x = self.fc(x)
        x = F.gelu(x)
        x = self.proj(x)
        return x

model = MyModel()
input = random.normal((10, 10))
# for development and debugging
max_output = model(input)
# for production
compiled_model = model.compile(input.type)
max_output = compile_model(input)

The latest API dramatically speeds up model development, debugging, and customization – especially for teams extending frontier-scale architectures. To get started using the new model APIs, check out our new online book to build an LLM from scratch with MAX. This is a guided lesson on building LLMs starting with GPT-2 that explains each component of the transformer model along the way. We’d love your feedback on forum.modular.com, and report bugs or feature requests on github.com/modular/modular/issues.

Expanded Model and Hardware Support

25.7 significantly broadens MAX’s reach across both models and hardware, ensuring high-performance inference on the newest accelerators and system architectures.

  • MAX now supports bfloat16 models running on GPUs attached to ARM CPU hosts, including Grace Hopper (GH200) and Grace Blackwell (GB200) systems. This unlocks higher performance and lower power consumption on next-generation NVIDIA platforms built around the Grace superchip architecture.
  • MAX delivers outsized performance wins on vision models, unlocking +30-80% additional throughput on Qwen2.5-VL compared to 25.6 and over 2x performance compared to vLLM.
  • Early support for GPT-OSS has landed, with MXFP4 integration coming soon to unlock major 4-bit performance gains.

Together, these updates deepen MAX’s position as the fastest fully portable inference engine. Stay tuned for very significant model and performance updates in our next release.

Dynamic LoRA for Real-Time Specialization

25.7 introduces support for Dynamic LoRA, initially tuned for speech and low-latency real-time workloads. This allows developers to:

  • Hot-swap LoRA adapters during runtime
  • Personalize models without restarting the server
  • Use LoRA for speaker-specific or domain-specific behaviors
  • Achieve near-zero-overhead tuning for production workloads

This is essential for those that require rapid, high-fidelity model adaptation with minimal cost. Today, dynamic LoRA in MAX powers real-time applications like Inworld AI’s voice cloning, but we’ll make this broadly available for more models shortly.

Mojo: Safer, better Apple GPU support, and more

Mojo is evolving into the simplest, most portable way to write high-performance GPU code. In the process, we are enhancing our hardware support, adding new language features, and improving the overall safety of the language.

Enhanced Apple Silicon GPU Support

We launched initial support for Apple silicon GPUs in 25.6, and in 25.7 we’ve been expanding the coverage of fundamental intrinsics and capabilities in these GPUs. This has included

  • Support for synchronization primitives
  • Access to shared memory
  • Mapping of warp operations to Metal equivalents

As a result, you can now run 20 of our popular Mojo GPU puzzles on Apple silicon GPUs in 25.7, up from 5 GPU puzzles in our last release.

Mojo now provides one of the cleanest paths to developing AI kernels on Apple hardware — no vendor-specific APIs required.

Safer GPU Programming

Correctness and illegal memory access errors in GPU code can be subtle and extremely hard to diagnose. Catching these earlier and making them more obvious saves an enormous amount of time when writing new kernels or AI models. In 25.7, some significant improvements have been made to the Mojo language and libraries to catch common GPU programming issues:

  • GPU functions now have strong type checking using enqueue_function_checked(), identifying many potential crashes and bad memory accesses before they happen.
  • A newly-reworked UnsafePointer type that fixes some frequent issues with lifetimes and accidentally unsafe changes.
  • No more implicit conversions between Int and UInt types, a source of common numerical correctness issues in GPU kernels.
  • Much better error messages around constraint failures: traces, line numbers, and printed parameter values.
  • Expanded support for using Address Sanitizer with Mojo code to identify memory leaks and illegal accesses.

Beyond language and library features, a new TestSuite module lays the foundation for better unit testing in Mojo. It provides a clean interface for test results, automatic test discovery, and more.

These safety improvements align with our long-term goal: Mojo makes GPU programming feel like writing high-level, safe systems code.

Try the Latest Updates and Join the Community!

The best way to experience everything in 25.7 is to try it yourself.

You can get everything you need to deploy an LLM with MAX, write GPU kernels with Mojo, and build models with the experimental API today by installing the modular package with pip, uv, pixi, or conda. For more details, see our quickstart guide.

shell

pip install modular

Once you’re set up, you can:

  • Explore the fully open-source MAX Python API on GitHub
  • Walk through our guided lessons to to build a transformer LLM from scratch
  • Review the complete list of changes in the MAX and Mojo changelogs
  • Browse the updated Guides section on docs.modular.com, now including all developer guides and tutorials for the Modular Platform

25.7 is a major step toward a unified, fully portable compute layer for AI – and we’re building it in the open. Your questions, feedback, and contributions directly shape the platform, so join the discussion and report any issues or request features.

We’re excited to see what you build. Come experiment, contribute, and help define the future of high-performance, hardware-portable AI with us.