惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Jina AI
Jina AI
V
Visual Studio Blog
博客园 - 司徒正美
TaoSecurity Blog
TaoSecurity Blog
博客园 - 聂微东
IT之家
IT之家
博客园_首页
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - Franky
雷峰网
雷峰网
罗磊的独立博客
S
Schneier on Security
C
Cybersecurity and Infrastructure Security Agency CISA
The Cloudflare Blog
T
Tailwind CSS Blog
B
Blog RSS Feed
H
Help Net Security
T
The Blog of Author Tim Ferriss
C
CXSECURITY Database RSS Feed - CXSecurity.com
T
Threatpost
C
CERT Recently Published Vulnerability Notes
博客园 - 三生石上(FineUI控件)
P
Palo Alto Networks Blog
I
Intezer
G
GRAHAM CLULEY
Engineering at Meta
Engineering at Meta
S
Securelist
J
Java Code Geeks
V
V2EX
Y
Y Combinator Blog
Simon Willison's Weblog
Simon Willison's Weblog
L
LINUX DO - 热门话题
云风的 BLOG
云风的 BLOG
Spread Privacy
Spread Privacy
MongoDB | Blog
MongoDB | Blog
P
Privacy International News Feed
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
B
Blog
Forbes - Security
Forbes - Security
Google Online Security Blog
Google Online Security Blog
Help Net Security
Help Net Security
S
SegmentFault 最新的问题
N
Netflix TechBlog - Medium
Webroot Blog
Webroot Blog
Microsoft Security Blog
Microsoft Security Blog
SecWiki News
SecWiki News
Scott Helme
Scott Helme
aimingoo的专栏
aimingoo的专栏
N
News and Events Feed by Topic

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
MAX 25.2: Unleash the power of your H200's–without CUDA!
No items found. · 2025-03-25 · via Modular Blog

We’re excited to announce MAX 25.2, a major update that unlocks industry-leading performance on the largest language models–built from the ground up without CUDA*.

This release builds on the momentum from MAX 25.1 released just a month ago, and delivers critical features to power faster, more responsive, and more customizable GenAI deployments at scale:

  1. State of the art H100 and H200 performance: with support for more than 500 GenAI models.
  2. Multi-GPU support: seamlessly run large LLMs that exceed single-GPU memory limits.
  3. Enhanced LLM Serving: improved scheduling, batching, and caching further improves TCO and performance–MAX is now 12% faster than vLLM 0.8 on Sonnet benchmarks with the same numerics.
  4. Ultra-slim containers for rapid deployment “without CUDA:” 80% smaller than NVIDIA containers–ideal for faster deployments to production environments.
  5. Unlock GPU Programming with Mojo 🔥: MAX allows you to write custom, high-performance GPU code in Mojo with direct access to NVIDIA GPUs.
  6. Advanced features: including GPTQ quantization to run the biggest models efficiently.

MAX 25.2 is a milestone in our mission to build a CUDA-free GenAI model inference platform that gives you uncompromising performance and control on a single heterogeneous stack.

👉 Dive into what’s new and get started today!

*MAX Requires the NVIDIA hardware driver for GPU access.

Multi-GPU + H100/H200 support: run massive LLMs with industry-leading performance

This release's headline feature is full multi-GPU support on NVIDIA H100 and H200s, allowing you to run even larger language models with incredible performance. You can now deploy Llama-3.3-70B-Instruct across multiple GPUs directly from MAX Builds.

After installing max-pipelines, try running a 70B parameter model in bfloat16 on 4 GPUs with a simple command:

Bash

max-pipelines generate \ --model-path=meta-llama/Llama-3.3-70B-Instruct \ --quantization-encoding bfloat16 \ --devices gpu:0,1,2,3 \ --prompt="Design a self-sustaining colony on Neptune's moon Triton..."

We rebuilt the entire AI stack from the bottom up–including all of the software that runs on the GPU–in order to bring a simple “it just works” experience to AI, without removing your power and control over AI. To do this, we implemented all of the GenAI GPU algorithms for the complex H100/H200 architecture, with performance that meets and beats the NVIDIA libraries. You don’t need to worry about this, but it does mean you can forget about CUDA version mismatches and its legacy weight!

More than 500 preconfigured GenAI models

But that's not all! We've also expanded our supported model architectures to include:

The MAX Builds repository now hosts over 500 preconfigured MAX models ready for immediate deployment. Whatever your AI use case, we've got you covered!

__wf_reserved_inherit

Beyond model generality, MAX transparently supports many advanced features. One example is GPTQ quantization for models–just specify the quantized weights, and MAX handles the rest. For Llama 3.1 70B, this reduces the total memory consumption of this model from about 140 GB to 35 GB, making MAX even more accessible on memory-limited devices. You can use max-pipelines to run Llama 3.1 70B using int4-quantized GPTQ weights:

Bash

max-pipelines generate \ --model-path hugging-quants/Meta-Llama-3.1-70B-Instruct-GPTQ-INT4 \ --prompt "Why is the sky blue?" \ --max-batch-size 1 --max-length 10000

Please see the detailed release notes for more details on this and other cool technology.

LLM Serving improvements

This release builds on our LLM serving capabilities to add:

  • Prefix cache-aware batch scheduling: The scheduler can now create larger batches when many prompt tokens are cached, improving throughput by up to 10% in some benchmarks.
  • In-flight batching: allowing token generation requests to be scheduled alongside context encoding, reducing inter-token latency.
  • Copy-on-write KV blocks: integrated into Paged Attention with Prefix Caching, improving performance by optimizing cache hits.

These improvements give you better performance and capabilities that continue to "just work."

Slim Docker container: deploy faster than ever

Speed matters not just in inference but also in deployment. Because MAX models are built without CUDA, we don't have to carry its weight. Our new slim Docker container slashes the footprint of the serving container to just 1.3 GB compressed, enabling lightning-fast deployments even for your largest models. Get your AI applications up and running in record time without compromising performance.

__wf_reserved_inherit

Simplified GPU programming with Mojo🔥

Beyond the the Open-AI compatible endpoint, MAX is also great for AI researchers and developers who want to unlock the full power of the GPU for novel applications. Our mission is to make "easy things work" without taking power or control away from you–we want to give you superpowers over AI.

MAX and Mojo have fundamentally reimagined GPU programming. If you're familiar with CUDA C++, you'll feel right at home with a powerful GPU programming model that offers all the benefits of Mojo's modern language features.

Check out our examples on MAX Builds, which recreate the first few chapters from the popular textbook "Programming Massively Parallel Processors," but translated into Mojo. From the basic principles of vector addition to advanced patterns like high-performance matrix multiplication, we have examples to help you master device-independent GPU programming. Examples include:

You can download these recipes directly using Magic, making starting your GPU programming journey even easier. Also, don't miss the first chapter of our comprehensive guide to GPU programming in the MAX documentation!

__wf_reserved_inherit

Experience the future of AI inference today!

MAX 25.2 represents a significant leap forward in our mission to make high-performance AI accessible to everyone. Whether running billion-parameter models or building custom GPU-accelerated applications, MAX provides the tools and performance you need.

Ready to get started? Download the latest release, check out our expanded documentation, or join our community to share your experiences and learn from other MAX users.

AI doesn't need to be held back by CUDA–check out MAX and Mojo🔥 today!