惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

W
WeLiveSecurity
The Last Watchdog
The Last Watchdog
Application and Cybersecurity Blog
Application and Cybersecurity Blog
G
Google Developers Blog
博客园 - 叶小钗
雷峰网
雷峰网
人人都是产品经理
人人都是产品经理
博客园_首页
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 三生石上(FineUI控件)
Help Net Security
Help Net Security
Cloudbric
Cloudbric
AI
AI
N
News | PayPal Newsroom
博客园 - 聂微东
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 【当耐特】
Forbes - Security
Forbes - Security
美团技术团队
Stack Overflow Blog
Stack Overflow Blog
SecWiki News
SecWiki News
H
Heimdal Security Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
MyScale Blog
MyScale Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
P
Proofpoint News Feed
S
Security @ Cisco Blogs
Google DeepMind News
Google DeepMind News
V
V2EX
大猫的无限游戏
大猫的无限游戏
阮一峰的网络日志
阮一峰的网络日志
S
Security Affairs
L
LangChain Blog
The Hacker News
The Hacker News
F
Full Disclosure
aimingoo的专栏
aimingoo的专栏
Hacker News - Newest:
Hacker News - Newest: "LLM"
腾讯CDC
Webroot Blog
Webroot Blog
A
About on SuperTechFans
H
Hacker News: Front Page
Cyberwarzone
Cyberwarzone
WordPress大学
WordPress大学
L
LINUX DO - 热门话题
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Attack and Defense Labs
Attack and Defense Labs
M
MIT News - Artificial intelligence

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Introducing MAX 24.6: A GPU Native Generative AI Platform
No items found. · 2024-12-17 · via Modular Blog

Three years ago we set out to redefine how AI is developed and deployed. Our goal wasn’t simply to improve existing systems, but to rebuild AI infrastructure from the ground up to deliver a more performant, programmable, and portable infrastructure platform. We recognized that to address today’s challenges and stay ahead of rapid technological evolution, we needed to completely rethink the AI stack from first principles.

The arrival of large-scale Generative AI transformed the very nature of AI infrastructure. To meet its rapidly growing resource demands, GenAI requires innovations from the lowest levels of GPU programming all the way up to the serving layers, something that only Modular is positioned to do.

Today, we’re announcing the first step in this journey to meet these challenges with MAX 24.6, featuring a preview of MAX GPU. This GPU release showcases the power of the MAX platform, and is just the beginning of the advancements we’ll bring to AI infrastructure as we head into 2025. We’d love for you to follow along, and develop with us as we continue to advance AI infrastructure for the world in the coming months.

Introducing MAX GPU: A new GenAI native serving stack

At the heart of the MAX 24.6 release is MAX GPU– the first vertically integrated Generative AI serving stack that eliminates the dependency on vendor-specific computation libraries like NVIDIA’s CUDA.

MAX GPU is built on two groundbreaking technologies. The first is MAX Engine, a high-performance AI model compiler and runtime built with innovative Mojo GPU kernels for NVIDIA GPUs–free from CUDA kernel dependencies. The second is MAX Serve, a sophisticated Python-native serving layer specifically engineered for LLM applications. MAX Serve expertly handles complex request batching and scheduling, delivering consistent and reliable performance, even under heavy workloads.

Unlike existing tools that address only specific parts of the AI workflow, MAX is designed to support the entire development experience–from initial experimentation, through deployment, and to production. MAX offers a unified platform for exploring new models, testing and optimization, and delivery of high-performance inference capabilities that before now required piecing together legacy technology.

__wf_reserved_inherit

Comparison between vLLM and MAX architecture

We cannot overstate how important this diagram is to our mission and vision of simplifying and making AI infrastructure accessible to everyone. We have strived to vastly reduce complexity across the AI infrastructure stack, and as you can see above, MAX reduces the incredible fragmentation of today's ecosystem. As a developer, it has become enormously challenging to navigate this landscape with such a vast array of differing technologies. We have been building for a new approach, and want to ensure the whole world can build with us.

We can now ship an uncompressed Docker container, without CUDA toolkit for NVIDIA GPUs, that drops our total size to under 3.7 GB vs. vLLM container which is 10.6GB - a 65% reduction. For customers that don't need PyTorch, and want to use MAX Graphs only, it drops even further to just 2.83GB. When we compress this, it's under 1GB in size.

Enterprise-grade development to deployment flexibility

MAX Engine enables flexible inference deployments across multiple hardware platforms, allowing developers to experiment locally on laptops and scale seamlessly into production cloud environments. Combined with MAX Serve’s native Hugging Face model support, teams can rapidly develop, test, and deploy any PyTorch LLM. Custom weight support, including Llama Guard integration, further enables developers to tailor models for specific tasks.

When it’s time to put a model into production, MAX Serve provides an OpenAI-compatible client API, shipped in a compact Docker container that works on NVIDIA platforms. Teams can then deploy models across all major clouds, including AWS, GCP, and Azure, with options for both direct VM deployment and enterprise-scale Kubernetes orchestration. This flexibility ensures that you can host your own models securely, keeping your GenAI infrastructure fully under your control.

At the heart of this workflow is Magic, Modular’s command-line tool that streamlines the entire MAX lifecycle. Magic handles everything from installation and environment management, to development and deployment–making it easier to manage your AI infrastructure. Read more about Magic here.

High-performance GenAI models and hardware portability

We’re also expanding MAX’s power with new high-performance models. These models deliver optimized implementations of many popular LLMs like Llama and Mistral. When running out-of-the-box on NVIDIA GPUs, MAX matches the performance of the established AI serving framework vLLM in standard throughput benchmarks. These models also support a range of quantization approaches, and we are working incredibly hard for this collection of native MAX models to define SOTA performance in the coming weeks and months.

The industry-standard ShareGPTv3 benchmark demonstrates MAX GPU’s performance capabilities with Llama 3.1, achieving a throughput of 3860 output tokens per second on NVIDIA A100 GPUs using only MAX’s innovative NVIDIA kernels–with GPU utilization greater than 95%. We’re currently achieving this level of performance without optimizations like PagedAttention - which will land early next year. All this is to highlight we’re just getting started, our numbers will only continue to improve and we make it easy to run these benchmarks yourself.

MAX GPU launches with support for NVIDIA A100, L40, L4, and A10 accelerators, the industry standards for LLM inference and the most optimized GPUs on the market–with H100, H200, and AMD support landing early next year.

__wf_reserved_inherit

Built to enable hardware portability

Up next is bringing up support for AMD MI300X GPUs which we are currently refining, and expect to deliver it soon. MAX's NVIDIA and AMD kernels are built on the same underlying technology, making it possible for us to rapidly support new hardware platforms. We’ll share the exciting details about our AMD bring-up shortly, and will continue to expand our AMD support through early 2025.

Try MAX 24.6 today & our nightly releases

We’re excited to invite developers to explore this early technology preview of MAX GPU and see how it can transform your AI workflow. Despite being a preview, this release is still packed with lots of features and capabilities like the new high-performance models running on NVIDIA GPUs, compatibility with OpenAI APIs, and interoperability with Hugging Face models.

Get started running Llama 3 on MAX GPU now!

This is just the beginning–in 2025, we’ll continue to expand our GPU technology stack, delivering even greater performance across more Generative AI modalities, such as text-to-vision and multi-GPU support for larger models. We’re also focusing on enhancing portability to new hardware architectures, along with introducing a complete GPU programming framework for low-level control and customization.

To help you stay ahead of these advancements, we’ve released detailed documentation for our nightly builds, making it easier to install and leverage the latest GPU features straight from the development branch.

As we wrap up the year, we want to extend our warmest wishes to all of you. 2025 is shaping up to be a pivotal year for AI infrastructure, and we’re thrilled to be at the forefront of this transformation. We could not be more excited. We look forward to continuing this journey with you in the new year—see you in January!