惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

B
Blog RSS Feed
K
Kaspersky official blog
Forbes - Security
Forbes - Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
P
Proofpoint News Feed
G
GRAHAM CLULEY
V
Vulnerabilities – Threatpost
Security Latest
Security Latest
Scott Helme
Scott Helme
S
Securelist
美团技术团队
T
Threat Research - Cisco Blogs
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
SegmentFault 最新的问题
W
WeLiveSecurity
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Apple Machine Learning Research
Apple Machine Learning Research
The Cloudflare Blog
AI
AI
L
Lohrmann on Cybersecurity
S
Security Affairs
Cloudbric
Cloudbric
SecWiki News
SecWiki News
爱范儿
爱范儿
雷峰网
雷峰网
Engineering at Meta
Engineering at Meta
C
Cyber Attacks, Cyber Crime and Cyber Security
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
I
InfoQ
S
Secure Thoughts
AWS News Blog
AWS News Blog
A
About on SuperTechFans
Schneier on Security
Schneier on Security
酷 壳 – CoolShell
酷 壳 – CoolShell
The Last Watchdog
The Last Watchdog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
Check Point Blog
P
Palo Alto Networks Blog
博客园 - 【当耐特】
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Google DeepMind News
Google DeepMind News
Latest news
Latest news
I
Intezer
博客园_首页
C
CXSECURITY Database RSS Feed - CXSecurity.com
V
V2EX
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
L
LangChain Blog
D
Docker

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modular Platform 25.5: Introducing Large Scale Batch Inference
No items found. · 2025-08-05 · via Modular Blog

Modular Platform 25.5 is here, and introduces Large Scale Batch Inference: a highly asynchronous, at-scale batch API built on open standards and powered by Mammoth. We're launching this new capability through our partner SF Compute, enabling high-volume AI performance with a fast, accurate, and efficient platform that seamlessly scales workloads across any hardware.

This release also features the open source launch of the MAX Graph API and expanded support for writing custom PyTorch operators directly in MAX. In addition, we’ve made Modular Platform development and deployment easier with optimized Docker containers and new standalone Mojo Conda packages.

Large Scale Batch Inference, powered by Mammoth

As AI workloads grow in complexity and scale, infrastructure needs to be smarter, faster, and more efficient. That’s why we partnered with SF Compute to build the Large Scale Batch Inference API, a high-throughput serving product powered by Mammoth, Modular’s intelligent Kubernetes-native cluster orchestration layer. To maximize throughput and efficiency, Mammoth continuously distributes jobs across GPU clusters using an optimized scheduler to maintain over 90% utilization of cluster resources.

This offering supports more than 20 state-of-the-art models across language, vision, and multimodal domains, runs on both NVIDIA and AMD hardware, and includes an OpenAI-compatible API. Watch our joint launch video and check out the blog post to learn more about SF Compute’s Large Scale Inference API. Sign up to get access to Large Scale Batch Inference through SF Compute.

__wf_reserved_inherit

Lightweight, modular packaging

Standalone Mojo packages

Mojo has empowered a new generation of GPU developers to write portable, performant code that breaks free from vendor lock-in. With 25.5, we're making GPU programming with Mojo even more accessible through two new standalone Conda packages.

The mojo package delivers everything you need for AI-independent CPU and GPU kernel development: compiler, runtime, debugger, and LSP, without requiring additional MAX packages. For streamlined distribution of Mojo-only or hybrid Python/Mojo packages, the mojo-compiler package includes just the essentials: compiler and runtime. This lightweight option provides a solid foundation for distributing Mojo applications and libraries through the official Modular community package repository. A Python wheel distribution is currently in the works to further simplify integration workflows.

Whether you’re new to GPU development or looking to optimize the performance of your existing system, our Mojo GPU Puzzles and GPU programming docs will take you from writing basic operations to crafting high-performance matrix multiplication.

Lighter-weight MAX serving packages  

To further reduce deployment overheads and drastically reduce cold start time in production, we’ve eliminated PyTorch dependencies from our containers, making MAX serving packages leaner than ever. Now under 700 MB for NVIDIA GPUs, you can deploy and scale your production workloads even faster, with a container that is over 90% smaller than the vLLM equivalent.

__wf_reserved_inherit

Open source MAX Graph API

This release also includes the open source launch of the MAX Graph API – the Python-based engine behind our state-of-the-art models, powered by the Mojo kernel library. With it, you can now build open, portable, GPU-accelerated graphs directly in Python.

Freely available code means that AI coding assistants can leverage the full source as a reference, making it faster and easier to port your most critical models to the Modular Platform. We’ve also enhanced the Graph API to perform compile-time verification of custom graphs, reducing development cycles and ensuring model correctness.

Earlier this year, we open-sourced all MAX kernels. Now, you can explore the model-building interfaces themselves – complete with supporting code and unit tests that demonstrate each kernel in small, well-defined examples.

Want to learn more? Connect with our growing community of AI developers, share your breakthroughs, and get expert help on the Modular community forum.

Seamless MAX and PyTorch integration

New in 25.5, MAX graphs can now be seamlessly integrated into PyTorch workflows as custom operators.

With the new @graph_op decorator, any MAX graph can be automatically wrapped as a custom PyTorch operator – dramatically expanding how you can extend PyTorch with MAX. This makes it easy to replace slow, platform-specific PyTorch kernels with high-performance MAX implementations, or even swap out entire sections of a model for faster, portable MAX graphs.

Best of all, you can do this while preserving your existing PyTorch codebase, ensuring minimal disruption and maximum performance gains.

Explore the full update

Check out all of 25.5’s updates with the full MAX and Mojo changelogs. Mojo adds parametric aliases, improved string handling, the Iterator trait, and stronger type-checking for custom GPU functions. MAX improvements include faster prefix caching using Mojo, up to 12x faster support for the OpenAI API’s logprobs and echo request parameters, and Idefics3 model support.

Get started now!

Ready to experience Modular 25.5? Get started by installing Modular and running GPU-accelerated inference with our quickstart guide, then explore our tutorials for step-by-step production deployment. Tune into next week's livestream to get an inside look at what's new in Modular Platform 25.5 and a chance to ask the team your questions live.