惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

S
Security @ Cisco Blogs
罗磊的独立博客
宝玉的分享
宝玉的分享
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
美团技术团队
T
Tailwind CSS Blog
博客园 - 三生石上(FineUI控件)
博客园 - Franky
G
Google Developers Blog
Jina AI
Jina AI
Stack Overflow Blog
Stack Overflow Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
V
Visual Studio Blog
腾讯CDC
S
SegmentFault 最新的问题
Recent Announcements
Recent Announcements
博客园 - 叶小钗
Microsoft Security Blog
Microsoft Security Blog
雷峰网
雷峰网
L
LangChain Blog
Vercel News
Vercel News
Forbes - Security
Forbes - Security
PCI Perspectives
PCI Perspectives
N
News | PayPal Newsroom
S
Security Affairs
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
博客园 - 司徒正美
J
Java Code Geeks
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Hacker News: Ask HN
Hacker News: Ask HN
Schneier on Security
Schneier on Security
A
About on SuperTechFans
Attack and Defense Labs
Attack and Defense Labs
Google Online Security Blog
Google Online Security Blog
aimingoo的专栏
aimingoo的专栏
MongoDB | Blog
MongoDB | Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
酷 壳 – CoolShell
酷 壳 – CoolShell
Cloudbric
Cloudbric
B
Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
P
Proofpoint News Feed
D
DataBreaches.Net
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
B
Blog RSS Feed
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
N
News and Events Feed by Topic

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In
No items found. · 2025-06-18 · via Modular Blog

We're excited to announce Modular Platform 25.4, a major release that brings the full power of AMD GPUs to our entire platform. This release marks a major leap toward democratizing access to high-performance AI by enabling seamless portability to AMD GPUs. Developers can now build and deploy models optimized for peak performance, with zero reliance on any single hardware vendor—unlocking greater flexibility, lower costs, and broader access to compute.

🚀 AMD GPUs now officially supported

The headline feature of 25.4 is official support for AMD GPUs, backed by our newly announced partnership with AMD. You can now deploy Modular with full acceleration on AMD MI300X and MI325X GPUs using the exact same code and container as NVIDIA, with zero changes or workflow tweaks. For the first time, enterprises can build portable, high-performance GenAI deployments that run on any platform without vendor lock-in or platform-specific optimizations.

Compared to existing infrastructure, Modular delivers substantial performance gains on AMD GPUs with popular LLM workloads—achieving competitive results even when compared with running the same workloads on NVIDIA GPUs:

  • Up to 53% better throughput on prefill-heavy BF16 workloads across Llama-3.1-8B, Gemma-3-12B, Mistral-Small-24B, and other state-of-the-art language models when compared with vLLM on AMD MI300X.
  • Up to 32% better throughput for decode-heavy BF16 workloads when compared with vLLM on AMD MI300X.
  • Throughput parity or better on ShareGPT workloads running on MI325X when compared to vLLM on NVIDIA H200.

__wf_reserved_inherit

AMD MI300X and MI325X GPUs often provide superior price-performance ratios for many AI workloads, giving you the flexibility to optimize total cost of ownership based on real-world economics rather than being locked into a single vendor's pricing structure. You can find a deeper analysis of these performance benchmarks in our recent AMD partnership announcement.

Beyond the headline AMD GPU support, we're also introducing:

While AI model support is still in development for some architectures, GPU programming capabilities are fully functional—and open source—across this expanded hardware ecosystem. You can learn to program this entire line of GPUs today using Mojo with our ever-expanding collection of Mojo GPU Puzzles.

🤖 Expanded model support

Modular 25.4 significantly expands our model ecosystem, including:

  • GGUF quantized Llamas with support for q4_0, q4_k, and q6_k quantization using a paged KVCache strategy.
  • Qwen3 family of models, with advanced reasoning and multilingual capabilities.
  • OLMo2 family of models, designed for research and common tasks.
  • Gemma3 multimodal models, offering optimized performance and improved safety.

Head over to Code with Modular where you can find these releases, along with more than 500 additional generative AI models.

📚 Enhanced documentation and developer experience

We've completely redesigned our documentation ecosystem with a unified navigation system across docs and code. Finding the resources you need is now easier than ever.

New documentation includes:

🐍 Python–Mojo bindings

Mojo's simplicity and ease of use has always drawn from Python for inspiration, and Mojo 25.4 brings the languages even closer together with a new developer preview of Python–Mojo bindings. You can now call Mojo functions directly from Python code without the need to manage complex build systems or dependency chains. Develop in Python, and seamlessly replace your performance hot-spots with blazing-fast Mojo equivalents. It’s like having a turbo button for your Python apps! Learn more in the Modular forum, and try out the code examples on GitHub.

👩‍💻 Now open for contributions!

We've made history by open sourcing over 450k lines of production-grade Mojo kernel and serving code, and now we're inviting the developer community to help shape the future of AI infrastructure. The MAX AI kernel library is officially open for contributions! Whether you're missing a key operator for your groundbreaking model, need to extend support for a new hardware architecture, or want to optimize performance at the kernel level, there’s a place for your contributions. Join the Modular developer community in pushing the boundaries of GPU programming and help us build the foundation for the next generation of AI breakthroughs.

Get started now, and join us in person!

Modular 25.4 represents our commitment to giving you more choice, better performance, and seamless integration with your existing workflows. Whether you're:

  • Optimizing total cost of ownership by choosing the most cost-effective hardware for your specific workloads.
  • Building resilient infrastructure that isn't dependent on a single GPU vendor's supply chain.
  • Future-proofing your AI investments against hardware vendor lock-in.
  • Working with the latest language models like Qwen3 or OLMo2.

This release has something powerful to offer!

Ready to experience Modular 25.4? Get started right away with our quickstart guide, and dive into our latest tutorials to learn how to start setting up production workloads.

To celebrate this launch, we have a couple of special events lined up. First, we’re hosting Modular Hack Weekend on June 27-29, kicking off with a GPU Programming Workshop on June 27th. Join us in-person or via livestream on Friday for the workshop and lightning talks, and participate in the weekend-long hackathon virtually!

__wf_reserved_inherit

Second, we just launched our comic, GPU Whisperers, a new series that perfectly captures the beautiful chaos of living through the GenAI revolution. Got an AI horror story that you want to immortalize? Submit your own AI problems and we’ll make art from your pain.

A full list of changes is available in the MAX and Mojo changelogs. As always, we welcome your feedback and contributions to help make the Modular platform even better. Join the discussion on our community forum, and come build with us!