惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Vercel News
Vercel News
Simon Willison's Weblog
Simon Willison's Weblog
云风的 BLOG
云风的 BLOG
宝玉的分享
宝玉的分享
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
The Register - Security
The Register - Security
S
SegmentFault 最新的问题
博客园 - 司徒正美
The GitHub Blog
The GitHub Blog
量子位
SecWiki News
SecWiki News
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 【当耐特】
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
P
Palo Alto Networks Blog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
P
Privacy & Cybersecurity Law Blog
爱范儿
爱范儿
S
Secure Thoughts
G
Google Developers Blog
Microsoft Security Blog
Microsoft Security Blog
D
Docker
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
F
Fortinet All Blogs
T
Threat Research - Cisco Blogs
P
Proofpoint News Feed
Schneier on Security
Schneier on Security
Y
Y Combinator Blog
博客园 - 三生石上(FineUI控件)
Recent Announcements
Recent Announcements
Recent Commits to openclaw:main
Recent Commits to openclaw:main
G
GRAHAM CLULEY
Recorded Future
Recorded Future
罗磊的独立博客
Forbes - Security
Forbes - Security
Security Latest
Security Latest
NISL@THU
NISL@THU
T
The Exploit Database - CXSecurity.com
Hugging Face - Blog
Hugging Face - Blog
T
Tenable Blog
C
Cybersecurity and Infrastructure Security Agency CISA
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Apple Machine Learning Research
Apple Machine Learning Research
K
Kaspersky official blog
月光博客
月光博客
小众软件
小众软件
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
WordPress大学
WordPress大学
Security Archives - TechRepublic
Security Archives - TechRepublic

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days
No items found. · 2025-10-17 · via Modular Blog

In late August, AMD and TensorWave reached out to collaborate on a presentation for AMD’s Media Tech Day—they asked if we could demo MAX on AMD Instinct™ MI355 on September 16th. There was just one problem: no one at Modular had access to an MI355.

That gave us just two weeks from the point we’d first have access. We saw this as an opportunity: could we achieve state-of-the-art performance on MI355 in just 2 weeks? We were confident we could, because we had spent several years building a technology platform to enable rapid AI hardware bringup.

Why AI Hardware Enablement Is Hard

Bringing new AI hardware online isn’t supposed to happen in just two weeks.

Why? The modern AI software ecosystem is fragmented. Different companies own different layers — hardware vendors are chasing silicon sales, researchers are prototyping bleeding-edge ideas, and application developers are stitching things together. The result? Complex, brittle stacks that are hard to extend and even harder to optimize.

That’s why, from day one, we built the Modular stack to make AI hardware enablement fast, consistent, and maintainable.

We often joke that the hardware is no longer the hard part — it’s the software that slows everyone down.

A Foundation Built for Portability

The Modular software stack — from Mojo (our programming language), to MAX (our inference framework) and Mammoth (our distributed serving system) — is architected for portability, meaning it can quickly move onto the latest hardware architectures and get SOTA performance.

Here’s how our design pays off:

  • Architecture-agnostic design. Mojo and MAX do not hardcode GPU knowledge and optimizations. Instead, they delegate all hardware-specific details to libraries.
  • Library-directed execution. Offloading, scheduling, and instruction selection are controlled through abstractions rather than hard-coded pathways in the compiler.
  • Parametric operations. Instead of “magic constants” (like SIMD widths or tile shapes), we use parameterized kernels that can be retuned for new hardware.

Because 99.9% of the stack is architecture-agnostic, adding support for a new GPU mostly involves updating a few kernels. MAX and Mammoth just work — out of the box.

That design is what allowed us to move so quickly once MI355 hardware arrived. In fact, we only need to revise the parts that changed in the hardware itself. These all reside in specific kernels—so let's look at the new features in MI355.

What’s New in AMD MI355

To understand why MI355 requires only incremental kernel updates, let's examine what the MI355 architecture offers and how Modular kernels address these features:

MI355 feature Impact Our adjustment
New FP32→BF16 conversion instructions Faster casting between data types Updated cast dispatch in Mojo’s standard library
Larger tensor-core tile sizes Higher arithmetic intensity Expanded matmul tile selection space
Increased shared memory (64KB→160KB) Larger tiles, less memory pressure Tuned tile-size parameters
Transposed DRAM→shared-memory loads Removes extra in-place transpositions Updated LayoutTensor load logic
Larger HBM capacity Allows larger dynamic batch sizes Nothing — handled automatically by MAX’s scheduler

These new hardware features in MI355 are specifically designed to optimize matmul-like operations, so that’s the only category that required changes in the Modular stack.

Two Weeks to SOTA: A Day-by-Day Story

Thanks to our stack design as described above, our work on AMD's MI355 started well before we had hardware access—we simply architected the Modular stack to enable fast hardware bringup.

With only 2 weeks until the demo, we had to be selective about what we could develop in a reasonable timeframe. While Modular moves fast, we pride ourselves on maintaining a good work-life balance—so we wanted to accomplish our goals without going into crunch mode.

Day 0: Preparing Without Hardware

Before we even touched an MI355, we began testing code generation offline. Mojo’s hardware-agnostic backend allowed us to simulate instruction paths — verifying the emitted assembly matched expectations. Try it yourself with Compiler Explorer.

By understanding the features available in the new hardware and ensuring we emit the correct instructions, we can prototype optimizations without actually running the program on the new hardware 🤯.

Day 1: First Login, First Success

On September 1st, TensorWave provisioned MI355 systems for us. We logged in, ran amd-smi, and saw the new hardware come online:

Bash

modular@mia1-p1-g49:~$ amd-smi +------------------------------------------------------------------------------+ | AMD-SMI 26.0.0+d30a0afe amdgpu version: 6.14.14 ROCm version: 7.0.0 | |-------------------------------------+----------------------------------------| | BDF GPU-Name | Mem-Uti Temp UEC Power-Usage | | GPU HIP-ID OAM-ID Partition-Mode | GFX-Uti Fan Mem-Usage | |=====================================+========================================| | 0000:05:00.0 AMD Instinct MI355X | 0 % 44 °C 0 229/1400 W | | 0 1 6 SPX/NPS1 | 0 % N/A 283/294896 MB | |-------------------------------------+----------------------------------------| | 0000:15:00.0 AMD Instinct MI355X | 0 % 45 °C 0 226/1400 W | | 1 3 7 SPX/NPS1 | 0 % N/A 283/294896 MB | |-------------------------------------+----------------------------------------| | 0000:65:00.0 AMD Instinct MI355X | 0 % 43 °C 0 233/1400 W | | 2 2 5 SPX/NPS1 | 0 % N/A 283/294896 MB | |-------------------------------------+----------------------------------------| | 0000:75:00.0 AMD Instinct MI355X | 0 % 43 °C 0 225/1400 W | | 3 0 4 SPX/NPS1 | 0 % N/A 283/294896 MB | |-------------------------------------+----------------------------------------| | 0000:85:00.0 AMD Instinct MI355X | 0 % 40 °C 0 226/1400 W | | 4 5 2 SPX/NPS1 | 0 % N/A 283/294896 MB | |-------------------------------------+----------------------------------------| | 0000:95:00.0 AMD Instinct MI355X | 0 % 41 °C 0 226/1400 W | | 5 7 3 SPX/NPS1 | 0 % N/A 283/294896 MB | |-------------------------------------+----------------------------------------|

We ran pip install modular, launched a serving endpoint with MAX, and everything worked out of the box—you can do the same by following our quickstart guide.

We weren't leveraging MI355's new hardware features yet, but we could profile and identify execution bottlenecks. We also mapped our execution against an internal performance estimation tool (stay tuned for an upcoming blog post about this).

With the hardware and initial validation in hand, we could now form a plan.

Week 1: Finding the Levers

We benchmarked the MI355 GPUs against B200 GPUs, running models such as Llama, Gemma, and Mistral. We analyzed the results against theoretical upper bounds to identify optimization opportunities—we've developed tools to automate much of this analysis. The tools pointed to matmul optimization as the path to a significant performance gains and SOTA results.

The matmul implementation — only about 500 lines of well-commented code — was easy to adapt to run efficiently on MI355 hardware. We also wanted portability across AMD hardware—the same code running at SOTA on MI300, MI325, and MI355—because even with tight deadlines, we still value good software design.

After working on the matmul code for a few hours and fixing the M=N=K=8192, we proved that minor parameterization changes could deliver performance close to AMD’s hipBLASLt library (which we measured to be SOTA and faster than hipBLAS).

Matmul kernel GFlop/s
Mojo (day 0) 1202302.79
hipblaslt 1561446.68
Mojo (day 1) 1610514.29

With a few more tweaks the same day, the Mojo matmul kernel was now 3% faster than SOTA. This day-one experiment proved our goal was achievable.

The rest of our week 1 effort included:

  • Generalizing our implementation across the different shapes present in the models
  • Refining our heuristics for optimal kernel parameter selection
  • Setting up automation for benchmarking
  • Configuring our remote login system to make compute resources easy to access and share

Of course, the first week wasn’t without some issues. We encountered delays due to hardware misconfiguration and missing GPU operators in the Kubernetes integration. Despite the setbacks, we had strong performance numbers by the end of the week and we enjoyed the weekend.

Week 2: From Optimization to Demo

We had one week to go—the demo presentation was the following Tuesday.

We continued optimizing the matmul kernel and began exploring other opportunities, such as optimizing the Attention kernel’s performance. By the middle of week 2, we had strong numbers to share with our partners.

The demo preparation was also in full swing now. The effort required that the IT department make sure we could present live benchmark results, the design department make our presentation professional, and the product team help craft the message.

Meanwhile, our engineering team continued to optimize our stack for the new hardware, using the automation we enabled in week 1 to feed live numbers to the other teams.

By Friday, MAX nightlies were running smoothly on MI355, and our benchmarks were consistently leading AMD’s custom fork of vLLM. With these results we were ready to demo on Tuesday!

💡 Fun fact: The MI355 bring-up wasn’t just fast, it was done by only two engineers working normal working hours, no late nights, and zero crunch. Turns out, great architecture scales both performance and sanity.

Results: Outperforming the Field

By the end of two weeks, MAX outperformed AMD’s optimized vLLM fork by up to 2.2× across multiple workloads — all while maintaining full portability across MI300, MI325, and MI355. These gains came partly from the kernels, but also from the entire stack working together to deliver both performance and portability.

Even more remarkably, the entire bring-up effort involved just two engineers working on MI355 bringup for just two weeks, one of which had a pre-planned vacation, so in reality we had 1.5 engineers working for two weeks. In total, this effort resulted in 20 small PRs and zero late nights. This is a textbook case of software architecture enabling velocity.

The performance results speak for themselves.

At the AMD Media Tech Day, MAX was the only inference solution to demonstrate clear TCO advantages when compared to NVIDIA’s flagship Blackwell architecture.

Are We Done? Not Even Close.

This two-week sprint was just the beginning. Since that demo, we’ve expanded MI355 support, added early Apple silicon support, and achieved state-of-the-art performance on both NVIDIA Blackwell and AMD MI355X. Check out our recent 25.6 release.

Our mission at Modular remains the same: To make AI hardware enablement fast, portable, and universal — no matter whose silicon it runs on.

Thank you to TensorWave for partnering with us as well as providing the MI355X systems and to AMD for inviting us to their Media Tech Day event.

Stay tuned for more updates — and for upcoming deep-dives on the kernel optimizations that made this milestone possible. Or feel free to reach out for a demo and we can talk about your AI use case.