惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Blog — PlanetScale
Blog — PlanetScale
博客园_首页
WordPress大学
WordPress大学
博客园 - 聂微东
P
Privacy International News Feed
Forbes - Security
Forbes - Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Last Week in AI
Last Week in AI
C
CERT Recently Published Vulnerability Notes
月光博客
月光博客
NISL@THU
NISL@THU
美团技术团队
T
Tailwind CSS Blog
Jina AI
Jina AI
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Apple Machine Learning Research
Apple Machine Learning Research
C
Cisco Blogs
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Hacker News
The Hacker News
B
Blog
P
Palo Alto Networks Blog
L
Lohrmann on Cybersecurity
有赞技术团队
有赞技术团队
The Register - Security
The Register - Security
S
Securelist
A
Arctic Wolf
MyScale Blog
MyScale Blog
H
Help Net Security
N
Netflix TechBlog - Medium
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
T
Threatpost
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Security Latest
Security Latest
T
Tor Project blog
V
Vulnerabilities – Threatpost
V
V2EX
AI
AI
Hugging Face - Blog
Hugging Face - Blog
大猫的无限游戏
大猫的无限游戏
博客园 - Franky
Simon Willison's Weblog
Simon Willison's Weblog
小众软件
小众软件
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
H
Hackread – Cybersecurity News, Data Breaches, AI and More
T
Troy Hunt's Blog
Schneier on Security
Schneier on Security
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
H
Heimdal Security Blog
Google Online Security Blog
Google Online Security Blog
Know Your Adversary
Know Your Adversary

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend
No items found. · 2025-07-09 · via Modular Blog

July 9, 2025

Caroline Frasca

Between a global hackathon, a major release, and standout community projects, last month was full of progress across the Modular ecosystem!

Modular Platform 25.4 launched on June 18th, alongside the announcement of our official partnership with AMD, bringing full support for AMD Instinct™ MI300X and MI325X GPUs. You can now deploy the same container across both AMD and NVIDIA hardware with no code changes, no vendor lock-in, and no additional configuration!

Highlights from 25.4 include up to 53% better throughput on prefill-heavy BF16 workloads across Llama 3.1, Gemma 3, Mistral, and other state-of-the-art language models. The release also added support for AMD MI300/325, RDNA3/4, and NVIDIA RTX 2060–5090, along with expanded model coverage.

June also united builders from around the world for Modular Hack Weekend, where developers created everything from Fast Fourier Transform implementations and GPU-accelerated quantum simulators to high-performance bioinformatics libraries. We introduced Mammoth, our new system for scaling GenAI inference across any GPU, and rolled out new ways to integrate Mojo kernels directly into Python workflows. The community pushed the boundaries of kernel design, explored breakthroughs in scientific computing, and continued expanding what’s possible with Mojo and MAX.

Let’s take a look at everything the Modular universe made possible last month.

Blogs, Tutorials, and Videos

  • Developers from across the AI and systems programming communities recently came together for Modular Hack Weekend: a global, virtual hackathon focused on GPU programming and model implementation with Mojo and MAX.
  • We dropped a series of exciting announcements in our video premiere:
    • Modular Platform is now generally available on AMD Instinct™ MI300X and MI325 GPUs! Benchmarks show up to 53% better throughput on prefill-heavy BF16 workflows. Together with AMD, we’re combining best-in-class compute with developer-friendly software. Check out the full blog post.
    • Meet Mammoth: our new Kubernetes-native system for scaling GenAI inference across any GPU. Deploy Hugging Face models across AMD and NVIDIA from a single container, with no manual configuration. Join the public preview.
    • Mojo in Python: you can now drop Mojo kernels directly into your Python workflows. Available today in nightly builds, and backed by 450k+ lines of open source Mojo kernel code. Start here.
  • Now on YouTube: Chris Lattner's full talk from AMD AdvancingAI 2025! Learn how Mojo brings together Python’s simplicity and C++ performance to power a next-gen AI software stack. Plus, catch the post-talk Q&A with Chris.
  • Chris Lattner joined the Latent Space podcast to share an inside look at the history of Modular and Mojo, and the future of GPU programming.
  • Our June community meeting featured two in-depth presentations on how Mojo is being applied in scientific computing:
    • Bioinformatics with Mojo: Seth walked us through ish, a high-performance, index-free alignment tool built in Mojo. He shared insights on SIMD optimizations, GPU acceleration, and benchmarking against C++ libraries like Parasail.
    • Particle Physics with Mojo: Photon shared how Mojo is helping streamline complex particle physics simulations. He introduced two open-source libraries, newmojo and hepjo, and discussed porting a C++/Python research pipeline to Mojo with promising performance gains.
  • Simon Veitner published a deep dive on crafting a blazing-fast matrix transpose kernel for NVIDIA Hopper. He covers TMA, swizzling, thread coarsening, and shows how far you can push performance using pure Mojo.
  • We released our comic series, GPU Whisperers, that perfectly captures the beautiful chaos of living through the GenAI revolution! 🧑‍🚀
  • Vincent Warmerdam shared an excellent writeup on calling Mojo from Python.
  • Modular is now available on the Amazon Web Services (AWS) Marketplace! 500+ Pre-Optimized Models, with an OpenAI API Compatible endpoint, ready for you to run across NVIDIA B200, H200, H100, A100, A10, L40 and L4 GPUs, with intelligent batching and memory management.
  • Modular Tech Talks is an exclusive series featuring internal presentations from our engineering team, explaining the inner workings of the Modular technology stack. In our most recent edition, Kyle Caverly gives a tour of the MAX Pipelines architecture, covering its major interfaces and how they enable the Modular team to rapidly bring up state of the art models with high-performance features like KV Cache optimization and speculative decoding.
  • Vibe coding your next Mojo masterpiece? You’re in luck: check out our guide on using AI coding assistants like Cursor and Copilot to build faster with Mojo and MAX.
  • Our recent Democratizing AI Compute series by Chris Lattner offers a clear perspective on the challenges shaping the future of AI infra, and you can now explore the full series in one convenient place! 🔖 Bookmark for later or subscribe to the RSS feed.
    • In the latest installment, Chris Lattner plots a flight plan across the Modular stack: Mojo the molten inner world, MAX the mighty gas giant, all orbiting in the Mammoth cluster.
  • Building high-performance AI infrastructure doesn’t have to take months. Inworld proved that by launching a state-of-the-art speech pipeline into production in under 8 weeks with Modular. Their blog post explains how they used MAX and Mojo to run on NVIDIA Blackwell GPUs, meet real-time latency targets that were 70% faster than using the latest vLLM.

Awesome MAX + Mojo

Open-Source Contributions

If you’ve recently had your first PR merged, message Caroline Frasca in the forum to claim your epic Modular swag! Check out the recently merged contributions from our amazing community members:

  • Person with blonde hair using a laptop with an Apple logo.

    Sign up today

    Signup to our Cloud Platform today to get started easily.

    Sign Up

  • Magnifying glass emoji with black handle and round clear lens.

    Browse open models

    Browse our model catalog, or deploy your own custom model

    Browse models

Sign up for our newsletter

Get all our latest news, announcements and updates delivered directly to your inbox. Unsubscribe at anytime.

Thanks for signing up to our newsletter! 🚀

Thank you,

Modular Sales Team

Oops! Something went wrong while submitting the form.