惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 【当耐特】
Scott Helme
Scott Helme
Google Online Security Blog
Google Online Security Blog
L
LINUX DO - 最新话题
O
OpenAI News
S
Secure Thoughts
Cisco Talos Blog
Cisco Talos Blog
Forbes - Security
Forbes - Security
V
Visual Studio Blog
有赞技术团队
有赞技术团队
Jina AI
Jina AI
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Microsoft Azure Blog
Microsoft Azure Blog
博客园_首页
C
CXSECURITY Database RSS Feed - CXSecurity.com
T
Troy Hunt's Blog
Recent Commits to openclaw:main
Recent Commits to openclaw:main
S
Schneier on Security
雷峰网
雷峰网
The Cloudflare Blog
量子位
Last Week in AI
Last Week in AI
T
Tor Project blog
V
Vulnerabilities – Threatpost
C
Cisco Blogs
B
Blog RSS Feed
S
Security @ Cisco Blogs
T
Threat Research - Cisco Blogs
G
GRAHAM CLULEY
MyScale Blog
MyScale Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Schneier on Security
Schneier on Security
Martin Fowler
Martin Fowler
Help Net Security
Help Net Security
J
Java Code Geeks
人人都是产品经理
人人都是产品经理
G
Google Developers Blog
S
Security Affairs
A
Arctic Wolf
T
Tenable Blog
PCI Perspectives
PCI Perspectives
Spread Privacy
Spread Privacy
AI
AI
L
LangChain Blog
Latest news
Latest news
博客园 - 叶小钗
博客园 - Franky
T
Threatpost
MongoDB | Blog
MongoDB | Blog
Hugging Face - Blog
Hugging Face - Blog

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥
No items found. · 2025-05-20 · via Modular Blog

On May 10th, over 100 engineers and researchers from across the AI ecosystem gathered at AGI House in Hillsborough, CA for our very first hackathon: a fast-paced day of hacking, learning, and building with Mojo. The Modular GPU Kernel Hackathon brought together developers of all experience levels to experiment with Mojo on modern GPU hardware, collaborate in person, and accelerate the future of high-performance AI infrastructure.

Participants tackled a wide range of problems, from low-level kernel implementations to full model training frameworks. Many participants had no Mojo or GPU programming experience. By the end of the day, dozens of teams had working prototypes and new insights into what Mojo can do.

We're thrilled by what this community accomplished in just a single day!

Hackathon talks now available

Before the hacking began, we were lucky to hear from an all-star lineup of speakers:

  • Chris Lattner, CEO of Modular, opened the event with a look at the challenges AI developers face today and how Mojo aims to solve them.
  • Ramine Roane, Corporate VP of AI at AMD, shared how AMD is approaching AI infrastructure with the MI300X GPU and the ROCm software stack.
  • Mark Saroufim, cofounder of GPU MODE and software engineer on the PyTorch team at Meta, broke down the tradeoffs of rewriting versus compiling PyTorch backends.
  • Jeff Niu, member of technical staff at OpenAI, explored how he brought Triton-style abstractions into Mojo to prototype high-performance kernels.
  • Simon Boehm and Sasha Krassovsky, members of technical staff at Anthropic, shared their real-world experience running inference across NVIDIA GPUs, Google TPUs, and AWS Tranium.

Winning projects

First place: Monolithic Sup

__wf_reserved_inherit

Team: Marcel Roed (PhD Student at Stanford University), Herman Brunborg (PhD Student at Stanford University), and Rajat Vadiraj Dwaraknath (PhD Student at Stanford University)

Marcel and the team tackled one of the most ambitious challenges of the hackathon. They built a training framework in Mojo/MAX and implemented the kernels and backward passes needed to train a Transformer model from scratch. This required implementing backpropagation in MAX, along with gradient descent algorithms like AdamW, as well as figuring out how to work with FlashAttention in the MAX kernel library and implementing its derivative to perform gradient training. Although they didn’t have time to fully implement FlashAttention during the hackathon and used standard scaled dot product attention instead, they’ve continued working on the project and plan to complete the implementation soon.

“It often makes sense to build something that works rather than trying to make something fast immediately… It was super fun to debug in real-time and discuss errors and solutions with the Modular employees as we ran into problems as we were trying to use Mojo in ways not previously explored by the Modular team.” - Herman
“It was difficult to get our kernels to be correct, and figuring out how to use the FlashAttention implementations... was quite challenging. But we showed that we can build useful SoTA-level tools from bare-bones with Mojo in a short period of time.” - Marcel

Second place: Fast Implementation of Prefix Scan Algorithms for AMD MI300X Using Decoupled Lookback

__wf_reserved_inherit

Team: Kirill Bobyrev (Software Engineer at Waymo)

Kirill set out to implement the optimal parallel prefix scan algorithm for high-performance GPUs—and in the process, helped improve Mojo’s standard library. His contributions included new, corrected implementations of warp- and block-level scans, and an efficient, tunable full device-wide scan. These updates are now part of Mojo’s open-source repo.

“GPU programming is notoriously challenging, but Mojo makes it surprisingly pleasant. The language feels modern, its templates and meta-programming features enable rapid experimentation, and the code is highly readable—akin to Rust's standard library, a clear contrast to the often cumbersome C++ standard library.” - Kirill

Even more impressively, Kirill had never written a line of CUDA or Mojo before the week of the hackathon.

Third place: Gaussian Splatting in Mojo

__wf_reserved_inherit

Team: Sandeep Menon (Software Engineer, Deep Learning at Kodiak) and Owen Leather (Perception Software Engineer at Kodiak)

Gaussian splatting is a rendering technique that, until now, had only been implemented in CUDA. Sandeep’s team aimed to break new ground by porting this kernel to Mojo so it could run on GPUs like AMD’s MI300X. While they weren’t able to fully finish the implementation by the end of the event, they made strong headway and are continuing to build on it.

“Learning Mojo and meeting the amazing Modular team was a highlight. It is humbling to code in the era of AI programming. I would love to be invited to future events and join many more hackathons as a way to kickstart learning on new topics.” - Sandeep

More hackathon projects

__wf_reserved_inherit

The dedicated crew of hackers who stuck around until late night to present their work 💪

Beyond our winners, teams dove into a wide range of challenges using Mojo:

  • A heat diffusion kernel based on the finite difference stencil method
  • A GPU-accelerated BM25 ranking algorithm
  • An optimized Non-Maximum Suppression implementation for YOLO-style models
  • A benchmarking study to run Mojo kernels without barrier or synchronize calls
  • A Fast Fourier Transform implementation in Mojo
  • A matrix inversion and LU decomposition kernel for small matrices
  • A dot product primitive and related linear algebra building blocks

Many participants wrote Mojo or GPU code for the first time and were eager to build on their projects in the coming weeks. We’re incredibly proud of everyone who participated in the hackathon. Whether you were building compilers, training frameworks, or GPU kernels from scratch, your work is what makes this community so exciting!

Thank you again to our sponsors, AMD, Crusoe, and GPU MODE, for making the event possible. And if you weren’t able to join us in person, be sure to catch the talk recordings on the Modular YouTube channel.

Until next time, keep building!