惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Engineering at Meta
Engineering at Meta
P
Privacy International News Feed
W
WeLiveSecurity
Spread Privacy
Spread Privacy
S
Schneier on Security
Google Online Security Blog
Google Online Security Blog
N
News and Events Feed by Topic
Forbes - Security
Forbes - Security
Cisco Talos Blog
Cisco Talos Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
L
Lohrmann on Cybersecurity
P
Privacy & Cybersecurity Law Blog
T
The Exploit Database - CXSecurity.com
C
CXSECURITY Database RSS Feed - CXSecurity.com
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
小众软件
小众软件
人人都是产品经理
人人都是产品经理
SecWiki News
SecWiki News
Schneier on Security
Schneier on Security
月光博客
月光博客
博客园_首页
腾讯CDC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Google DeepMind News
Google DeepMind News
Cyberwarzone
Cyberwarzone
www.infosecurity-magazine.com
www.infosecurity-magazine.com
AWS News Blog
AWS News Blog
WordPress大学
WordPress大学
AI
AI
酷 壳 – CoolShell
酷 壳 – CoolShell
Hacker News: Ask HN
Hacker News: Ask HN
Attack and Defense Labs
Attack and Defense Labs
IT之家
IT之家
P
Proofpoint News Feed
The Hacker News
The Hacker News
The Cloudflare Blog
Vercel News
Vercel News
Application and Cybersecurity Blog
Application and Cybersecurity Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Cloudbric
Cloudbric
C
Cisco Blogs
TaoSecurity Blog
TaoSecurity Blog
I
Intezer
Jina AI
Jina AI
雷峰网
雷峰网
阮一峰的网络日志
阮一峰的网络日志
Microsoft Azure Blog
Microsoft Azure Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
A
About on SuperTechFans
B
Blog

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modular Raises $250M to scale AI's Unified Compute Layer
No items found. · 2025-09-24 · via Modular Blog

September 24, 2025

Modular Team

Modular has raised $250M in its third financing round to continue its mission to build AI’s unified compute layer – a hypervisor for AI. The round was led by Thomas Tull’s US Innovative Technology fund, with DFJ Growth joining and with participation from all existing investors including GV (Google Ventures), General Catalyst and Greylock Ventures. This brings its total capital raised to $380M across three rounds since its founding in 2022 and values Modular at $1.6 billion – almost tripling its valuation from its last raise. The investment reflects Modular’s incredible momentum and reinforces its position as the world’s only truly unified AI infrastructure platform to power the future of AI superintelligence.

“Strategic AI implementation is the most important competitive factor in today’s economy, and as the public and private sectors ramp up their efforts to remain competitive, the demand for compute power to handle these heavy workloads is greater than ever. Modular is foundational in this new era of diverse AI infrastructure, providing a unified AI compute layer that maximizes efficiency, resilience, and cost reduction - and their platform is already in high demand from enterprises, clouds and developers. We are proud to support the team and their vision for portable AI that will power both the U.S. and global economies.” 

- Thomas Tull, Chairman of USIT

"Modular is addressing the most urgent challenge in AI: unifying the compute layer by enabling diversified processing hardware and software to operate cohesively. Modular's platform is poised to become a defining pillar of AI systems, unlocking portability, performance, and efficiency that will accelerate the path to superintelligence."

- Sam Fort, DFJ Growth partner

Building the future of AI infrastructure for everyone

The world's appetite for compute is insatiable. CPUs yield to GPUs and ASICs as AI transforms everything, while data centers rise at unprecedented pace to feed the demand. Superintelligence won't just live in server farms - it's coming to every device, every chip becoming an AI-enabled agent. Inference costs plummet as reasoning models drive explosive usage, yet training costs climb relentlessly higher. The paradox deepens: amid this computational renaissance, massive underutilization haunts our existing capacity, fragmented by every hardware vendor's insistence on proprietary software stacks. The imperative is elegant but unforgiving: chase every flop, and make every one count - because software, not silicon, will determine whether this revolution soars or stalls.

Modular has spent the last 3+ years building foundational infrastructure to solve this for the world. Modular reinvented the world's accelerated compute programming model from the ground up, and is rapidly scaling to meet the enormous demand they are seeing from advanced enterprises and hardware partners. Modular has grown to more than 130 people today with their main headquarters in San Francisco Bay Area, along with a global footprint in North America, United Kingdom and Europe.

Since launching their platform in 2023, Modular has redefined what is possible with heterogeneous programming and AI across the world's CPU and GPU silicon. Its platform is being downloaded 10K’s of times per month and is growing at 75% m/m, has earned 24K+ GitHub stars, powers trillions of tokens served daily in production, and has 100K’s of developers in their ecosystem across more than 100 countries. Modular has now released 600K+ lines of open-source code, has thousands of contributions from developers globally, achieved state-of-the-art performance across NVIDIA and AMD on a single, unified stack, and delivered up to 70% latency reduction and 80% cost reductions for their partners and customers. Modular is building the future of AI alongside an emerging alliance of co-architects - including enterprises like Inworld, SF Compute and more, research teams like Jane Street; cloud providers like Oracle, AWS, Lambda Labs and TensorWave; and hardware leaders like AMD and NVIDIA - each rallying and driving towards the vision of a simpler, more open and innovative AI hardware ecosystem.

The Modular Platform

The Modular Platform is the first enterprise-grade AI inference stack that abstracts away hardware complexity. By replacing vendor-specific runtimes like CUDA and ROCm with a unified low-level layer, it eliminates the fragmentation that holds back existing AI frameworks that were never developed for modern Generative AI inference. The Platform includes the following components, that span from cloud orchestration layer down to hardware programming model (more details):

  • 🦣Mammoth: A Kubernetes-native control plane, router, and substrate specially-designed for large-scale distributed AI serving. It supports multi-model management, prefill-aware routing, disaggregated compute and cache, and other advanced, at-scale, AI optimizations.
  • 🧑🏻‍🚀MAX: A high-performance GenAI serving framework that delivers state-of-the-art optimizations – like speculative decoding and operator-level fusions – out of the box. It exposes an OpenAI-compatible endpoint, runs both MAX-native and PyTorch models seamlessly across GPUs and CPUs, and offers deep customization at the model and kernel level for maximum performance and flexibility.
  • 🔥Mojo: A kernel-focused systems programming language that enables high-performance GPU and CPU programming, blending Pythonic syntax with the performance of C/C++ and the safety of Rust. All the kernels in MAX are written with Mojo and it can be used to extend MAX Models with novel algorithms.

And with the latest release, 25.6, Modular delivers 20–50% performance gains over the latest vLLM and SGLang on next-generation hardware like NVIDIA’s B200 and AMD’s MI355. At the same time, Modular is expanding support to new silicon such as Apple GPUs and many other consumer grade GPUs, along with many upcoming ASIC’s.

Mission and the future

This next round of funding will enable Modular to aggressively scale the Modular Platform natively in the cloud, extend support across cloud and edge hardware platforms, and power the world’s most advanced AI workloads – delivering throughput, latency, cost, and accuracy gains unmatched by any other inference company. All of this is in service of empowering developers with a single, unified AI infrastructure they can build on and trust.

“When we founded Modular, we believed that the world needed a unified platform for AI, and today, that vision is more important than ever. This funding will enable us to realize that vision for developers, enterprises and hardware companies around the world.” said Chris Lattner, CEO of Modular.

Modular is hiring in North America and Europe across a wide range of roles. If you want to shape the future of AI infrastructure and collaborate with some of the brightest minds in the field, visit Modular’s career page. The future of AI infrastructure is bright! Download, contribute and help Modular power the future of AI in the open today.

__wf_reserved_inherit

The Modular Team - Join us, we're currently hiring
  • Person with blonde hair using a laptop with an Apple logo.

    Sign up today

    Signup to our Cloud Platform today to get started easily.

    Sign Up

  • Magnifying glass emoji with black handle and round clear lens.

    Browse open models

    Browse our model catalog, or deploy your own custom model

    Browse models

Sign up for our newsletter

Get all our latest news, announcements and updates delivered directly to your inbox. Unsubscribe at anytime.

Thanks for signing up to our newsletter! 🚀

Thank you,

Modular Sales Team

Oops! Something went wrong while submitting the form.