惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
爱范儿
爱范儿
V
Visual Studio Blog
The Register - Security
The Register - Security
P
Proofpoint News Feed
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
H
Hackread – Cybersecurity News, Data Breaches, AI and More
GbyAI
GbyAI
Y
Y Combinator Blog
M
MIT News - Artificial intelligence
大猫的无限游戏
大猫的无限游戏
L
LangChain Blog
The Cloudflare Blog
Hugging Face - Blog
Hugging Face - Blog
Microsoft Azure Blog
Microsoft Azure Blog
T
Threatpost
P
Proofpoint News Feed
美团技术团队
A
About on SuperTechFans
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
MongoDB | Blog
MongoDB | Blog
C
Check Point Blog
Vercel News
Vercel News
L
Lohrmann on Cybersecurity
N
News and Events Feed by Topic
宝玉的分享
宝玉的分享
T
Tor Project blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Spread Privacy
Spread Privacy
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
C
Cisco Blogs
博客园 - 司徒正美
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Cyberwarzone
Cyberwarzone
C
Cybersecurity and Infrastructure Security Agency CISA
S
Security @ Cisco Blogs
AWS News Blog
AWS News Blog
SecWiki News
SecWiki News
I
InfoQ
PCI Perspectives
PCI Perspectives
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Hacker News - Newest:
Hacker News - Newest: "LLM"
Latest news
Latest news
Stack Overflow Blog
Stack Overflow Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
H
Help Net Security
B
Blog RSS Feed
H
Hacker News: Front Page
雷峰网
雷峰网
Know Your Adversary
Know Your Adversary

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute
No items found. · 2025-02-27 · via Modular Blog

February 27, 2025

Caroline Frasca

We recently introduced MAX 25.1, a major leap forward in AI development. This release enhances agentic and LLM workflows, introduces MAX Builds as a central hub for GenAI models and application recipes, and debuts a new GPU programming interface. Developers can now take advantage of GPU-accelerated embeddings, OpenAI-compatible function calling, structured output generation, and high-performance LLM optimizations like paged attention and prefix caching for improved efficiency.

MAX 25.1 also marks a shift to a nightly release model, giving developers early access to new features, real-time community-driven improvements, and continuous innovation. With offline batch inference, Mojo-powered GPU programming via MAX Graphs, and streamlined deployment from local to cloud, MAX 25.1 delivers improved performance and flexibility. Get started today by exploring MAX Builds, diving into the docs, and joining the community forum!

Blogs, Tutorials, and Videos

  • Experience faster, smarter AI with MAX 25.1! Prefix caching and paged attention boost LLM performance, offline batch inference improves latency and load time, and MAX Builds is your go-to hub for GenAI models, recipes, and packages. Get the full details in our recent blog post.
  • Our MAX 25.1 livestream was packed with insights, demos, powerful updates to MAX and Mojo, and a live audience Q&A with Chris Lattner and team.
  • In Community Meeting #13, we discussed Owen's structured async Mojo proposal and highlighted two standout community projects: Brian's EmberJSON for Mojo JSON parsing and Martin's Modo for generating Mojo docs.
  • In part 1 of our “Democratizing AI Compute” series, Chris Lattner explored how novel ideas, backed by highly focused teams, can unlock efficiency breakthroughs.
    • In part 2, he tackled the question, “What exactly is CUDA?”, revealing why DeepSeek and others are bypassing it entirely.
    • Part 3 examined how CUDA has become the dominant force in GPU computing, dissecting the layers of NVIDIA’s strategy.
    • In part 4, Chris addressed the question, “Is CUDA any good?”, digging into the perspectives of frequent users in the Gen AI ecosystem.
  • We're thrilled to introduce Paged Attention and Prefix Caching in MAX Serve, delivering cutting-edge optimizations for LLM inference.
  • Chris Lattner delivered a keynote on MAX and spoke on an expert panel on compute and at the Democratize Intelligence conference in San Francisco.
  • MAX Builds is now your go-to hub to get started building with MAX, featuring the latest GenAI models supported on both CPU and GPU, community-created packages, and application recipes.
  • Check out all of our new recipes, which are step-by-step guides to deploy GenAI using MAX:
    • Continuous Chat App With MAX Serve: build a chat application using Llama 3 and MAX by implementing efficient token management with rolling context windows, handling concurrent requests for optimal performance, and containerizing and deploying your application with Docker Compose.
    • Generate Embeddings with MAX Serve: run an OpenAI-compatible embeddings endpoint on MAX Serve with Docker, and generate embeddings with MPNet using the OpenAI Python client.
    • MAX Serve OpenAI Function Calling: explore LLM function calling with MAX Serve and llama3-8B on both CPU and GPU, demo OpenAI's function calling for interacting with external tools, and run a working example locally.
    • Offline Inference With MAX: use MAX to run inference with models from Hugging Face and generate text completions using the Llama-3.1 8B model.
    • Build Your Own AI Weather Agent: integrate OpenAI’s function calling with FastAPI and Llama 3.1 to build an interactive app that retrieves real-time data based on user queries.
    • Use Open WebUI With MAX Serve: use MAX Serve to create an OpenAI-compatible endpoint for Llama 3.1, set up Open WebUI for a robust chat interface, explore its RAG and web search capabilities, and configure the setup for multiple users.
    • MAX Serve Multi-Modal Structured Output: run a multimodal vision model with Llama 3.2 Vision and MAX Serve, implement structured output parsing with Pydantic models, convert image analysis into strongly-typed JSON, and leverage MAX Serve’s capabilities for multimodal models, type-safe parsing, and simple deployment with the magic CLI.

Awesome MAX + Mojo

Open-Source Contributions

If you’ve recently had your first PR merged, message Caroline Frasca in the forum to claim your epic Mojo swag!

Check out the recently merged contributions from our valuable community members:

Coming Up

Beyond CUDA: Accelerating GenAI Workloads with Modular’s MAX Engine

Join us on March 4th in San Francisco for an exclusive in-person talk with Chris Lattner, exploring how the MAX Engine accelerates GenAI workloads on both CPUs and GPUs—without relying on CUDA. Chris will break down the next-gen graph compiler and runtime behind MAX, designed to make high-performance AI more accessible across diverse hardware.

Modular at NVIDIA GTC

Planning to attend NVIDIA GTC? Stop by the Modular booth to connect with the team in-person! Modular's booth is #2315, and you can sign up for updates on Modular at GTC. We'll send you an email update as we get closer to the conference with a full schedule of the live demos at our booth, instructions to find our booth (#2315), and details on claiming your epic swag. Live demos at our booth will include programming GPUs with Mojo and deploying agent workflows on MAX.

  • Person with blonde hair using a laptop with an Apple logo.

    Sign up today

    Signup to our Cloud Platform today to get started easily.

    Sign Up

  • Magnifying glass emoji with black handle and round clear lens.

    Browse open models

    Browse our model catalog, or deploy your own custom model

    Browse models

Sign up for our newsletter

Get all our latest news, announcements and updates delivered directly to your inbox. Unsubscribe at anytime.

Thanks for signing up to our newsletter! 🚀

Thank you,

Modular Sales Team

Oops! Something went wrong while submitting the form.