惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Threat Research - Cisco Blogs
C
Cybersecurity and Infrastructure Security Agency CISA
T
Tenable Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cisco Blogs
I
Intezer
Hacker News - Newest:
Hacker News - Newest: "LLM"
Hacker News: Ask HN
Hacker News: Ask HN
Schneier on Security
Schneier on Security
H
Heimdal Security Blog
Simon Willison's Weblog
Simon Willison's Weblog
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Cyberwarzone
Cyberwarzone
V2EX - 技术
V2EX - 技术
W
WeLiveSecurity
Help Net Security
Help Net Security
S
Secure Thoughts
P
Privacy & Cybersecurity Law Blog
S
Securelist
SecWiki News
SecWiki News
P
Palo Alto Networks Blog
C
CERT Recently Published Vulnerability Notes
Know Your Adversary
Know Your Adversary
The Last Watchdog
The Last Watchdog
N
News | PayPal Newsroom
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
V
Vulnerabilities – Threatpost
H
Hacker News: Front Page
NISL@THU
NISL@THU
Scott Helme
Scott Helme
L
LINUX DO - 热门话题
Attack and Defense Labs
Attack and Defense Labs
Security Archives - TechRepublic
Security Archives - TechRepublic
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
Google Online Security Blog
Google Online Security Blog
The Hacker News
The Hacker News
Cloudbric
Cloudbric
G
Google Developers Blog
Google DeepMind News
Google DeepMind News
N
News and Events Feed by Topic
A
Arctic Wolf
Latest news
Latest news
S
Schneier on Security
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
V
Visual Studio Blog
Project Zero
Project Zero
P
Privacy International News Feed
B
Blog
云风的 BLOG
云风的 BLOG

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Take control of your AI Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
SF Compute and Modular Partner to Revolutionize AI Inference Economics
No items found. · 2025-07-31 · via Modular Blog

Modular has partnered with SF Compute to address a fundamental asymmetry in the AI ecosystem: while model capabilities advance exponentially, the economic structures governing compute costs remain anchored in legacy paradigms. 

We’re excited to launch the Large Scale Inference Batch API – a high-throughput, asynchronous interface built for large-scale offline inference tasks like data labeling, summarization, and synthetic generation. At launch, it supports 20+ state-of-the-art models across language, vision, and multimodal domains, from efficient 1B models to 600B+ frontier systems. Powered by Modular’s high-performance inference stack and SF Compute’s real-time spot market, the API delivers up to 80% lower cost than typical market alternatives.

Try it today - we’re offering 10M’s of batch inference tokens for free to the first 100 new customers that get started now.

The Best Price-Performance for the Rest of Us

The economics of AI inference are fundamentally broken - characterized by underutilized hardware, rigid pricing, and infrastructure built for traditional AI workloads. The collaboration between SF Compute and Modular rethinks this from first principles, combining a real-time spot market for GPUs with an industry leading AI serving stack to unlock entirely new token economics for deploying AI inference at scale.

Model ID Hugging Face Name Size
DeepSeek‑R1 deepseek-ai/DeepSeek-R1 671B
DeepSeek‑V3 deepseek-ai/DeepSeek-V3 671B
DeepSeek‑R1‑Distill‑Llama‑70B deepseek-ai/DeepSeek-R1-Distill-Llama-70B 70B
Llama‑3‑70B‑chat meta-llama/Llama-3-70b-chat-hf 70B
Llama‑3.1‑405B‑Instruct meta-llama/Meta-Llama-3.1-405B-Instruct 405B
Llama‑3.1‑70B‑Instruct meta-llama/Meta-Llama-3.1-70B-Instruct 70B
Llama‑3.1‑8B‑Instruct meta-llama/Meta-Llama-3.1-8B-Instruct 8B
Llama‑3.3‑70B‑Instruct meta-llama/Meta-Llama-3.3-70B-Instruct 70B
Llama‑4‑Scout‑17B‑Instruct meta-llama/Llama-4-Scout-17B-16E-Instruct 109B
Llama‑4‑Maverick‑17B‑128E‑Instruct meta-llama/Llama-4-Maverick-17B-128E-Instruct 400B
Llama 3.2 Vision meta-llama/Llama-3.2-11B-Vision-Instruct 11B
Mistral‑7B‑Instruct mistralai/Mistral-7B-Instruct-v0.1 7B
Mixtral‑8x7B‑Instruct mistralai/Mixtral-8x7B-Instruct-v0.1 56B
Mistral‑Small‑24B‑Instruct mistralai/Mistral-Small-24B-Instruct-2501 24B
Qwen‑2.5‑72B‑Instruct Qwen/Qwen2.5-72B-Instruct 72.7B
Qwen‑2.5‑7B‑Instruct Qwen/Qwen2.5-7B-Instruct 7B
Qwen 3‑14B Qwen/Qwen3-14B 14.8B
Qwen 3‑8B Qwen/Qwen3-8B 8.2B
QwQ‑32B Qwen/QwQ-32B 32.5B
InternVL3‑9B OpenGVLab/InternVL3-9B 9B
InternVL3‑14B OpenGVLab/InternVL3-14B 14B
InternVL3‑38B OpenGVLab/InternVL3-38B 38B
InternVL3‑78B OpenGVLab/InternVL3-78B 78B
Gemma‑3‑12B‑in‑chat google/gemma-3-12b-it 12B
Gemma‑3‑27B‑in‑chat google/gemma-3-27b-it 27B

Under the hood, SF Compute provides real-time access to thousands of NVIDIA H100, H200, and AMD MI300/325X GPUs (coming soon) via its dynamic pricing marketplace. Spot rates are often below $1.40/hour–far below the current $6-$8/on-demand standard–and is even lower than the 1-3 year locked-in reserve pricing rate of traditional clouds. The Modular Platform complements this with compiler-native execution, GenAI-specific serving optimizations, and the world’s most performant AI kernels - achieving up to 60% higher throughput relative to existing industry infrastructure (NVIDIA, AMD).

Together, we are able to deliver a structural shift in how inference is monetized. While incumbents rely on fixed, over-provisioned infrastructure to preserve margins, we optimize for volume, efficiency, and developer value - collapsing the cost stack and returning the gains to users.

Modular + SF Compute: The Power of Unification

For years, AI development has been constrained by rigid hardware silos and inflexible cloud infrastructure. NVIDIA’s CUDA stack dominated the ecosystem, while AMD’s ROCm struggled to gain adoption despite competitive hardware. Meanwhile, traditional cloud platforms enforced fixed provisioning, long-term contracts, and static pricing models. The result: artificial scarcity, vendor lock-in, and concentrated pricing power that stifles innovation and inflates the true cost of AI deployment.

Our architecture treats this as an engineering problem rather than market reality. By combining SF Compute’s unified cloud marketplace with Modular’s hardware abstraction platform, we’ve built true fungibility across compute vendors. This required solving several uniquely challenging technical challenges:

  • Hardware unification: The Modular Platform provides unified model cluster, serving and kernel APIs  - delivering industry leading performance without sacrificing portability across heterogeneous hardware platforms.
  • Cloud unification: SF Compute’s platform abstracts physical infrastructure behind a programmable spot market and dynamic scheduler, enabling seamless allocation across heterogeneous compute backends.
  • Intelligent placement & routing: Together they automatically place workloads based on a multidimensional optimizer - factoring latency, bandwidth, model characteristics, and real-time market pricing.

The result: H100s, H200s, MI300/MI325Xs (coming soon), and next-gen accelerators compete in a single market on pure price-performance. For developers, hardware complexity disappears - models are routed to the best-fit resources, transparently. This isn’t just cost reduction – it’s a fundamental redefinition of how AI gets built and deployed.

Unlocking AI Innovation

The launch of the Large Scale Inference Batch API marks the first step in a larger transformation of AI infrastructure economics. Our roadmap targets key inefficiencies across the stack with innovations designed to unlock dramatically better cost-performance:

  • Long-running, online inference support for persistent, low-latency applications
  • Expanded model and hardware compatibility across the API and marketplace
  • Additional multi-cluster-level optimizations to further reduce compute overhead for large workloads

Start saving money today!

Ready to operate in a world where inference economics no longer constrain your ambitions? Reach out to the SF Compute team to get a quote for your use case.