惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
G
GRAHAM CLULEY
www.infosecurity-magazine.com
www.infosecurity-magazine.com
V2EX - 技术
V2EX - 技术
The Last Watchdog
The Last Watchdog
S
Secure Thoughts
Webroot Blog
Webroot Blog
PCI Perspectives
PCI Perspectives
L
LINUX DO - 最新话题
Hacker News: Ask HN
Hacker News: Ask HN
N
News and Events Feed by Topic
H
Heimdal Security Blog
H
Help Net Security
T
The Blog of Author Tim Ferriss
P
Proofpoint News Feed
The GitHub Blog
The GitHub Blog
Jina AI
Jina AI
Recent Commits to openclaw:main
Recent Commits to openclaw:main
F
Full Disclosure
小众软件
小众软件
S
Securelist
罗磊的独立博客
NISL@THU
NISL@THU
D
Darknet – Hacking Tools, Hacker News & Cyber Security
C
Cisco Blogs
云风的 BLOG
云风的 BLOG
C
CERT Recently Published Vulnerability Notes
Cisco Talos Blog
Cisco Talos Blog
Know Your Adversary
Know Your Adversary
S
Schneier on Security
D
DataBreaches.Net
M
MIT News - Artificial intelligence
V
Vulnerabilities – Threatpost
N
News and Events Feed by Topic
有赞技术团队
有赞技术团队
F
Fortinet All Blogs
T
Tenable Blog
The Register - Security
The Register - Security
C
Check Point Blog
AWS News Blog
AWS News Blog
Cloudbric
Cloudbric
C
CXSECURITY Database RSS Feed - CXSecurity.com
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
Cyber Attacks, Cyber Crime and Cyber Security
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Google Online Security Blog
Google Online Security Blog
博客园 - 叶小钗
Hacker News - Newest:
Hacker News - Newest: "LLM"
博客园 - 司徒正美

Modular Blog

Qualcomm to Acquire Modular Modular 26.4: SOTA MoE Serving, Model Bringup via Agent Skills, Mojo 1.0 Beta 2 and More ModCon 2026: Modular’s Developer Conference Day Zero: MiniMax M3 Open Weights on Modular Cloud Modverse #55: Mojo 1.0 Beta, Community Mojo Libraries, and Real-Time Patient Conversations Powered by MAX What about OpenCL and CUDA C++ alternatives? (Democratizing AI Compute, Part 5) Why LLM Inference Needs a New Kind of Router - Part 3 Three trends from MLSys 2026 Why LLM Inference Needs a New Kind of Router - Part 2 How I built a pure Mojo app (and 10 libraries) with AI agents Hippocratic AI partners with Modular to power flexible, high-quality inference for real-time patient conversations Translating to Mojo via AI Agents Inkwell: Why Your Inference Platform Matters As Much As Your Model Why LLM Inference Needs a New Kind of Router - Part 1 Modular 26.3: Mojo 1.0 Beta, MAX Video Gen, and more Modverse #54: AMD AI DevDay, New Modular Offices, and a Community That Keeps Shipping How Frontier Coding Agents Built a Video Diffusion Pipeline on MAX TileTensor Part 1 - Safer, More Efficient GPU Kernels Modular Opens Edinburgh & San Francisco Offices Structured Mojo Kernels Part 4 - Portability and the Road Ahead Day Zero Launch: Fastest Performance for Gemma 4 on NVIDIA and AMD Modverse #54: From GTC to Edinburgh, a Community Building Momentum Software Pipelining for GPU Kernels: Part 1 - The Pipeline Problem Structured Mojo Kernels Part 3 - Composition in Practice Modular 26.2: State-of-the-Art Image Generation and Upgraded AI Coding with Mojo Modular at NVIDIA GTC 2026: MAX on Blackwell, Mojo Kernel Porting, and DeepSeek V3 on B200 Structured Mojo Kernels Part 2 - The Three Pillars Modverse #53: Community Builds, Research Milestones, and a Growing Ecosystem Structured Mojo Kernels Part 1 - Peak Performance, Half the Code The Claude C Compiler: What It Reveals About the Future of Software BentoML Joins Modular The Five Eras of KVCache Modular 26.1: A Big Step Towards More Programmable and Portable AI Infrastructure How to Beat Unsloth's CUDA Kernel Using Mojo—With Zero GPU Experience 🔥 Modular 2025 Year in Review The path to Mojo 1.0 Modverse #52: Advancing AI Together — Community Projects & Platform Milestones Modular 25.7: Faster Inference, Safer GPU Programming, and a More Unified Developer Experience "TTS 1 Max" (powered by Modular Platform) Ranked #1 Speech Model on Artificial Analysis PyTorch and LLVM in 2025 — Keeping up With AI Innovation Achieving State-of-the-Art Performance on AMD MI355 — in Just 14 Days Modular Raises $250M to scale AI's Unified Compute Layer Modular 25.6: Unifying the latest GPUs from NVIDIA, AMD, and Apple Matrix Multiplication on Blackwell: Part 4 - Breaking SOTA Modverse #51: Modular x Inworld x Oracle, Modular Meetup Recap and Community Projects Matrix Multiplication on Blackwell: Part 3 - The Optimizations Behind 85% of SOTA Performance Matrix Multiplication on Blackwell: Part 2 - Using Hardware Features to Optimize Matmul Matrix Multiplication on Blackwell: Part 1 - Introduction Modverse #50: Modular Platform 25.5, Community Meetups, and Mojo's Debut in the Stack Overflow Developer Survey Modular Platform 25.5: Introducing Large Scale Batch Inference SF Compute and Modular Partner to Revolutionize AI Inference Economics AI Agents for AWS Marketplace Modverse #49: Modular Platform 25.4, Modular 🤝 AMD, and Modular Hack Weekend Inside Modular Hack Weekend: Top Projects and Community Highlights How is Modular Democratizing AI Compute? (Democratizing AI Compute, Part 11) Modular 25.4: One Container, AMD and NVIDIA GPUs, No Lock-In Introducing Mammoth: Enterprise-Scale GenAI Deployments Made Simple Modular + AMD: Unleashing AI performance on AMD GPUs Modverse #48: Modular Platform 25.3, MAX AI Kernels, and the Modular GPU Kernel Hackathon Exploring Metaprogramming in Mojo Modular GPU Kernel Hackathon Highlights: Innovation, Community, & Mojo🔥 Modular’s bet to break out of the Matrix (Democratizing AI Compute, Part 10) Modular Platform 25.3: 450K+ Lines of Open Source Code and pip Packaging A New, Simpler License for MAX and Mojo Why do HW companies struggle to build AI software? (Democratizing AI Compute, Part 9) Modverse #47: MAX 25.2 and an evening of GPU programming at Modular HQ What about the MLIR compiler infrastructure? (Democratizing AI Compute, Part 8) What about Triton and Python eDSLs? (Democratizing AI Compute, Part 7) MAX 25.2: Unleash the power of your H200's–without CUDA! What about TVM, XLA, and AI compilers? (Democratizing AI Compute, Part 6) Modverse #46: MAX 25.1, MAX Builds, and Democratizing AI Compute CUDA is the incumbent, but is it any good? (Democratizing AI Compute, Part 4) MAX 25.1 - Introducing MAX Builds How did CUDA succeed? (Democratizing AI Compute, Part 3) Paged Attention & Prefix Caching Now Available in MAX Serve What exactly is “CUDA”? (Democratizing AI Compute, Part 2) Modular DeepSeek's Impact on AI (Democratizing AI Compute, Part 1) Modular Hands-on with Mojo 24.6 Evaluating Llama Guard with MAX 24.6 and Hugging Face Modular Introducing MAX 24.6: A GPU Native Generative AI Platform MAX GPU: State of the Art Throughput on a New GenAI platform Understanding SIMD: Infinite Complexity of Trivial Problems Community Spotlight: Writing Mojo with Cursor Hands-on with Mojo 24.5 MAX 24.5 - With SOTA CPU Performance for Llama 3.1 Announcing stack-pr: an open source tool for managing stacked PRs on GitHub Debugging in Mojo🔥 Write hardware-agnostic custom ops for PyTorch | Modular Develop locally, deploy globally A brief guide to the Mojo n-body example What's new in MAX 24.4? MAX on macOS, fast local Llama3, native quantization and GGUF support What’s new in Mojo 24.4? Improved collections, new traits, os module features and core language enhancements MAX 24.4 - Introducing quantization APIs and MAX on macOS Deep dive into ownership in Mojo What ownership is really about: a mental model approach Fast⚡k-means clustering in Mojo🔥: a guide to porting Python to Mojo🔥 for accelerated k-means clustering
Take control of your AI
No items found. · 2024-07-09 · via Modular Blog

In today’s rapidly evolving technology landscape, adopting and rolling out AI to enhance your enterprise is critical to improving your organization’s productivity and ensuring that you are delivering a world-class product and service experience to your customers. AI is without question, the single most important technological revolution of our time—representing a new technology super-cycle that your enterprise cannot be left behind on.

Today, enterprises are trying to adopt AI at an unprecedented pace — in fact, the latest research offered by Bain & Company suggests:

87% of companies surveyed by Bain said that they were already developing, piloting, or have deployed generative AI in some capacity, with most of these early deployments in software code development, customer service, marketing and sales, and product differentiation.

__wf_reserved_inherit

Bain & Company: AI Survey: Four Themes Emerging

While the benefits of AI are increasingly clear, a notable trend is also starting to emerge: enterprise teams are now assessing how to own and control their AI systems rather than relying on third-party providers like OpenAI.

Bringing AI in-house

As the race to deploy AI into enterprise exploded rapidly, many enterprises tried to adopt AI as fast as possible without considering the broader impacts to their organizations and the impact on their employees, products, and customers. Now, as enterprises see their AI efforts maturing, and with more proof of concepts (POCs) moving to production, they are increasingly asking what they need to do to scale AI inside their organizations.

We regularly talk to enterprises about their AI needs, and here is a list of what enterprise customers tell us is important to them:

Customization and flexibility

Product and engineering teams need the ability to customize their AI deployments — ensuring insight into data processing methods, model formats, AI pipelines, training methods, model serving, production monitoring and more. This level of control and customization is unattainable with off-the-shelf, or with third-party AI solutions. Product and engineering teams must innovate without constraints, deploying AI and customizing the implementation for your organizational needs.

Intellectual property protection

You wouldn’t trust your most critical IP to a third party, and AI will increasingly become some of your most important IP. Over time, enterprise teams will create an increasing number of proprietary algorithms and approaches. By controlling enterprise AI, you can own and scale this important intellectual property, maintaining or growing your enterprise's competitive edge.

Innovation and agility

The ability to experiment, iterate, and adapt quickly is crucial for determining where AI can help your organization. Owning your AI systems ensures that you can foster an environment of continuous innovation, where organizational teams can explore new applications and rapidly respond to market changes. Third-party solutions can be short-lived, as they constrain innovation to third-party update cycles and schedules. Empowering your organization to own its AI enables your enterprise to drive innovation forward quickly.

Resource allocation & cost efficiency

While establishing in-house AI capabilities requires upfront investment, rapidly deploying AI across an enterprise is expensive. Third-party services often come with recurring fees and scaling costs that can escalate unpredictably — particularly on a token or context basis. Owning your AI future means that you can manage cost and scale AI within resource allocation and budgetary control standards. Escalating costs from a lack of control of your AI infrastructure systems often means months or years of challenges as you race to reduce your third-party AI dependency.

Data privacy and security

Safeguarding data privacy and security is paramount. AI systems thrive on vast datasets, often containing sensitive and proprietary information. By maintaining control over your AI infrastructure, your teams can implement robust security measures tailored to their unique needs. This control mitigates the risk of data breaches and unauthorized access — issues that can be magnified when outsourcing AI solutions. Ensuring that you control AI in-house allows your product and engineering teams to build a fortress around their data, ensuring its integrity and confidentiality.

Compliance and regulatory alignment

For enterprises in finance, healthcare, and telecommunications - there are real and stringent regulations governing data handling and processing. In these sectors, enterprise teams face the difficult challenge of ensuring AI systems not only scale correctly, but also comply with strict state and federal legal requirements. By owning your AI, you can directly implement and monitor compliance measures - reducing the risk of regulatory breaches and associated penalties.

Data quality and bias mitigation

AI systems are deeply reliant on data and its quality. Enterprises with control over their AI infrastructure have greater oversight of data preprocessing and cleaning procedures, ensuring high data integrity standards which improves AI model quality. Further, this enables active identification and mitigation of biases in AI models, delivering fairer and more accurate outcomes. The lack of transparency into third-party services, indicates that you do not have deep insight into what is actually being done with your data.

Integration challenges

Many enterprises operate within complex IT ecosystems that blend legacy systems, databases, and applications with modern code bases. A lack of control over AI workloads means its can be hard to scale AI into these environments. Controlling your AI deployment infrastructure enables you minimize integration challenges, building a more cohesive technology stack.

Building internal expertise

Without question, growing and scaling your AI talent ensures you can constantly scale with the latest AI breakthroughs. Building an AI-first enterprise requires deep investment in product and engineering teams and a commitment to developing expertise to scale AI systems. This will only make your organization more self-sufficient and resilient. All your teams must be AI teams, and fostering a strong internal culture of innovation and talent development ensures you can own your AI future.

How does MAX help?

At Modular, we have been at the forefront of building infrastructure that hands back ownership and control to enterprises seeking to own their AI future. As leaders in AI infrastructure who helped build the infrastructure that shaped today’s AI industry, we have taken a new approach to help developers and enterprises answer the critical question: What if deploying AI workloads into production was so simple you could do it yourself?

That’s why we built MAX - the Modular Accelerated Xecution Platform.

__wf_reserved_inherit

MAX gives you everything you need to deploy low-latency, high-throughput AI applications into production with minimal effort. More specifically, it provides the following:

Industry standard protocols

MAX works with the industry-standard APIs and protocols that have become ubiquitous with the mass adoption of Generative AI and LLMs, including the OpenAI’s completion and chat APIs. Importantly, this significantly reduces the cost of migrating your application code to your AI systems, giving you the flexibility to choose the best engine for your needs.

Portable, performant model execution

MAX provides a state-of-the-art AI compiler and runtime that optimizes the latency and throughput of popular open source AI models like Llama3-8B and Gemma2 across a wide range of AI hardware, from local laptops to common cloud instances. MAX enables you to seamlessly move the same model, without code changes, across a wide range of CPU architectures—Intel, AMD, ARM—and GPUs, allowing you to take advantage of the breadth and depth of different cloud instances at the best price, and always get the best inference cost-performance ratio across cloud environments.

Compatible with what you already use

For most organizations experimenting with GenAI and LLMs, this isn’t their first foray into AI. Many of them already have traditional AI models scaled in production. These organizations have been forced to split their infrastructure efforts, despite already having well established standard infrastructure. MAX standardizes this existing infrastructure and extends it to GenAI and LLMs, replacing the parts of the stack that matter most. MAX is compatible with all PyTorch and ONNX models, including object detection, recommenders, and much more. It integrates with industry standard technologies such as Triton Inference Server, Docker, Prometheus, and Grafana.

Composable abstractions that allow for extensibility

For Enterprises looking to get their hands dirty and adapt GenAI and LLMs to their needs, MAX provides clean, composable abstractions, including the Serve, Engine, Drive, Graph, and Extensibility APIs. These abstractions enable users to go beyond stock LLMs and build a moat for their business in areas of AI that have yet to become mainstream.

Deploy to your own VPC or data center

Finally, MAX is deployable into any cloud or on-prem environment. Enterprises maintain data sovereignty, privacy, and control over how their data is used and how their AI services are scaled. When they’re ready, users can adopt the Enterprise Edition and get world-class support from the team that scaled Google’s AI.

MAX is free! Download now

The decision for an enterprise to control its own AI is more than a strategic one — it's a call to action for the entire organization to become AI first. It signifies a commitment to data security, customization, and innovation. By taking charge of AI, enterprises empower their entire organization to craft sophisticated solutions that drive long-term success. This approach protects valuable data and intellectual property. It positions their team at the forefront of technological advancement—ready and able to drive innovation and seize the AI opportunities that lie ahead.

By adopting MAX in your enterprise, you can drive this innovation across your entire organization, ensuring that you are in a position to control for your AI future. Learn more here, and contact us if you need help.