惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Cybersecurity and Infrastructure Security Agency CISA
N
News and Events Feed by Topic
S
Securelist
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
Spread Privacy
Spread Privacy
T
Threat Research - Cisco Blogs
T
Tor Project blog
C
Cyber Attacks, Cyber Crime and Cyber Security
K
Kaspersky official blog
L
LINUX DO - 热门话题
T
The Exploit Database - CXSecurity.com
S
Schneier on Security
A
Arctic Wolf
Security Latest
Security Latest
T
Threatpost
P
Palo Alto Networks Blog
Simon Willison's Weblog
Simon Willison's Weblog
AWS News Blog
AWS News Blog
Cyberwarzone
Cyberwarzone
L
Lohrmann on Cybersecurity
P
Privacy International News Feed
V
Vulnerabilities – Threatpost
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Cisco Talos Blog
Cisco Talos Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
G
GRAHAM CLULEY
The Hacker News
The Hacker News
C
CERT Recently Published Vulnerability Notes
Know Your Adversary
Know Your Adversary
I
Intezer
Scott Helme
Scott Helme
T
Tenable Blog
NISL@THU
NISL@THU
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
C
Cisco Blogs
N
News and Events Feed by Topic
P
Proofpoint News Feed
P
Privacy & Cybersecurity Law Blog
Project Zero
Project Zero
Latest news
Latest news
Hacker News: Ask HN
Hacker News: Ask HN
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Forbes - Security
Forbes - Security
Security Archives - TechRepublic
Security Archives - TechRepublic
AI
AI
S
Security Affairs
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
D
Docker
P
Proofpoint News Feed
博客园 - Franky

IBM Research

All of AI benchmarking at your fingertips What are spin qubits? | IBM Quantum Computing Blog IBM to acquire HRL Laboratories IBM commits $50M in quantum access for US Genesis Mission It’s time for cryptography to get its own abstraction layer It’s time for cryptography to get its own abstraction layer This could be the largest synthetic code dataset yet How to measure the performance of a quantum computer | IBM Quantum Computing Blog Release News: Qiskit v2.5 is here! | IBM Quantum Computing Blog CoFrGeNets replace the ‘bones’ of transformer-based models How training environments can teach AI models to misbehave What’s new at IBM Quantum - Q2 2026 | IBM Quantum Computing Blog Modeling the chemistry of fusion reactor material | IBM Quantum Computing Blog Ponder This Challenge - July 2026 - Return of the Superheroes Apply to IBM Quantum Developer Conference 2026 | IBM Quantum Computing Blog Qiskit Paulice: postselected quantum error correction | IBM Quantum Computing Blog What is IBM’s nanostack chip architecture? IBM introduces the smallest computer chip in the world A new playbook for quantum optimization benchmarking Running AI on mixed hardware for speed and affordability Explore next-gen quantum algorithms with IBM Quantum Credits | IBM Quantum Computing Blog Allstate explores quantum computing for insurance portfolios | IBM Quantum Computing Blog Can LLMs discover quantum error correction codes? Prototype and validate fermionic circuits faster with ffsim | IBM Quantum Computing Blog Bringing the power of semantic AI to IBM Db2 The fast Fourier transform, how and why it works Building AI more like software The future of quantum takes center stage at NY Tech Week Qiskit Fall Fest 2026: Applications open | IBM Quantum Computing Blog IBM to invest $10 billion in quantum computing | IBM Quantum Computing Blog Renowned mathematician Subhash Khot joins IBM Research Ponder This Challenge - June 2026 - The Superhero Team Movies New Classroom Accounts expand quantum access for educators | IBM Quantum Computing Blog Qiskit Global Summer School 2026: Registration now open | IBM Quantum Computing Blog How researchers built a record-setting quantum circuit | IBM Quantum Computing Blog IBM charts a new research path with MIT How IBM is using quantum computing to understand the operating system of the universe How to use sample-based quantum diagonalization on IBM hardware Quantum-centric supercomputing simulates 12,635-atom protein | IBM Quantum Computing Blog A decade of quantum on the cloud | IBM Quantum Computing Blog Ponder This Challenge - May 2026 - The Powers of a Binary Matrix Where the frontiers of high-speed racing and computing meet Introducing the IBM Granite 4.1 family of models Building the future of computing, together Next-generation algorithms could move fusion from the lab to the grid Bringing quantum-centric supercomputing to Illinois What’s new at IBM Quantum - Q1 2026 | IBM Quantum Computing Blog Release News: Qiskit v2.4 is here! | IBM Quantum Computing Blog How IBM Quantum is enabling healthcare and biology research | IBM Quantum Computing Blog How an extra training step can unlock AI’s reasoning power IBM demonstrates extreme scale for content-aware storage with a 100-billion vector database Ponder This Challenge - April 2026 - The Unlabeled Clock IBM Research and ETH Zurich open a new era of innovation IBM’s newest time-series models cover a full range of enterprise prediction tasks Toward a transparent supply chain for AI Quantum computers take a step into real materials science Cleveland Clinic & IBM debut new quantum simulation workflow | IBM Quantum Computing Blog Turning turbulence into transcripts Like the information in a dream: IBM’s Charles H. Bennett receives ACM Turing award Doubling down on open-access quantum computing | IBM Quantum Computing Blog Unveiling the first reference architecture for quantum-centric supercomputing Realizing Feynman’s vision for the future of simulation | IBM Quantum Computing Blog IBM is working today to secure communication from tomorrow’s quantum risks Building PyTorch-native support for the IBM Spyre Accelerator Quantum simulates properties of the first-ever half-Möbius molecule, designed by IBM and researchers A look back at the International Year of Quantum | IBM Quantum Computing Blog TerraStackAI: Bringing Earth and space AI to Red Hat and the world Ponder This Challenge - March 2026 - Path game on a hole-riddled chessboard IBM demonstrates High NA EUV process capability on track for insertion below 2 nm nodes at SPIE 2026 Quantum Advantage Tracker: the race to advantage | IBM Quantum Computing Blog
Donating llm-d to the Cloud Native Computing Foundation
2026-03-24 · via IBM Research

Operationalizing AI inference is hard, especially with cutting-edge models and the infrastructure they require. New workloads are variable, and APIs don’t always make it possible to orchestrate inference. The cloud‑native world is racing to keep up with the demands of modern AI, and large language model (LLM) inference is one place where that pressure is felt most intensely.

As organizations push models into production, they’re discovering that serving LLMs at scale presents a new class of distributed systems challenges. That’s exactly the gap llm‑d was created to fill. llm-d addresses the limitations of traditional routing and autoscaling by offering a Kubernetes‑native distributed inference framework.

Today at KubeCon Europe, IBM Research, Red Hat, and Google Cloud announced the contribution of llm-d to the CNCF as a sandbox project. Launched as a collaborative effort with founding contributors NVIDIA and CoreWeave, and joined by industry leaders AMD, Cisco, Hugging Face, Intel, Lambda, and Mistral AI, alongside university supporters at the University of California, Berkeley, and the University of Chicago, the project has rapidly evolved into state-of-the-art AI infrastructure. This move marks a major milestone in IBM’s mission to make high‑performance, vendor‑neutral, Kubernetes‑native LLM inference accessible to everyone. By aligning with the CNCF, we’re doubling down on open governance, community‑driven development, and the belief that scalable generative AI should be a core feature of the cloud‑native ecosystem.

“llm-d bridges the gap between traditional distributed systems and the emerging AI inference stack, making large-scale model serving a first-class, cloud-native workload,” said Carlos Costa, a Distinguished Engineer at IBM Research who specializes in hybrid cloud platform for AI. “This donation helps establish the CNCF as a home for AI inference infrastructure, catalyzing a broader ecosystem of composable systems and projects.”

Any model, any accelerator, any cloud

The mission of llm-d from the outset was to build a vendor-neutral inference serving stack that can be used with any combination of hardware and software. llm‑d, which was launched in 2025, is a Kubernetes‑native, high‑performance distributed inference framework designed to make serving LLMs at scale both predictable and efficient. To support modern generative AI workloads, it provides a modular architecture that turns inference engines like vLLM into production-ready, distributed, cloud-native inference systems capable of sustaining low latency and high throughput under real-world traffic.

“In the most fundamental sense, we’re taking inference from just standing up something and playing with models, to running them in production at scale with multiple users and models,” said Priya Nagpurkar, vice president of AI platform at IBM Research. “You need the scale, distribution, and reliability of what Kubernetes provided for the previous era, while also recognizing that this is a very different workload,” she added.

At its core, llm‑d addresses the most challenging aspects of LLM inference, including KV‑cache locality management, balancing prefill and decode phases, coordinating multi‑node deployments, maintaining low latency, and efficiently utilizing heterogeneous accelerator hardware. “To deliver efficient inference, llm‑d introduces intelligent inference scheduling and prefix‑cache‑aware routing,” said Vita Bortnikov, IBM Fellow specializing in distributed AI inferencing at IBM Research. “This ensures that each request is routed to the most optimal replica based on cache state, traffic patterns, and hardware topology.”

“Another key capability of llm‑d is hierarchical KV‑cache offloading across GPU, CPU, and storage tiers,” she added. “This significantly improves performance, particularly for long‑context workloads and high levels of concurrency.”

Through prefill/decode disaggregation, llm‑d allows these two fundamentally different phases of inference to scale independently, dramatically improving efficiency for variable workloads. Autoscaling is traffic‑ and hardware‑aware, adapting to real‑time workload characteristics rather than relying on generic CPU and GPU metrics. llm‑d integrates deeply with emerging Kubernetes standards, including the Kubernetes Gateway API Inference Extension (GAIE) and LeaderWorkerSet (LWS), making distributed inference a first‑class Kubernetes workload.

A central promise of llm-d is that it will turn AI infrastructure from a black box into a replicable blueprint for manageable, cloud-native microservices. “This is a well-lit path,” Costa said. “We tested this for you. We benchmarked it. We went through the pain, and this is a path that we provide the community a clear path from experimentation to production.” With reproducible benchmarks, validated deployment patterns, and vendor‑neutral design, llm‑d provides a well‑lit path for organizations seeking production‑grade generative AI infrastructure across NVIDIA, AMD, Intel, and Google TPU accelerators.

A well-lit path

The CNCF is a natural venue for approaching this varied landscape. llm‑d was contributed to the CNCF as a sandbox project to accelerate the standardization, openness, and interoperability of distributed LLM inference across the cloud‑native ecosystem. As organizations race to operationalize generative AI, they’re discovering that LLM inference introduces challenges — stateful scheduling, KV cache locality, multi‑phase execution, heterogeneous accelerators — that expose limitations in the original workload model Kubernetes was designed around. These challenges are too fundamental and too shared to be solved inside a single company’s product roadmap. They require a neutral, community‑driven approach.

By contributing llm‑d to the CNCF, the project’s maintainers aim to establish a vendor‑agnostic, Kubernetes‑native blueprint for high‑performance inference that any organization can adopt. CNCF provides the governance model, IP clarity, and community trust needed for llm‑d to evolve from a promising framework into a widely accepted standard. While IBM, Red Hat, and Google are driving core contributions and early adoption, a growing ecosystem of collaborators is actively exploring integrations with the stack. CNCF stewardship ensures that no single vendor controls the project’s direction and that it remains aligned with upstream Kubernetes APIs such as the GAIE and LWS.

Joining the CNCF also strengthens llm‑d’s mission to create well‑lit paths for production‑grade AI infrastructure. The foundation’s ecosystem provides the ideal environment for building interoperable, standards‑driven components. Ultimately, contributing llm‑d to the CNCF is about ensuring that scalable, efficient, and portable LLM inference becomes a core capability of the cloud‑native stack, not a proprietary feature locked behind closed platforms.

What’s next?

Following the announcement of llm-d’s contribution to the CNCF, the project’s next phase will focus on deepening adoption, expanding technical capabilities, and strengthening its position as the neutral, open-governance inference stack for the AI ecosystem. The donation formalizes llm-d as a community project that will grow as more collaborators join, Costa said.

A key next step is collaborating to support next-generation AI architectures. For instance, Mistral AI is currently contributing features to the llm-d ecosystem to help advance open standards around disaggregated serving. "Creating a common foundation stack has already proven its value," said Costa. "It allows the entire ecosystem to focus on pushing the boundaries of the AI platform rather than rebuilding the basic building blocks."

At the same time, IBM Research will continue driving innovation, especially in areas where the industry lacks proven solutions. This includes work at the intersection of inference and training — reinforcement learning — as well as advancing self‑managing, AI‑guided optimization across caching, scaling, and configuration. The border between scale inference and model adaptation jobs is becoming blurrier, and an inference platform needs to be adapted to suit this reality.

As the project matures, the broader community is actively tackling the next generation of AI infrastructure challenges. The technical roadmap introduces standardized support for multi-modal workloads, expands integration to additional inference engines, and optimizes scheduling for multi-LoRA environments alongside advanced multi-tier KV cache offloading, ensuring llm-d meets the ecosystem’s evolving table-stakes expectations while pushing into new frontiers. Together, these steps position llm-d to evolve rapidly under CNCF governance and accelerate its role as the operating layer for distributed inference.