惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
J
Java Code Geeks
I
InfoQ
V
Visual Studio Blog
M
MIT News - Artificial intelligence
H
Help Net Security
博客园_首页
Blog — PlanetScale
Blog — PlanetScale
F
Fortinet All Blogs
Apple Machine Learning Research
Apple Machine Learning Research
人人都是产品经理
人人都是产品经理
G
Google Developers Blog
A
About on SuperTechFans
腾讯CDC
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Last Week in AI
Last Week in AI
小众软件
小众软件
aimingoo的专栏
aimingoo的专栏
罗磊的独立博客
大猫的无限游戏
大猫的无限游戏
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
云风的 BLOG
云风的 BLOG
S
SegmentFault 最新的问题
WordPress大学
WordPress大学

SiliconANGLE

Will agentic AI governance run amok? The lesson of Asimov’s Three Laws - SiliconANGLE AI + quantum, Amazon vs. Starlink and the wide-open US-China internet battle - SiliconANGLE Team Cymru launches Total Insights Feed to replace legacy threat intelligence lists - SiliconANGLE AI Mode in Chrome adds split-screen view to enhance the web search experience - SiliconANGLE Resolve AI raises $40M at $1.5B valuation to optimize production environments - SiliconANGLE How Zscaler and OpenAI turn zero-trust security into an AI accelerator - SiliconANGLE OpenAI ratchets up Codex's agentic capabilities to rival Claude Code - SiliconANGLE Anthropic launches Claude Opus 4.7 with coding, visual reasoning improvements - SiliconANGLE Slash raises $100M at a $1.4B valuation to expand AI-powered banking platform for online businesses - SiliconANGLE Canva unveils Canva AI 2.0, recasting its platform as an agentic system for work - SiliconANGLE Data center, consumer device chips boost TSMC’s revenue - SiliconANGLE Mission-critical security cannot be bolted on, says Oracle - SiliconANGLE Agentic infrastructure reshapes enterprise AI - SiliconANGLE Data quality, and data freedom, foundational for AI success - SiliconANGLE Data trust is a bedrock in successful, scalable AI outcomes - SiliconANGLE Google introduces new agentic AI-ready tools and resources for Android developers  - SiliconANGLE Agentic AI orchestration separates winners from laggards - SiliconANGLE Data-driven tools turning the tide against human trafficking - SiliconANGLE Achieving trusted AI development goes beyond 'vibes' - SiliconANGLE Impinj boosts edge computing power in updated R700 RAIN RFID reader - SiliconANGLE Certinia powers professional services with AI - SiliconANGLE Antioch prepares to accelerate simulated testing for autonomous robots after raising $8.5M - SiliconANGLE Developer tooling startup Expo nabs $45M investment - SiliconANGLE Solidroad lands $25M to bring AI to customer support interactions - SiliconANGLE DuploCloud lands compliance and AI governance certifications as enterprise buyers tighten scrutiny - SiliconANGLE Lua lands $5.8M to help businesses build and manage AI agent workforces - SiliconANGLE Best of frenemies: Oracle's and AWS' clouds unite with dedicated, private connectivity - SiliconANGLE NIST shifts National Vulnerability Database to risk-based triage as CVE submissions hit record levels - SiliconANGLE Cisco goes to the races with new Churchill Downs multiyear partnership - SiliconANGLE Susecon 2026 will tackle the future of open-source platforms - SiliconANGLE
Scalable AI inference in focus amid move beyond GPUs - Si...
Ryan Stevens · 2026-05-13 · via SiliconANGLE

Red Hat and Intel spotlight scalable AI inference as enterprises move beyond the GPU gold rush

As companies move from testing AI to broader adoption, the biggest challenge is building scalable AI inference systems that perform without breaking the budget. The next wave of AI won’t be won on raw power alone — it will be decided by who can do more with less.

When AI inference first took off, the focus was on deploying the largest possible models across massive GPU clusters following the rise of ChatGPT and open-weight models. That’s when customers turned to Red Hat Inc., looking for ways to scale those models across platforms like Red Hat Enterprise Linux and OpenShift without sacrificing control or cost efficiency, according to Taneem Ibrahim (pictured, right), director of engineering for AI inference at Red Hat.

“That’s when the friction moment came in for us, like, ‘How do I take this project — called vLLM, [which] we’re the largest commercial contributor to — and work it at scale with a project like llm-d?’” Ibrahim said. “How you drive the cost per token down so that you can operationalize your AI, you can govern your AI [and] you can deploy it at scale?”

Ibrahim and Bill Pearson (left), vice president of data center and AI at Intel Corp., spoke with theCUBE’s Rob Strechay and Rebecca Knight at Red Hat Summit 2026, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed the development of scalable AI inference systems and the growing role of open-source, CPU-driven AI deployments. (* Disclosure below.)

Scalable AI inference shifts infrastructure priorities

As agentic AI reshapes infrastructure demands, CPUs are playing a bigger role than they did during the earlier GPU-heavy phase of adoption, Pearson explained. Companies are now focused on finding the right balance of both to meet performance needs efficiently, underpinning Red Hat and Intel’s latest collaboration in bringing full vLLM support for Intel Xeon to Red Hat AI 3.4.

“It isn’t a one-size-fits-all approach, but rather, ‘What’s my workload? What’s the outcome I’m looking for?’” he said. “’How do I put together the right combination of hardware and software to go and deliver that outcome?’”

Part of that calculus is recognizing the hardware companies already have. CPUs are already deployed across most data centers, and a growing share of inference workloads — particularly agentic tasks like tool calling and data orchestration — don’t require GPUs at all. That frees up GPU capacity for the heavy lifting, according to Pearson.

“As we’ve gone through this with our customers in the industry, we’ve seen that people often have just assumed, ‘I’ve got a hammer. I need the nail to hit it with,'” he said. “Once they take a step back to say, ‘Wait a minute. I have these CPUs in my data center’ — or, ‘I need to figure out how to balance the right number of CPUs with the right number of GPUs to achieve that outcome I’m looking for’ — they’re actually going to get better results at a better price point for delivering lower-cost tokens.”

Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of Red Hat Summit 2026:

(* Disclosure: Red Hat sponsored this segment of theCUBE. Neither Red Hat nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)

Photo: SiliconANGLE

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

About SiliconANGLE Media

SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.