惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Schneier on Security
Schneier on Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Tenable Blog
P
Proofpoint News Feed
V
Vulnerabilities – Threatpost
Project Zero
Project Zero
Latest news
Latest news
S
Schneier on Security
C
Cyber Attacks, Cyber Crime and Cyber Security
T
Threatpost
Simon Willison's Weblog
Simon Willison's Weblog
Cyberwarzone
Cyberwarzone
T
The Exploit Database - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
Spread Privacy
Spread Privacy
Cisco Talos Blog
Cisco Talos Blog
T
Troy Hunt's Blog
S
Secure Thoughts
C
Cisco Blogs
Application and Cybersecurity Blog
Application and Cybersecurity Blog
V2EX - 技术
V2EX - 技术
Hacker News: Ask HN
Hacker News: Ask HN
O
OpenAI News
L
LINUX DO - 最新话题
T
Threat Research - Cisco Blogs
Recent Commits to openclaw:main
Recent Commits to openclaw:main
P
Palo Alto Networks Blog
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
博客园_首页
月光博客
月光博客
博客园 - 【当耐特】
雷峰网
雷峰网
C
CXSECURITY Database RSS Feed - CXSecurity.com
博客园 - 叶小钗
aimingoo的专栏
aimingoo的专栏
L
Lohrmann on Cybersecurity
D
DataBreaches.Net
美团技术团队
B
Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
有赞技术团队
有赞技术团队
D
Docker
Jina AI
Jina AI
The GitHub Blog
The GitHub Blog
H
Hacker News: Front Page
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
AI
AI
Martin Fowler
Martin Fowler
Attack and Defense Labs
Attack and Defense Labs
小众软件
小众软件

The Next Platform: In-depth coverage of high end computing

Oak Ridge Starts Weaving Together A Quantum, Classical HPC, And AI System Stack Dell Bulks Up Hardware As AI Infrastructure Shifts To On-Premises Cisco Wins Over AI Customers With Merchant Silicon And Optics With Its IPO Done, Cerebras Can Get Back To Pushing The AI Envelope HPE Throws VM Users A Lifeline, Unifying Containers And VM Management In Cloud Stack OpenAI, Microsoft And Friends Build A Better, More Scalable Ethernet Compute And Memory Price Hikes Drive IT Spending Way Higher Sometimes, Air Is The Only Way For AI Systems To Keep Their Cool Arista Rides AI Scale Out Networks, Moves Into Scale Across, And Awaits Scale Up If You Can Make A Compute Engine, You Can Sell A Compute Engine Cleveland Clinic Simulates Large Proteins With Quantum-Centric Supercomputing Broadcom Helps CPU And XPU Makers Go Vertical With Compute Microsoft Committed To Doubling AI Infrastructure In Two Years Google Is A Full Stack AI Player, And Is Playing Well AWS Will Be An OEM, Just Like Google And Maybe Microsoft New Google Networks Tuned Up For GenAI Inference And Training Microsoft And OpenAI Remain Friends, Are Looking To Hook Up With Others AI-Driven CPU Shortage Saves Intel’s Financial Cookies The GenAI Battle Shifts From Frontier Models To Agentic Platforms With TPU 8, Google Makes GenAI Systems Much Better, Not Just Bigger Cisco Scales Out Quantum Systems With A Quantum Network Switch The Second Time Will Be The IPO Charm For Cerebras Imagine An Army Of AI Minions Handling Incident Response AI Will Soon Drive A Third Of TSMC’s Business Bechtolsheim & Friends Breathe Life Into Pluggable Optics One Last Time How HPC And AI Digital Twins Accelerate Quantum Error Correction The Embrace Of AI In Design Transforms Cadence And Its Customers Nvidia Brings The Power Of Open Source AI Models To Quantum Computing Building The Imperfect Beast For Enterprises, GPUs Need Virtualization As Much As CPUs Ever Did CoreWeave Takes As Much Financial Engineering As It Does Datacenter Design Contemplating Meta’s Homegrown MTIA Compute Engine Roadmap Most Neoclouds, Sovereigns, And Enterprises Will Buy, Not Build, Their AI Stacks Broadcom And Google Benefit Mightily From Anthropic’s Meteoric Growth Rebellions AI Rings Up The Money To Rack Up AI Inference Systems Nvidia Software Pushes MLPerf Inference Benchmarks To New Highs Broadcom Makes Its Pitch To Run Kubernetes On VMware VCF The $2 Billion Nvidia Deal With Marvell Is About A Lot More Than NVLink Fusion Classiq Says Quantum Is On Its Way, But Patience Is Needed Demonstrating The Scientific Usefulness Of Quantum Systems We Need Servers – Lots Of Servers. . . . Arm Comes Full Circle With Homegrown, AI-Tuned Server CPU Riding The Memory Boom And Trying To Avoid The Bust Data Analytics Helps Make The Mighty Lionesses Roar Driving Down The AI System Roadmap With Nvidia The Open Agentic AI World According To Nvidia Nvidia Finally Admits Why It Shelled Out $20 Billion For Groq Nvidia Says OpenClaw Is To Agentic AI What GPT Was To Chattybots IBM Unrolls Blueprint For Quantum-Classical HPC Computing Women Get Data-Driven Health Boost As The FA Tackles Sports Science Four Months Into Its Comeback, Zapata Stakes Its Claim In Quantum Software Eridu Cuts To The AI Networking Chase With High Radix Switch System HPE Works Harder And Smarter To Chase Datacenter Profits We Need A Proper AI Inference Benchmark Test How AI Is Boosting Gender Equality In High Performance Racing Custom Compute Engine Biz Growing More Than Marvell Ever Hoped Broadcom May Become The Biggest Counterbalance To Nvidia Ayar Labs Gets $500 Million To Ramp Photonics Into 2028 AI Systems With Cisco Outshift, Agentic AI Is Teed Up For the Internet Of Cognition Nvidia Sees The Light On Silicon Photonics And Maybe Optical Switching AI Servers Finally Dominate Dell’s Systems Business VAST Data: What Controls The Data Is More Important Than What Stores It So Far, Nobody Turns Tokens Into Money Like Nvidia SambaNova Pits Its Engineering Against Nvidia For Agentic AI Some More Game Theory, This Time On The AMD-Meta Platforms Deal AMD Says “Helios” Racks And MI400 Series GPUs On Track For 2H 2026 CPU-Only Compute Still Matters To A Lot Of HPC Centers Taalas Etches AI Models Onto Transistors To Rocket Boost Inference Some Game Theory On That Nvidia-Meta Platforms Partnership AI Eats The World, And Most Of Its Flash Storage The Current AI Networking Wave Will Be A Tsunami Of Money By 2027 The Memory Crunch Pinches Cisco’s Profits Only A Few AI Platforms Can Survive The Greatest AI Show On Earth Cisco Doubles Up The Switch Bandwidth To Take On AI Scale Out And Eventually Scale Up Datacenter Spending Forecast Revised Upwards – Yet Again The Twin Engine Strategy That Propels AWS Is Working Well With GenAI Turbochargers, Google Is Shifting Its Cloud Into A Higher Gear AMD Finally Makes More Money On GPUs Than CPUs In A Quarter Dassault And Nvidia Bring Industrial World Models To Physical AI TACC Explores Mixed Precision And FP64 Emulation For HPC With Horizon Robotics Will Break AI infrastructure: Here's What Comes Next Oracle’s Financing Primes The OpenAI Pump Gartner Takes Another Stab At Forecasting AI Spending Microsoft Is More Dependent On OpenAI Than The Converse Big Blue Poised To Peddle Lots Of On Premises GenAI Microsoft Takes On Other Clouds With “Braga” Maia 200 AI Compute Engines Nvidia’s $2 Billion Investment In CoreWeave Is A Drop In A $250 Billion Bucket Intel Is Still Struggling In The Datacenter, But It Could Get Better Is Nvidia Assembling The Parts For Its Next Inference Platform? TSMC Has No Choice But To Trust The Sunny AI Forecasts Of Its Customers Cerebras Inks Transformative $10 Billion Inference Deal With OpenAI By Decade’s End, AI Will Drive More Than Half Of All Chip Sales Startup Quantum Elements Brings AI, Digital Twins To Quantum Computing D-Wave Makes Gate-Model Power Move With Quantum Circuits Buy Building The Future Of Software In The AI-Native Era Arista Modular Switches Aim At Scale Across Networks, Hit Scale Out, Too NextSilicon Takes Aim At CPUs And GPUs With “Maverick-2” Dataflow Engine How HPC Is Igniting Discoveries In Dinosaur Locomotion – And Beyond Oracle First In Line For AMD “Altair” MI450 GPUs, “Helios” Racks
HPE Rides The Agentic AI Wave Back Into The Datacenter
Jeff Burt · 2026-06-22 · via The Next Platform: In-depth coverage of high end computing

During the AI era, the support systems of Hewlett Packard Enterprise have regularly processed billions of operational signals coming in from customers every day, and as those operational environments became more autonomous, executives for the IT gear supplier watched as the number of tokens those systems consumed scaled with the amount of signals.

Along with this, the per-token cost to HPE also steadily mounted, the result of what has become known as “tokenomics.”

In response to the rapidly rising costs, HPE engineers built an AI-first support platform on premises with the company’s GreenLake Intelligence – a framework of AI agents – and Private Cloud AI, an on-prem AI infrastructure engineered with Nvidia. Running the AI workloads on its own infrastructure gave HPE control over the economics of the work, according to Fidelma Russo, executive vice president, president and general manager of HPE’s hybrid cloud business unit and the company's chief technology officer.

“It allowed us to govern that really important customer data and it gave us better performance,” Russo said from the keynote stage at the recent HPE Discovery 2026 event in Las Vegas. “It also helped us significantly minimize the token spend associated with operating AI at scale. We stopped being consumers of AI and we became producers of intelligence.”

As a result, HPE lowered the cost by more than 30 times, saving nearly $100,000 a month, which she said gave the company the capacity needed to scale more quickly.

This is an example of the shift in the industry away from enterprises running their AI workloads primarily on the big clouds to building AI datacenters that run on-premises – and even reach out to the edge – and create more hybrid inferencing environments. There are several reasons for this, including data sovereignty and security, as well as latency. However, key among them are the rising costs associated with expansion of AI inferencing and the emergence of agentic AI.

“Once agents have continuous access to data, every interaction consumes a token,” Russo said. “That includes every decision, includes some validating the decision, and includes them taking the action. Unlike traditional AI, agents don't stop after one response. They continuously reason, they continuously coordinate, and they continuously interact with other systems. What that means is that inference is a continuous operational workload, not a one-time request. This brings us back to economics, because every time an AI system reasons, validates, acts, or takes action, it's consuming a token.”

The consumption piles up the costs very quickly, she said. What seems like a simple prompt can becomes thousands and millions of model interactions. For example, public data shows that OpenClaw – the widely popular virtual personal AI agent – has processed more than 600 billion tokens in a single month to support roughly 100 continuously operating coding agents, Russo said, which comes to about $13,000 per agent per month.

“Suddenly, AI economics looks a lot like infrastructure economics,” she said. “It comes down to utilization, efficiency, and scale, and how well we operate the full system and not just the model.”

According to validation firm Signal65, agents can utilize 4X to 15X times as many tokens than standard AI chat interactions, and as agentic workloads evolve, autonomous agents could push 1,000 times the inference demand than reasoning AI. All of this will drive the cost of inferencing and agents even higher.

“Training might happen in the cloud, and that part of the story remains largely true,” wrote Steve McDowell, founder and chief analyst with NAND Research. “But something unexpected is happening with inference. It's moving back on-prem and out to the edge, a quiet reversal that's forcing a fundamental rethink of enterprise AI architecture. The ‘cloud for everything’ approach that seemed inevitable just two years ago is proving impractical for production AI workloads. IT organizations are discovering that while cloud infrastructure excels at certain AI tasks, inference often works better closer to home.”

The trend has driven traditional OEMs to create infrastructure that can run these AI workloads on-premises and software tools that can manage the resulting hybrid environments that are arising, including in the form of AI factories created by the likes of not only HPE but also Dell Technologies and Cisco Systems in partnership with Nvidia. Executives speaking at Dell Technologies World 2026 in May highlighted their efforts to expand the vendor’s on-premises AI infrastructure capabilities, with founder and chief executive officer Michael Dell pointing to a study by the company that said 67 percent of AI workloads run outside the cloud and 88 percent of those surveyed said they are running at least one AI workload in their own datacenter.

And at its Cisco Live 2026 show earlier this month, Cisco Systems gave a similar assessment, putting a networking-heavy emphasis on the initiative while outlining the advantages of its large infrastructure, services, and software.

As we noted, HPE also put a focus on its networking as it continues to cross-pollinate its Aruba branch and campus networking lineup with the technology inherited when it bought Juniper Networks for $14 billion last year. However, the reached deeper into a range of other areas, from the cloud and the edge to software and security.

In software, a focus is on GreenLake Intelligence, with a central agent registry to ensure that organizations know not only what agents they have, but also where they are and what they’re permitted to do, and OpsRamp Copilot for AI agents and large language models to monitor utilization, token-based consumption, and keeping tabs on costs related to not only agents but also AI factories and workloads.

Morpheus 9, the latest iteration of the software for infrastructure automation that now includes Morpheus Central for federated manage of multiple sites, the Morpheus Orchestration Copilot that uses natural language for provisioning, and integrated software-defined networking.

“Morpheus Central provides a single operational layer across every distributed Morpheus instance, regardless of where it runs or what it manages,” Russo said.

HPE is integrating its Alletra MPX 10000 storage with the two-year-old Private Cloud AI to automatically apply policies for governance and metadata.

Russo said that every agent, inference, and AI workflow depends on context, the background information that makes up the systems working memory and situational awareness to generate relative and accurate responses. Systems that need to rebuild context every time it’s used, tokens are burned and processes are slowed down, with Russo noting that “in AI, memory is no longer a technical detail or a supply chain challenge. It is a strategic resource.”

KV cache removes the need to rebuild the context, meaning that storage becomes active memory for AI, which for HPE means its HPE Alletra Storage X10000 system becomes the repository for that information.

“Not simply to store data, but to keep data, context, and intelligence available wherever AI needs it,” she said. “It reduces the time GPUs spend waiting for data. It increases the amount of useful work your infrastructure can perform. And the X10K turns storage into an active part of AI efficiency and is a key contributor to the economic value of AI.”

In tests, Russo said the MX 10000 provides 20 times faster time-to-first-token and 17 higher throughput.

Private Cloud AI servers will come with Nvidia’s Agent Toolkit, which includes the GPU maker’s OpenShell secure runtime, NemoClaw blueprints, and Nemotron models to give developers the tools they need to design, build, run, and orchestrate large-scale multi-agent environments.

HPE’s upcoming ProLiant DL394 Gen12 server, due early next year, will include Nvidia’s new Vera CPUs that are used to support agents with rapid tool calls, orchestration, and real-time data processing.

These points and others address some of the ways HPE is trying to ease the path for enterprises to bringing AI infrastructure to internal datacenters, because that’s where the workloads are heading, according to Cheri Williams, senior vice president and general manager of HPE’s private cloud and flex solutions.

“There's still a place for experimentation and model training in the public cloud, and you still see customers doing that,” Williams said during a panel discussion. “But when it comes to production, on-prem is where most of the enterprise customers are going. The economics don't make sense to be in production in the public cloud.”