惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cisco Talos Blog
Cisco Talos Blog
K
Kaspersky official blog
T
The Exploit Database - CXSecurity.com
NISL@THU
NISL@THU
AWS News Blog
AWS News Blog
V2EX - 技术
V2EX - 技术
Google DeepMind News
Google DeepMind News
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
S
Security @ Cisco Blogs
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Commits to openclaw:main
Recent Commits to openclaw:main
J
Java Code Geeks
Microsoft Azure Blog
Microsoft Azure Blog
Attack and Defense Labs
Attack and Defense Labs
Jina AI
Jina AI
The Last Watchdog
The Last Watchdog
W
WeLiveSecurity
H
Help Net Security
V
Visual Studio Blog
宝玉的分享
宝玉的分享
C
Cybersecurity and Infrastructure Security Agency CISA
T
Threat Research - Cisco Blogs
IT之家
IT之家
Hugging Face - Blog
Hugging Face - Blog
Latest news
Latest news
T
Tor Project blog
I
Intezer
美团技术团队
GbyAI
GbyAI
T
Tailwind CSS Blog
Last Week in AI
Last Week in AI
博客园 - 三生石上(FineUI控件)
Google DeepMind News
Google DeepMind News
Scott Helme
Scott Helme
Y
Y Combinator Blog
博客园 - 司徒正美
T
Tenable Blog
O
OpenAI News
N
News and Events Feed by Topic
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
V
Vulnerabilities – Threatpost
P
Palo Alto Networks Blog
博客园 - 聂微东
酷 壳 – CoolShell
酷 壳 – CoolShell
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threatpost
Google Online Security Blog
Google Online Security Blog
Apple Machine Learning Research
Apple Machine Learning Research
云风的 BLOG
云风的 BLOG
Help Net Security
Help Net Security

The Register - Special Features: Supercomputing Month

GPUs aren't worth their weight in gold HPC won't be x86 forever – and it's starting to show Norway's new supercomputer to use waste heat to raise salmon The exascale offensive: America's race to rule AI HPC India satisfies its supercomputing needs, not its ambitions UK lines up £250M cloud procurement for AI research How supercomputer filesystem DAOS breaks out of its niche Eviden to build France's first exascale rig with AMD chips Power: The answer to and source of all your DC dilemmas GPU monsters eat supercomputing, legacy storage starves HPE details Vera Rubin blades for next-gen Cray Battery trade war hits booming datacenter market AI isn't throttling HPC. It is HPC Oak Ridge lab gets $125M to combine HPCs with quantum Power crunch threatens to derail AI datacenter construction $10B + spent on liquid cooling this week – it's only Tuesday Nvidia, OpenAI, and the trillion-dollar loop Nvidia will help build 7 AI supercomputers for for DoE NextSilicon eyes HPC market with Maverick-2 accelerators UK waves £750M supercomputer contract at HPC builders Tsunami forecasting to get faster thanks to El Capitan
HPE to build Discovery exascale successor for Oak Ridge
2025-10-28 · via The Register - Special Features: Supercomputing Month

REG AD

Supercomputing Month

HPE's Discovery to succeed Frontier supercomputer with next-gen Cray tech

Oak Ridge's $500M system due in 2028, paired with a separate Lux AI cluster arriving two years earlier

HPE is set to build a successor to the Frontier exascale system for America's Oak Ridge National Laboratory, based on the next generation of its Cray supercomputer platform, plus a separate AI cluster to advance machine learning with a multi-tenant cloud-like platform.

The Discovery system will "bolster productivity up to 10x," according to HPE, and like many other supercomputers will be used for scientific research into various areas including medicine, cancer research, nuclear energy, and aerospace.

ORNL GX 3D system mock-up Discovery

Mock-up of HPE's forthcoming Discovery GX5000 system

Oak Ridge issued a request for proposals (RFP) for a successor to Frontier last year, with an expected delivery date of late 2027 to early 2028 and anticipated budget of $500 million.

REG AD

HPE now says delivery of Discovery is expected in 2028, with user operations set to begin in 2029.

REG AD

The national laboratory will also receive a second HPE-built system, Lux, the AI cluster intended to support both training and inference work at the site. This is expected to be installed early in 2026.

Discovery will be based on HPE's Cray Supercomputing GX5000, the next iteration of its supercomputing architecture, and will also feature a new Cray Storage Systems K3000 running the DAOS object storage platform, plus the next generation of Cray's Slingshot high-performance networking.

HPE says the Discovery nodes will be built with AMD's "Venice" (a code name) server processors, which are not due to be launched until next year, plus Instinct MI430X GPUs – also due next year – for the level of performance required for modeling, simulation, and AI projects.

However, HPE did not disclose how many nodes or CPUs and GPUs will go into building Discovery, or how much memory the system will have.

For interconnect, it will use the next generation of Slingshot networking HPE gained when it acquired Cray, although this has yet to launch and the company didn't give a date as to when it will. The current Slingshot 11 supports 200 Gbps per port, and can be regarded as a superset of Ethernet.

Discovery will be supported by Cray Storage Systems K3000, which HPE claims will support up to 75 million input/output operations per second per storage rack, 4x more performance than the next 30 storage systems on the IO 500 list, according to the firm.

This will be based on the open source DAOS (Distributed Asynchronous Object Storage) platform, but will complement rather than replace the Lustre file system-based Cray Storage Systems E2000, which will also be included in Discovery.

DAOS was developed by Intel, but farmed out to an independent foundation after the chipmaker canceled its Optane memory technology in 2022 and lost interest. HPE then hired Intel's DAOS engineers and brought them into its own storage team.

REG AD

Lux, meanwhile, is set to be an all-AMD affair, based on liquid-cooled HPE ProLiant Compute XD685 nodes with Epyc CPUs, Instinct MI355X GPUs, and linked together using AMD's Pensando SmartNIC networking.

Liquid cooling innovations

Crosshead text

Trish Damkroger, HPE's senior VP for HPC and AI Infrastructure Solutions, told The Register that the GX5000 had been in the works for years, but the company had "made some pivots over the last year and a half, as we've seen the growth of TDPs (thermal design points), the growth of different silicon coming out from all the vendors, and the need to be able to support all of these different workloads."

She said the racks will be able to accommodate up to 25 kilowatts per compute slot, 127 percent higher than before. But she seemed prouder of the liquid cooling for the GX5000 infrastructure, which now supports 40°C (104°F) water to meet new energy requirements for a lot of customers in Europe.

This means additional chillers and refrigerators are not needed, which cuts power, so it is a much more energy-efficient system for upcoming deployments.

"It is a bookend design," she said. "So basically, the cooling pump is designed to be more compact. And can be placed on the side of the system instead of in the middle. And each pump is going to have redundancy to ensure that there's always-on operation."

HPE next-gen cooling

HPE next-gen cooling

Damkroger added that users can now control the water flow rate, so instead of every single blade having the same, it can be optimized for each blade and its workloads.

REG AD

HPE said there will be an opportunity to see the new GX5000 infrastructure at the SC 25 high-performance compute conference in St. Louis, Missouri, next month, though the platform is not expected to be available to customers until early 2027. ®