惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Threat Research - Cisco Blogs
H
Hacker News: Front Page
IT之家
IT之家
I
Intezer
GbyAI
GbyAI
MongoDB | Blog
MongoDB | Blog
博客园_首页
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
S
SegmentFault 最新的问题
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threatpost
Cisco Talos Blog
Cisco Talos Blog
C
Check Point Blog
P
Proofpoint News Feed
P
Privacy International News Feed
有赞技术团队
有赞技术团队
T
Tailwind CSS Blog
Scott Helme
Scott Helme
U
Unit 42
J
Java Code Geeks
W
WeLiveSecurity
H
Hackread – Cybersecurity News, Data Breaches, AI and More
C
CERT Recently Published Vulnerability Notes
小众软件
小众软件
The Hacker News
The Hacker News
L
LINUX DO - 热门话题
博客园 - 【当耐特】
G
Google Developers Blog
Latest news
Latest news
AWS News Blog
AWS News Blog
NISL@THU
NISL@THU
S
Secure Thoughts
P
Proofpoint News Feed
L
Lohrmann on Cybersecurity
F
Full Disclosure
S
Securelist
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Engineering at Meta
Engineering at Meta
Security Archives - TechRepublic
Security Archives - TechRepublic
人人都是产品经理
人人都是产品经理
T
Tor Project blog
Recent Announcements
Recent Announcements
Security Latest
Security Latest
N
News | PayPal Newsroom
A
About on SuperTechFans
Hugging Face - Blog
Hugging Face - Blog
Y
Y Combinator Blog
大猫的无限游戏
大猫的无限游戏
博客园 - Franky
T
The Blog of Author Tim Ferriss

The Register - Special Features: Supercomputing Month

GPUs aren't worth their weight in gold HPC won't be x86 forever – and it's starting to show Norway's new supercomputer to use waste heat to raise salmon The exascale offensive: America's race to rule AI HPC India satisfies its supercomputing needs, not its ambitions UK lines up £250M cloud procurement for AI research Eviden to build France's first exascale rig with AMD chips Power: The answer to and source of all your DC dilemmas GPU monsters eat supercomputing, legacy storage starves HPE details Vera Rubin blades for next-gen Cray Battery trade war hits booming datacenter market AI isn't throttling HPC. It is HPC Oak Ridge lab gets $125M to combine HPCs with quantum Power crunch threatens to derail AI datacenter construction $10B + spent on liquid cooling this week – it's only Tuesday Nvidia, OpenAI, and the trillion-dollar loop Nvidia will help build 7 AI supercomputers for for DoE HPE to build Discovery exascale successor for Oak Ridge NextSilicon eyes HPC market with Maverick-2 accelerators UK waves £750M supercomputer contract at HPC builders Tsunami forecasting to get faster thanks to El Capitan
How supercomputer filesystem DAOS breaks out of its niche
Chris Mellor Chris Mellor · 2025-11-25 · via The Register - Special Features: Supercomputing Month

DAOS has been a great success in the traditional HPC/supercomputing world, but is nowhere in the new, AI-focused, GPU supercomputing arena. What will it take for DAOS to find customers outside its high-end, legacy supercomputing niche?

The DAOS parallel filesystem has a strong IO500 presence, holding positions 1 (Argonne) and 2 (LRZ) in the current Production SC25 list. The two, according to HPE, combined have four times the storage benchmark score of the next 30 storage systems. DAOS also appears at number 13 (Zuse Institute, Berlin), and 17 (China Telecom Research Institute). The software appears more often in the full IO500 list, with 16 of the top 30 submissions using DAOS, and 26 of the top 45 being DAOS devotees.

DAOS has stronger still representation in the IO500 10-Node production list: systems with just 10 clients. It holds the top 3 positions plus number 6.

Top 10 IO500 10-Node production list entries

Top 10 IO500 10-Node production list entries

But DAOS is widespread, with 15 to 20+ production systems in active use. For its use to spread, it has to demonstrate, we understand, significantly better storage IO performance than competing software, meaning supporting more processing cores and delivering higher bandwidth. DAOS is open-source code and no single parallel processing storage system supplier is reliant on it. HPE has its ClusterStor as well as DAOS. DDN has its Lustre software, and VAST and WEKA each have their own software.

The IO500 measures storage IO performance while the TOP500 rates pure supercomputer power. Nvidia GPU systems are appearing in it, with the number 17 position held by CHIE-4, a DGX B200 system. Nvidia-based systems also appear at positions 17, 22, 24, 29, 30, and 32. AMD GPU systems are also appearing.

Enakta Labs co-founder Denis Nuja reckons that in supercomputers, “Lustre is still pretty much number one. … I haven't seen a GPFS (Storage Scale) system in a long time.”

Modern supercomputers, in his view, are "usually a two storage system." Nuja cites the recommended requirements for storage from AMD, based on a 10,000 GPU specification: "They have two specs. One is a production spec and the other one is a high resolution spec. … For production you need like five terabytes per second of reads, and two and a half terabytes per second of writes, which is fine. And then you look at the high resolution spec, they mentioned 40 terabytes per second of read and 20 terabytes a second of write," which is much higher.

“If you look at 40 terabytes per second reads and 20 [in a] commercially viable way,” suppliers may need to over-provision the system enormously, to meet the write target, Nuja said.

A Lustre system would have to deploy a lot of extra capacity to hit these numbers, for example, whereas DAOS can match these numbers.

Nuja’s view is that, in the supercomputing world, Lustre is by far the most popular filesystem in use, but it’s weak at the top end, where DAOS reigns, and getting nibbled away at the lower end and in the GPU supercomputer area by VAST Data and WEKA. He says that DAOS is great with single huge systems. Lustre, VAST and WEKA, he thinks, are excellent with smaller partitions: “They are effectively building a lot of smaller systems. So you can actually segregate and have 15 VAST clusters or Lustre or WEKA or whatever, and it's a very different story.”

DAOS would have an easier path to growth if it supported GPUs better, Nvidia GPUs specifically, but there is no GPUDirect support, for example.

Nuja thinks Nvidia, by moving to object storage, is “effectively democratizing storage … anybody can pretty much plug in, whereas the GPUDirect story and all that stuff was very much under the full control of Nvidia.”

Our thinking is that Nvidia has primarily and effectively supported, and helped, parallel filesystem storage suppliers who can help it sell lots of GPUs and its software. Market presence wins. From that point of view, DAOS, with 25 or fewer systems in deployment, represents small potatoes.

But object world interface requirements are different. Nuja tells us: “We have developed an S3 interface. We can plug it into everything that consumes S3. Nvidia AIStore consumes S3. So that should work [but] we haven't tested.”

PyTorch is another GPU system data access option in his view: “if you are using PyTorch for example; a lot of people are using PyTorch for computing on their GPUs, we built the whole integration for PyTorch. We natively integrate into PyTorch. So there's nothing preventing you from doing that. It has nothing to do with the GPU itself. It has to do with the framework that's on top of it.”

How could DAOS grow outside its top-end supercomputing storage niche?

Nuja said: “We need to build on more integrations and we need to get DAOS to a position where it's easily manageable, deployed, and everything, which is exactly what Enakta has been doing for two years.

“The third thing, which is the most problematic: it needs to become a choice in the minds of the users, end users. Because at the moment, all of them, when they hear about DAOS, say 'We heard about it, but it's something obscure.' …We need to educate the market, the end users, that DAOS is an option."

He added, “It's not only some kind of a science experiment, that one or two supercomputers use, and that takes time. It takes time, it takes effort. Let's be honest, the whole Optane debacle didn't really help. That's behind us, thankfully.”

Nuja continued: "Now there's a clear path forward. We're doing all we can to develop a lot of these interfaces that will help users be able to consume DAOS in a more sensible way. And we think that DAOS will continue having an edge versus other systems for the foreseeable future from a performance standpoint.” ®