惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LINUX DO - 最新话题
Microsoft Azure Blog
Microsoft Azure Blog
The Register - Security
The Register - Security
Vercel News
Vercel News
Cloudbric
Cloudbric
Recent Announcements
Recent Announcements
O
OpenAI News
Y
Y Combinator Blog
Hacker News: Ask HN
Hacker News: Ask HN
I
InfoQ
WordPress大学
WordPress大学
S
Secure Thoughts
P
Proofpoint News Feed
Hacker News - Newest:
Hacker News - Newest: "LLM"
小众软件
小众软件
I
Intezer
博客园 - 司徒正美
Google DeepMind News
Google DeepMind News
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
G
GRAHAM CLULEY
P
Palo Alto Networks Blog
Attack and Defense Labs
Attack and Defense Labs
MyScale Blog
MyScale Blog
P
Proofpoint News Feed
美团技术团队
V
Vulnerabilities – Threatpost
Scott Helme
Scott Helme
C
Check Point Blog
H
Help Net Security
博客园 - Franky
AWS News Blog
AWS News Blog
MongoDB | Blog
MongoDB | Blog
Project Zero
Project Zero
N
Netflix TechBlog - Medium
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
V2EX - 技术
V2EX - 技术
Cisco Talos Blog
Cisco Talos Blog
Latest news
Latest news
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Last Week in AI
Last Week in AI
V
V2EX
C
Cybersecurity and Infrastructure Security Agency CISA
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
Hugging Face - Blog
Hugging Face - Blog
The Hacker News
The Hacker News
Schneier on Security
Schneier on Security
PCI Perspectives
PCI Perspectives
Apple Machine Learning Research
Apple Machine Learning Research

The Next Platform: In-depth coverage of high end computing

Uncle Sam Awards $2 Billion-Plus To Quantum Companies, But Wants A Cut Oak Ridge Starts Weaving Together A Quantum, Classical HPC, And AI System Stack Dell Bulks Up Hardware As AI Infrastructure Shifts To On-Premises Cisco Wins Over AI Customers With Merchant Silicon And Optics With Its IPO Done, Cerebras Can Get Back To Pushing The AI Envelope HPE Throws VM Users A Lifeline, Unifying Containers And VM Management In Cloud Stack Compute And Memory Price Hikes Drive IT Spending Way Higher Sometimes, Air Is The Only Way For AI Systems To Keep Their Cool Arista Rides AI Scale Out Networks, Moves Into Scale Across, And Awaits Scale Up If You Can Make A Compute Engine, You Can Sell A Compute Engine Cleveland Clinic Simulates Large Proteins With Quantum-Centric Supercomputing Broadcom Helps CPU And XPU Makers Go Vertical With Compute Microsoft Committed To Doubling AI Infrastructure In Two Years Google Is A Full Stack AI Player, And Is Playing Well AWS Will Be An OEM, Just Like Google And Maybe Microsoft New Google Networks Tuned Up For GenAI Inference And Training Microsoft And OpenAI Remain Friends, Are Looking To Hook Up With Others AI-Driven CPU Shortage Saves Intel’s Financial Cookies The GenAI Battle Shifts From Frontier Models To Agentic Platforms With TPU 8, Google Makes GenAI Systems Much Better, Not Just Bigger Cisco Scales Out Quantum Systems With A Quantum Network Switch The Second Time Will Be The IPO Charm For Cerebras Imagine An Army Of AI Minions Handling Incident Response AI Will Soon Drive A Third Of TSMC’s Business Bechtolsheim & Friends Breathe Life Into Pluggable Optics One Last Time How HPC And AI Digital Twins Accelerate Quantum Error Correction The Embrace Of AI In Design Transforms Cadence And Its Customers Nvidia Brings The Power Of Open Source AI Models To Quantum Computing Building The Imperfect Beast For Enterprises, GPUs Need Virtualization As Much As CPUs Ever Did CoreWeave Takes As Much Financial Engineering As It Does Datacenter Design Contemplating Meta’s Homegrown MTIA Compute Engine Roadmap Most Neoclouds, Sovereigns, And Enterprises Will Buy, Not Build, Their AI Stacks Broadcom And Google Benefit Mightily From Anthropic’s Meteoric Growth Rebellions AI Rings Up The Money To Rack Up AI Inference Systems Nvidia Software Pushes MLPerf Inference Benchmarks To New Highs Broadcom Makes Its Pitch To Run Kubernetes On VMware VCF The $2 Billion Nvidia Deal With Marvell Is About A Lot More Than NVLink Fusion Classiq Says Quantum Is On Its Way, But Patience Is Needed Demonstrating The Scientific Usefulness Of Quantum Systems We Need Servers – Lots Of Servers. . . . Arm Comes Full Circle With Homegrown, AI-Tuned Server CPU Riding The Memory Boom And Trying To Avoid The Bust Data Analytics Helps Make The Mighty Lionesses Roar Driving Down The AI System Roadmap With Nvidia The Open Agentic AI World According To Nvidia Nvidia Finally Admits Why It Shelled Out $20 Billion For Groq Nvidia Says OpenClaw Is To Agentic AI What GPT Was To Chattybots IBM Unrolls Blueprint For Quantum-Classical HPC Computing Women Get Data-Driven Health Boost As The FA Tackles Sports Science Four Months Into Its Comeback, Zapata Stakes Its Claim In Quantum Software Eridu Cuts To The AI Networking Chase With High Radix Switch System HPE Works Harder And Smarter To Chase Datacenter Profits We Need A Proper AI Inference Benchmark Test How AI Is Boosting Gender Equality In High Performance Racing Custom Compute Engine Biz Growing More Than Marvell Ever Hoped Broadcom May Become The Biggest Counterbalance To Nvidia Ayar Labs Gets $500 Million To Ramp Photonics Into 2028 AI Systems With Cisco Outshift, Agentic AI Is Teed Up For the Internet Of Cognition Nvidia Sees The Light On Silicon Photonics And Maybe Optical Switching AI Servers Finally Dominate Dell’s Systems Business VAST Data: What Controls The Data Is More Important Than What Stores It So Far, Nobody Turns Tokens Into Money Like Nvidia SambaNova Pits Its Engineering Against Nvidia For Agentic AI Some More Game Theory, This Time On The AMD-Meta Platforms Deal AMD Says “Helios” Racks And MI400 Series GPUs On Track For 2H 2026 CPU-Only Compute Still Matters To A Lot Of HPC Centers Taalas Etches AI Models Onto Transistors To Rocket Boost Inference Some Game Theory On That Nvidia-Meta Platforms Partnership AI Eats The World, And Most Of Its Flash Storage The Current AI Networking Wave Will Be A Tsunami Of Money By 2027 The Memory Crunch Pinches Cisco’s Profits Only A Few AI Platforms Can Survive The Greatest AI Show On Earth Cisco Doubles Up The Switch Bandwidth To Take On AI Scale Out And Eventually Scale Up Datacenter Spending Forecast Revised Upwards – Yet Again The Twin Engine Strategy That Propels AWS Is Working Well With GenAI Turbochargers, Google Is Shifting Its Cloud Into A Higher Gear AMD Finally Makes More Money On GPUs Than CPUs In A Quarter Dassault And Nvidia Bring Industrial World Models To Physical AI TACC Explores Mixed Precision And FP64 Emulation For HPC With Horizon Robotics Will Break AI infrastructure: Here's What Comes Next Oracle’s Financing Primes The OpenAI Pump Gartner Takes Another Stab At Forecasting AI Spending Microsoft Is More Dependent On OpenAI Than The Converse Big Blue Poised To Peddle Lots Of On Premises GenAI Microsoft Takes On Other Clouds With “Braga” Maia 200 AI Compute Engines Nvidia’s $2 Billion Investment In CoreWeave Is A Drop In A $250 Billion Bucket Intel Is Still Struggling In The Datacenter, But It Could Get Better Is Nvidia Assembling The Parts For Its Next Inference Platform? TSMC Has No Choice But To Trust The Sunny AI Forecasts Of Its Customers Cerebras Inks Transformative $10 Billion Inference Deal With OpenAI By Decade’s End, AI Will Drive More Than Half Of All Chip Sales Startup Quantum Elements Brings AI, Digital Twins To Quantum Computing D-Wave Makes Gate-Model Power Move With Quantum Circuits Buy Building The Future Of Software In The AI-Native Era Arista Modular Switches Aim At Scale Across Networks, Hit Scale Out, Too NextSilicon Takes Aim At CPUs And GPUs With “Maverick-2” Dataflow Engine How HPC Is Igniting Discoveries In Dinosaur Locomotion – And Beyond Oracle First In Line For AMD “Altair” MI450 GPUs, “Helios” Racks
OpenAI, Microsoft And Friends Build A Better, More Scalable Ethernet
Timothy Prickett Morgan · 2026-05-13 · via The Next Platform: In-depth coverage of high end computing

Sometimes, to solve a particular system architecture problem, you have to invent a new technology. And sometimes, you just need to squint at the problem a little and look at what you already have and use the parts in a different way.

The latter approach is what has happened as researchers at OpenAI, Microsoft, Broadcom, AMD, and Nvidia took a hard look at how ever-embiggening bandwidth on network ports is not necessarily a valuable thing compared to have scale out networks that have higher radix switches – meaning a lot more network links between devices – and also flatter networks with fewer switches. Lowering the switch count means the scale out network lashing together AI system nodes has lower latency (fewer hops across the network between any two endpoints), lower cost (which lowers the total cost of acquisition), and lower power consumption (which further lowers the total cost of ownership).

With most great engineering ideas, when you look at it, the new approach is intuitively obvious and you have to wonder why it wasn’t always done that way. Such is the case with Multipath Reliable Connection, a new network protocol that lays down atop Ethernet switch ASICs and that borrows many of the ideas of the Ultra Ethernet specification put forth by the Ultra Ethernet Consortium, which was founded back in July 2023 for the express purpose of scaling Ethernet to more than 1 million endpoints as well as making it as good as the InfiniBand low latency network for AI clusters.

The MRC effort was started two years ago. OpenAI did a lot of the talking for the new MRC protocol as it was unveiled last week, but we strongly suspect that Microsoft did a lot of the work based on its extensive experience with both RoCE Ethernet and InfiniBand networks. You can read the OpenAI blog about MRC here, download the paper the five companies release there, and see the Open Compute Project spec for the effort at this link.

In essence, what MRC does is stop chasing ports with higher and higher bandwidth and start using the same aggregate bandwidth of a given switch ASIC to increase the number of links between devices. I know what you’re thinking: Won’t increasing the number of ports and the number of links mean increasing the number of potential failures in those links, thereby making it more the absolutely synchronous work like an AI training run comes to a crashing halt more often? No, it won’t, if you radically increase the number of links between endpoints. If you have enough links, as it turns out, and the right protocol, you an heal around link failures and while the AI training job slows down, there are enough ways to reroute traffic that the network can heal around the link failure. And at your convenience, without having to stop the AI training job, you can repair the link.

Endpoint failures – meaning GPUs and XPUs – will still crash the training run, of course. To which we say: Why not locally snapshot checkpoints on each server node, stream them out to network storage or a shared memory appliance, keep a few spare GPUs or XPUs in the network, and restore that one failed compute engine and then resume the calculation? Perhaps this is harder than it sounds. . . . It might be better to have an out of band compute engine monitor that predicts a failure for a compute engine, freezes the training run before it crashes, takes the failing compute engine offline, loads up data on the spare compute engine, and resumes processing. Why let it crash at all?

Anyway, back to MRC. While the Ultra Ethernet protocol is a brand new protocol that starts from a blank sheet of paper to make Ethernet more like InfiniBand in terms of low latency, traffic shaping, and adaptive load balancing, the MRC protocol is much less drastic of a change and is, in fact, a superset extension of the current RDMA over Converged Ethernet (RoCE) protocol that hyperscalers, cloud builders, and supercomputing centers have been complaining about for more than a decade.

The adaptive load balancing is based on Explicit Congestion Notification, and like Ultra Ethernet, MRC supports out of order delivery of packets, packet spraying across multiple links, selective retransmission, and packet trimming to help deal with congestion.

Packet trimming is neat in that it only retransmits packets that have been dropped due to switch ASIC buffer overflows, and it does so without invoking the global ECN mechanism. (Nvidia has a good explanation of packet trimming, which was implemented in the Cumulus Linux network operating system it acquired shortly after buying Mellanox, here.) While ECN tells packet senders to slow down when they are flooding a switch or endpoint, packet trimming keeps track of the headers, drops the packet payload, and asks the network to retransmit only missing data when a packet is dropped due to congestion. Packet trimming requires acceleration and processing inside the switch ASIC or the network interface card.

The new MRC protocol is paired with IPv6 segment routing, which is used to route packets in a static fashion around the network and, ironically, the dynamic routing mechanisms in the protocol stack were turned off. The combination of the adaptive load balancing and static routing across eight links per endpoint is relatively easy to do and means adaptive routing is not necessary because links and ports are not so scare and there are eight different ways to get to any endpoint.

This is accomplished by parallelizing the data plane, and this is the topology magic that makes MRC is really powerful. This is one of those cases where two pictures are literally worth a thousand words, so let’s compare and contrast how an AI cluster is built today using 51.2 Tb/sec switches today and how you do it with MRC.

For the past decade or so, if you wanted to build an AI cluster, you used a three-tier network – leaf switches in the rack, spine switches linking them together to make pods, and superspine switches for linking the pods together. This is how you can lash together an AI supercomputer with traditional RoCE Ethernet using 800 Gb/sec ports on the 51.2 Tb/sec switches, which have 64 ports each:

Each pod has 64 GPUs or XPUs, each with an 800 Gb/sec network interface. Those 64 compute engines feed into 32 Tier 0 top of rack leaf switches, which are cross-coupled with 32 Tier 1 spine switches. Each Tier 1 spine switch has 32 ports pointing up to the Tier 2 superspine switches and 32 ports pointing down to the Tier 0 leaf switches. There are 1,024 Tier 2 superspines that interlink 64 pods together, providing connectivity for a total of 65,536 GPUs or XPUs. There is a total of 5,120 64-port 51.2 Tb/sec switches in this three tier network, and they present a single Clos topology data plane.

If you want to do more GPUs than this, you have to wait for 102.4 Tb/sec switches or you have to add a fourth tier to the network. This will add cost as well as latency. With a three tier network, any GPU or XPU can be as much as five to seven switch hops away.

Now, shift to a higher radix point of view. Instead of craving 800 Gb/sec ports, split that 51.2 Tb/sec switch ASIC so it supports 512 ports running at 100 Gb/sec. And then instead of having one Clos data plane, have eight unique Clos data planes (eight unique paths between any two devices in the network). How many devices can you connect in a two tier network? Look at this:

With the same bandwidth going into each GPU or XPU (spread over eight planes instead of one), you can lash together 131,072 compute engines in a two tier network. Crazy, right? Each pod in the Ethernet MRC cluster has 256 XPUs, and they feed up to a Tier 0 leaf complex with 512 switches per plane, for a total of 4,096 switches. The Tier 1 spine layer has 256 switches per plane, for a total of 2,048 switches. Add it up, you need 6,144 switches to do the MRC network. And no GPU or XPU is more than three hops away from another.

That is 20 percent more switches to get double the compte engine capacity with the same bandwidth per compute engine. This seems like a fair trade.

There is a point in the MRC paper where it says “for full-bisection bandwidth, we require 2/3s of the optics and 3/5s the number of switches compared to a three tier network.” That is not for the comparison above, with 65,536 compute engines in a three tier network versus 131,072 compute engines in a two tier network. Those ratios are only true for the same number of compute engines in each cluster.

By the way, the three tier RoCE Ethernet network with 65,536 endpoints has a total of 196,608 links between all of the switches and compute engines, and the two tier design has an incredible 1,179,648 links. But for a two tier network with only 65,536 compute engines would have 1,048,576 links. You do copper DAC cables within the racks and optical links to the spines and superspines (if you need that layer). Given that DACs and optical links are not free, it could turn out that even though you save on the switch budget or be able to double the scale with an MRC network while only spending 20 percent more on the switch budget, you get killed on the link budget.

But, ironically, that may not matter, and here is why: Nothing, but nothing, in the whole wide world is more expensive than a GPU or XPU compute engine that can’t do work because a link fails, and if a link fails in “normal” Clos switch architectures, the whole AI supercomputer stops and you have to go back to a checkpoint and start over. So tens of thousands to more than a hundred thousand compute engines are just sitting there.

But with MRC, if you lose one of the eight links going into a GPU or XPU, you only lose 12 percent of that 800 Gb/sec bandwidth – and the AI training job keeps running. And you can go change that dead link and it will come back online; the switches will activate that link and start giving you the bandwidth back into that particular endpoint.

You can even reboot a Tier 1 switch in an MRC network, and the system heals around it, like this:

In its blog post, OpenAI said that aside from the increased scale of the AI cluster, “MRC’s adaptive packet spraying load-balances well enough that we see essentially no congestion in the core of the network. This greatly reduces variation in throughput between flows during synchronous training, where eliminating outliers is central to performance. It also means that when multiple jobs share the cluster, they do not impact one another’s performance.”

OpenAI has run MRC on clusters at the Oracle Stargate datacenter in Abilene, Texas as well as the Microsoft Azure AI datacenter in Fairwater, Wisconsin. The MRC protocol was implemented in Nvidia ConnectX-8 SmartNICs, AMD “Pollara” and “Vulcano” DPUs, and Broadcom Thor Ultra SmartNICs. The SRv6 static routing was implemented on Nvidia Spectrum 4 and Spectrum 5 switches running both the Cumulus Linux and SONiC network operating systems as well as on Arista Networks switches based on the EOS variant of Linux running atop Broadcom Tomahawk 5 ASICs.