惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

腾讯CDC
N
Netflix TechBlog - Medium
aimingoo的专栏
aimingoo的专栏
P
Proofpoint News Feed
F
Fortinet All Blogs
大猫的无限游戏
大猫的无限游戏
I
InfoQ
V
V2EX
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
有赞技术团队
有赞技术团队
G
Google Developers Blog
L
LangChain Blog
博客园_首页
M
MIT News - Artificial intelligence
H
Hackread – Cybersecurity News, Data Breaches, AI and More
月光博客
月光博客
IT之家
IT之家
量子位
宝玉的分享
宝玉的分享
S
SegmentFault 最新的问题
Stack Overflow Blog
Stack Overflow Blog
V
Visual Studio Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
雷峰网
雷峰网

AI demand is so high, AWS customers are trying to buy out its entire capacity | Network World

Cisco: Latest news and insights 2026 network outage report and internet health check Selector targets the network visibility gap in multi-cloud infrastructure AI reshapes cybersecurity workforce priorities as IT teams brace for new risks Top network and data center events of 2026 How AI is transforming network incident response (and where it still falls short) Google opens TPUs to enterprises beyond its own cloud via Blackstone JV AI, cybersecurity skills top IT pay premiums Startup Bolt Graphics promises 5x performance over Nvidia’s best GPU Wireless security is a battle of AI vs. AI NetOps teams look to AI to automate Day 2 operations Digital twins reshape network and data center management Network outages, power failures strain data center resiliency Five takeaways from Cisco's blowout quarter and what it means to customers Cisco to cut nearly 4,000 jobs despite strong growth in AI, enterprise networking Startup SPAN teams with Nvidia to put data center nodes in your backyard Hard drive shortage affecting enterprise storage needs Wi-Fi 8 is closer than you think. Here’s what you need to know Cisco open-sources agentic AI security spec HPE revamps private cloud stack for enterprises rethinking VMware Versa takes aim at fragmented enterprise security with CSPM, orchestration update, and AI agent controls Red Hat opens Ansible to AI agents, within limits Red Hat offers endless Linux support — for a fee Red Hat: Sovereignty is more than just compliance Tech job postings hit three-year high as AI demand fuels hiring rebound HPE memory server targets compute-heavy and agentic AI workloads PCI group begins work on new spec to support bandwidth-hungry apps like AI, HPC Q&A: Quantum physicist Sonia Fernández-Vidal on why classical computing isn't going anywhere AWS hit by US-East-1 outage after data center thermal event Gluware's Titan rises to meet Mythos network vulnerability challenge
OpenAI-led consortium seeks to address AI processing bott...
2026-05-08 · via AI demand is so high, AWS customers are trying to buy out its entire capacity | Network World

An OpenAI-led consortium of tech giants including AMD, Broadcom, Intel, Microsoft, and Nvidia have unveiled a new networking protocol designed to address network congestion, a problem that has always existed but has been exacerbated by the massive amounts of data required for AI processing.

The new protocol, called Multipath Reliable Connection (MRC), is for training models on 100,000+ GPUs by distributing traffic across hundreds of network paths simultaneously rather than forcing it down a few lanes that can get easily congested.

“Network congestion, link, and device failures are the most common sources of delay and jitter in transfers,” OpenAI wrote in a blog post announcing the project. “These problems get more frequent, and harder to solve, as the size of the cluster increases.”

It went on to note that a single failure could often cause a training job to crash, forcing a restart from a saved checkpoint, or stall progress for many seconds while the network recomputed routes. Such interruptions are costly in both GPU cycles and time.

“The larger the job we run, the greater the impact of any single link flap or failure. These workloads act as a form of ‘failure amplifier,’ so preventing this has become critical,” the company said.

OpenAI led the development of the protocol and worked with AMD, Broadcom, Intel, Microsoft, and Nvidia, all of whom made significant technical contributions. The project is hosted and coordinated by the Open Compute Platform (OCP) consortium.

Nvidia is making its presence felt with the use of its Spectrum-X Ethernet as a part of MRC. The company says it is running MRC in production at some of the world’s largest AI training clusters, including OpenAI, for training frontier LLM models like ChatGPT and Codex.

Spectrum-X is also used in Microsoft’s Fairwater and Oracle Cloud Infrastructure (OCI’s) Abilene data center (a part of Project Stargate), two of the largest AI factories purpose-built for training and deploying leading-edge frontier LLMs.

MRC delivers the best GPU utilization possible by load-balancing traffic across all available paths, avoiding congestion by dynamically avoiding overloaded paths in real time. Conventional network fabrics can take seconds or even tens of seconds to stabilize after failures, according to OpenAI.

This helps keep maximum GPU utilization while training runs through network slowdowns, congestion, or failures or other events that would ordinarily disrupt or stall the training process. Administrators also gain fine-grained visibility and control over traffic paths, monitoring network traffic from a simple, single pane of glass.

SUBSCRIBE TO OUR NEWSLETTER

From our editors straight to your inbox

Get started by entering your email address below.