惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

MongoDB | Blog
MongoDB | Blog
宝玉的分享
宝玉的分享
博客园 - 三生石上(FineUI控件)
小众软件
小众软件
罗磊的独立博客
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
SegmentFault 最新的问题
Last Week in AI
Last Week in AI
人人都是产品经理
人人都是产品经理
博客园 - 聂微东
博客园 - 司徒正美
博客园 - 叶小钗
T
Tailwind CSS Blog
博客园 - Franky
V
V2EX
有赞技术团队
有赞技术团队
美团技术团队
雷峰网
雷峰网
爱范儿
爱范儿
Jina AI
Jina AI
D
DataBreaches.Net
H
Help Net Security
酷 壳 – CoolShell
酷 壳 – CoolShell

The Next Platform: In-depth coverage of high end computing

Uncle Sam Awards $2 Billion-Plus To Quantum Companies, But Wants A Cut Oak Ridge Starts Weaving Together A Quantum, Classical HPC, And AI System Stack Dell Bulks Up Hardware As AI Infrastructure Shifts To On-Premises Cisco Wins Over AI Customers With Merchant Silicon And Optics With Its IPO Done, Cerebras Can Get Back To Pushing The AI Envelope HPE Throws VM Users A Lifeline, Unifying Containers And VM Management In Cloud Stack OpenAI, Microsoft And Friends Build A Better, More Scalable Ethernet Compute And Memory Price Hikes Drive IT Spending Way Higher Sometimes, Air Is The Only Way For AI Systems To Keep Their Cool Arista Rides AI Scale Out Networks, Moves Into Scale Across, And Awaits Scale Up If You Can Make A Compute Engine, You Can Sell A Compute Engine Cleveland Clinic Simulates Large Proteins With Quantum-Centric Supercomputing Broadcom Helps CPU And XPU Makers Go Vertical With Compute Microsoft Committed To Doubling AI Infrastructure In Two Years Google Is A Full Stack AI Player, And Is Playing Well AWS Will Be An OEM, Just Like Google And Maybe Microsoft New Google Networks Tuned Up For GenAI Inference And Training Microsoft And OpenAI Remain Friends, Are Looking To Hook Up With Others AI-Driven CPU Shortage Saves Intel’s Financial Cookies The GenAI Battle Shifts From Frontier Models To Agentic Platforms With TPU 8, Google Makes GenAI Systems Much Better, Not Just Bigger Cisco Scales Out Quantum Systems With A Quantum Network Switch The Second Time Will Be The IPO Charm For Cerebras Imagine An Army Of AI Minions Handling Incident Response AI Will Soon Drive A Third Of TSMC’s Business Bechtolsheim & Friends Breathe Life Into Pluggable Optics One Last Time How HPC And AI Digital Twins Accelerate Quantum Error Correction The Embrace Of AI In Design Transforms Cadence And Its Customers Nvidia Brings The Power Of Open Source AI Models To Quantum Computing Building The Imperfect Beast
Robotics Will Break AI infrastructure: Here's What Comes ...
Evan Helda Evan Helda · 2026-02-03 · via The Next Platform: In-depth coverage of high end computing

SPONSORED CONTENT Physical AI and robotics are moving from the lab to the real world – and the cost of getting it wrong is no longer theoretical. With robots deployed in factories, warehouses, and public settings, large-scale simulation has become tightly coupled with real-world operations.

Physical AI companies need new types of infrastructure to continuously build, train, simulate, and deploy models that operate in dynamic, physical environments. With the cloud’s current limitations, the next wave of physical AI won’t scale.

Here are three reasons why the infrastructure stack needs to be purpose-built for physical AI.

The Need For – And Scarcity Of – Training Data

Physical AI can’t be trained on internet text, like an LLM. It requires context-specific data – from images and video to LiDAR, sensor streams, and motion data – that maps directly to actions and outcomes. With variation across environments, tasks, and hardware configurations, this data is not easy to obtain.

Collecting training data exclusively in the real world is slow and expensive. Virtual environments allow teams to generate synthetic data, test edge cases, and iterate faster than real-world deployment alone.

Simulation has become a critical way to bootstrap training, but scaling it is a heavy lift. It requires orchestrating large GPU fleets, parallelizing simulations, preparing “sim-ready” 3D assets, and often using different classes of GPUs than training or inference. Inference inside simulation mirrors the forward pass on real robots, but must run at massive scale, optimized for throughput rather than latency, which creates a distinct infrastructure requirement of its own.

Hardware reliability matters here: when simulations run across thousands of GPUs, interruptions or failures can derail entire training cycles. Price-performance ratio and mean time to failure become first-order concerns when choosing a cloud for simulations.

Big data, High Stakes, Low Latency

Data usability presents another challenge. Once physical AI systems are deployed, teams are suddenly faced with massive volumes of data, including simulation output alongside photos, video, LiDAR, and sensor data from active robots.

Simply dumping multimodal training data into object storage won’t work. Unlike curated training datasets, this data is noisy, contextual, and time sensitive. To be useful, it must be indexed, synchronized, and organized (ideally, through automated pipelines) so teams can search, segment, and select the right data for each training run.

Latency raises the stakes further. Physical systems must react in milliseconds, which rules out centralized, batch-style processing. As a result, physical AI increasingly relies on fast inference at the edge paired with higher-level planning and coordination models in the cloud, operating together as a single system.

Sophisticated platforms must be purpose-built for multimodal ingestion and querying. Without them, more data does not translate to better models.

Data Movement Becomes The Constraint

In physical AI, the hardest problem is often not model size – it’s moving data. Robotics systems generate continuous streams of video, sensor readings, and motion data that must be processed and acted on in real time.

In these systems, infrastructure breaks in unexpected ways. Many existing platforms were designed for batch-style workloads; they struggle when faced with sustained, high throughput multimodal data. Scaling GPUs alone is not enough if data cannot move quickly and efficiently between devices, local systems, and the cloud.

The expense of moving this data adds up quickly. Transferring large volumes across systems can cost more than storing it, making naive scaling inefficient. Supporting physical AI at scale requires infrastructure optimized for fast read and write performance, high-bandwidth pipelines, and predictable throughput – not just more memory or more compute.

The New Requirements For A Physical AI Stack

Physical AI is pushing AI out of controlled, digital environments and into the real world, where failure modes are physical, rather than theoretical. These systems place new demands on compute, networks, and data infrastructure, and there is no single blueprint yet for how to build them.

Coordinating a single robot is difficult. Scaling that to fleets operating in dynamic environments – continuously learning from simulation and real-world feedback – raises the bar higher. Data becomes more valuable, latency more consequential, and infrastructure decisions more tightly coupled to system behavior.

Progress in physical AI depends not just on better models, but on infrastructure that can support continuous learning, real-time response, and coordination across edge and cloud systems. Failing to meet these requirements risks stalled deployments, unreliable systems, and real-world consequences.

The challenges are clear. By necessity, a robust physical AI stack will be a hybrid of large-scale simulation and training in the cloud paired with fast, on-device inference and continuous learning at the edge. The question now is who will build it first.

How Nebius Is Building Robotics Solutions

The AI stack of the future isn’t defined by raw compute alone. It’s shaped by speed, data movement, orchestration, and the ability to operate seamlessly across virtual and physical worlds.

At Nebius, we are obsessed with solving the unique constraints of the physical world. We are engineering the infrastructure specifically for this next phase of AI, combining optimal price/performance GPUs and high-throughput storage with flexible, managed orchestration designed to handle the dynamic nature of robotics workloads.

Whether you are bursting massive simulation workloads via Slurm or training foundation models on reliable large-scale clusters, Nebius provides the foundation to move faster, scale reliably, and operate with confidence.

The best way to understand the difference is to experience it. Sign up today to start building on Nebius, or contact our Physical AI team to discuss how Nebius can support your architecture.

Evan Helda is Head of Physical AI at Nebius.