惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
C
CERT Recently Published Vulnerability Notes
Know Your Adversary
Know Your Adversary
Security Archives - TechRepublic
Security Archives - TechRepublic
Security Latest
Security Latest
P
Privacy & Cybersecurity Law Blog
P
Privacy International News Feed
月光博客
月光博客
Stack Overflow Blog
Stack Overflow Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
H
Help Net Security
Recent Commits to openclaw:main
Recent Commits to openclaw:main
AI
AI
O
OpenAI News
M
MIT News - Artificial intelligence
Scott Helme
Scott Helme
U
Unit 42
P
Proofpoint News Feed
罗磊的独立博客
C
Check Point Blog
MongoDB | Blog
MongoDB | Blog
Engineering at Meta
Engineering at Meta
博客园 - 三生石上(FineUI控件)
阮一峰的网络日志
阮一峰的网络日志
Apple Machine Learning Research
Apple Machine Learning Research
T
The Exploit Database - CXSecurity.com
I
InfoQ
云风的 BLOG
云风的 BLOG
aimingoo的专栏
aimingoo的专栏
Google DeepMind News
Google DeepMind News
W
WeLiveSecurity
Webroot Blog
Webroot Blog
P
Palo Alto Networks Blog
C
Cybersecurity and Infrastructure Security Agency CISA
N
News and Events Feed by Topic
Cisco Talos Blog
Cisco Talos Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - Franky
A
About on SuperTechFans
美团技术团队
J
Java Code Geeks
T
Tenable Blog
L
LINUX DO - 最新话题
NISL@THU
NISL@THU
G
Google Developers Blog
Forbes - Security
Forbes - Security
爱范儿
爱范儿
S
Security @ Cisco Blogs
Project Zero
Project Zero
有赞技术团队
有赞技术团队

The Register - On-Prem: Systems

Qualcomm teases agentic CPUs and smartphones Fujitsu says quantum and AI will replace mainframes in 2035 ZTE & China's NCRC partner for smart interventional medicine Core Scientific accelerates crypto-to-AI pivot Meta to use millions of AWS Graviton cores AI now gobbling up power and management chips for servers Tesla stakes AI dreams on Intel's unfinished AI chip SK Hynix breaks ground on Indiana advanced packaging plant Datacenter boom keeps dirty coal plants alive in the US AMD's Ryzen 9 9950X3D2 Dual Edition tested World's blandest man steps down from CEO job to spend more time in tastefully appointed home World's blandest man steps down from CEO job Intel eases reliance on TSMC with 'Merica-made Core Series 3 processors Intel eases reliance on TSMC with Core Series 3 CPUs Guide to GPU virtualization: passthrough, vGPU, and MIG Guide to GPU virtualization: passthrough, vGPU, and MIG Orbital datacenter startup admits launch economics don't fly AI-powered mainframe exits are a bubble set to pop Cloud-smart strategy helps Interactive meet GenAI demands Oracle taps Bloom for fuel cells to support datacenter binge Oracle taps Bloom for fuel cells to support datacenter binge UK signs Rolls-Royce SMR design deal Japan going back to the future by reviving its chip industry Japan going back to the future by reviving its chip industry AWS ponders selling its home-grown chips by the rack-load Supply chain challenges risk delaying Nvidia's Rubin GPUs Supply chain challenges risk delaying Nvidia's Rubin GPUs Supermicro investigating alleged China chip smuggling Intel trapped in Elon's reality distortion field UALink delivers 2.0 spec before v. 1.0 silicon ships OpenInfra General Manager on sovereignty and kill switches Anthropic reveals $30bn run rate, plan to use new Google TPU Nvidia embraces optical scale-up as copper reaches limits Nvidia embraces optical scale-up as copper reaches limits IBM wants Arm software on its mainframes for AI support AI datacenters create heat islands around them, paper finds Arm says AI agents need a new CPU. Intel doesn't buy it Memory-makers' shares are down. Don't blame Google Memory-makers' shares are down. Don't blame Google US PC shipments to fall 13% as memory and storage crunch hits budget systems US PC shipments to fall 13% as memory costs surge ZTE showcases intelligent computing at CloudFest 2026 Rebellions eyes global expansion with rack-scale AI platform Rebellions eyes global expansion with rack-scale AI platform Enterprise infrastructure is entering an economic reset AMD doubles up on V-Cache with 9950X3D2 Dual Edition Apple's making more iPhone bits in US, but not the iPhone Three more charged with trying to smuggle GPUs to China Three more charged with trying to smuggle GPUs to China Dell slims Pro laptops, boosts battery and cooling Alibaba delivers RISC-V server chip optimized for Chinese AI Alibaba delivers RISC-V server chip optimized for Chinese AI AI-pilled Arm CEO teases mystery products for $1T TAM Arm rolls its own 136-core AGI CPU to chase AI hype train SoftBank builds AI mega-datacenter on nuke site SoftBank builds AI mega-datacenter on nuke site Chip tester shrugged off ransomware – then came the leak Explainer: AI-ready servers Elon Musk proposes 'Terafab' to level up chip production Australia to datacenter operators: BYO energy or stay home Australia to datacenter operators: BYO energy or stay home Supermicro co-founder charged over $2.5B GPU sales to China Blue Origin applies to launch 51,000 datacenter satellites Blue Origin applies to launch 51,000 datacenter satellites Alibaba has made 470,000 AI chips, admits they’re inferior Decoding Nvidia's Groq-powered LPX and the rest of its new rack systems Your next car might need 300 GB of RAM, and so will autonomous robots Your next car might need 300 GB of RAM, and so will robots It's not a binary choice: Boffin builds ternary CPU Nvidia H200 back on in China, production ramping: Huang Nvidia slaps $20B Groq tech into massive new LPX racks to speed AI response time Meta reveals four Broadcom-built custom AI chips, claims some outperform commercial silicon Meta reveals custom AI chips it says beat Nvidia ZTE and Orange Morocco launch Livebox 7 for smart homes Ayar Labs taps Wiwynn to cram 1,024 GPUs into a photonic rack system Ayar Labs, Wiwynn to cram 1,024 GPUs into photonic system ZTE and Whale Cloud Showcase Digital Transformation at MWC AI datacenters may gulp NYC's daily water supply at peak Mystery outage behind JetBlue's request for grounding HPE tweaks T&Cs so it can change quotes as RAM prices rise Supermicro launches probe after staff charged with China export violations
AI Burning Man happens next week – what to expect at Nvidia GTC 2026
2026-03-13 · via The Register - On-Prem: Systems

Nvidia has a bit of a problem. Popular generative AI workloads like code assistants and agentic systems generate massive quantities of tokens and need to move them at speed. But the GPU giant's chips currently struggle to deliver.

That will start to change next week when Nvidia CEO Jensen Huang uses his company’s GPU Technology Conference (better known as GTC) to explain how he will use the token-spewing accelerator tech he acquired with upstart Groq late last year.

Market-watching firm SemiAnalysis' latest InferenceX benchmarks shows how Groq's tech helps to fill the gap in Nvidia’s current portfolio.

InferenceX's efficiency Pareto curve can be broken down into three main categories. Bulk tokens on the left, expensive low-latency tokens on the right, and the so called "goldilocks zone" in the middle.

InferenceX's efficiency Pareto curve can be broken down into three main categories. Bulk tokens on the left, expensive low-latency tokens on the right, and the so called "goldilocks zone" in the middle. - Click to enlarge

While Nvidia's NVL72 rack systems scale well at lower per-user token generation rates, they become progressively less efficient as user interactivity increases.

By contrast, SRAM-heavy architectures, like those championed by Groq and Cerebras, excel in latency sensitive scenarios and can achieve token generation rates often exceeding 500 or even 1,000 tokens a second. That’s many more tokens than GPU-based architectures can deliver.

In fact, this capability is how Cerebras won OpenAI's business earlier this year to power its Codex model. Nvidia didn’t own anything to match Cerebras until it acquired Groq's intellectual property and talent for a staggering $20 billion in December.

By combining its GPU tech and CUDA software libraries with Groq's dataflow architecture, Nvidia has the opportunity to raise the Pareto curve dramatically, reducing the cost per token, while at the same time bolstering output speeds.

Extending Nvidia's CUDA hardware stack to include Groq's dataflow architecture will not be easy. At GTC, Nvidia might announce it will add limited support for Groq's existing architecture relatively quickly.

More silicon

This GTC already feels a bit different as Nvidia has spilled the beans on its Rubin GPUs back at CES in January.

To recap, Rubin packs up to 288 GB of HBM4 memory good for 22 TB/s of bandwidth and 35-50 petaFLOPS of dense NVFP4 performance depending on the use case.

The launch represents a major performance uplift over Nvidia’s current Blackwell-generation parts, delivering 5x the dense floating point throughput. So far, Nvidia has announced the chips will be available in both an eight-way HGX platform or its NVL72 rack system, which as the name suggests, crams 72 Rubin SXM modules into a single system.

There's also Rubin GPX, which was announced back at Computex in June 2025, which will slot into select NVL racks to provide additional compute capacity for large context and video processing workflows.

We expect to see Huang hammer on the performance optimizations and efficiency gains delivered by its growing portfolio of GPUs. But with those GPUs growing ever hotter – estimates put Rubin’s thermal design power at 1.8kW or perhaps even higher – liquid cooling isn’t optional. Some buyers may balk at that requirement, which would benefit AMD and its air-cooled kit.

However, given the generation gains delivered by the Rubin architecture, there's nothing stopping Nvidia from releasing a single-die, air cooled version of the chip with five or six HBM stacks rather than eight. Such a chip would still deliver a 2.5x uplift in performance over Blackwell – without requiring liquid cooling.

That's just speculation, but we have a sneaking suspicion we might see something along these lines during next week’s festivities.

Some Vera, Vera powerful cores

Alongside its latest datacenter GPUs, we anticipate more details on Nvidia's standalone Vera CPU.

First teased at last year's GTC, Vera features 88 custom-Arm cores which add support for simultaneous multithreading and a slew of confidential computing features previously only available on x86 platforms.

So far, we've only seen the CPU packaged as part of Nvidia's Vera-Rubin superchip. However, we've since learned Nvidia will offer the chip as a standalone processor that will compete with Intel and AMD for some mainstream applications.

Previously, Nvidia had offered Grace CPU superchips, but those were primarily for use in supercomputers and other HPC applications. However, last month the GPU giant revealed Meta would be its first partner to deploy Grace at scale and that the Social Network was already evaluating Vera CPUs for use in its datacenters as well.

Setting expectations

Alongside new datacenter silicon, we also anticipate Huang will share more details about Nvidia's next-gen Kyber racks and Feynman GPUs, which should debut in 2027 and 2028.

We first saw the Kyber at last year's GTC. The 600 kW behemoth is set to cram 144 GPU sockets, each with four Rubin Ultra GPU dies into a standard rack form factor.

Nvidia disclosed the existence of Kyber in part because datacenter operations were already struggling with the 120kW NVL72 systems announced the year before. By revealing Kyber, Nvidia lit a fire under datacenter physical infrastructure providers so they could provision the power supplies and cooling kit necessary to support such a system by 2027. With a yearly release cadence, Nvidia can't wait for the rest of the industry to catch up – it must telegraph its next move years in advance.

With Feynman just two years out, we suspect Huang may repeat the exercise, setting new power and cooling targets, likely exceeding a megawatt per rack.

Will Nvidia throw gamers a bone?

Nvidia has long rumored to have been working on an Arm-based system on chip for PCs.

A part capable of doing that job arrived last year in form of the DGX Spark and GB10 partner systems that put it to work. So far, however, OEMs have only used the chip in workstation class mini-PCs running Linux. Recent reports indicate Nvidia is working with the likes of Lenovo and Dell to bring a similar product to the Windows PC market.

As we previously reported, Nvidia is also working with Intel to integrate its GPU dies into Chipzilla's next-gen processors.

GTC seems like as good a time as any to throw gamers a bone and give Nvidia a new market to chase beyond its side hustles in the pro-visualization markets.

Integrated Nvidia graphics might not be the RTX 50 Super series cards that many had hoped to see at CES, but given the state of the memory market, it seems unlikely we'll see them make an appearance at GTC.

The Claw, robotics, and everything else

Beyond big iron and the remote possibility of some consumer hardware, you can bet on OpenClaw being a major talking point at GTC.

Jensen Huang is apparently quite fond of the agentic framework in spite of its many security vulnerabilities, reportedly describing it as the "most important software release probably ever."

The company is reportedly working on its own, presumably safer, version of the platform called NemoClaw.

Speaking of claws, we also expect to see a fair few more robots take the stage. Since announcing its Isaac GR00T robotics platform nearly two years ago, Nvidia has launched a steady supply of new toolkits, frameworks, and hardware development platforms aimed at giving generative AI physical form.

And to teach them to function in an unpredictable world, you can count on Nvidia's Omniverse digital twin platform to make another appearance. Introduced in 2019 at a time of rising Metaverse hype, the platform aimed to create a virtual environment in which physical processes could be simulated in the digital world before real-life implementation.

Developers have since integrated Omniverse in a variety of simulation platforms, including those used to design and build AI bit barns.

El Reg will be on the ground in San Jose next week for GTC to bring you the latest news from what has become one of the world’s most-watched tech conferences. ®