惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Privacy International News Feed
A
Arctic Wolf
Security Latest
Security Latest
雷峰网
雷峰网
V2EX - 技术
V2EX - 技术
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
Schneier on Security
WordPress大学
WordPress大学
J
Java Code Geeks
宝玉的分享
宝玉的分享
T
The Exploit Database - CXSecurity.com
T
Troy Hunt's Blog
Scott Helme
Scott Helme
爱范儿
爱范儿
罗磊的独立博客
Apple Machine Learning Research
Apple Machine Learning Research
Application and Cybersecurity Blog
Application and Cybersecurity Blog
I
Intezer
博客园 - 【当耐特】
T
Threat Research - Cisco Blogs
L
LINUX DO - 最新话题
W
WeLiveSecurity
K
Kaspersky official blog
Google Online Security Blog
Google Online Security Blog
IT之家
IT之家
N
News and Events Feed by Topic
The Hacker News
The Hacker News
Know Your Adversary
Know Your Adversary
小众软件
小众软件
博客园 - 叶小钗
Latest news
Latest news
P
Proofpoint News Feed
G
GRAHAM CLULEY
Schneier on Security
Schneier on Security
T
Tor Project blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
T
Tailwind CSS Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Project Zero
Project Zero
The Cloudflare Blog
美团技术团队
大猫的无限游戏
大猫的无限游戏
S
Secure Thoughts
量子位
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
博客园_首页
Jina AI
Jina AI
V
Visual Studio Blog
Help Net Security
Help Net Security

The Next Platform: In-depth coverage of high end computing

Uncle Sam Awards $2 Billion-Plus To Quantum Companies, But Wants A Cut Oak Ridge Starts Weaving Together A Quantum, Classical HPC, And AI System Stack Dell Bulks Up Hardware As AI Infrastructure Shifts To On-Premises Cisco Wins Over AI Customers With Merchant Silicon And Optics With Its IPO Done, Cerebras Can Get Back To Pushing The AI Envelope HPE Throws VM Users A Lifeline, Unifying Containers And VM Management In Cloud Stack OpenAI, Microsoft And Friends Build A Better, More Scalable Ethernet Compute And Memory Price Hikes Drive IT Spending Way Higher Sometimes, Air Is The Only Way For AI Systems To Keep Their Cool Arista Rides AI Scale Out Networks, Moves Into Scale Across, And Awaits Scale Up If You Can Make A Compute Engine, You Can Sell A Compute Engine Cleveland Clinic Simulates Large Proteins With Quantum-Centric Supercomputing Broadcom Helps CPU And XPU Makers Go Vertical With Compute Microsoft Committed To Doubling AI Infrastructure In Two Years Google Is A Full Stack AI Player, And Is Playing Well AWS Will Be An OEM, Just Like Google And Maybe Microsoft New Google Networks Tuned Up For GenAI Inference And Training Microsoft And OpenAI Remain Friends, Are Looking To Hook Up With Others AI-Driven CPU Shortage Saves Intel’s Financial Cookies The GenAI Battle Shifts From Frontier Models To Agentic Platforms With TPU 8, Google Makes GenAI Systems Much Better, Not Just Bigger Cisco Scales Out Quantum Systems With A Quantum Network Switch The Second Time Will Be The IPO Charm For Cerebras Imagine An Army Of AI Minions Handling Incident Response AI Will Soon Drive A Third Of TSMC’s Business Bechtolsheim & Friends Breathe Life Into Pluggable Optics One Last Time How HPC And AI Digital Twins Accelerate Quantum Error Correction The Embrace Of AI In Design Transforms Cadence And Its Customers Nvidia Brings The Power Of Open Source AI Models To Quantum Computing Building The Imperfect Beast For Enterprises, GPUs Need Virtualization As Much As CPUs Ever Did CoreWeave Takes As Much Financial Engineering As It Does Datacenter Design Contemplating Meta’s Homegrown MTIA Compute Engine Roadmap Most Neoclouds, Sovereigns, And Enterprises Will Buy, Not Build, Their AI Stacks Broadcom And Google Benefit Mightily From Anthropic’s Meteoric Growth Rebellions AI Rings Up The Money To Rack Up AI Inference Systems Nvidia Software Pushes MLPerf Inference Benchmarks To New Highs Broadcom Makes Its Pitch To Run Kubernetes On VMware VCF The $2 Billion Nvidia Deal With Marvell Is About A Lot More Than NVLink Fusion Classiq Says Quantum Is On Its Way, But Patience Is Needed Demonstrating The Scientific Usefulness Of Quantum Systems We Need Servers – Lots Of Servers. . . . Arm Comes Full Circle With Homegrown, AI-Tuned Server CPU Riding The Memory Boom And Trying To Avoid The Bust Data Analytics Helps Make The Mighty Lionesses Roar Driving Down The AI System Roadmap With Nvidia The Open Agentic AI World According To Nvidia Nvidia Says OpenClaw Is To Agentic AI What GPT Was To Chattybots IBM Unrolls Blueprint For Quantum-Classical HPC Computing Women Get Data-Driven Health Boost As The FA Tackles Sports Science Four Months Into Its Comeback, Zapata Stakes Its Claim In Quantum Software Eridu Cuts To The AI Networking Chase With High Radix Switch System HPE Works Harder And Smarter To Chase Datacenter Profits We Need A Proper AI Inference Benchmark Test How AI Is Boosting Gender Equality In High Performance Racing Custom Compute Engine Biz Growing More Than Marvell Ever Hoped Broadcom May Become The Biggest Counterbalance To Nvidia Ayar Labs Gets $500 Million To Ramp Photonics Into 2028 AI Systems With Cisco Outshift, Agentic AI Is Teed Up For the Internet Of Cognition Nvidia Sees The Light On Silicon Photonics And Maybe Optical Switching AI Servers Finally Dominate Dell’s Systems Business VAST Data: What Controls The Data Is More Important Than What Stores It So Far, Nobody Turns Tokens Into Money Like Nvidia SambaNova Pits Its Engineering Against Nvidia For Agentic AI Some More Game Theory, This Time On The AMD-Meta Platforms Deal AMD Says “Helios” Racks And MI400 Series GPUs On Track For 2H 2026 CPU-Only Compute Still Matters To A Lot Of HPC Centers Taalas Etches AI Models Onto Transistors To Rocket Boost Inference Some Game Theory On That Nvidia-Meta Platforms Partnership AI Eats The World, And Most Of Its Flash Storage The Current AI Networking Wave Will Be A Tsunami Of Money By 2027 The Memory Crunch Pinches Cisco’s Profits Only A Few AI Platforms Can Survive The Greatest AI Show On Earth Cisco Doubles Up The Switch Bandwidth To Take On AI Scale Out And Eventually Scale Up Datacenter Spending Forecast Revised Upwards – Yet Again The Twin Engine Strategy That Propels AWS Is Working Well With GenAI Turbochargers, Google Is Shifting Its Cloud Into A Higher Gear AMD Finally Makes More Money On GPUs Than CPUs In A Quarter Dassault And Nvidia Bring Industrial World Models To Physical AI TACC Explores Mixed Precision And FP64 Emulation For HPC With Horizon Robotics Will Break AI infrastructure: Here's What Comes Next Oracle’s Financing Primes The OpenAI Pump Gartner Takes Another Stab At Forecasting AI Spending Microsoft Is More Dependent On OpenAI Than The Converse Big Blue Poised To Peddle Lots Of On Premises GenAI Microsoft Takes On Other Clouds With “Braga” Maia 200 AI Compute Engines Nvidia’s $2 Billion Investment In CoreWeave Is A Drop In A $250 Billion Bucket Intel Is Still Struggling In The Datacenter, But It Could Get Better Is Nvidia Assembling The Parts For Its Next Inference Platform? TSMC Has No Choice But To Trust The Sunny AI Forecasts Of Its Customers Cerebras Inks Transformative $10 Billion Inference Deal With OpenAI By Decade’s End, AI Will Drive More Than Half Of All Chip Sales Startup Quantum Elements Brings AI, Digital Twins To Quantum Computing D-Wave Makes Gate-Model Power Move With Quantum Circuits Buy Building The Future Of Software In The AI-Native Era Arista Modular Switches Aim At Scale Across Networks, Hit Scale Out, Too NextSilicon Takes Aim At CPUs And GPUs With “Maverick-2” Dataflow Engine How HPC Is Igniting Discoveries In Dinosaur Locomotion – And Beyond Oracle First In Line For AMD “Altair” MI450 GPUs, “Helios” Racks
Nvidia Finally Admits Why It Shelled Out $20 Billion For Groq
Timothy Prickett Morgan · 2026-03-18 · via The Next Platform: In-depth coverage of high end computing

Back in late December, Nvidia did a $20 billion “acquihire” of most of the development team at Groq and licensed the technology underlying its LPU dataflow engines for doing AI inference. We expected for Nvidia to move fast to deploy the tensor streaming processors created by Jonathan Ross, the ex-Googler who created a fully-scheduled, programmable tensor processing unit after he left the search engine giant. When the GenAI boom took off, these were renamed Language Processing Units, but the architecture did not change. Now, Nvidia is working with Samsung to bring the third generation LP30 chips to market, which Nvidia co-founder and chief executive officer Jensen Huang said in his opening keynote presentation at the GTC 2026 conference would happen in the second half of this year, and very likely in the third quarter.

Nvidia is not wasting any time, and that is because it does not have time to waste. Groq was going to start getting traction in low latency inference, just as Cerebras Systems has and that SambaNova Systems can do given their focus on ultra-high bandwidth SRAM memory against more modest compute to have zippy inference across a large number of compute engines. Where speed matters, these system makers and the dozens of upstarts who are trying to tackle inference at scale are so many piranhas swarming towards a fat cow standing in the Amazon (the river, not the bookseller and cloud utility). So Nvidia had to moooooooove. . . .

Hence, the dramatic $20 billion acquihire of Groq, which could not be an outright acquisition because that might take a year or two and might not pass muster with the world’s antitrust regulators. And hence its immediate absorption into the Vera-Rubin platform. Which arguably should be called the Vera-Rubin-Groq platform, given that Huang said during his keynote that low latency, premium priced token generation should represent somewhere on the order of 25 percent of the compute in an AI cluster.

Remember that Rubin CPX large context compute engine that Nvidia preview back in September 2025? The one based on a variant of the Rubin architecture and equipped with cheaper and more available GDDR7 graphics memory?

“We discovered a great idea,” Ian Buck, vice president of AI and HPC at Nvidia, said on a call ahead of GTC 2026 going over the systems announcements. “Integrating the LPU and LPX into our Rubin platform to optimize the decode. That's where we're focused right now, and we're excited to be bringing that to market.”

In other words, scratch Rubin CPX.

Huang stacked up what we presume is the “Rubin” R200 GPU accelerator beside what we presume was called the “Alan-3” Groq LP30 inference accelerator. One is a general purpose, dynamically scheduled compute engine that is pretty good at batching up lots of inferences and pipelining them through HBM stacked memory with reasonable latency and supporting many concurrent users. (That would be the GPU.) And the other is a rack or more of fairly modest, inference-specific, statically scheduled, deterministic compute engines that work in concert to support a small number of users – that number is likely one most of the time – and distribute model weights (not data) across their aggregate SRAM in such a way that the response time for token generation scales down as you add more machines. The GPU is a thresher, the LPU is a speed demon. They can work together with the Dynamo inference stack to provide a more balanced pareto curve for inference performance across a range of throughout and latency.

Here are the feeds and speeds of the R200 and the LP30 chips:

A fuller comparison would take into account the full memory hierarchy of these systems, including flash and main memory in host processors, but you get the idea. Also we would normalize to FP8 flops, which shows the performance gap is 21X at the same data precision and if the decode part of your AI workload can take advantage of FP4 processing – which is a fairly large if – then you can get 42X more peak theoretical oomph out of the R200 than out of the LP30.

But look at the complexity of the GPU, which is directly proportional to its costs – and most of the bill of materials for the R200 will cover the cost of the HBM4 stacked memory and the interposer it requires to link it to the GPU. So what one must consider is that not only will the latency of the speed demon be lower than that of the thresher, the cost per token for a reasonable level of interactivity could also be lower.

The most important thing to consider as we move from humans interacting with chattybots to agentic AI systems speaking to each other to perform tasks at much higher speeds and with much more reasoning and therefore orders of magnitude more tokens, is that the architectures of the world that look like those of Groq, Cerebras, SambaNova are going to be more important. There will have to be variants of Google TPUs and Amazon Trainiums aimed specifically at agentic AI inference and presenting a better balance between memory bandwidth and compute while also not sacrificing memory capacity.

We will do a deeper dive down deeper into the hardware. Fear not. Right now, we are just reviewing the strategy that Huang and Buck have elucidated, and the main thing you need to see are two pareto performance curves showing prior, current, and future coherent GPU memory domain systems and then what happens when you add the LP30 designed by Groq into the mix. The goal is to span from free to premium tiers with the inference iron in Huang’s conception of the inference universe, which is a reasonable way to look at it.

Here is how the Hopper NVL8, Grace-Blackwell NVL72, and Vera-Rubin NVL72 systems stack up in terms of throughput (tokens per second per megawatt) and interactivity (tokens per second per user):

It is immediately obvious that the larger shared GPU memory domain enabled by NVSwitch helps stretch the curves out from Hopper to Blackwell, but it is the memory, memory bandwidth, and compute that can only shift the curve up – but not stretch it to the right – with the move to the Rubin GPU. Nvidia will eventually increase this memory domain, but no in the 2026 hardware generation.

Now here is what happens when you add the Groq LP30 to the system mix, targeting the medium and premium tiers and driving out to a very profitable ultra tier as more and more LP30s are added to do the inference:

So what does that amazing curve tell you? Let me sum it up in plain American for you. 

If you are doing cheapass inference where response time is not the issue, like with a chattybot talking to slow-speaking humans or a couple of agents helping automate various kinds of human work, Vera-Rubin is fine for you. You will probably also need Vera-Rubin for training. But in a world of agentic AI, where the number of tokens needed to be generated is truly enormous and the latency of token generation has to be low so that huge collections of agents can complete their tasks – any delay is lost money that you might as well light on fire on the floor of the datacenter, or the New York Stock Exchange – then there is no one, and I mean no one, that will choose a hybrid CPU-GPU system to do this decoding work.

Which is why Nvidia paid $20 billion to take the best of Groq for itself.

AMD knows the co-founders of Cerebras really well is all that I am saying for now.

With the Vera-Rubin architecture, which refers to the 88-core “Vera” CV100 Arm server processor with custom “Olympus” cores paired to the “Rubin” R200 GPU accelerator, there are seven different chips that comprise five different styles of rackscale systems that can be mixed and matched in a Vera-Rubin AI supercomputer.

Huang showed off a comparison of 1 GW of “Hopper” H100 GPU capacity paired with X86 processors and embodied in HGX NVL8 systems (eight GPUs sharing memory on a scale up network, scaling out using InfiniBand) to what we presume is a cluster of VR200 NVL72 rackscale systems (72-way memory sharing for the GPUs).

In this comparison, it takes half as many GPUs to deliver 13.3X more AI processing performance. To be fair, the H100 could only shrink precision to FP8, while the R200 will have FP4 formats (just like the prior “Blackwell” GPUs did). So 2X of that 13.3X comes from the precision shrink. And the FP4 formats are not just a benchmark game, either – models are being tweaked to get the precision of answers to within a point or two of FP8 while cutting the data and therefore the processing precision in half. People are making that trade with production workloads.

But here’s the thing. If you need half as many GPUs, but they cost three or four times as much each, Nvidia gets to radically increase its revenues by selling at least twice as many devices, but your IT budget doesn’t go down and if your AI workloads are scaling – and they most certainly will be – then your IT budget increases. But so does that of every other IT organization deploying AI, and now the demand once again far outstrips supply, compelling prices to rise even more, driving Nvidia revenues and profits even higher than they might otherwise be in a non-constrained environment.

It’s good to be the Inference King.

But it was almost Jonathan Ross, creator of the Google TPU and the arguably much better Groq architecture, that was inference king. Ross just got an offer he could not refuse, and I think there is a very good chance Cerebras will get one, too. Intel missed its chance with SambaNova Systems – but perhaps there is still time and money to get a deal done.