惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
The GitHub Blog
The GitHub Blog
Y
Y Combinator Blog
博客园 - 三生石上(FineUI控件)
T
The Blog of Author Tim Ferriss
宝玉的分享
宝玉的分享
Hugging Face - Blog
Hugging Face - Blog
WordPress大学
WordPress大学
V
Visual Studio Blog
小众软件
小众软件
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
MongoDB | Blog
MongoDB | Blog
V
V2EX
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 【当耐特】
Microsoft Azure Blog
Microsoft Azure Blog
The Cloudflare Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Engineering at Meta
Engineering at Meta
L
LangChain Blog
Martin Fowler
Martin Fowler
GbyAI
GbyAI
博客园 - 司徒正美

Latest from Tom's Hardware in News-analysis

Kimi K3 rocks the AI industry as Moonshot AI undercuts closed-source American competitors on price — but the huge 2.8T open-weight model still needs serious hardware to deploy at scale Tower Semiconductor revives shuttered Panasonic-era fab in $3 billion Japan photonics expansion — METI-backed plan targets $3.6 billion revenue by 2028 Intel Intel Micron commits $500 million to GlobalWafers Anthropic says it can read Claude Elon Musk receives FTC greenlight to buy Mesh Optical as interconnects emerge as AI SiPearl South Korea Inside the history of DRAM price-fixing lawsuits — how HBM allocations could make a difference after two decades of failed cases U.S. PC shipments drop 7%, market isn Chinese Z.ai The AI tokenmaxxing party is crashing over spiraling costs — leaked consulting firm audio suggests no one is sure how to measure AI effectiveness US Secures Netherlands for Pax Silica Alliance in key win for strategic chip alliance — tension remains over MATCH Act restrictions Arm servers capture over 45% of data center market revenue — GPU clusters and high-end AI infrastructure fuel a tectonic shift away from x86 Post-silicon era gets closer as industry giants crack the 2D transistor scaling bottleneck with breakthrough tech — imec, ASML, and TSMC fab complementary 2D-material transistors at 50nm pitch on a 300mm wafer US pulls the Marvell details vision of optically-interconnected data centers spanning across thousands of kilometers — new interconnects sampling later this year would allow CSPs to pool resources based on workload Nvidia's high-speed AI data center storage servers break cover, touting 2.9 petabytes of storage and extreme PCIe 6.0 performance — Wiwynn shows off SCADA server with GPU-accelerated storage AI is set to consume up to 600 billion gallons of water by 2030 — rising energy consumption primarily to blame as… Google reportedly books Intel for packaging more than 3 million TPUs in 2028 — SK hynix is testing Intel's… Anthropic's warning over AI self-improvement has a hidden message — accelerating development requires more compute before companies ever risk losing control of frontier AI models Executives are cutting jobs for an AI future that hasn't fully arrived yet, even as productivity gains remain difficult to prove — data neither confirms nor refutes an AI unemployment apocalypse Jensen Huang says 'every edge device will become autonomous' — Nvidia maps one computing pattern from… AMD's Helios MI455X AI platform breaks cover, initial systems use UALink-over-Ethernet interconnects — AMD's Vera Rubin rival surfaces, but the downsides of Ethernet could hamstring performance Frore shows off LiquidJet Nexus coldplate for Nvidia Vera Rubin, other AI accelerators — offers up claimed 10% token generation boost over rival liquid-cooling solutions The rise of local agentic computing faces a brutal reality: rising DRAM prices —  RTX Spark, Gorgon Halo chips subject to 63% DRAM contract price hike this quarter Astera Labs showcases 320-lane PCIe 6.0 switch for vendor-agnostic scaling in data centers — up to 80 accelerators… AI costs begin to bite as agents may increase token demand by 24 times, says Goldman Sachs report — Uber and Microsoft among companies feeling the bite of tokenized billing IBM spins off America's first quantum chip foundry with $2 billion in federal and private funding — newly-minted 'Anderon' foundry to offer 300mm quantum wafer fab and manufacturing services
How Nvidia's $20 billion Groq 3 LPU deal reshapes the Nvi...
Luke James · 2026-03-19 · via Latest from Tom's Hardware in News-analysis

Nvidia unveiled the Groq 3 language processing unit at GTC 2026 in San Jose on Monday, marking the first chip to emerge from its $20 billion licensing and talent deal with AI inference startup Groq, which was struck on Christmas Eve last year. The SRAM-based inference accelerator slots into the Vera Rubin platform as a dedicated decode-phase co-processor, and Nvidia plans to ship it in Q3 2026, manufactured by Samsung on a 4nm process. It is the company's first rack-scale product built around non-GPU silicon — and its arrival has already displaced a homegrown Nvidia chip from the roadmap.

The LP30 chip at the heart of the Groq 3 LPX rack carries 512 MB of on-chip SRAM per die, delivering 150 TB/s of memory bandwidth. That figure dwarfs the 22 TB/s available from the 288 GB of HBM4 on each Rubin GPU. A full LPX rack houses 256 LPUs for a total of 128GB of SRAM and 40 PB/s of aggregate bandwidth. Nvidia claims the LPX rack, paired with a Vera Rubin NVL72, delivers 35 times higher throughput per megawatt than Blackwell NVL72 alone for trillion-parameter models, at a target price point of $45 per million tokens.