惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

有赞技术团队
有赞技术团队
小众软件
小众软件
The Cloudflare Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园_首页
酷 壳 – CoolShell
酷 壳 – CoolShell
爱范儿
爱范儿
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
雷峰网
雷峰网
Jina AI
Jina AI
博客园 - 【当耐特】
V
Visual Studio Blog
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
罗磊的独立博客
大猫的无限游戏
大猫的无限游戏
量子位
IT之家
IT之家
G
Google Developers Blog
V
V2EX
The GitHub Blog
The GitHub Blog
月光博客
月光博客
GbyAI
GbyAI

Latest from Tom's Hardware in News-analysis

Kimi K3 rocks the AI industry as Moonshot AI undercuts closed-source American competitors on price — but the huge 2.8T open-weight model still needs serious hardware to deploy at scale Tower Semiconductor revives shuttered Panasonic-era fab in $3 billion Japan photonics expansion — METI-backed plan targets $3.6 billion revenue by 2028 Intel Intel Micron commits $500 million to GlobalWafers Anthropic says it can read Claude Elon Musk receives FTC greenlight to buy Mesh Optical as interconnects emerge as AI SiPearl South Korea Inside the history of DRAM price-fixing lawsuits — how HBM allocations could make a difference after two decades of failed cases U.S. PC shipments drop 7%, market isn Chinese Z.ai The AI tokenmaxxing party is crashing over spiraling costs — leaked consulting firm audio suggests no one is sure how to measure AI effectiveness US Secures Netherlands for Pax Silica Alliance in key win for strategic chip alliance — tension remains over MATCH Act restrictions Arm servers capture over 45% of data center market revenue — GPU clusters and high-end AI infrastructure fuel a tectonic shift away from x86 Post-silicon era gets closer as industry giants crack the 2D transistor scaling bottleneck with breakthrough tech — imec, ASML, and TSMC fab complementary 2D-material transistors at 50nm pitch on a 300mm wafer US pulls the Marvell details vision of optically-interconnected data centers spanning across thousands of kilometers — new interconnects sampling later this year would allow CSPs to pool resources based on workload Nvidia's high-speed AI data center storage servers break cover, touting 2.9 petabytes of storage and extreme PCIe 6.0 performance — Wiwynn shows off SCADA server with GPU-accelerated storage AI is set to consume up to 600 billion gallons of water by 2030 — rising energy consumption primarily to blame as… Google reportedly books Intel for packaging more than 3 million TPUs in 2028 — SK hynix is testing Intel's… Anthropic's warning over AI self-improvement has a hidden message — accelerating development requires more compute before companies ever risk losing control of frontier AI models Executives are cutting jobs for an AI future that hasn't fully arrived yet, even as productivity gains remain difficult to prove — data neither confirms nor refutes an AI unemployment apocalypse Jensen Huang says 'every edge device will become autonomous' — Nvidia maps one computing pattern from… AMD's Helios MI455X AI platform breaks cover, initial systems use UALink-over-Ethernet interconnects — AMD's Vera Rubin rival surfaces, but the downsides of Ethernet could hamstring performance Frore shows off LiquidJet Nexus coldplate for Nvidia Vera Rubin, other AI accelerators — offers up claimed 10% token generation boost over rival liquid-cooling solutions The rise of local agentic computing faces a brutal reality: rising DRAM prices —  RTX Spark, Gorgon Halo chips subject to 63% DRAM contract price hike this quarter Astera Labs showcases 320-lane PCIe 6.0 switch for vendor-agnostic scaling in data centers — up to 80 accelerators… AI costs begin to bite as agents may increase token demand by 24 times, says Goldman Sachs report — Uber and Microsoft among companies feeling the bite of tokenized billing IBM spins off America's first quantum chip foundry with $2 billion in federal and private funding — newly-minted 'Anderon' foundry to offer 300mm quantum wafer fab and manufacturing services
How Nvidia's $20 billion Groq 3 LPU deal reshapes the Nvi...
Luke James · 2026-03-19 · via Latest from Tom's Hardware in News-analysis

Nvidia unveiled the Groq 3 language processing unit at GTC 2026 in San Jose on Monday, marking the first chip to emerge from its $20 billion licensing and talent deal with AI inference startup Groq, which was struck on Christmas Eve last year. The SRAM-based inference accelerator slots into the Vera Rubin platform as a dedicated decode-phase co-processor, and Nvidia plans to ship it in Q3 2026, manufactured by Samsung on a 4nm process. It is the company's first rack-scale product built around non-GPU silicon — and its arrival has already displaced a homegrown Nvidia chip from the roadmap.

The LP30 chip at the heart of the Groq 3 LPX rack carries 512 MB of on-chip SRAM per die, delivering 150 TB/s of memory bandwidth. That figure dwarfs the 22 TB/s available from the 288 GB of HBM4 on each Rubin GPU. A full LPX rack houses 256 LPUs for a total of 128GB of SRAM and 40 PB/s of aggregate bandwidth. Nvidia claims the LPX rack, paired with a Vera Rubin NVL72, delivers 35 times higher throughput per megawatt than Blackwell NVL72 alone for trillion-parameter models, at a target price point of $45 per million tokens.