惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
V
V2EX
博客园_首页
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Recent Announcements
Recent Announcements
博客园 - 司徒正美
Microsoft Security Blog
Microsoft Security Blog
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Latest news
Latest news
Vercel News
Vercel News
The Register - Security
The Register - Security
T
The Exploit Database - CXSecurity.com
S
Schneier on Security
N
Netflix TechBlog - Medium
WordPress大学
WordPress大学
小众软件
小众软件
L
Lohrmann on Cybersecurity
GbyAI
GbyAI
P
Privacy & Cybersecurity Law Blog
T
Tor Project blog
AWS News Blog
AWS News Blog
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
K
Kaspersky official blog
B
Blog RSS Feed
G
Google Developers Blog
量子位
大猫的无限游戏
大猫的无限游戏
Google DeepMind News
Google DeepMind News
Scott Helme
Scott Helme
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
I
Intezer
雷峰网
雷峰网
Martin Fowler
Martin Fowler
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Blog — PlanetScale
Blog — PlanetScale
IT之家
IT之家
F
Full Disclosure
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 【当耐特】
The Hacker News
The Hacker News
U
Unit 42
S
SegmentFault 最新的问题
I
InfoQ
aimingoo的专栏
aimingoo的专栏
Y
Y Combinator Blog
宝玉的分享
宝玉的分享
罗磊的独立博客
Spread Privacy
Spread Privacy
C
CERT Recently Published Vulnerability Notes

NVIDIA Newsroom

Claude Meets Blackwell Ultra: Anthropic’s Models Now Run on NVIDIA GB300 in Azure Firefly Aerospace Operates NVIDIA Jetson in Lunar Orbit for the First Time Open Models, Closed Environments: Palantir Brings Secure AI to US Agencies With NVIDIA Nemotron The Ultimate Summer Sale Pairing: Steam Sale Meets GeForce NOW Discounts NVIDIA and AWS Collaborate to Bring AI to Production at Scale How Businesses Are Building Specialized AI They Can Trust NVIDIA Announces BioNeMo Agent Toolkit — Tools for Agents to Accelerate Scientific Discovery NVIDIA Powers Over 400 of the World’s 500 Fastest Supercomputers NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations At ISC, JUPITER Shows What Exascale Science Looks Like NAIRR Science Program Reshapes Scientific Research, Powered by NVIDIA AI Infrastructure From Materials Simulation to Experimental Astronomy, New NVIDIA AI Software Unlocks Scientific Discoveries NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory Eco Wave Power Turns Waves Into Watts With NVIDIA AI Infrastructure and Digital Twins NVIDIA Vera Rubin Delivers World-Class Supercomputers for Science Europe Unveils a Record 35 New NVIDIA AI Supercomputers NVIDIA Announces Halos for Robotics, the Industry’s First Full-Stack Safety System for Physical AI Hotter Than a Hot Tub: The 45°C Breakthrough to Cool AI’s Biggest Machines How FERC’s Large-Load Interconnection Actions Help Address Grid Stress, Improve Affordability At Cannes Lions, NVIDIA Partners Reshape Advertising and Marketing With AI Sync and Stream: GeForce NOW Connects to Members’ Game Libraries Across Devices France Advances Europe’s AI Future With NVIDIA Technologies Hands Free, AIs Forward: NVIDIA XR AI Brings Agents to AR Glasses Coherent Breaks Ground on Expanded Texas Facility, Scaling AI’s Optical Backbone HPE AI Factory With NVIDIA Expands for the Era of Agents Fastest, Largest, Strongest: NVIDIA Blackwell Sweeps MLPerf Training 6.0 NVIDIA Blackwell Leads on First Agentic AI Infrastructure Benchmark NVIDIA Stockholder Meeting Set for June 24; Individuals Can Participate Online Save Big and Play Bigger: GeForce NOW Summer Sale Brings Major Membership Savings For Robotaxis, Safety Must Be Built In, Not Bolted On NVIDIA Confidential Computing to Help Expand Apple’s Private Cloud Compute How the UK Is Turning Sovereign AI Ambition Into Action With NVIDIA Technologies NVIDIA and LG Group Build an AI Factory to Advance Physical AI, Mobility and AI Infrastructure NVIDIA and Doosan Group Collaborate to Advance Physical AI and AI Factory Infrastructure NVIDIA and SK hynix Announce Multiyear Technology Partnership to Advance Memory for AI Factories SK Telecom and NVIDIA Build AI Infrastructure to Power Korea’s AI Innovation NAVER Expands AI Infrastructure With NVIDIA to Serve Surging Global AI Demand NVIDIA, KRAFTON, NC and Reigning ‘League of Legends’ Champions T1 Celebrate RTX Spark at Korea’s PC Bangs Seoul Purpose: How NVIDIA and South Korea Are Building the Future of AI Forecast: Fun Ahead — 18 Games Join in June to Stream on GeForce NOW NVIDIA Research Unlocks Advanced Grasping, Smarter Autonomous Driving and Agent Training at Scale NVIDIA Enables the Next Era Of Physical AI Research With Agent Skills For Autonomous Vehicles, Robotics And Vision AI Industrial Software Leaders Build Secure, Autonomous AI Engineers With NVIDIA NemoClaw NVIDIA Partners With Microsoft on Unified Stack for Agentic AI Deployment, From Windows Devices to Cloud to Local Why Financial Institutions Are Converging on Transaction Foundation Models to Build Their Own Intelligence NVIDIA Jetson Brings Agentic AI to the Physical World NVIDIA AI Cloud Ecosystem Expands Worldwide to Meet Global AI Compute Demand NVIDIA Factory Operations Blueprint Gives Factories a New AI Brain Taiwan’s Industry Titans Turbocharge World’s AI Infrastructure Buildout With NVIDIA NVIDIA and TSMC Bring AI Into Fabs to Advance Semiconductor Design and Manufacturing NVIDIA, Foxconn and Taiwan Medical Centers Bring Agentic and Physical AI to ‘Healthy Taiwan’ NVIDIA Releases Major Collection of Open Source Agent Tools and Skills for Physical AI NVIDIA Announces NVIDIA Isaac GR00T Reference Humanoid Robot for Academic Research NVIDIA DRIVE Hyperion Becomes the Global Platform for a Robotaxi-Ready World NVIDIA Launches Alpamayo 2 Super Open Reasoning Model for Robotaxis How Cosmos 3 Helps Physical AI Think Before It Acts NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI NVIDIA DGX Station for Windows Puts a Trillion-Parameter AI Supercomputer on Every Enterprise Desk NVIDIA Levels Up Local AI Agents Across RTX PCs and DGX Spark NVIDIA and Microsoft Reinvent Windows PCs for the Age of Personal AI Enterprise Software Leaders Build AI Agents With NVIDIA NVIDIA Unveils Vera, the CPU for Agents NVIDIA Vera BlueField-4 STX Brings Agentic AI Storage Processing With In-Silicon Security NVIDIA Vera Rubin Ramps Into Full Production to Power Agentic AI Factories Worldwide NVIDIA DSX Gives Infrastructure Builders the Playbook for AI Factories NVIDIA Research Advances Robotics From Simulation to the Real World The Name’s Gaming … Cloud Gaming: ‘007 First Light’ Launches on GeForce NOW NVIDIA GTC Taipei at COMPUTEX: Live Updates on What’s Next in AI NVIDIA CEO Jensen Huang at Dell Technologies World: ‘Demand Is Going Parabolic, Utterly Parabolic’ Linked and Loaded: Gaijin Single Sign-On Now Available on GeForce NOW NVIDIA and ServiceNow Partner on New Autonomous AI Agents for Enterprises It’s Gonna Be May: 16 Games Hit the Cloud This Month, With More NVIDIA GeForce RTX 5080 Power NVIDIA Launches Nemotron 3 Nano Omni Model, Unifying Vision, Audio and Language for up to 9x More Efficient AI Agents Into the Omniverse: Manufacturing’s Simulation-First Era Has Arrived Tag, You’re It: GeForce NOW Levels Up Game Discovery With Xbox Game Pass and Ubisoft+ Labels Making Sense of the Early Universe From Rainforests to Recycling Plants: 5 Ways NVIDIA AI Is Protecting the Planet NVIDIA and Google Cloud Collaborate to Advance Agentic and Physical AI Autonomous AI at Scale: Adobe Agents Unlock Breakthrough Creative Intelligence With NVIDIA and WPP No Need for Space Gear — Capcom’s ‘PRAGMATA’ Joins GeForce NOW on Launch Day Rethinking AI TCO: Why Cost per Token Is the Only Metric That Matters New Adobe Premiere Color Grading Mode Accelerated on NVIDIA GPUs Strength and Destiny Collide: ‘Samson: A Tyndalston Story’ Arrives in the Cloud National Robotics Week — Latest Physical AI Research, Breakthroughs and Resources From RTX to Spark: NVIDIA Accelerates Gemma 4 for Local Agentic AI Press Start on April: GeForce NOW Brings 10 Games to the Cloud Efficiency at Scale: NVIDIA, Energy Leaders Accelerating Power‑Flexible AI Factories to Fortify the Grid Into the Omniverse: NVIDIA GTC Showcases Virtual Worlds Powering the Physical AI Era Game On: Five New Titles Now Streaming on GeForce NOW The Future of AI Is Open and Proprietary Blowing Off Steam: How Power-Flexible AI Factories Can Stabilize the Global Energy Grid Advancing Open Source AI, NVIDIA Donates Dynamic Resource Allocation Driver for GPUs to Kubernetes Community How Autonomous AI Agents Become Secure by Design With NVIDIA OpenShell NVIDIA's CEO Projects $1 Trillion in AI Chip Sales as New Computing Era Begins Nvidia CEO: We have the most energy efficient architecture in the world An Interview with Nvidia CEO Jensen Huang About Accelerated Computing NVIDIA GTC 2026: Live Updates on What’s Next in AI Smooth Moves: 90 Frames-Per-Second Virtual Reality Arrives on GeForce NOW From Simulation to Production: How to Build Robots With AI More Than Meets the Eye: NVIDIA RTX-Accelerated Computers Now Connect Directly to Apple Vision Pro
NVIDIA Accelerates Google DeepMind’s DiffusionGemma for Local AI
Michael Fukuyama · 2026-06-11 · via NVIDIA Newsroom

Today, Google DeepMind released DiffusionGemma — an experimental open model built for exceptionally fast text generation. NVIDIA has optimized DiffusionGemma to run even faster across NVIDIA GeForce RTX GPUs, the NVIDIA RTX PRO platform and NVIDIA DGX Spark systems, from local PCs to the cloud. 

Rather than generating text one word at a time, DiffusionGemma generates multiple words in parallel to output whole blocks of text, opening a new, low-latency frontier for the kind of single-user workloads that developers, researchers and AI enthusiasts run every day. 

Features of the new model include: 

  • Parallel generation: DiffusionGemma denoises up to 256 tokens per step instead of predicting one at a time. 
  • Built on Gemma 4: DiffusionGemma is built on Gemma 4, a 26-billion-parameter mixture-of-experts model that activates just 3.8 billion parameters per step, pairing a diffusion head with Google’s Gemma 4 architecture. 
  • Up to 4x faster performance: The boost means fast text generation, where single-user generation usually stalls — on local hardware. 
  • Open and local: DiffusionGemma is open weights under a permissive Apache 2.0 license and runs entirely on RTX and DGX Spark — no cloud, no per-token cost — with day-zero support in Hugging Face Transformers, vLLM and Unsloth. 

A Different Way to Generate Text 

Almost every large language model (LLM) in wide use today is autoregressive — meaning it generates text one token at a time, with each new word depending on the one before it. That sequential process is what makes interactive AI feel like it’s typing. 

DiffusionGemma takes a different path. Built on the Gemma 4 26B mixture-of-experts architecture, it generates text the way diffusion models generate images: by starting from noise and refining a whole block of text at once. Each step denoises up to 256 tokens in parallel rather than emitting a single token and waiting to compute the next. 

The result is a model that thinks in blocks instead of sequentially. For latency-sensitive, single-user work — such as interactive chat, agentic loops or on-device assistants that plan and act — that parallelism translates into responses fast enough to keep pace with how developers think and iterate.

DiffusionGemma Flies on NVIDIA GPUs 

Generating one token at a time is fundamentally a memory-bound problem — a traditional LLM spends most of its time waiting on memory bandwidth, not doing math, which leaves a lot of compute on the table. 

Diffusion flips the equation. Pulling a full 256-token block through the transformer in parallel is a compute-bound workload — exactly what NVIDIA GPUs are built for. NVIDIA Tensor Cores accelerate the dense parallel math, and the CUDA software stack lets the model run efficiently from day one without bespoke tuning. In short, the model’s design plays directly to the GPUs strengths. 

That shows up in the numbers. DiffusionGemma delivers 1,000 tokens/sec on a single NVIDIA H100 Tensor Core GPU, 150 tokens/sec on NVIDIA DGX Spark and up to 2,000 tokens/sec on NVIDIA DGX Station — roughly 4x faster than an equivalent autoregressive model running in the same single-user regime.

That advantage holds across NVIDIA’s full lineup, running: 

  • Locally on the NVIDIA DGX Spark deskside personal AI supercomputer — powered by the NVIDIA GB10 Grace Blackwell Superchip with 128GB of unified memory — with the preinstalled NVIDIA AI software stack ready for prototyping, fine-tuning and fully local agent workflows. 
  • On NVIDIA RTX PRO 6000 workstations, providing developers, researchers and AI professionals with the headroom to run local low-latency generation and agentic loops as part of a professional workflow. 
  • On DGX Station, delivering best-in-class, local high-speed inference with up to 2,000 tokens/sec for low-latency text generation and agentic loops with 748GB of coherent memory.
  • On GeForce RTX GPUs, with llama.cpp support coming soon. 

The fastest way to start testing and prototyping the model is through Hugging Face Transformers, which runs DiffusionGemma on a GeForce RTX 5090 or DGX Spark out of the box. For higher-throughput inference, vLLM provides day-zero serving support.  

For adapting the model to a specific task or domain, fine-tuning is available through Unsloth and NVIDIA NeMo framework, with ready-made DGX Spark playbooks to get a local environment running quickly. Check out the vLLM playbooks for DGX Spark , RTX PRO and DGX Station. 

Try Diffusion Gemma on Hugging Face or test it for free using NVIDIA-hosted application programming interfaces at build.nvidia.com. 

Go deeper on the architecture and local deployment by reading the NVIDIA technical blog and the Google DeepMind announcement.

#ICYMI: The Latest From RTX AI Garage 

🎬 NVIDIA researchers released SANA-WM, an open source world model that turns a single image and a camera path into a minute-long, 720p video with precise 6-DoF control. At just 2.6 billion parameters, its distilled version generates a full 60-second clip in 34 seconds on a single NVIDIA GeForce RTX 5090 GPU using the NVFP4 format — delivering up to 36x higher throughput than comparable open models while running on one GPU. Read the paper. 

🛠️ Building Windows agents just got a full toolset — NVIDIA and Microsoft rolled out turnkey agent sandboxing on native Windows — Microsoft eXecution Containers plus the NVIDIA OpenShell runtime — alongside up to 2x faster agentic inference and native Windows support for Hermes Agent. 

🤖DGX Spark goes from unboxing to a running agent in minutes — A streamlined NVIDIA NemoClaw install gets developers to a working local agent fast, with Qwen3.6-35B running up to 2.6x faster on vLLM. And the new cluster assistant in NVIDIA Sync links up to four DGX Spark units into one 512GB pool — enough for ~400-billion-parameter models. 

Plug in to RTX Spark on FacebookInstagramTikTok and X — and stay informed by subscribing to the RTX Spark newsletter. 

See notice regarding software product information.