惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
T
The Exploit Database - CXSecurity.com
IT之家
IT之家
酷 壳 – CoolShell
酷 壳 – CoolShell
T
Tailwind CSS Blog
宝玉的分享
宝玉的分享
Last Week in AI
Last Week in AI
阮一峰的网络日志
阮一峰的网络日志
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Hugging Face - Blog
Hugging Face - Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
博客园 - 聂微东
博客园 - Franky
美团技术团队
WordPress大学
WordPress大学
博客园 - 司徒正美
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
D
DataBreaches.Net
腾讯CDC
大猫的无限游戏
大猫的无限游戏
人人都是产品经理
人人都是产品经理
Microsoft Azure Blog
Microsoft Azure Blog
D
Docker
Security Archives - TechRepublic
Security Archives - TechRepublic
F
Fortinet All Blogs
T
Tor Project blog
G
GRAHAM CLULEY
Simon Willison's Weblog
Simon Willison's Weblog
I
InfoQ
Cyberwarzone
Cyberwarzone
V
V2EX
T
Tenable Blog
NISL@THU
NISL@THU
Scott Helme
Scott Helme
K
Kaspersky official blog
Latest news
Latest news
S
Schneier on Security
Martin Fowler
Martin Fowler
博客园 - 三生石上(FineUI控件)
Know Your Adversary
Know Your Adversary
Microsoft Security Blog
Microsoft Security Blog
S
Securelist
M
MIT News - Artificial intelligence
V
Vulnerabilities – Threatpost
P
Proofpoint News Feed
L
LangChain Blog
T
Threat Research - Cisco Blogs
Spread Privacy
Spread Privacy
T
Threatpost
有赞技术团队
有赞技术团队

Towards AI

Building AI Agents in Rust — part 4 | Towards AI The Verified Identity Agent Bridge | Towards AI You Can’t Prompt Your Away Your LLM Problems | Towards AI The Free Agent Trap | Towards AI Your Agentic Loop Will Drift. Here Is the KL Divergence Equation That Measures How Far It Has Wandered From Its Original Instruction. | Towards AI Beyond Chat: Processing Images, PDFs, and Documents with the OpenAI Adapter in Oracle Integration Cloud | Towards AI Building AI Agents in Rust — part 3 | Towards AI Self-Hosting Airflow at Home: Automating Stock Price Data Collection | Towards AI The 76-Hour Frontier: How the Takedown of Claude Fable 5 Birthed the Military-Industrial-AI Complex | Towards AI I Trained a Markdown File to Boost GPT-5.5 by 23 Points — It Shouldn't Work | Towards AI We Replaced ChatGPT With a Local AI Server. Six Months of Honest Data. | Towards AI What Really Makes Cars Pollute? A Data Science Deep Dive into CO₂ Emissions | Towards AI Training GPT-2 From Scratch on a GTX1050 | Towards AI Principal Component Analysis (PCA): Theory, Mathematics, and Applications Build a Zero-Cost Web Automation Pipeline With OpenRouter, OpenClaw, and MediaUse I Gave Qwen3.7-Plus a Screenshot and It Found the Exact Pixel to Click for $0.40 Beyond the Prompt: Why Autonomous AI Agents Are Replacing the Chatbot Moonshot Cracked Claude Code’s Playbook with an MIT Terminal Agent and a $0.60 Model Connections, Roles, and Warehouses: Getting CoCo Desktop Production-Ready from Day One My First $5,000 Month Writing About AI Engineering on Medium Google Shrank Gemma 4 by 72% and Unsloth Fixed the 4-Bit Bug Nobody Else Caught on One 4090, and 4-Bit Shouldn’t Be This Good LangChain Explained: Understanding Models, Prompts, Chains, Memory, Indexes, and Agents TOON: Beyond JSON for LLMs Claude Code Casual, Pro, Elite: The Three Working Personas of Claude Code Mastery MiniMax M3 Decodes 1M Tokens 15x Faster — and It Shouldn’t Be This Cheap Using Amazon SQS for AI Agent Orchestration I Ran a 1.5B-Active Model on My Laptop That Embarrassed a 26B by 46 Points How to Build a Self-Improving Company with AI Part 3 — Implementation/Engine-Level: Choosing the Runtime That Gives You These for Free Part 2 — Serve-Level Speed: System Design That Stabilizes P95/P99 3-Part Series: LLM Latency in Production (Part 1) Claude Code: The AI Coding Partner Changing How Developers Build Software Claude Code Pitfalls: Claude Code Won’t Do What You Told It: A Troubleshooting Catalog Full-Stack Data Scientists for the Agentic Coding World Building Production-Grade AI Skills with Snowflake Cortex AI Function Studio I Tried 10 AI Agent Frameworks in 2026 — Here’s the Honest Guide I Wish I Had Earlier How One Spring Boot Optimization Saved Our Startup $30,000 a Year Inside Palantir AIP: How the World’s Most Controversial AI Platform Actually Works What Is a Reverse Proxy? (And Why Every Backend Developer Should Care) What Claude Opus 4.8 Actually Changes If You’re Building Agents QWEN 3.7 Max Worked For 35 Hrs Straight And The Results Were Mind-blowing When LLMs Meet Knowledge Graphs on the Battlefield Fine-Tuning is Dead: Why Context Orchestration Won in 2026 5 Things Broke When I Shipped a RAG + MCP Agent to Production. Google Co-Scientist: Hyper Scaling Research and Discovery Microsoft Just Embarrassed Browser Web Agents — 1,000 Lines Made GPT-5.4 Beat Opus 4.6 on 200 Web Tasks The Modern Data Stack Is Broken — Here’s How to Fix It With AI, Governance, and Real Architecture Building Production MCP Servers: What the Spec Won’t Tell You When Should an Agent Stop? The Anatomy of Termination Harness Engineering: The Layer That Matters More Than the Model AI Engineers Who Can’t Debug Are Getting Fired (Here’s How I Debug with Claude Code) Claude Code Memory: Why You Keep Explaining the Same Thing to Claude (and the Five Layers That Fix It) Claude Code Subagents: The Claude Code Feature You Skip Every Day (And Why It Quietly Wrecks Your Sessions) Agentic AI and the SMB Banking Advantage Claude Code: Spec-Driven Development — Why Your AI Coding Sessions Fall Apart at Hour Three The Real Cost of Agentic AI Nobody Budgets For SVM : 40 must visit Interview Questions (Part 2) Your AI Agent Works Perfectly in the Demo. Here Are the 6 Ways It Dies in Production. Unleashing the Power of ONNX for Speedier SBERT Inference Terraform vs CI/CD for Serverless Deployments Merve Noyan Stopped Writing Training Scripts — Her Agent Just Fine-Tuned 18 Models Solo for $11.40 Why Your Sales Forecast Is Always 20% Wrong (And How To Make It 12% Wrong) Genetic Cubic n{C/A} Ratios For Elementary Robotics Design Top 20 AdaBoost Interview Questions & Answers (Part 2 of 2) Agentic AI Vs AI Agents — What Are the Key Differences? LAI #127: The Infrastructure Layer of AI Is Becoming the Product Anthropic Caught Its Own AI Planning to Blackmail Engineers RNNs Cannot Think What Transformers Think Cheaply. ICLR 2026 Proved the Gap Is Exponential. Time Series Made So Easy My Aunt Got It on the Second Read Claude Cowork 101 | Towards AI Is 3-Bit KV Cache the Holy Grail? A Reality Check on Google’s TurboQuant LangGraph Multi-Agent Architecture: Building a Self-Critiquing AI Debate System AutoML on Autopilot | Towards AI I Ran This Open-Source AI Tool on a Messy Codebase and Got 71x Fewer Tokens — Here Is Exactly What Happened Month in 4 Papers (April 2026) AI Kept Forgetting My Notes. Fixing That Taught Me How It Actually Works. How ChatGPT Makes You Addicted Crack ML Interviews with Confidence: K-Nearest Neighbors (KNN 20 Q&A) The Event-Driven Blueprint: How I Scaled a Spring Boot System to 10 Million Kafka Messages/Day Building Vector Search? Why FAISS Alone Isn’t Enough TAI #202: GPT-5.5 Moves Codex Into Real Work Machine Learning System Design -The Model Serving Triangle, With One Forward Pass Flowing Through Every Trade-off (Part3) AI Orchestration in Action: How MuleSoft and LLMs Fuel the Future of Enterprise AI GPT-4 Has 1.8 Trillion Parameters. It Uses 2% of Them Per Token. Part 20: Data Manipulation in Multi-Dimensional Aggregation A Fundamental Introduction to Genetic Algorithm -Part Two TAI #200: Anthropic’s Mythos Capability Step Change and Gated Release From Notebook to Production: Running ML in the Real World (Part 4) Sqribble’s Template‑Driven Document Automation Anthropic Just Shipped the Layer That’s Already Going to Zero Long-Term vs Short-Term Memory for AI Agents: A Practical Guide Without the Hype The L1 Loss Gradient, Explained From Scratch Your Postcode Is Deciding Your Care. I Built a Pipeline to Prove It. I Directed AI Agents to Build a Tool That Stress-Tests Incentive Designs. Here’s What It Found. Your System Prompt Is the Product — Not the Feature The LLM Wiki Trend Has a Retention Problem Nobody Mentions Top 20 Data Preparation Interview Questions and Answers (Part 2 of 2) LAI #122: Word Embeddings Started in 1948, Not With Word2Vec Top 15 Computer Vision Datasets [2026] 40 Generative AI Interview Questions That Actually Get Asked in 2026 (With Answers)
You Do Not Need 50 Diffusion Steps. Here Is What Nvidia Proved at GTC. | Towards AI
Siddhant Nitin Patil · 2026-06-25 · via Towards AI

Originally published on Towards AI.

You Do Not Need 50 Diffusion Steps. Here Is What Nvidia Proved at GTC.

The video diffusion industry has had the same conversation for two years.

Better model. More parameters. Higher resolution. Longer clips. Richer motion. And underneath all of it, the same silent constraint that nobody advertises: generating a single second of 720p video still takes long enough to make most real-time use cases a fantasy.

At GTC 2026 in San Jose, Nvidia’s Ziv Ilan from the AI Labs team in Paris gave a 20-minute talk that reframed the problem entirely. The title: You Might Not Need 50 Diffusion Steps.

The argument was not about a new model. It was about what happens when you stop treating the step count as a fixed constraint and start treating it as an engineering variable.

Why Step Count Is the Real Bottleneck

Diffusion models generate images and videos through iterative denoising. Random noise gets progressively cleaned up across a series of steps, each step moving the output closer to the final result. Standard production models run 20 to 50 denoising steps. Each step is a full forward pass through a model that, in the case of modern video diffusion architectures, can have 20 to 40 billion parameters.

The math compounds fast. A single 1,328 x 1,328 image generated with Qwen-Image involves approximately 12,900 TFLOPs of computation, producing a latency of up to 127 seconds per image on an Nvidia H20 GPU. For video, where you need consistent quality across frames with temporal coherence, the compute demand grows faster than linearly with resolution and duration.

This is why Adobe’s Firefly video generation model, before optimization, was architecturally capable but commercially constrained. State-of-the-art image diffusion already took tens of seconds per image. Video diffusion with a 50-step process at production resolution was simply not viable for interactive or real-time applications.

The path forward was not a bigger model. It was a smarter inference stack.

The Three-Technique Stack

Ilan’s talk organized the solution space into three composable techniques: quantization, caching, and distillation. Critically, these are not alternatives. They are stackable. You deploy them in combination, and each one adds a multiplier to the performance gains of the others.

Quantization: Making Each Step Cheaper

Quantization reduces the numerical precision of the model’s weights and activations from 16-bit or 32-bit floating point to lower-precision formats: INT8, FP8, or even FP4 in the latest research.

For LLMs, the impact of quantization is well understood and well documented. Diffusion models present a more complex picture because they are attention-heavy in ways that LLMs are not. The multi-head attention mechanisms in transformer-based diffusion architectures (DiT models) are more sensitive to precision loss than the feed-forward layers in autoregressive models. This means that naive quantization approaches developed for LLMs often produce measurable quality degradation in diffusion models even at INT8 precision.

The solution Nvidia has deployed in production, demonstrated through their collaboration with Black Forest Labs on Flux 2, uses dynamic quantization rather than static quantization. Static quantization pre-computes the activation range across a calibration dataset and applies fixed scaling factors at inference time. Dynamic quantization computes activation ranges on the fly per batch, adapting to the actual data distribution being processed. For diffusion models where the latent space evolves significantly across denoising steps, dynamic quantization maintains quality that static approaches cannot match.

The hardware layer amplifies this further. Nvidia’s Blackwell architecture introduced NVFP4 support, a 4-bit floating point format that, combined with Blackwell’s dedicated FP4 tensor cores, delivers performance gains that dwarf what FP8 achieved on Hopper. In ComfyUI benchmarks, NVFP4 optimizations on RTX 50-series cards delivered up to 3x performance boosts over FP16 baselines. For Stable Diffusion 3.5 Large, FP8 quantization alone cuts the VRAM requirement from 18GB to 11GB, opening up mid-range 12GB GPUs for a model that previously required 24GB.

The Adobe Firefly case is the most concrete enterprise data point. Using TensorRT with mixed FP8 and BF16 precision on Hopper GPUs via AWS EC2 P5 instances: 60% latency reduction, 40% total cost of ownership reduction, serving more users with fewer GPUs. This is not a research result. It is a production deployment that is live today.

One important note from Ilan on diffusion-specific quantization considerations: because these models are more attention-heavy than LLMs, the memory savings from quantization are less dramatic than in the LLM world. The performance gains still matter, but the ratio of memory benefit to compute benefit is different. Quantization should be treated as the entry-point optimization, the lowest-friction gain available, rather than the primary strategy.

Quantization gets you into the field. Caching and distillation win the game.

Caching: Skipping the Computation You Already Did

The second technique exploits a property of diffusion that is counterintuitive until you see it: adjacent denoising steps are highly redundant.

When a diffusion model runs 50 steps to generate a video frame, the feature representations in the model’s internal layers do not change dramatically between step 23 and step 24. The high-level structure, the composition, the semantic layout, these are largely determined in the early steps. The middle steps refine. The late steps clean up residual noise and adjust texture. Large swaths of the computation happening in steps 24 through 48 are recalculating values that changed very little from the previous step.

This is the same insight that motivated KV caching in LLMs: if you have already computed something and it has not changed meaningfully, do not recompute it. In the autoregressive case, KV cache is straightforward because you are generating one token at a time and the previously computed keys and values are definitionally unchanged. In diffusion, the cache mechanics are more complex because you are denoising across a full latent space simultaneously, but the redundancy is real and measurable.

T-cache, the approach Ilan referenced in his talk, operates at the full pixel or latent space level. It computes a similarity metric between the current denoising step’s output and the previous step’s output. If the change falls below a configurable threshold, the next step reuses the cached computation rather than running the full forward pass. The threshold is the key tuning parameter: set it too aggressively and you see visible quality degradation, set it too conservatively and the speed gains are marginal.

More recent caching techniques have moved from this global approach to chunk-based spatial caching. The insight, illustrated vividly in Ilan’s classroom analogy: if most of a video frame is static (the audience sitting still) but one region is dynamic (the presenter moving), you do not need to recompute the entire frame. You recompute only the dynamic region and reuse the cached computation for everything else.

BWCache, published in February 2026, implements this on top of HunyuanVideo and Wan 2.1 by using a lightweight similarity indicator to dynamically determine cache reuse per spatial block, without requiring additional training or architectural modifications. The approach is training-free, meaning it can be applied directly to existing pretrained models. The tradeoff: BWCache deliberately skips caching in the first few denoising steps where feature changes are most pronounced, sacrificing some acceleration in the early steps to protect generation quality where it matters most.

AdaCache adapts the caching interval based on motion complexity: videos with low spatial complexity and slow motion need far fewer recomputed steps than videos with fast motion and complex feature changes. This variable-rate approach more accurately matches compute to necessity, but its aggressive step-skipping policy can introduce visible artifacts in high-motion sequences.

The key practical guidance from the research: aggressive caching buys speed at the cost of quality, and the tradeoff is nonlinear. Modest threshold settings preserve quality nearly completely while delivering meaningful speedups. Extreme settings push performance further but require careful evaluation against your specific quality requirements.

For production deployments, Nvidia’s TRT-LLM visual gen repository exposes caching as a flag with a configurable threshold. The same capability is available in vLLM Omni and GLM Diffusion. For most teams, the practical starting point is: enable caching with a conservative threshold, benchmark your quality metrics against your uncached baseline, then tighten the threshold until quality begins to move.

Distillation: Eliminating the Steps You Do Not Need

The third technique is the most impactful and the most operationally complex. It is also the one most commonly misunderstood.

When DeepSeek released its R1 distillation results in early 2025, the industry learned what model distillation means in the LLM context: train a smaller student model to replicate the outputs of a larger teacher model, trading some quality for a dramatic reduction in compute requirements. The student model is smaller. The step count stays the same.

Diffusion distillation is different. The student model keeps the same number of parameters as the teacher. The goal is not to shrink the model. The goal is to train the student model to produce equivalent quality output in 4 to 8 steps instead of 50. In some architectures, in a single step.

The mechanism is conceptually elegant. You have a teacher model that runs 50 denoising steps. You train a student model to match the teacher’s output quality while compressing the trajectory. Two main approaches have emerged:

Trajectory-based distillation trains the student to follow the same denoising path as the teacher, step by step, but at a compressed rate. The student learns to mimic not just the final output but the intermediate states the teacher passes through. This produces stable training but constrains the student to a specific trajectory, limiting how much compression is possible before quality degrades.

Distribution-based distillation is less constrained. The student is trained only to match the teacher’s output distribution, not the specific trajectory it follows to get there. The student can discover its own path to the same destination. This approach generally achieves higher quality at greater compression ratios, and represents the current state of the art. Nvidia’s own FastGen library, released in February 2026, uses distribution-based methods as its primary technique, specifically DMD2 (Distribution Matching Distillation 2) and LADD (Latent Adversarial Diffusion Distillation).

FastGen supports models up to 14B parameters with full multi-GPU sharding via FSDP2, dynamic batching for gradient accumulation on constrained hardware, and a Hydra-style configuration system that separates experiment design from method-specific hyperparameters. The benchmarks Nvidia has published show 10x to 100x sampling speedups with maintained quality, varying by model architecture and compression target.

The GTC 2026 demonstration combined FastGen distillation with quantization on a single Blackwell B200 GPU and achieved near real-time video generation at production resolution. That is the end-to-end result of the full three-technique stack applied together.

The operational complexity Ilan flagged honestly: distillation is a post-training technique. It requires data, compute, and iteration. For a general-purpose use case, Nvidia’s open-source data can get you most of the way there. For domain-specific applications, whether that is medical imaging, satellite imagery, protein structure visualization, or industrial inspection, you will need data from your target distribution. The model trained on general video distributions will not generalize perfectly to your specific use case without fine-tuning on representative samples.

The compute requirement scales with model size but does not require the hardware needed for pre-training. Hopper GPUs (H100, H200) handle distillation for most production model sizes without requiring Blackwell. For models in the 2B to 4B parameter range, a single H100 node is sufficient. For 14B models, multi-node setups with FSDP2 sharding are necessary but available through standard cloud providers.

What Real-Time Actually Unlocks

Here is where the engineering story becomes a product story.

Real-time video diffusion means generating frames faster than they are displayed, at sustainable compute costs. At GTC 2026, Nvidia demonstrated this on a single B200 at production resolution. Hybrid Forcing, the approach published in April 2026 combining linear temporal attention with block-sparse attention and decoupled distillation, achieved 29.5 FPS at 832×480 on a single H100 without quantization or model compression, demonstrating that the distillation alone can cross the real-time threshold for certain architectures.

The use cases that cross from research demo to engineering problem once real-time is achieved:

World models for robotics. Nvidia’s Cosmos platform, which passed 2 million downloads by January 2026, generates physics-aware synthetic video for robot training. The bottleneck for sim-to-real transfer has been the speed of synthetic data generation. Real-time diffusion means generating training environments at the speed of the training loop itself, rather than pre-generating datasets as a separate pipeline stage. This changes the economics of robotic policy training fundamentally.

Interactive gaming environments. Google DeepMind’s Genie 3, the first real-time interactive world model generating persistent 3D environments at 24 fps, demonstrated what becomes possible. The gap between Genie 3’s architecture and a production-deployable version of the same idea is precisely the optimization stack Ilan described: distilled models, cached denoising, quantized weights. Microsoft Research’s Muse, built with Ninja Theory, represents the same direction applied to action-conditioned game generation.

Streaming personalized content. The bottleneck for personalized video generation at streaming scale has been per-user latency. If generating a 10-second clip takes 40 seconds of compute, you cannot serve it interactively. Real-time diffusion changes that constraint entirely.

Augmented reality overlays. The highest-demand use case for low-latency video generation. Generating consistent, physics-aware video overlays at 30+ FPS on consumer hardware requires exactly the combination of distillation, caching, and quantization that the stack delivers.

The Practical Deployment Order

Ilan was explicit on this in his talk, and it is worth preserving exactly: the order in which you implement these techniques matters for how much friction you encounter.

Start with quantization. It is the lowest-friction intervention. Pre-quantized checkpoints for Flux 2, LTX 2, and the Wan model family are available on HuggingFace today. For teams not doing fine-tuning, you can deploy a quantized model without any training infrastructure. The TRT-LLM visual gen repository has working examples. The gains are real. The risk is low.

Add caching next. Enable it as a flag in your serving library with a conservative threshold. Measure quality impact against your specific use case’s requirements. Tighten the threshold until quality begins to move, then back off one step. This is two days of work for a team that is already running quantized inference.

Distillation last. It is the most impactful technique and the most operationally demanding. If quantization and caching get you to an acceptable performance level, stay there. If your use case requires real-time or near-real-time generation and the first two techniques are not enough, distillation is the path. Plan for data curation, multi-GPU training infrastructure, and iteration. FastGen provides the scaffolding. Your domain-specific data and quality evaluation framework are the inputs only you can provide.

The stack is additive. Each layer compounds the gains of the previous one. A team that applies all three to a 14B video diffusion model can realistically expect the output that would otherwise require a rack of GPUs to run on a single Blackwell node.

That is not a research projection. That is what Nvidia demonstrated at GTC with real hardware, real models, and a public benchmark.

The question is no longer whether real-time video diffusion is achievable. The question is how quickly your team’s production stack gets there.

Published via Towards AI

Towards AI Academy

We Build Enterprise-Grade AI. We'll Teach You to Master It Too.

15 engineers. 100,000+ students. Towards AI Academy teaches what actually survives production.

Start free — no commitment:

6-Day Agentic AI Engineering Email Guide — one practical lesson per day

Agents Architecture Cheatsheet — 3 years of architecture decisions in 6 pages

Our courses:

AI Engineering Certification — 90+ lessons from project selection to deployed product. The most comprehensive practical LLM course out there.

Agent Engineering Course — Hands on with production agent architectures, memory, routing, and eval frameworks — built from real enterprise engagements.

AI for Work — Understand, evaluate, and apply AI for complex work tasks.

Note: Article content contains the views of the contributing authors and not Towards AI.