惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
IT之家
IT之家
博客园_首页
量子位
博客园 - 三生石上(FineUI控件)
小众软件
小众软件
博客园 - 聂微东
罗磊的独立博客
酷 壳 – CoolShell
酷 壳 – CoolShell
Hugging Face - Blog
Hugging Face - Blog
V
V2EX
爱范儿
爱范儿
大猫的无限游戏
大猫的无限游戏
宝玉的分享
宝玉的分享
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
雷峰网
雷峰网
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Google DeepMind News
Google DeepMind News
Microsoft Azure Blog
Microsoft Azure Blog
有赞技术团队
有赞技术团队
S
SegmentFault 最新的问题
Engineering at Meta
Engineering at Meta
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
Retrieval‑Augmented Memory Reduces Sliding‑Window Limitat...
Papers Mache · 2026-06-17 · via DEV Community

Papers Mache

VideoMLA’s low‑rank latent KV cache cuts KV‑cache demand by roughly 90 % and LongLive‑RAG’s retrieval‑augmented memory helps mitigate the temporal drift introduced by sliding‑window attention. The KV‑cache reduction comes from replacing per‑head keys and values with a shared low‑rank latent, shaving 92.7 % off per‑token cache size; separately, the retrieval module lets the generator attend to non‑local history instead of a stale recent window, helping prevent error accumulation across thousands of frames.

Before these advances, long‑horizon video diffusion relied on a fixed‑size sliding window that constantly overwrites the KV cache. As the window slides, any appearance error that slips in becomes permanent, and because the model can only look at the most recent tokens, identity drift compounds unchecked. Researchers tried shuffling token order or tweaking positional encodings, but the fundamental bottleneck—a growing KV cache that forces either truncation or out‑of‑memory failures—remained.

VideoMLA achieves a 92.7 % reduction in per‑token KV cache memory while preserving compatibility with standard chunk‑causal generation. The paper shows that “VideoMLA reduces per-token KV cache memory by 92.7 % while preserving compatibility with standard chunk‑causal generation” and Figure 4 confirms that subject identity, scene structure, and visual fidelity stay intact over 30‑second rollouts despite the compact latent cache [1].

LongLive‑RAG adds a lightweight retrieval step that draws from the entire self‑generated latent history, so the generator can condition on truly relevant frames instead of a narrow window. The authors note that “this lightweight retrieval step adds only a small overhead relative to generation and lets the generator condition on non‑local context instead of only the recent window,” and experiments across multiple AR backbones report “improved long‑video quality and the best average VBench‑Long rank” [2].

Both methods leave open questions about scaling to truly minute‑scale generation without additional latency. Retrieval introduces a query‑embedding cost and a memory‑search latency that can become noticeable as the history grows, a limitation the authors acknowledge but do not quantify. VideoMLA’s low‑rank cache has only been validated up to 30‑second rollouts; whether the same rank budget suffices for several‑minute sequences remains speculative.

If these techniques hold, the default architecture for autoregressive video diffusion should drop the sliding‑window KV cache in favor of a shared low‑rank latent cache plus a history‑retrieval module, enabling minute‑scale synthesis on a single high‑end GPU without exceeding memory limits.

References

  1. VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion
  2. LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation