惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

I
InfoQ
H
Heimdal Security Blog
罗磊的独立博客
B
Blog RSS Feed
WordPress大学
WordPress大学
The Register - Security
The Register - Security
N
Netflix TechBlog - Medium
美团技术团队
量子位
GbyAI
GbyAI
Recent Announcements
Recent Announcements
博客园 - 叶小钗
D
DataBreaches.Net
S
SegmentFault 最新的问题
Hacker News - Newest:
Hacker News - Newest: "LLM"
T
Troy Hunt's Blog
The Last Watchdog
The Last Watchdog
O
OpenAI News
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
酷 壳 – CoolShell
酷 壳 – CoolShell
Webroot Blog
Webroot Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Last Week in AI
Last Week in AI
V
V2EX
N
News and Events Feed by Topic
Jina AI
Jina AI
Y
Y Combinator Blog
T
The Blog of Author Tim Ferriss
IT之家
IT之家
C
Check Point Blog
H
Hacker News: Front Page
爱范儿
爱范儿
Schneier on Security
Schneier on Security
Apple Machine Learning Research
Apple Machine Learning Research
P
Privacy & Cybersecurity Law Blog
L
LINUX DO - 最新话题
Forbes - Security
Forbes - Security
人人都是产品经理
人人都是产品经理
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Microsoft Azure Blog
Microsoft Azure Blog
C
Cyber Attacks, Cyber Crime and Cyber Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
S
Secure Thoughts
The Cloudflare Blog
Simon Willison's Weblog
Simon Willison's Weblog
Stack Overflow Blog
Stack Overflow Blog
腾讯CDC
MongoDB | Blog
MongoDB | Blog
V2EX - 技术
V2EX - 技术
AI
AI

The Decoder

The AI industry's platform trap is starting to look a lot like Microsoft's OpenAI buys Ona to push Codex toward long-running, autonomous coding tasks Jeff Bezos' AI startup Prometheus closes $12 billion round at a $41 billion valuation Free Deezer tool lets users on any streaming service check their playlists for AI music OpenAI vs. Anthropic: A price war over API tokens is brewing Dario Amodei's new essay reads like a Cold War playbook for the AI age Claude Fable 5: Anthropic admits "wrong tradeoff" after invisibly throttling rival AI researchers Google's new open model DiffusionGemma generates text from noise instead of word by word OpenAI's IPO slips as Altman tells staff to expect a public offering "within the next year" Anthropic study shows AI needs hours, not weeks, to build exploits from security patches OpenAI wants its biggest data center yet, and Nvidia would back the bill Claude Fable 5: The first Mythos model is powerful, expensive, and heavily filtered Germany's National Security Council greenights an AI Safety Institute modeled after the UK's AISI Google's NotebookLM now runs its own cloud computer with code execution and agent-based research Anthropic releases Claude Fable 5 and Mythos 5 with major gains in coding and science Google's Gemini 3.5 Live Translate delivers real-time voice translation across 70+ languages SpaceX wants to put data centers in orbit, and Musk says it's no big deal Landmark German ruling declares Google's AI Overviews are Google's own words and makes it liable for false answers Beijing's $295 billion AI buildout would require 80 percent domestic chips, locking out US suppliers Apple Intelligence gets a second shot with help from Google and Nvidia OpenAI now says "entirely automating everything is not the future we want" OpenAI says going public is "a complicated set of tradeoffs" and is unsure about the timing Microsoft Research's Lens proves detailed captions matter more than raw scale for training efficient image generators Intel gets a second life as Google and Nvidia explore it as a TSMC backup for AI chips Most companies are flying blind on AI spending Frontier Radar #3: How agentic AI is turning tokens into a business metric Instagram AI chatbot breach may have affected over to 20,000 accounts, Meta discloses Microsoft tightens rules for conflict zones after investigation into Israel's military use of Azure Moonshot AI targets a $30 billion valuation, more than six times its late-2025 worth Deepseek topped Ramp's trending software vendors in June 2026 as US companies chase cheaper AI OpenAI says "chat is dead" and plans to rebuild ChatGPT as a full-blown agent app Perplexity's "Search as Code" lets AI models write their own search pipelines instead of calling fixed APIs ChatGPT's new Lockdown Mode lets you disable web access and more to protect sensitive data from prompt injection Anthropic poaches OpenAI's second-ever chip engineer as both companies race toward IPOs Researchers pinpoint why larger language models pick up skills that small ones miss Sakana AI bets AI that improves itself can break the compute arms race of frontier labs Meta's Hatch AI agent could cost up to $200 a month and marks its first paid AI product Elon Musk's xAI reportedly trained its coding models on Claude outputs for months before getting cut off New open-source voice model listens nonstop and decides every 0.4 seconds whether to speak or stay silent SpaceX signs $920 million per month deal with Google for 110,000 Nvidia AI chips ahead of IPO OpenAI and the Trump administration are negotiating a government stake in the AI startup Qwen3.7-Plus is Alibaba's bid to turn multimodal AI into a full-blown autonomous agent Florida's lawsuit against OpenAI and CEO Altman treats ChatGPT as a defective product and public nuisance Satya Nadella publicly torches a VP's plan to make Microsoft's AI agent deliberately addictive Microsoft trained its MAI models on unlicensed web data despite promising "enterprise grade, clean and commercially licensed data" Anthropic's Mythos model is reportedly powering NSA offensive cyber ops against China and Iran Anthropic says Claude now writes over 90% of its code and wants the world to have an AI pause button Cloudflare CEO says the web's future is "pay to crawl" as bots overtake human traffic ChatGPT now saves narrative dossiers about you sorted by work, hobbies, and travel preferences Bain study finds companies miss AI savings targets because humans keep getting in the way OpenAI CEO Sam Altman sees "proactive AI" as the next big phase after chatbots and agents AI can now coach amateur virologists, and top tech leaders want Congress to act on DNA security xAI updates Grok Imagine to 1.5 with image-to-video generation at 720p resolution Google Deepmind's Gemma 4 12B squeezes multimodal AI onto a laptop with just 16 GB of RAM Google lets sites opt out of AI search results, knowing most have nowhere else to go Ideogram 4.0 drops as an open-weight model with native 2K resolution and improved text rendering Trump's new executive order wants AI companies to voluntarily submit models for government safety reviews Perplexity announces hybrid AI system that decides what runs locally or in the cloud AI music startup Suno doubles its valuation to $5.4 billion while fighting major record labels in court Nous Research releases Hermes Desktop, an open-source AI agent for every platform Build 2026: Microsoft tops Google in image generation while playing catch-up on reasoning OpenAI expands Codex with role-specific plugins to build a general-purpose app for non-developers Anthropic scales Project Glasswing to 150 partners across 15 countries to hunt critical software flaws Hackers hijacked high-profile Instagram accounts by simply asking Meta's AI chatbot to change the email OpenAI turns ChatGPT into a career platform with job search and CV editor Warren Buffett's Berkshire Hathaway bets $10 billion on Alphabet's AI infrastructure buildout OpenAI models now available on Amazon Web Services Claude maker Anthropic files for IPO with the SEC Turing Award winner Richard Sutton says pure generative AI can't do real science MiniMax M3: Open-weight model with a million-token context challenges proprietary leaders Nvidia's Nemotron 3 Ultra becomes the smartest open US model, but China still leads Nvidia bets big on physical AI at GTC Taipei with a new world model, driving brain, and open humanoid robot Nvidia pitches RTX Spark as the chip that finally makes local AI agents practical on Windows devices OpenAI starts with infrastructure robots but aims for "everyone having a personal robot doing anything they need" Ask AI what goes with chicken and the answer depends on whether it learned from recipes or molecules Anthropic bans AI tools during job interviews to see how candidates actually think Anthropic study finds men use AI coding agents more than twice as often as women in social science research SoftBank plans 75 billion euro AI data center buildout in France AI search agents often confirm what they already know instead of actually researching the web Microsoft and Nvidia reportedly team up on AI PCs that run actual agents instead of Copilot Making AI chatbots helpful weakens their ability to simulate human behavior, large-scale study finds Terence Tao argues AI could bring division of labor to math for the first time in history Attackers abuse shared ChatGPT and Claude chats to spread malware OpenAI's Codex can now operate your Windows PC autonomously, hunting bugs and testing apps on its own Salesforce claims AI agents cut a 231-day migration to 13 days with fewer incidents Meta's leaked memo reveals AI pendant, supersensing glasses, and enterprise wearables strategy OpenAI gives GPT-5.5 Instant a readability upgrade while phasing out two older models Google fixes several bugs in Gemini usage limits that burned through quotas too fast One company reportedly spent $500 million on Claude in one month after failing to cap AI usage OpenAI is giving away its life sciences AI model to help governments prepare for the next pandemic New review paper argues code is how AI agents think and act, not just what they produce Amazon kills internal AI leaderboard after employees gamed it with pointless tasks Claude company Anthropic nears a trillion-dollar valuation after raising $65 billion in Series H Anthropic ships Claude Opus 4.8 as a "modest but tangible improvement" that tops GPT-5.5 in most benchmarks Google Cloud responds to AI-accelerated cyberattacks with a platform that aims to close security gaps in minutes Google launches a tiny board that runs Gemma 3 locally Mistral rebrands LeChat as Vibe, betting its chatbot's future is as a full-blown work agent Meta One: Zuckerberg finally puts a price tag on all that AI spending Amazon builds its own AI production platform and greenlights three AI animated series for Prime Video ElevenLabs Music v2 promises opera-to-metal transitions without losing musical coherence
Microsoft Research's Mirage gives video generation a persistent spatial memory that doesn't forget what's around the corner
Jonathan Kemper · 2026-06-14 · via The Decoder

Mirage is a new video world model that skips the costly detour through pixel-based memory. That speeds up generation and keeps a scene's spatial structure stable even during long camera moves. Researchers from several universities built it with Microsoft Research.

Video world models turn a starting frame and a camera path into plausible moving images, handy for simulations or as world simulators. But without some kind of memory, even strong generators lose track of space over time. A corner of a room you've already passed looks different when the camera swings back. Furniture shifts, and textures change.

Systems like Voyager, WonderWorld, and Spatia try to fix this with a 3D point cloud that gets fed a steady stream of color data. Every new generation step has to render that cloud and then translate the result back into the model's internal feature space. Microsoft's new paper calls this a double bottleneck: It eats compute, and information leaks out every time the data passes through pixel space.

Mirage takes a different approach. Rather than holding onto visible color points, it stores the internal image features the diffusion model already uses. Each feature gets a spot in 3D space, which turns it into an entry in spatial memory.

Comparison diagram of two video world model pipelines. Top: an RGB point cloud memory with a render-and-encode loop. Bottom: Mirage's latent spatial memory, built and read directly in latent space.
Two video world model pipelines side by side. Top: an RGB point cloud memory with a render-and-encode loop. Bottom: Mirage's latent spatial memory, built and read directly in latent space. | Image: Wang et al.

To generate a new viewpoint, the model projects this store straight onto the target camera and hands the result to the generator, skipping the step of rendering a point cloud and re-encoding it. The authors say this also slashes memory use, since the data sits in the model's compact internal resolution instead of at full image size.

How the memory grows with each step

Mirage builds videos in segments, seeding the spatial memory from the starting image. For every later segment, the system pulls the relevant data from memory, generates the new frames, then writes their contents back to the cache. The memory keeps growing as it goes.

Mirage pipeline in which a VAE plus depth estimation builds the latent cache from the first frame. Each generation chunk reads from it via readout and updates it via write, while the latent 3D representation grows over time from t0 to tN.
Mirage seeds the latent cache from the starting image, then reads from it and writes to it chunk by chunk, keeping static scene content intact across the whole run. | Image: Wang et al.

A filter keeps the system from tripping over itself by stripping out moving objects and the sky before writing, so only stable geometry lands in long-term memory. The researchers built on Alibaba's open-source video model Wan2.2, bolting on a small add-on module that teaches the model to use the new memory, then fine-tuning the whole thing with LoRA adapters.

Faster and lighter than color-based rivals

On the WorldScore benchmark, Mirage beats its closest rival Spatia, which still keeps memory as color points, and leaves general video generators like Wan2.1 and CogVideoX far behind. It shines at holding a scene's spatial structure together and keeping surfaces looking consistent across many frames.

It also leads two of three metrics on the RealEstate10K dataset in the closed-loop test. Here the camera circles back to its starting point, a brutal stress test because every tiny error piles up over the full path.

Two bar charts across five generation chunks. Left: average generation time per frame. Right: peak cache VRAM. Mirage stays consistently low on both metrics, while Spatia, VMem, and Gen3C climb sharply.
Mirage holds compute time and memory nearly flat across the whole run, while rival models get hungrier with every chunk. | Image: Wang et al.

Efficiency is Mirage's strongest point. Color-based memory scales badly on longer runs and keeps demanding more graphics memory. Mirage's compute cost per frame barely moves after the first segment. The researchers put the total gain at up to 10.57x faster generation and up to 55x less memory than color-based systems.

They're upfront about one catch. Moving objects get dropped at segment boundaries because their geometry can't be trusted, and the filter deliberately tosses them out. Busy scenes gain less from spatial memory than quiet interiors do. The team points to storing dynamic content as the obvious next problem to solve.

You can find more on Mirage on the project page. Microsoft also runs a GitHub repository for Latent Spatial Memory.

Video world models are one of the hottest research areas in AI video right now. Models like Veo mostly produce single, internally consistent clips, while world models try to make a scene navigable and keep it consistent over time. Google Deepmind showed this off recently with Genie 3, which spins up interactive environments in real time and holds them for several minutes. At I/O, Google also pitched Gemini Omni as a world model and the potential successor to its text-to-video model Veo.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

Subscribe now