惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Stack Overflow Blog
Stack Overflow Blog
S
SegmentFault 最新的问题
大猫的无限游戏
大猫的无限游戏
The GitHub Blog
The GitHub Blog
M
MIT News - Artificial intelligence
T
Tailwind CSS Blog
aimingoo的专栏
aimingoo的专栏
Last Week in AI
Last Week in AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
罗磊的独立博客
H
Help Net Security
Engineering at Meta
Engineering at Meta
Microsoft Security Blog
Microsoft Security Blog
阮一峰的网络日志
阮一峰的网络日志
J
Java Code Geeks
T
The Blog of Author Tim Ferriss
Hugging Face - Blog
Hugging Face - Blog
C
Check Point Blog
F
Fortinet All Blogs
腾讯CDC
博客园 - Franky
WordPress大学
WordPress大学
U
Unit 42

Interesting Engineering

US firm to scale laser-based nuclear fusion ‘breakthrough’ with new partnership Military Archives - Interesting Engineering World’s first non-nuclear lead-cooled reactor to generate electricity begins installation US scientists devise new process to turn sewage sludge into 99% pure natural gas US firm unveils submarine-hunting drone with 9,200-mile-range, 35 mph top speed Military Archives - Interesting Engineering Supercomputer finds lithium-titanium tweak to boost sodium-ion batteries for grids Lockheed Martin demonstrates vertical launch missile system for mobile drone defense China’s 1116 MWe Taipingling Unit 1 reactor goes online, set to generate 9bn kWh yearly ChatGPT Images 2.0 update combines reasoning, research, and design with 2K output US Navy tests plug-and-play laser system on USS Bush carrier, downs drones at sea China’s CATL reveals 621-mile EV battery, under-7-minute charging to challenge BYD US uses world’s first exascale supercomputer to model supernovae, fusion reactors AI and Robotics Archives - Interesting Engineering First-in-human study confirms safety of graphene-based brain interface Tesla’s Optimus humanoid robot greets runners, poses for photos at Boston Marathon Interlocking materials offer high strength and flexibility for robotics, infrastructure US redeploys 100,000-ton nuclear-powered aircraft carrier in Red Sea after repairs US scientists unveil concept for ‘world’s first neutrino laser’ to unlock breakthroughs New military tech can maintain communication in contested electronic warfare environments Got a dark personality? Psychologists can help you choose your career wisely Humidity boosts performance of 3D-printed nanogenerator instead of degrading it China demonstrates microwave beam that recharges drones in flight, continues power delivery Scientists run compact free-electron laser for eight hours, cracks FEL stability problem China’s PLA considers to use minelaying underwater drones to enforce Taiwan blockade: Report 1-ton sharks may struggle for survival in waters exceeding 62.6°F, study suggests US firm’s thorium nuclear fuel bundles move to manufacturing for commercial reactors Tesla hits 0% charge in remote Chilean desert as YouTuber uses hood-mounted solar Humanoid robot surpasses human world record in Beijing half-marathon, clocking 50:26 mins New method extracts maximum work from unknown quantum states using symmetry tricks
Google rolls out Gemini Omni Flash for autonomous video c...
Neetika Walt · 2026-05-23 · via Interesting Engineering

Google has started rolling out Gemini Omni Flash, its new multimodal AI model that can generate and edit videos using text, images, audio and video inputs. The rollout follows the model’s announcement during Google I/O 2026 and marks the point where users can now actively use the system inside the Gemini app, Google Flow and YouTube Shorts.

The company says the model is designed to combine reasoning and creative generation in a single system, allowing users to build and modify video content through natural conversation.

With Gemini Omni Flash, users can prompt the model to create videos from scratch or modify existing clips step by step. Each instruction builds on the previous one, allowing continuous refinement of scenes without breaking continuity. Google says this helps maintain consistency in characters, objects and environments across edits, even as the video changes through multiple iterations.

The model also supports multi-input workflows, where users can combine different types of inputs such as text prompts, images, video clips and audio references. This allows a single output video to be shaped using multiple reference points instead of relying on a single prompt. Google says the system is built to understand how these inputs relate to each other and produce a coherent final scene.

The rollout is part of Google’s broader push to integrate generative AI into its consumer ecosystem, especially platforms focused on short-form video creation. YouTube Shorts and the YouTube Create app are among the first platforms where Omni Flash capabilities are being introduced, signalling a tighter connection between AI generation tools and content creation pipelines.

The company also says all outputs generated through the system will include SynthID watermarking for identification of AI-generated content.

Conversational video editing

Gemini Omni Flash allows users to edit videos using natural language commands instead of traditional editing tools. Users can describe changes such as altering environments, adding objects or changing actions within a scene, and the model updates the video accordingly while preserving overall structure.

The system is designed to maintain visual continuity across edits, ensuring that characters and objects remain consistent as changes are made over multiple steps. Google says this makes the editing process more iterative and flexible compared to conventional video production tools.

The model also draws on Gemini’s broader world knowledge to improve realism in generated content. It uses this understanding to simulate physical interactions such as motion, lighting and environmental effects more accurately, according to Google.

From prompts to production

Google has positioned Gemini Omni Flash as part of a wider shift toward multimodal AI systems that can handle creation and reasoning together. The model is designed to process multiple input formats and generate output video that reflects combined instructions rather than isolated prompts.

The company says the goal is to reduce the gap between idea and execution, allowing users to move from concept to finished video using a single conversational interface. Over time, Google plans to expand output formats beyond video, with support for images and audio also planned for future updates.

The rollout of Gemini Omni Flash is currently limited to select subscription tiers in the Gemini app, with broader access expected as the deployment expands.

The Blueprint

Get the latest in engineering, tech, space & science - delivered daily to your inbox.

With over a decade-long career in journalism, Neetika Walter has worked with The Economic Times, ANI, and Hindustan Times, covering politics, business, technology, and the clean energy sector. Passionate about contemporary culture, books, poetry, and storytelling, she brings depth and insight to her writing. When she isn’t chasing stories, she’s likely lost in a book or enjoying the company of her dogs.