惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
雷峰网
雷峰网
Hugging Face - Blog
Hugging Face - Blog
IT之家
IT之家
H
Help Net Security
腾讯CDC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
The GitHub Blog
The GitHub Blog
V
V2EX
M
MIT News - Artificial intelligence
Vercel News
Vercel News
WordPress大学
WordPress大学
博客园 - 三生石上(FineUI控件)
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
阮一峰的网络日志
阮一峰的网络日志
B
Blog RSS Feed
D
Docker
V
Visual Studio Blog
博客园 - 叶小钗
美团技术团队
S
SegmentFault 最新的问题
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

CNET

Valve's Steam Machine: Summer Release Planned, Still No Price Apple TV: 28 of the Best Shows You're Probably Not Watching YouTube TV vs. DirecTV vs. Hulu Live and More: Which Has the Most Must-Have Channels Out of 100? If You Want to Be a Better Pet Parent, AI Can Help I Was Shocked by How Good These Budget TVs Were Trump Phone Looks Different, Has No Launch Date, Isn't Made in America The Apple Watch Series 12 Is Rumored to Revive a Retired iPhone Feature Best Projector of 2026: Tested by Experts Best Home Theater Systems of 2026 How to Use Apple's Clean Up Tool to Remove Unwanted People and Things From Your Photos Today's NYT Strands Hints, Answers and Help for April 12 #770 Today's NYT Connections Hints, Answers and Help for April 12, #1036 Today's Wordle Hints, Answer and Help for April 12, #1758 Today's NYT Mini Crossword Answers for Sunday, April 12 Today's NYT Connections: Sports Edition Hints and Answers for April 12, #566 Watch a Robot Stuff Cash Into a Wallet Just Like You Do This Animation Startup Wants to Make It Easier to Tell Open-Ended Stories The 23 Best Graduation Gifts for 2026 Grand National 2026 Livestream: How to Watch Aintree Horse Racing From Anywhere Amazon Luna to Drop Support for Third-Party Games and Subscriptions in June YouTube Premium Is the Latest Streaming Service to Hike Prices Today's NYT Mini Crossword Answers for Saturday, April 11 Elden Ring: Tarnished Edition for Switch 2 Reignites Controversy Over Game-Key Cards Comcast Adds New StreamSaver Bundles: HBO Max, Disney Plus, Hulu Now Part of the Lineup Samsung's Galaxy Z Fold 7 Just Got a Price Hike, 9 Months After Its Release Microsoft Is Scrubbing the Copilot Name From Some Windows 11 Apps These $299 Glasses Are Like an HDR TV on Your Face Today's NYT Connections: Sports Edition Hints and Answers for April 11, #565 How to Make Sure Your Private Signal Messages Aren't Still Lurking on Your Phone Apple AirPods Max 2 Review: Seemingly Small Changes Make a Substantial Difference
Google Introduces Gemini Omni, a Multimodal AI That Knows...
Blake Stimac · 2026-05-20 · via CNET

Google announced its latest AI product, Gemini Omni, during its I/O conference on Tuesday. Unlike existing text-to-video products such as Veo, Omni can take in virtually any input to create realistic, lifelike videos. 

Built on Gemini modeling architecture, Omni is a true multimodal input and output system, allowing you to create videos from text, images and existing videos. At launch, you'll be able to create videos with the aforementioned inputs, but image; text generations will be available in a future update. 

With Gemini at its core, Omni can process and interpret multiple types of inputs to produce a consistent, sophisticated final product. Omni builds on Google's existing products by integrating Gemini Intelligence.

The rise of AI-created videos comes at a paradoxical time as companies such as Google make incredible advances with the technology, while social media feeds become more filled with AI slop. Google considers Omni the "next big step" toward building AI that can model and simulate the real world. It's a world model with advanced reasoning, capable of generating videos grounded in the world we know today. Omni demonstrates advanced physics capabilities, enabling it to create realistic video outputs. Here's what's coming in Gemini Omni from Google I/O.

Powerful (and scary) editing capabilities

As with its powerful video generation, Omni also has advanced video editing capabilities. If you create a video with Omni, you can feed it back into the tool, make impressive changes with just a prompt or incorporate additional media. You can even upload your own videos and change or swap out individual elements, allowing for a new way to edit videos that has essentially never been available before. 

That ability to fully replace elements in a person's video could lead to some dark outcomes, making Omni's advanced editing abilities as alarming as they are impressive. But Google has built-in guardrails. First, any output from Omni will automatically include Google's SynthID watermark, so you know that what you're viewing has been altered in some way by AI. This is a big deal, as Omni essentially lets you change how reality is perceived. 

Multiple access points

People will be able to play with Gemini Omni in a variety of ways. It's a prominent feature within the newly redesigned Gemini app, where you can add built-in templates to your camera roll with a single click. Additionally, you'll be able to create a custom avatar that looks and sounds like you and add it to videos. 

For some paid subscribers, Omni will be available on Google Flow and YouTube Shorts, starting on Tuesday. Omni will roll out to developers and enterprise customers via APIs in the coming weeks, allowing for custom integrations. 

Omni Flash and Omni Pro

Like most Gemini models, Omni will be split into Flash and Pro versions, though the former will be available initially. Google is working on an even more powerful model, Omni Pro, which will become available in the future.