惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The GitHub Blog
The GitHub Blog
I
InfoQ
U
Unit 42
WordPress大学
WordPress大学
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Apple Machine Learning Research
Apple Machine Learning Research
J
Java Code Geeks
月光博客
月光博客
D
Docker
Stack Overflow Blog
Stack Overflow Blog
D
DataBreaches.Net
阮一峰的网络日志
阮一峰的网络日志
Blog — PlanetScale
Blog — PlanetScale
V
Visual Studio Blog
博客园 - 聂微东
A
About on SuperTechFans
腾讯CDC
Jina AI
Jina AI
Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
博客园 - 【当耐特】
罗磊的独立博客
博客园 - 三生石上(FineUI控件)
M
MIT News - Artificial intelligence

Gemini Models

Proactive cyber defense for governments and enterprises Introducing Gemini 3.8 Flash and 3.8 Flash Cyber The latest AI news we announced in August 2026 Introducing agentic video understanding with Gemini Gemini Omni 1.1 Flash lets you build with more control Intelligent transcription with Gemini 3.5 Transcribe What does “full-stack” AI actually mean? Introducing Gemini 3.7 Flash Omni experts share what excites them most about the model. See what 5 builders are making with Gemini Omni The latest AI news we announced in July 2026 Inside our 353,000-person vibe coding course Simplify your morning with this vibe-coded schedule app. Introducing Gemini Robotics ER 2 How Gemini Flash agents are helping a Michigan dairy farmer Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber The latest AI news we announced in June 2026 Start building with Nano Banana 2 Lite and Gemini Omni Flash Introducing computer use in Gemini 3.5 Flash Fluid, natural voice translation with Gemini 3.5 Live Translate The latest AI news we announced in May 2026 How we used Gemini to build Google I/O 2026 9 demos of Gemini Omni and Gemini 3.5 in action Catch up on 12 major I/O 2026 moments I/O 2026 I/O 2026: Welcome to the agentic Gemini era Gemini 3.5: frontier intelligence with action Introducing Gemini Omni The latest AI news we announced in April 2026 Join the new AI Agents Vibe Coding Course from Google and Kaggle
Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking
Tom Ouyang · 2026-09-16 · via Gemini Models

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. Major upgrades in intelligence and parallel reasoning make them more intuitive to collaborate with and use to execute complex tasks using your voice.


Malini Jaganathan

Member of Technical Staff, on behalf of the Gemini Audio Team


Text "Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking" with the Gemini Spark, all on a light blue background

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Today, we’re introducing two new models that bring advancements in near real-time reasoning to more effectively enable voice agents and make conversing with AI feel more intuitive and intelligent.

  • Gemini 3.8 Live: Built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding.
  • Gemini 3.8 Live Extended Thinking: Built for high-complexity tasks, with increased intelligence and multi-step reasoning.

For developers and enterprises, these models deliver the building blocks for reliable, production-ready voice agents. They also make speaking with Gemini across the Gemini app, Google Workspace, and Search more fluid and collaborative — helping you tackle complex tasks using just your voice.

Experience more fluid, intelligent conversations

Gemini 3.8 Live Extended Thinking provides enterprise-grade task completion and intelligence, capturing the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index (82.6), and leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. It also provides strong reasoning capabilities, scoring 97.7% on Big Bench Audio, while maintaining a highly competitive price point compared to other frontier models.

Gemini 3.8 Live has shown a high preference among users, securing a second place in the Speech Agent Arena. In addition to this performance, it remains highly cost-effective — providing developers and enterprises with a capable and efficient model built for scale.

a chart showing Artificial Analysis Speech to Speech Index

A chart showing Artificial Analysis agentic performance

A chart showing Sierra

A chart showing Artificial Analysis cost per hour of input audio

On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, our models push the Pareto Frontier for complex workflows by successfully balancing accuracy with conversational quality.

Note: This was run on the Live API on Gemini Enterprise Agent Platform.

A chart showing EVA Bench Experience to Task Completion

Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It automatically detects and transitions between 97 supported languages mid-conversation. It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge requests and keep chatting while tasks finish in the background.

For tasks that require deeper reasoning, 3.8 Live Extended Thinking reasons and speaks simultaneously. It delivers increased intelligence for complex workflows while maintaining an uninterrupted conversational flow — using early verbal cues like “Let me check that…” to acknowledge prompts naturally, and live progress narration to walk users through multi-step background tasks as they progress.

Across Google Workspace and Search, our Live models deliver more intuitive, collaborative experiences — especially when tackling your most complex tasks.

Get step-by-step, real-time troubleshooting help powered by Gemini 3.8 Live — right inside Search Live.

Empowering the developer and enterprise voice ecosystem

By using the Gemini Live API, developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.

We’re also partnering with companies like Salesforce, Genspark, and Lumeris who are excited about 3.8 Live and 3.8 Live Extended Thinking, highlighting its impressive latency, fluidity, and tool-calling capabilities.

Salesforce quote

11Sight Quote

Equal AI Quote

ServiceNow quote

Genspark quote

Lenskart quote

Lumeris quote

Agora quote

Ambr AI quote

LiveKit Quote

Casuu quote

Ensure transparency with SynthID watermarking

All audio generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated content remains detectable to help prevent misinformation. For details on our approach to safety and responsibility, review the model card.

Start using our latest Gemini Audio models:

3.8 Live is rolling out starting today:

3.8 Live Extended Thinking is rolling out starting today:

Get the latest news from Google in your inbox

Sign up for our newsletters with product updates, event information, special offers, and more.

Your information will be used in accordance with Google's privacy policy. You may opt out at any time.