惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 三生石上(FineUI控件)
Blog — PlanetScale
Blog — PlanetScale
B
Blog
GbyAI
GbyAI
爱范儿
爱范儿
月光博客
月光博客
N
Netflix TechBlog - Medium
T
Tailwind CSS Blog
G
Google Developers Blog
大猫的无限游戏
大猫的无限游戏
Vercel News
Vercel News
H
Hackread – Cybersecurity News, Data Breaches, AI and More
WordPress大学
WordPress大学
The GitHub Blog
The GitHub Blog
Recent Announcements
Recent Announcements
腾讯CDC
MyScale Blog
MyScale Blog
V
Visual Studio Blog
The Cloudflare Blog
Microsoft Security Blog
Microsoft Security Blog
A
About on SuperTechFans
Google DeepMind News
Google DeepMind News
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

Developer tools

DevFest is back The latest AI news we announced in August 2026 Pairing Google Antigravity with Gemini 3.7 Flash solves notable multi-agent math and engineering problems. Gemini Omni 1.1 Flash lets you build with more control How developers build AI for good with Gemma 4 Inside the Gemmaverse: Celebrating one billion Gemma downloads Inside our 353,000-person vibe coding course Introducing Gemini Robotics ER 2 Gemini API Managed Agents: 3.6 Flash, hooks, and more We're rolling out AlphaEvolve widely to solve Google Cloud customers' hardest problems. Expanding Managed Agents in Gemini API: background tasks, remote MCP and more The latest AI news we announced in June 2026 Ask an AI expert: What exactly is the full stack? Interactions API: our primary interface for Gemini models and agents DiffusionGemma: 4x faster text generation See what 3 builders are making with Gemma 4 Bringing the latest Gemini models to Apple developers Gemma 4 QAT models: Optimizing model compression for mobile and laptop efficiency Kaggle is making AI benchmark creation effortless Introducing Gemma 4 12B: a unified, encoder-free multimodal model How we used Gemini to build Google I/O 2026 Take our I/O 2026 quiz, vibe coded in Google AI Studio. Here's what developers can do with the latest Google Play updates. Building the agentic future: Developer highlights from I/O 2026 I/O 2026 Introducing Managed Agents in the Gemini API Bring any idea to life: Google AI Studio at I/O 2026 Gemini API File Search is now multimodal: build efficient, verifiable RAG Accelerating Gemma 4: faster inference with multi-token prediction drafters The latest AI news we announced in April 2026
Build real-time voice applications with Gemini 3.8 Live a...
Alisa Fortin · 2026-09-16 · via Developer tools

New Gemini Audio models are available for developers to build more intelligent conversational experiences via the Gemini API and Google AI Studio.


Thor Schaeff

Member of the Technical Staff (DevX), Google DeepMind


Try the models today in ai.studio/live

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Today, we released new Gemini Live models in the Gemini API and Google AI Studio, expanding our developer suite for building real-time, voice-first product experiences:

Gemini 3.8 Live and 3.8 Live Extended Thinking: Gemini 3.8 Live brings a step change to our native speech-to-speech models, capable of performing tasks while maintaining dialogue. For complex requests, 3.8 Live Extended Thinking delivers deeper reasoning, ranking #1 on Artificial Analysis’ Speech-to-Speech leaderboard.

Gemini 3.5 Transcribe: Our dedicated speech-to-text model brings highly precise transcription across 85+ languages. Released last month, it achieved an average Word Error Rate (WER) of 4.0% (streaming) and 2.6% (non-streaming).

Gemini 3.8 Live & 3.8 Live Extended Thinking: Build more intelligent conversational agents

Our new models, Gemini 3.8 Live and 3.8 Live Extended Thinking enable developers to build voice agents that can reason and execute tasks while maintaining the flow of conversations. Key capabilities include:

  • Asynchronous function calling: Execute API and tool calls in the background while continuing to stream audio responses to the user
  • Visual context: Ground dialogue in live visual inputs to help enable agents that can understand what users say and see
  • Alphanumeric precision: Accurately parse confirmation codes, claim numbers, and technical data
  • Multilingual support: Reach global audiences with coverage for 97+ languages and accent consistency
  • Incremental content updates: Seamlessly merge real-time audio with structured data to return context-aware responses

3.8 Live Extended Thinking also supports configurable thinking to help handle complex, multi-step reasoning in the background, while responding or narrating its progress in the main conversation. These models represent a step-change from our previous live models and provide a more streamlined alternative to cascaded architectures.

Jamie Wood, cofounder and Chief of Technology & Products, says," Ambr AI trains enterprise teams to negotiate, lead, and resolve difficult customer conversations through realistic AI simulations. Switching to the Gemini 3.8 Live Extended Thinking made our simulations faster, more expressive, and better at handling complex conversations. It’s also enabled us to deliver training in over 70 languages through a single integration, helping our customers develop these skills across their global workforce."

Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are available via the Live API. Competitively priced at $0.005/min for audio input and $0.018/min 1 for audio output, they allow developers to scale voice applications with industry-leading performance.

Benchmarks on Artificial Analysis Speech-to-Speech Index

Benchmark on Agentic Performance

Benchmark of Tau-bench leaderboard

Artificial Analysis Overall Cost of Audio

Developers can also access the models through Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel, and Vision Agents, our Live API integration partners that handle media streaming infrastructure for real-world deployment:

Live API partners include Pipecat, LiveKit, LangChain, Agora, Fishjam, Voximplant, Vercel, and VisionAgents.

Gemini 3.5 Transcribe: Convert streamed speech to text

Real-time speech understanding is critical for voice-first interfaces. Last month, we released Gemini 3.5 Transcribe for low-latency transcription with high precision, achieving a 4.0% WER, and useful features:

  • Automatic code-switching: Handle intra-sentence and inter-sentential code- and language-switching without manual configuration
  • Custom vocabulary biasing: Steer speech recognition toward domain-specific terms, uncommon jargon, company names, and proper nouns by passing a custom_vocabulary list of up to 1,000 terms
  • Smart transcription mode: Deliver polished, reader-ready transcripts with structured formatting, self-corrections, and disfluency removal that eliminates filler words

3.5 Transcribe supports 85+ languages and provides a strong listening engine for voice experiences and stateless tasks like sub-second captioning, call center agents, and real-time audio analytics. You can also access the model via the Interactions API to transcribe audio files up to 1 hour long with structured timestamps and speaker diarization. Read our developer guide to learn more.

Our complete audio suite for developers

To get started, try out the models in ai.studio/live, clone example apps from GitHub, or equip your agent with our live api skill.

You can also create audio experiences with our speech and music generation models, all available in the Gemini API:

The mic is yours, and we can’t wait to hear what you build!

Get the latest news from Google in your inbox

Sign up for our newsletters with product updates, event information, special offers, and more.

Your information will be used in accordance with Google's privacy policy. You may opt out at any time.