惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

U
Unit 42
Vercel News
Vercel News
博客园 - 叶小钗
大猫的无限游戏
大猫的无限游戏
MyScale Blog
MyScale Blog
P
Proofpoint News Feed
量子位
Engineering at Meta
Engineering at Meta
B
Blog RSS Feed
博客园 - 【当耐特】
Recent Announcements
Recent Announcements
Google DeepMind News
Google DeepMind News
D
DataBreaches.Net
Stack Overflow Blog
Stack Overflow Blog
博客园 - 聂微东
小众软件
小众软件
Hugging Face - Blog
Hugging Face - Blog
人人都是产品经理
人人都是产品经理
IT之家
IT之家
T
The Blog of Author Tim Ferriss
Last Week in AI
Last Week in AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Jina AI
Jina AI
博客园 - 三生石上(FineUI控件)

Interesting Engineering

US firm to scale laser-based nuclear fusion ‘breakthrough’ with new partnership Military Archives - Interesting Engineering World’s first non-nuclear lead-cooled reactor to generate electricity begins installation US scientists devise new process to turn sewage sludge into 99% pure natural gas US firm unveils submarine-hunting drone with 9,200-mile-range, 35 mph top speed Military Archives - Interesting Engineering Supercomputer finds lithium-titanium tweak to boost sodium-ion batteries for grids Lockheed Martin demonstrates vertical launch missile system for mobile drone defense China’s 1116 MWe Taipingling Unit 1 reactor goes online, set to generate 9bn kWh yearly ChatGPT Images 2.0 update combines reasoning, research, and design with 2K output US Navy tests plug-and-play laser system on USS Bush carrier, downs drones at sea China’s CATL reveals 621-mile EV battery, under-7-minute charging to challenge BYD US uses world’s first exascale supercomputer to model supernovae, fusion reactors AI and Robotics Archives - Interesting Engineering First-in-human study confirms safety of graphene-based brain interface Tesla’s Optimus humanoid robot greets runners, poses for photos at Boston Marathon Interlocking materials offer high strength and flexibility for robotics, infrastructure US redeploys 100,000-ton nuclear-powered aircraft carrier in Red Sea after repairs US scientists unveil concept for ‘world’s first neutrino laser’ to unlock breakthroughs New military tech can maintain communication in contested electronic warfare environments Got a dark personality? Psychologists can help you choose your career wisely Humidity boosts performance of 3D-printed nanogenerator instead of degrading it China demonstrates microwave beam that recharges drones in flight, continues power delivery Scientists run compact free-electron laser for eight hours, cracks FEL stability problem China’s PLA considers to use minelaying underwater drones to enforce Taiwan blockade: Report 1-ton sharks may struggle for survival in waters exceeding 62.6°F, study suggests US firm’s thorium nuclear fuel bundles move to manufacturing for commercial reactors Tesla hits 0% charge in remote Chilean desert as YouTuber uses hood-mounted solar Humanoid robot surpasses human world record in Beijing half-marathon, clocking 50:26 mins New method extracts maximum work from unknown quantum states using symmetry tricks
OpenAI launches next-gen voice AI models built for realti...
Aamir Kholla · 2026-05-08 · via Interesting Engineering

OpenAI has introduced three new audio models through its API, expanding its push into real-time voice AI for developers. The launch includes GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper, each targeting a different part of live voice interaction.

The company said the new models aim to make voice software more useful in everyday situations. That includes handling conversations while driving, navigating airports, or getting customer support without typing.

OpenAI framed the launch around a broader shift in computing interfaces. “Voice is becoming one of the most natural ways for people to use software,” the company said.

Smarter voice interactions

GPT-Realtime-2 serves as the flagship model in the release. OpenAI described it as its first voice model with GPT-5-class reasoning capabilities. The system can process harder requests, manage interruptions, and continue conversations naturally.

The model also supports live tool usage. Developers can let the AI access calendars, search systems, or other tools while speaking with users. OpenAI said the model can explain those actions in real time using phrases like “checking your calendar” or “looking that up now.”

OpenAI also expanded the model’s context window from 32K to 128K. That allows longer conversations and more complex tasks without losing context.

The company said GPT-Realtime-2 can recover more smoothly when something fails. It also better understands industry-specific terminology, including healthcare vocabulary and proper nouns.

OpenAI shared benchmark improvements tied to live voice performance. GPT-Realtime-2 (high) scored 15.2% higher on Big Bench Audio than GPT-Realtime-1.5. The xhigh version improved instruction-following scores by 13.8% on Audio MultiChallenge tests.

OpenAI’s new audio models place the company in close competition with Google’s Gemini Live. However, the latter still stands out for fast responses and stronger language support. OpenAI’s approach seems to be focusing more on making conversations feel natural during longer interactions. The new models can handle interruptions, use tools during calls, and “keep pace with the speaker,” according to the company.

Live translation features

OpenAI also launched GPT-Realtime-Translate, a real-time translation model designed for multilingual conversations.

The model translates speech from more than 70 input languages into 13 output languages while keeping pace with the speaker. OpenAI positioned the model for customer support, travel, and cross-language communication systems.

The company pointed to examples already in development. Deutsche Telekom is building voice support tools that let customers speak in their preferred language while the AI translates conversations live.

Speech-to-Text expansion

The third release, GPT-Realtime-Whisper, focuses on live transcription. The model converts speech into text while a person speaks, supporting streaming speech-to-text use cases.

OpenAI said the broader goal involves moving beyond simple voice assistants toward systems that can actively complete tasks during conversations.

For example, Zillow is developing a voice assistant that can search for homes, filter preferences, and schedule tours from spoken requests alone.

OpenAI said these models push real-time audio systems closer to agents that can “listen, reason, translate, transcribe, and take action as a conversation unfolds.”

The Blueprint

Get the latest in engineering, tech, space & science - delivered daily to your inbox.

Aamir is a seasoned tech journalist with experience at Exhibit Magazine, Republic World, and PR Newswire. With a deep love for all things tech and science, he has spent years decoding the latest innovations and exploring how they shape industries, lifestyles, and the future of humanity.