惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Docker
IT之家
IT之家
Microsoft Security Blog
Microsoft Security Blog
博客园 - 司徒正美
云风的 BLOG
云风的 BLOG
P
Proofpoint News Feed
D
DataBreaches.Net
B
Blog RSS Feed
博客园_首页
The GitHub Blog
The GitHub Blog
I
InfoQ
L
LangChain Blog
G
Google Developers Blog
M
MIT News - Artificial intelligence
美团技术团队
腾讯CDC
V
Visual Studio Blog
aimingoo的专栏
aimingoo的专栏
博客园 - 聂微东
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Apple Machine Learning Research
Apple Machine Learning Research
A
About on SuperTechFans
博客园 - 三生石上(FineUI控件)
博客园 - 叶小钗

TestingCatalog

SpaceXAI gearing up for upcoming Grok 4.5 release ByteDance set to launch Seedance 2.5 with 3-minute output Meta prepares Scheduled tasks for Meta AI users on web Google tests new Gemini Inbox section for Workspace triage OpenAI might be preparing GPT-5.6 for next week's release Mistral releases Leanstral 1.5 model for proof engineering xAI debuts Grok Voice Agent Builder for Enterprises Condense launches proxy to cut AI coding agent bills by 66% Vellum adds agent-to-agent AI collaboration for Slack Early look at Anthropic's Claude Science app for researchers Google launches Nano Banana 2 Lite and Gemini Omni Flash Google might be testing Gemini Flash upgrade on LM Arena Anthropic may impose KYC restrictions for Fable 5 access Anthropic launches Claude Sonnet 5 model on Claude and APIs NoimosAI launches Creative Agent for brand assets Apify lets AI Agents pay via Coinbase x402 for web tools Bloome launches chat platform for AI agent teams Meituan launches LongCat-2.0 1.6T parameter model Cursor releases its iOS app for vibe coding on the go OpenAI prepares upgraded Office controls for Codex OpenAI tests gifting Codex credits as new growth strategy Microsoft launches MAI-Code-1-Flash on GitHub Copilot Google adds Computer Use to Gemini 3.5 Flash Google tests notebook collections for NotebookLM OpenAI launches GPT-5.6 Sol preview for select partners Microsoft adds Copilot finance tools to Excel for M365 users DeepReinforce releases Ornith-1.0 open-source coding models Gemini to get voice dictation and Magic Pointer on desktop Meta launches AI glasses with three new styles from $299 Anthropic launches Claude Tag on Team and Enterprise plans
Microsoft released MAI voice and image models for Build 2026
https://www.facebook.com/testingcatalog · 2026-05-31 · via TestingCatalog

UPDATE: Microsoft has announced 7 new AI models during Microsoft Build 2026 - MAI Image 2.5, MAI Image 2.5 Flash, MAI Voice 2, MAI Voice 2 Flash, MAI Transcribe 1.5, MAI Code 1 Flash, and MAI Thinking 1.

Seven new models launching at Build: let’s go!
Reasoning. Code. Image. Transcribe. Voice.

Built from scratch on a clean data lineage, designed for efficiency, working seamlessly as a family of models

Thread 🧵 #MSBuild pic.twitter.com/g3WQIcIQ24

— Microsoft AI (@MicrosoftAI) June 2, 2026

The Story

Microsoft heads into its Build conference on June 2 in San Francisco with more in its model pipeline than the MAI-Image-2.5 that it has already shown on Arena, where the text-to-image system landed third behind OpenAI’s gpt-image-2 and Google’s Nano Banana 2. That release is lined up for the MAI Playground and Foundry, but three additional models are taking shape within the company’s stack, none of which are publicly available yet.

— Arena.ai (@arena) May 26, 2026

The first, MAI-Transcribe-1.5, is a modest step up from the speech-to-text model launched in April, which already claimed the lowest word error rate across 25 languages. The image side draws more attention: MAI-Image-2.5 looks set to ship in two variants, a high-quality version and a faster one labeled MAI-Image-2.5e, mirroring the split seen with MAI-Image-2. It would also accept image uploads, opening the model to editing as well as generation, putting it on par with rivals from Google and OpenAI.

The most striking find is MAI-Voice-2, a multilingual successor to the company’s text-to-speech model. While MAI-Voice-1 began in English, the new version adds German, Australian and US English, Spanish, French, Hindi, Indonesian, Italian, Japanese, Korean, Dutch, Portuguese, Turkish, Vietnamese, and Chinese, with a wider emotional range that covers tones such as angry, confused, and embarrassed. Early samples suggest it can whisper, too.

audio-thumbnail

audio-thumbnail

audio-thumbnail

All three would feed Copilot, Teams, and Azure Speech, and fit the developer crowd that Build is made for. The timing matches a broader push, as Mustafa Suleyman’s team weans the company off OpenAI following April’s renegotiation. Reports point to a homegrown coding model for GitHub Copilot at the show, too, while a Copilot “super app” that integrates chat, coding, and agents into a single hub is expected later in the summer.