惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Y
Y Combinator Blog
GbyAI
GbyAI
爱范儿
爱范儿
H
Hackread – Cybersecurity News, Data Breaches, AI and More
C
Check Point Blog
M
MIT News - Artificial intelligence
量子位
宝玉的分享
宝玉的分享
MongoDB | Blog
MongoDB | Blog
V
Visual Studio Blog
罗磊的独立博客
F
Fortinet All Blogs
美团技术团队
博客园_首页
博客园 - 【当耐特】
L
LangChain Blog
月光博客
月光博客
腾讯CDC
The Cloudflare Blog
D
Docker
博客园 - 聂微东
Stack Overflow Blog
Stack Overflow Blog
WordPress大学
WordPress大学
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

Recent Commits to openclaw:main

test: merge chat side-result checks · openclaw/openclaw@ddd2c2a test: merge cron history checks · openclaw/openclaw@f7eb746 test: merge responsive navigation shell checks · openclaw/openclaw@c2e4b47 docs(changelog): add codex oauth fixes · openclaw/openclaw@628e6cd test: merge navigation routing cases · openclaw/openclaw@5d8cecb Tests: mock channel registry bundled fallback · openclaw/openclaw@2b08233 Secrets: avoid broad web search discovery for single plugin config · openclaw/openclaw@a464f59 test: merge config view browser checks · openclaw/openclaw@20cf511 fix(status): align oauth health with runtime · openclaw/openclaw@eed7116 feat: add macOS screen snapshots for monitor preview (#67954) thanks … · openclaw/openclaw@f377db1 fix: report shared auth scopes in hello-ok (#67810) thanks @BunsDev · openclaw/openclaw@0b6c39b Auto-reply: avoid eager bundled route fallback · openclaw/openclaw@3ea1bf4 Tests: narrow session binding contract setup · openclaw/openclaw@54e4e16 fix(macOS): enable undo/redo in webchat composer text input (#34962) · openclaw/openclaw@00951dc Tests: speed up channel setup promotion · openclaw/openclaw@82b529a Docs: refresh agent instructions · openclaw/openclaw@5775fe2 fix(auth): serialize OAuth refresh across agents to fix #26322 (#67876) · openclaw/openclaw@8e79080 test: allow ollama public surface boundary test · openclaw/openclaw@7d4f1a6 Docs: add test performance guardrails · openclaw/openclaw@89706d3 Tests: restore context-engine usage proof · openclaw/openclaw@e4c4f95 Tests: slim context engine runtime coverage · openclaw/openclaw@74c198f ci: retry failed custom checkouts · openclaw/openclaw@0ee5baf test: trim duplicate provider auth onboarding cases · openclaw/openclaw@1ffc02e matrix: fix sessions_spawn --thread subagent session spawning (#67643) · openclaw/openclaw@1ce2596 test: reduce auth choice fixture churn · openclaw/openclaw@857b9cd test: mock health status config boundaries · openclaw/openclaw@9d5ab4a test: mock onboard config io boundary · openclaw/openclaw@299694d test: mock legacy state plugin boundaries · openclaw/openclaw@2713089 test: mock channel install boundaries · openclaw/openclaw@b945248 test: mock doctor preview channel boundaries · openclaw/openclaw@b1a3ad4
docs: document Google realtime voice support · openclaw/o...
steipete · 2026-04-24 · via Recent Commits to openclaw:main

@@ -18,31 +18,31 @@ OpenClaw generates images, videos, and music, understands inbound media (images,

1818

| Image generation | `image_generate` | ComfyUI, fal, Google, MiniMax, OpenAI, Vydra, xAI | Creates or edits images from text prompts or references |

1919

| Video generation | `video_generate` | Alibaba, BytePlus, ComfyUI, fal, Google, MiniMax, OpenAI, Qwen, Runway, Together, Vydra, xAI | Creates videos from text, images, or existing videos |

2020

| Music generation | `music_generate` | ComfyUI, Google, MiniMax | Creates music or audio tracks from text prompts |

21-

| Text-to-speech (TTS) | `tts` | ElevenLabs, Microsoft, MiniMax, OpenAI, xAI | Converts outbound replies to spoken audio |

21+

| Text-to-speech (TTS) | `tts` | ElevenLabs, Google, Microsoft, MiniMax, OpenAI, xAI | Converts outbound replies to spoken audio |

2222

| Media understanding | (automatic) | Any vision/audio-capable model provider, plus CLI fallbacks | Summarizes inbound images, audio, and video |

23232424

## Provider capability matrix

25252626

This table shows which providers support which media capabilities across the platform.

272728-

| Provider | Image | Video | Music | TTS | STT / Transcription | Media Understanding |

29-

| ---------- | ----- | ----- | ----- | --- | ------------------- | ------------------- |

30-

| Alibaba | | Yes | | | | |

31-

| BytePlus | | Yes | | | | |

32-

| ComfyUI | Yes | Yes | Yes | | | |

33-

| Deepgram | | | | | Yes | |

34-

| ElevenLabs | | | | Yes | Yes | |

35-

| fal | Yes | Yes | | | | |

36-

| Google | Yes | Yes | Yes | | | Yes |

37-

| Microsoft | | | | Yes | | |

38-

| MiniMax | Yes | Yes | Yes | Yes | | |

39-

| Mistral | | | | | Yes | |

40-

| OpenAI | Yes | Yes | | Yes | Yes | Yes |

41-

| Qwen | | Yes | | | | |

42-

| Runway | | Yes | | | | |

43-

| Together | | Yes | | | | |

44-

| Vydra | Yes | Yes | | | | |

45-

| xAI | Yes | Yes | | Yes | Yes | Yes |

28+

| Provider | Image | Video | Music | TTS | STT / Transcription | Realtime Voice | Media Understanding |

29+

| ---------- | ----- | ----- | ----- | --- | ------------------- | -------------- | ------------------- |

30+

| Alibaba | | Yes | | | | | |

31+

| BytePlus | | Yes | | | | | |

32+

| ComfyUI | Yes | Yes | Yes | | | | |

33+

| Deepgram | | | | | Yes | | |

34+

| ElevenLabs | | | | Yes | Yes | | |

35+

| fal | Yes | Yes | | | | | |

36+

| Google | Yes | Yes | Yes | Yes | | Yes | Yes |

37+

| Microsoft | | | | Yes | | | |

38+

| MiniMax | Yes | Yes | Yes | Yes | | | |

39+

| Mistral | | | | | Yes | | |

40+

| OpenAI | Yes | Yes | | Yes | Yes | Yes | Yes |

41+

| Qwen | | Yes | | | | | |

42+

| Runway | | Yes | | | | | |

43+

| Together | | Yes | | | | | |

44+

| Vydra | Yes | Yes | | | | | |

45+

| xAI | Yes | Yes | | Yes | Yes | | Yes |

46464747

<Note>

4848

Media understanding uses any vision-capable or audio-capable model registered in your provider config. The table above highlights providers with dedicated media-understanding support; most LLM providers with multimodal models (Anthropic, Google, OpenAI, etc.) can also understand inbound media when configured as the active reply model.

@@ -58,12 +58,14 @@ ElevenLabs, Mistral, OpenAI, and xAI also register Voice Call streaming STT

5858

providers, so live phone audio can be forwarded to the selected vendor

5959

without waiting for a completed recording.

606061-

OpenAI maps to OpenClaw's image, video, batch TTS, batch STT, Voice Call

62-

streaming STT, realtime voice, and memory embedding surfaces. xAI currently

63-

maps to OpenClaw's image, video, search, code-execution, batch TTS, batch STT,

64-

and Voice Call streaming STT surfaces. xAI Realtime voice is an upstream

65-

capability, but it is not registered in OpenClaw until the shared realtime

66-

voice contract can represent it.

61+

Google maps to OpenClaw's image, video, music, batch TTS, backend realtime

62+

voice, and media-understanding surfaces. OpenAI maps to OpenClaw's image,

63+

video, batch TTS, batch STT, Voice Call streaming STT, backend realtime voice,

64+

and memory embedding surfaces. xAI currently maps to OpenClaw's image, video,

65+

search, code-execution, batch TTS, batch STT, and Voice Call streaming STT

66+

surfaces. xAI Realtime voice is an upstream capability, but it is not

67+

registered in OpenClaw until the shared realtime voice contract can represent

68+

it.

67696870

## Quick links

6971