惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

有赞技术团队
有赞技术团队
Martin Fowler
Martin Fowler
N
Netflix TechBlog - Medium
WordPress大学
WordPress大学
罗磊的独立博客
H
Help Net Security
MongoDB | Blog
MongoDB | Blog
A
About on SuperTechFans
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
D
Docker
云风的 BLOG
云风的 BLOG
Microsoft Security Blog
Microsoft Security Blog
Blog — PlanetScale
Blog — PlanetScale
P
Proofpoint News Feed
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
I
InfoQ
J
Java Code Geeks
博客园 - 聂微东
大猫的无限游戏
大猫的无限游戏
Engineering at Meta
Engineering at Meta
美团技术团队
小众软件
小众软件
Stack Overflow Blog
Stack Overflow Blog
C
Check Point Blog

Bing: albert einstein

Laptop Screen Flickering or Glitching? 10 Proven Fixes for 2026 Troubleshoot screen flickering in Windows How to Fix Windows 11 Screen Flicker: Hardware vs Software Guide PC Screen Flickering: Causes, Tests, and Effective Solutions GitHub - nexu-io/open-design: 🎨 Local-first, open-source alternative to Anthropic's Claude Design. ⚡ 19 Skills · ✨ 71 brand-grade Design Systems 🖼 Generate web · desktop · mobile prototypes · slides · images · videos · HyperFrames 📦 Sandboxed preview · HTML/PDF/PPTX/MP4 export 🤖 Runs on Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen / Copilot / Hermes / Kimi CLI. Pune - Compare our ride options - BlaBlaCar New-delhi - Compare our ride options - BlaBlaCar Bengaluru - Compare our ride options - BlaBlaCar Carpool in India - BlaBlaCar Models and pricing for GitHub Copilot - GitHub Docs B12全合成大师:三次被提名,等到97岁 | #一起聊诺贝尔奖 苏黎世联邦理工学院有机化学教授Albert Eschen… 人民币货币符号是「Y」加一横还是两横? - 知乎 ¥、$、€、£等符号是怎么来的? - 知乎 数学符号 ŷ 的中文和英文读法分别是什么? - 知乎 为什么《生化危机》系列里的威斯克(Albert Wesker)总给人一种亦正亦邪的感觉? - 知乎 如何评价特斯拉新出的焕新版 model Y? - 知乎 数学公式中,y上面有个^是什么意思,怎么读,如何在WORD中打出来_百度知道 川A、川B、川C、D、E、F..........U、V、W、川X、川Y、川Z开头的车牌号分别代表四川哪个市的?_百度知道 y开头的单词大全集_百度知道 【ラグビー】試合時間は約80分|年代別一覧もご紹介! - スポスルマガジン|様々なスポーツ情報を配信 About Classroom - Classroom Help ラグビーの試合時間は何分?国際試合・中学・高校・大学・社会人まで解説 拼音y和w为什么不是声母,为什么?_百度知道 OpenClassrooms Get started with Classroom for students - Computer Classroom Help - OpenClassrooms How do I sign in to Classroom? - Computer 第5条 試合時間 | みんなでラグビー|日本ラグビーフットボール ... ラグビーの試合時間は約80分!ワールドカップ観戦のためのラグビー解説!
ChatGPT can now see, hear, and speak
2023-09-25 · via Bing: albert einstein
OpenAI

We are beginning to roll out new voice and image capabilities in ChatGPT. They offer a new, more intuitive type of interface by allowing you to have a voice conversation or show ChatGPT what you’re talking about.

Voice and image give you more ways to use ChatGPT in your life. Snap a picture of a landmark while traveling and have a live conversation about what’s interesting about it. When you’re home, snap pictures of your fridge and pantry to figure out what’s for dinner (and ask follow up questions for a step by step recipe). After dinner, help your child with a math problem by taking a photo, circling the problem set, and having it share hints with both of you.

We’re rolling out voice and images in ChatGPT to Plus and Enterprise users over the next two weeks. Voice is coming on iOS and Android (opt-in in your settings) and images will be available on all platforms.

You can now use voice to engage in a back-and-forth conversation with your assistant. Speak with it on the go, request a bedtime story for your family, or settle a dinner table debate.

Use voice to engage in a back-and-forth conversation with your assistant.

To get started with voice, head to Settings → New Features on the mobile app and opt into voice conversations. Then, tap the headphone button located in the top-right corner of the home screen and choose your preferred voice out of five different voices.

The new voice capability is powered by a new text-to-speech model, capable of generating human-like audio from just text and a few seconds of sample speech. We collaborated with professional voice actors to create each of the voices. We also use Whisper, our open-source speech recognition system, to transcribe your spoken words into text.

You can now show ChatGPT one or more images. Troubleshoot why your grill won’t start, explore the contents of your fridge to plan a meal, or analyze a complex graph for work-related data. To focus on a specific part of the image, you can use the drawing tool in our mobile app.

Show ChatGPT one or more images.

To get started, tap the photo button to capture or choose an image. If you’re on iOS or Android, tap the plus button first. You can also discuss multiple images or use our drawing tool to guide your assistant.

Image understanding is powered by multimodal GPT‑3.5 and GPT‑4. These models apply their language reasoning skills to a wide range of images, such as photographs, screenshots, and documents containing both text and images.

OpenAI’s goal is to build AGI that is safe and beneficial. We believe in making our tools available gradually, which allows us to make improvements and refine risk mitigations over time while also preparing everyone for more powerful systems in the future. This strategy becomes even more important with advanced models involving voice and vision.

The new voice technology—capable of crafting realistic synthetic voices from just a few seconds of real speech—opens doors to many creative and accessibility-focused applications. However, these capabilities also present new risks, such as the potential for malicious actors to impersonate public figures or commit fraud.

This is why we are using this technology to power a specific use case—voice chat. Voice chat was created with voice actors we have directly worked with. We’re also collaborating in a similar way with others. For example, Spotify is using the power of this technology for the pilot of their Voice Translation(opens in a new window) feature, which helps podcasters expand the reach of their storytelling by translating podcasts into additional languages in the podcasters’ own voices.

Vision-based models also present new challenges, ranging from hallucinations about people to relying on the model’s interpretation of images in high-stakes domains. Prior to broader deployment, we tested the model with red teamers for risk in domains such as extremism and scientific proficiency, and a diverse set of alpha testers. Our research enabled us to align on a few key details for responsible usage.

Like other ChatGPT features, vision is about assisting you with your daily life. It does that best when it can see what you see. 

This approach has been informed directly by our work with Be My Eyes, a free mobile app for blind and low-vision people, to understand uses and limitations. Users have told us they find it valuable to have general conversations about images that happen to contain people in the background, like if someone appears on TV while you’re trying to figure out your remote control settings.

We’ve also taken technical measures to significantly limit ChatGPT’s ability to analyze and make direct statements about people since ChatGPT is not always accurate and these systems should respect individuals’ privacy.

Real world usage and feedback will help us make these safeguards even better while keeping the tool useful.

Users might depend on ChatGPT for specialized topics, for example in fields like research. We are transparent about the model's limitations and discourage higher risk use cases without proper verification. Furthermore, the model is proficient at transcribing English text but performs poorly with some other languages, especially those with non-roman script. We advise our non-English users against using ChatGPT for this purpose.

You can read more about our approach to safety and our work with Be My Eyes in the system card for image input.

Plus and Enterprise users will get to experience voice and images in the next two weeks. We’re excited to roll out these capabilities to other groups of users, including developers, soon after.