ChatGPT Voice and Image Capabilities Update
OpenAI has integrated voice and image capabilities into ChatGPT, transforming the interface from text-only to a multimodal experience. This update allows users to have real-time voice conversations and provide visual inputs for the model to analyze and discuss.
Voice Interaction Capabilities
ChatGPT now supports back-and-forth voice conversations, allowing users to interact with the assistant hands-free. This feature is available on iOS and Android via an opt-in setting in the "New Features" menu.
Technical Implementation of Voice
The voice experience is powered by a combination of two primary technologies:
- Text-to-Speech (TTS): A new model capable of generating human-like audio from text and a few seconds of sample speech. The available voices were created in collaboration with professional voice actors.
- Speech Recognition: OpenAI's open-source Whisper system is used to transcribe spoken words into text for the model to process.
Image Understanding and Visual Input
Users can now provide one or more images to ChatGPT to perform tasks such as troubleshooting hardware, planning meals based on pantry contents, or analyzing complex data graphs. On mobile devices, a drawing tool is available to help the user focus the model's attention on specific parts of an image.
Multimodal Model Architecture
Image understanding is powered by multimodal versions of GPT-3.5 and GPT-4. These models apply language reasoning skills to various visual formats, including photographs, screenshots, and documents that combine text and images.
Safety and Deployment Strategy
OpenAI is deploying these features gradually to Plus and Enterprise users over a two-week period, with plans to expand access to developers and other user groups later. This phased rollout is intended to refine risk mitigations and prepare users for more powerful systems.
Voice Safety and Risk Mitigation
Because the ability to create realistic synthetic voices can be used for impersonation or fraud, OpenAI has limited the technology's application to a specific use case: voice chat with voices created from professional actors. Outside of ChatGPT, OpenAI is collaborating with partners like Spotify to pilot voice translation features for podcasters.
Vision Safety and Privacy
To address risks such as hallucinations regarding people or misuse in high-stakes domains, OpenAI implemented the following measures:
- Red Teaming: The model was tested by red teamers and alpha testers for risks related to scientific proficiency and extremism.
- Privacy Restrictions: Technical measures have been implemented to significantly limit the ChatGPT's ability to analyze and make direct statements about people to protect privacy and ensure accuracy.
- Accessibility Collaboration: The development of vision features was informed by work with Be My Eyes, an app for blind and low-vision individuals.
Model Limitations
OpenAI notes specific limitations regarding the vision capabilities:
- Verification Required: Users are discouraged from using the model for high-risk use cases without proper verification.
- Language Constraints: While the model is proficient at transcribing English text, it performs poorly with other languages, particularly those using non-roman scripts.