Introducing GPT‑Live‑1 and GPT‑Live‑1 mini: OpenAI’s New Full‑Duplex Voice Models
Introducing GPT‑Live‑1 and GPT‑Live‑1 mini: OpenAI’s New Full‑Duplex Voice Models
TL;DR
OpenAI launches GPT‑Live‑1 and GPT‑Live‑1 mini, a new generation of voice models that use a full‑duplex architecture to listen and speak simultaneously, delegate heavy reasoning to GPT‑5.5, and provide a more natural ChatGPT Voice experience.
Overview
We’re launching GPT‑Live, a new generation of voice models that make talking with AI feel much more like having a real conversation.
GPT‑Live is built on a full‑duplex architecture, allowing it to listen and speak at the same time. It can show attentiveness with phrases like “mhmm” or “yeah”, engage in quick back‑and‑forth, or stay quiet when the user needs a moment to think.
At launch, GPT‑Live uses GPT‑5.5 in the background for tasks that require web search, deeper reasoning, or more complex work, while keeping the conversation flowing.
Two versions are released today: GPT‑Live‑1 and GPT‑Live‑1 mini, with plans to bring them to the API soon.
Technical Architecture
GPT‑Live addresses these limitations through two architectural changes.
Continuous interaction via full‑duplex – Instead of processing discrete turns, GPT‑Live continuously processes input while generating output, enabling it to decide many times per second whether to speak, listen, pause, interrupt, or invoke a tool. This supports natural back‑and‑forth, better timing awareness, and live translation.
Delegation for deeper work – The interaction layer (GPT‑Live) is decoupled from a separate model that handles search, reasoning, or agentic tasks. When a question needs such capabilities, GPT‑Live delegates to a frontier model like GPT‑5.5, maintains the conversation, and returns the result when ready. This design also lets GPT‑Live continuously use the latest frontier models as they are released.
Capabilities and User Experience
Over time, we believe this research will also unlock the ability to use voice for increasingly complex, longer‑running, and more agentic work.
- Users can interrupt, pause, or ask ChatGPT to slow down; the model acknowledges with phrases like “mhmm” or “got it”.
- Nine distinct voices have been remastered for GPT‑Live.
- While speaking, ChatGPT can display rich visual cards for topics such as weather, stocks, and sports.
- Voice continues to support search, memory, images, and file uploads.
- Reasoning levels can be chosen: Instant for fast responses, or Medium and High for more thinking time.
Evaluation
We built new human evaluations to measure pleasantness and the flow of conversation.
In head‑to‑head 5–10 minute conversations, GPT‑Live‑1 and GPT‑Live‑1 mini are strongly preferred over Advanced Voice Mode on overall preference, turn‑taking, interruptions, conversational flow, and naturalness.
Specific benchmark results mentioned in the source:
- GPQA: GPT‑Live‑1 substantially outperforms Advanced Voice Mode on expert‑level scientific reasoning across biology, chemistry, and physics.
- BrowseComp: GPT‑Live‑1 shows strong gains over Advanced Voice Mode on agentic web search and finding difficult‑to‑locate information.
- τ³‑Voice Telecom (internal variant): GPT‑Live‑1 outperforms Advanced Voice Mode on realistic, multi‑turn telecom support tasks.
Safety
GPT‑Live was designed to be safe by default.
Safety measures include:
- Expanded safety testing with new audio‑native evaluations and synthetic evaluations focusing on self‑harm, psychosis and mania, emotional reliance on AI, violence, and sexual content.
- Internal red‑teaming for voice‑specific risks.
- Real‑time safeguards that can steer the model toward safer responses, surface safety messaging, or end the conversation in higher‑risk cases.
- Adapted support flows for self‑harm, offering expert‑vetted crisis helpline.
- Age‑appropriate behavior trained directly into the model to reduce inappropriate responses for teens; parental controls allow parents to enable or disable ChatGPT Voice for teens and receive notifications in higher‑risk situations.
- Ongoing long‑term measurement and post‑launch monitoring focused on emotional reliance.
- Use of a predefined set of voices with safeguards to prevent voice impersonation of real individuals.
The source states that GPT‑Live performed comparably or better performance than Advanced Voice Mode across nearly all evaluated safety areas.
Rollout and Availability
We’re beginning to roll out two versions of GPT‑Live – GPT‑Live‑1 and GPT‑Live‑1 mini – to ChatGPT users globally today.
- GPT‑Live‑1 becomes the default model powering ChatGPT Voice for Go, Plus, and Pro users.
- GPT‑Live‑1 mini becomes the default for Free users.
- The rollout covers iOS, Android, and ChatGPT.com.
- Developers and enterprises can sign up to be notified when the models are available via the API using the provided form.
- Legacy versions of ChatGPT Voice (Standard and Advanced Voice Mode) remain accessible where features like voice with video or screen sharing are needed.
Limitations and Future Work
At launch, GPT‑Live will not support voice with video or screen sharing in ChatGPT, but we’re working to introduce these capabilities soon.
- For certain languages the model may have a non‑native accent or gaps in fluency; improvements are ongoing.
- The vision is to enable truly natural human‑AI interaction, extending voice to increasingly complex, longer‑running, and more agentic work over time.
All statements above are drawn directly from the source; no external facts or numbers have been added.
Sources
- OriginalIntroducing GPT-Live