OpenAI is releasing GPT-Live, a new generation of voice models designed to make talking with AI feel significantly more like a real conversation.
GPT-Live is built on a full-duplex architecture, meaning it can listen and speak simultaneously. During conversations, GPT-Live can signal that it's paying attention with phrases like "mhmm" or "yeah," engage in rapid back-and-forth, or remain quiet when the user needs a moment to think. The result is a voice experience that feels refreshingly easy to talk to.
GPT-Live is also OpenAI's smartest voice model to date. For questions that require web search, deeper reasoning, or more complex work, it delegates to OpenAI's latest frontier model behind the scenes and brings the result back into the conversation when ready. While it works, GPT-Live can keep talking with the user and maintain conversational flow. At launch, GPT-Live uses GPT-5.5 in the background. As OpenAI releases new frontier models, the model used by GPT-Live will be continuously updated.
These advances power a new ChatGPT Voice experience that is more intelligent and natural to use. Over time, OpenAI believes this research will also unlock the ability to use voice for increasingly complex, longer-running, and more agentic work.
OpenAI is beginning to roll out two versions of GPT-Live - GPT-Live-1 and GPT-Live-1 mini - to ChatGPT users globally. There are also plans to bring them to the API soon, and developers and enterprises can sign up to be notified via a form on OpenAI's website.
A New Era of Human-AI Interaction
OpenAI's vision is to enable truly natural human–AI interaction: a world where collaborating with AI feels as fluid and responsive as working with another person, while reasoning and complex task execution happen seamlessly in the background.
Previous Approaches
Older generations of voice AI systems moved closer to that vision but came with important tradeoffs.
Cascaded voice systems relied on a series of models acting one after another to process each turn. The original ChatGPT Voice chained three models together: a speech-to-text model to transcribe speech, a large language model to produce a response, and a text-to-speech model to convert it back into speech. This approach enabled users to talk to frontier AI models for the first time, but the complexity came at a cost: information could be lost across models, and responses were slow and stilted.
Turn-based voice models like ChatGPT Advanced Voice Mode processed and generated audio within a single model, reducing latency and making conversations smoother - but they still operated through discrete turns. The model had to wait for the user to stop speaking before responding, resulting in rigid back-and-forth. Additionally, because turn detection was based on silence, even a brief pause or background noise could be mistaken for the end of a turn - causing the model to interrupt at unnatural times.
OpenAI's New Approach
GPT-Live addresses these limitations through two architectural changes.
Continuous interaction: GPT-Live was built for continuous interaction using a full-duplex architecture. Instead of processing a sequence of separate messages, GPT-Live continuously processes input while generating output. The model can therefore make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool. This allows the model to engage in more natural back-and-forth, maintain a better sense of time, and even perform live translation.
Delegation for deeper work: OpenAI decoupled GPT-Live - which handles continuous interaction - from deeper work. When a question requires search, reasoning, or more agentic capabilities, GPT-Live can delegate the task to another model like GPT-5.5. This allows it to keep the conversation going even as it handles multiple tasks in the background. This architectural change also allows GPT-Live to continuously use the latest models and agents, combining frontier intelligence with natural interaction.
Evaluations
OpenAI built new human evaluations to measure pleasantness and the flow of conversation. In head-to-head comparisons, GPT-Live-1 and GPT-Live-1 mini are strongly preferred over Advanced Voice Mode in matched 5–10 minute conversations that measure overall preference, turn-taking, interruptions, conversational flow, and how natural each interaction felt.
- GPQA: GPT-Live-1 substantially outperforms Advanced Voice Mode on GPQA, which tests expert-level scientific reasoning across biology, chemistry, and physics.
- BrowseComp: GPT-Live-1 shows strong gains over Advanced Voice Mode on BrowseComp, which tests agentic web search and the ability to find difficult-to-locate information.
- τ³-Voice Telecom (internal variant): GPT-Live-1 outperforms Advanced Voice Mode on τ³-Voice Telecom, which tests voice agents on realistic, multi-turn telecom support tasks.
An Upgraded ChatGPT Voice Experience
Each week, more than 150 million people talk to ChatGPT using features like Voice and Dictation. They use it for hands-free everyday help, language practice, bedtime stories, or just chatting during a commute.
When users tap the Voice button to talk with ChatGPT, they now get an improved experience powered by GPT-Live - with more natural conversations, smarter answers, better listening, and visual responses.
More Natural Conversations
Talking with ChatGPT now feels much more like a real conversation. Users can interrupt with a question, pause to gather thoughts, or ask ChatGPT to slow down. It naturally acknowledges what users are saying with phrases like "mhmm" or "got it," so they know it's following along. OpenAI has also remastered the nine distinct voices in ChatGPT for GPT-Live.
Smarter Answers
ChatGPT Voice can now draw on OpenAI's latest frontier models, providing smarter answers when needed. Users can also choose the level of reasoning that fits their needs: Instant for fast responses, or Medium and High when they want ChatGPT to spend more time thinking.
Better Listening
If a user takes a moment to think, ChatGPT Voice now waits instead of jumping in and interrupting. If asked to stay quiet and listen, it will. And when there's background noise, like passing traffic or nearby conversations, ChatGPT is better at focusing on the user's voice instead of getting distracted.
Visual Answers at a Glance
Some answers are more useful when they can be seen. While talking, ChatGPT can now show rich visual cards for topics like weather, stocks, sports, and more. Voice also continues to support search, memory, images, and file uploads.
Safety Designed for Voice
GPT-Live was designed to be safe by default. It builds on the safety advances from OpenAI's latest models while adding dedicated safety training across key risk areas and new safeguards designed specifically for voice.
Expanded Safety Testing
To better reflect how people use voice in real-life settings, OpenAI expanded safety testing to include new audio-native evaluations. Synthetic evaluations using generated audio were also created to focus more intensively on key safety areas, drawing on lessons learned from Advanced Voice Mode. Those areas include self-harm, psychosis and mania, emotional reliance on AI, violence, and sexual content. Internal experts also red-teamed the model for risks unique to voice.
In testing, GPT-Live performed comparably to or better than Advanced Voice Mode across nearly all evaluated areas. More details are available in the GPT-Live system card.
Built-in Safeguards
Because voice conversations unfold in real time, OpenAI built safeguards that can act while the model is speaking. When the system detects potentially unsafe output, it can steer the model toward a safer response, surface additional safety messaging or resources, or end the voice conversation in higher-risk cases. For conversations involving self-harm, OpenAI adapted ChatGPT's support flows for voice, including offering expert-vetted crisis helpline support.
Additional protections were designed to support teen users, and age-appropriate behavior was trained directly into the model to reduce the risk of inappropriate responses. Parents can choose whether their teen can use ChatGPT Voice through Parental Controls, and linked parents may be notified in higher-risk situations involving signs of potential self-harm or suicidal intent.
Learning from Real-World Use
OpenAI is also rolling out longer-term measurement and post-launch monitoring focused on emotional reliance to continue improving understanding and refining safeguards. Building on previous research into affective use and emotional well-being, this will help identify emerging patterns and improve how the system responds in emotionally sensitive interactions.
GPT-Live is designed for conversation, not voice impersonation. It uses a set of predefined voices in ChatGPT, with safeguards to prevent it from imitating a real person's voice.
Availability and Limitations
GPT-Live is rolling out to ChatGPT users globally across iOS, Android, and ChatGPT.com. GPT-Live-1 becomes the default model powering ChatGPT Voice for Go, Plus, and Pro users, and GPT-Live-1 mini becomes the default for Free users.
GPT-Live has been optimized for some of the most popular languages in ChatGPT. For certain languages, the model may have a non-native accent or gaps in fluency. OpenAI is actively working to improve the experience across languages.
At launch, GPT-Live does not support voice with video or screen sharing in ChatGPT, but OpenAI is working to introduce these capabilities soon. Users can still access legacy versions of ChatGPT Voice, including Standard and Advanced Voice Mode, where these features are available.