OpenAI releases new voice models for more natural live conversations
OpenAI says its new voice mode can speak and listen at the same time, a key ability for live translation. OpenAI today released new conversational models, called GPT-Live-1 and GPT-Live-1 mini, claiming that they sound more natural and can handle turn-taking better. These are full-duplex models, meaning they can speak and listen at the same time, allowing users to interrupt naturally and enabling features like live translation.
Key Takeaways
- The company is also replacing its current Advanced Voice Mode in ChatGPT with GPT-Live-1 mini by default.
The previous model combined a speech-to-text model to transcribe speech, a large language model to generate responses, and a text-to-speech model to deliver the final answer.
- Other startups like Monogram, which raised $40 million in seed funding from DST and Lux Capital , are also leaning into visual responses to make assistants more interactive.
The company said the new voice mode in ChatGPT is designed to have longer conversations.
- "Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work.
The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work," Eleti said.
- Startups like Sesame , founded by Oculus co-founder Brendan Iribe and Ankit Kumar, also launched AI assistants with more natural conversation while completing tasks in the background.
OpenAI is moving in the same direction, aiming to let users talk to its assistant hands-free for a longer time.
- Ivan Mehta Ivan covers global consumer tech developments at TechCrunch.
Stats & Key Facts
- #Other startups like Monogram, which raised $40 million in seed funding from DST and Lux Capital , are also leaning into visual responses to make assistants more interactive.
- #The company said that more than 150 million people talk to ChatGPT using features like Voice and Dictation.
The company said in a press briefing that the new models solve issues like interrupting users while they're talking and not having enough intelligence to answer questions. OpenAI's new models will send the query to its latest text models like GPT-5.5 for search, reasoning, or agentic capabilities while continuing the conversation. OpenAI also showed that the model can stay silent for a long time and absorb the context of the conversation until it's called upon.
Plus, as the new voice mode has access to newer GPT models, it can also present some information in a visual format. Other startups like Monogram, which raised $40 million in seed funding from DST and Lux Capital , are also leaning into visual responses to make assistants more interactive. The company said the new voice mode in ChatGPT is designed to have longer conversations.
During the briefing, ChatGPT Voice's product lead, Atty Eleti, said he has had 30- to 40-minute-long conversations with the voice feature during walks. OpenAI thinks that voice could be the primary interface to computing for complex work. Reports have suggested that it could launch a pair of earbuds with AI capabilities this year .
For more details please read the original article at TechCrunch AI.
Continue Learning
Comments
Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.
No approved comments yet.