OpenAI launches full-duplex GPT-Live-1 voice models

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- OpenAI released GPT-Live-1 and GPT-Live-1 mini, full-duplex conversational models claiming more natural sound and better turn-taking, enabling natural interruptions and live translation
- ChatGPT will replace its Advanced Voice Mode with GPT-Live-1 mini by default; paid-tier users gain access to the larger GPT-Live-1, which can query GPT-5.5 mid-conversation for search, reasoning, or agentic tasks
- The architecture shifted from a three-model pipeline (speech-to-text, then a large language model, then text-to-speech) to a single full-duplex model, a consolidation OpenAI says resolves issues like cutting users off mid-sentence
- Adoption is already large — more than 150 million people use ChatGPT's Voice and Dictation features, and product lead Atty Eleti said he personally runs 30- to 40-minute voice conversations during walks
- Rivals are pursuing similar conversational upgrades: Apple and Amazon updated their assistants for better context handling, while startups like Sesame (founded by Oculus co-founder Brendan Iribe and Ankit Kumar) launched expressive AI assistants that complete tasks in the background
- The demo stuttered — during a Hindi live-translation showcase, the assistant spoke with a heavy American accent and produced unnatural, bookish Hindi, and OpenAI declined to specify which languages the model is optimized for
- Safety guardrails built into the new models include age-appropriate responses for teens and resource referrals when conversations turn to self-harm, though OpenAI emphasized it is not building an AI companion
Why it matters: By collapsing its three-stage voice pipeline into a single full-duplex model and tying it to GPT-5.5, OpenAI is making voice the front door to its most capable AI — potentially reshaping how 150 million existing ChatGPT voice users interact with agentic and reasoning tools. The architectural shift lets the model stay silent and absorb context before responding, a capability competitors like Apple, Amazon, and Sesame are also racing to match.
Ask SkimNews


