OpenAI Launches GPT-Live Voice Model: Full-Duplex Conversations Feel Natural, but "Over-Enthusiasm" Draws User Complaints

In the early hours of July 9 Beijing time, OpenAI officially launched its next-generation GPT-Live voice model series, aiming to completely shed the rigid turn-based "you say something, I say something" feel of AI voice assistants and achieve natural interaction more akin to chatting with a real person. However, after the first wave of user experiences, social media has been flooded with complaints labeling the model "annoying," with many finding that the model's excessive short feedback—such as "mhmm" and "got it"—has become a new form of interruption.
According to OpenAI's official introduction, the core upgrade of GPT-Live lies in its adoption of a "full-duplex" architecture. Previously, whether it was the cascaded system of the original ChatGPT Voice or the later Advanced Voice Mode based on GPT-4o, both were essentially "turn-based" conversations. The model had to determine whether the user had finished speaking before it could begin generating a response, which meant a brief pause by the user could be interrupted by the model, or the system could be falsely triggered by background noise. GPT-Live, by contrast, can listen and speak simultaneously, making multiple interaction decisions per second to judge whether it should talk, continue listening, pause, or interrupt. This allows it to naturally insert short feedback like "mhmm" or "got it" into the conversation, or remain silent and wait when the user is thinking.
OpenAI CEO Sam Altman described this update as "both magical and real." From a technical architecture standpoint, GPT-Live also introduces a mechanism that separates "front-end interaction" from "back-end reasoning." GPT-Live itself is primarily responsible for maintaining the fluency and natural rhythm of the conversation. When it encounters tasks requiring web searches, deep reasoning, or complex agent operations, it offloads those tasks to a more powerful frontier model running in the background. At launch, this backend model is GPT-5.5, and OpenAI has stated it will continue updating the underlying support as new models are released. This decoupled design allows the voice assistant to maintain natural voice interaction with the user on the front end even while performing high-intensity computation in the background, avoiding awkward long silences.
In officially released evaluation data, GPT-Live-1 achieved an overall preference rate of 75.7% in 5-to-10-minute conversation tests, far surpassing the previous generation Advanced Voice Mode. On conversation fluency scoring, GPT-Live-1 received 4.96 points (out of 7), compared to just 3.80 for Advanced Voice Mode. In the challenging expert-level scientific reasoning benchmark GPQA, GPT-Live-1 at the high-intensity reasoning tier achieved an accuracy of 84.2%, nearly double the 45.3% of Advanced Voice Mode. In the complex agent web search test BrowseComp, Advanced Voice Mode scored a mere 0.7% accuracy, while GPT-Live-1 reached as high as 75.2%—a leap from "unusable" to "capable of getting work done."
That said, these significant performance gains are largely attributable to the backend calling the latest GPT-5.5 model, rather than a fundamental qualitative change in the voice interaction layer itself. GPT-Live is now rolling out to global users on iOS, Android, and the web. Paid Go, Plus, and Pro users will default to the full-powered GPT-Live-1, while free users will experience the GPT-Live-1 mini version. OpenAI has also re-recorded the nine exclusive voices in ChatGPT and offers three reasoning levels for responses: Instant, Medium, and High.
However, it is precisely these short feedback mechanisms—designed to be "more natural"—that have become the primary target of early user complaints. Some users feel that GPT-Live's frequent use of "mhmm" and "yeah" as filler during conversations, while intended to provide psychological reassurance, comes across as overly verbose and even irritating in practice. Commentators have pointed out that the biggest difference between voice assistants and text assistants is that sound directly intrudes on a person's attention. Once the assistant jumps in too eagerly or provides too much short feedback, what the user perceives is not intelligence, but annoyance.
Additionally, GPT-Live has some functional limitations at launch. It currently does not support voice calling, video calling, or screen sharing within ChatGPT; users needing these features must switch back to the old version. In real-time mode, it also cannot access ChatGPT's memory. OpenAI says it is working to add these features as quickly as possible.
On the safety front, OpenAI conducted specialized training targeting the real-time nature of voice. Internal red team testing shows that GPT-Live-1's scores in preventing illegal activities, self-harm, hate speech, and other areas have all improved significantly. The system also introduces real-time protection mechanisms: once a potentially unsafe output is detected, it can intervene and guide the model while it is speaking, or in extremely high-risk situations, directly cut off the voice. For teenage users, the system includes built-in parental controls and has injected age-appropriate behavioral norms into the model's foundation. Notably, to prevent voice cloning scams, GPT-Live is strictly limited to using only preset voices and absolutely does not offer the ability to mimic real people's voices. An important backdrop to this is the major controversy in 2024 when GPT-4o launched with a "Sky" voice that sounded highly similar to actress Scarlett Johansson's voice; OpenAI subsequently removed that voice and apologized.
From an industry competition perspective, full-duplex voice interaction is rapidly becoming a standard feature for consumer AI products. Google's Gemini Live already supports full-duplex conversation along with camera and screen sharing—the latter two being features GPT-Live currently lacks. Nvidia also released PersonaPlex, a full-duplex model with customizable voices, earlier this year. The launch of GPT-Live appears more like OpenAI integrating capabilities already emerging in the industry into ChatGPT as a mainstream entry point, while deeply binding them with the advanced models within its ecosystem.
OpenAI disclosed that over 150 million people currently use ChatGPT's voice and dictation features each week, primarily for practicing foreign languages, chatting during commutes, or hands-free queries for everyday information. With the rollout of GPT-Live, OpenAI is attempting to upgrade voice interaction from a simple "text substitute" into a super real-time entry point that commands search, memory, and complex tasks. But striking the right balance between "natural" and "intrusive"—truly teaching AI when it's appropriate to stay quiet—will be the key to determining whether this generation of voice models can retain users.
Once added, BigGo Finance appears first in Google Search Top Stories, so you get the broadest, most up-to-the-minute, and most comprehensive global financial news first.