I do not fully understand the complexity behind achieving full-duplex but I hope this sets the bar for Anthropic to follow. Turn-based simplex is yesterday.
You have always been able to pick between voices of many kinds though? Do you find them all sultry? From the British woman to the 17th century pirate soundalike?
What I’m missing from this announcement is the capability to use connectors and tools. I don’t really get it - NONE of the frontier assistants can use tools / connectors while in voice mode - Claude, ChatGPT, Gemini, Grok. It seems so obvious: I want to be able to research stuff, pull up documents, jot down notes and do productive work while I’m talking to it, and not end voice mode whenever I need to connect to an app or service.
It’s weird. The old Claude voice mode WAS able to use tools but when they revamped it, it lost that capability and is now pinned to Haiku ![]()
So, yay for finally a voice mode that’s powered by a frontier model and hopefully as good as Grok voice, but sad to still not see tool use while in voice mode.
(I haven’t tried it yet, only read the announcement)
(Atty from OpenAI here)
GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.
Would love to hear your feedback!
watched the live translation video very impressive
Seems like a shift from previous voice models where it sequentially processes voice to text then feeds it to LLM and then back which cant escape the clunky lag
not sure how pipecat stands now, gpt live seems like it takes audio tokens and does inference on it directly
I’m so mad that this might make me re-subscribe to ChatGPT. I wouldn’t have believed how much I use the voice feature before LLMs and ChatGPT currently has the best voice interface. I think Grok’s interface is the next best, then Claude.
I for one am greatly looking forward to the day these kind of voice models can be run locally. It seems like the gap between open-weight and frontier is way larger for voice models than coding/language models.
Is it a dumb-down version of GPT like the current voice model? At least in french, I find the current GPT voice mode to be useless, to the point I only use the dication mode. I would ask a question and it would answer something along “That’s a interesting question. I can help you with that. Anything you want to know about X?” I would ask again and it would answer the same kind of non answer.
Thank you for testing and the feedback Simon!
Is it responsive to personality settings? I actively don’t want fake AI girlfriend, but I do get a ton of value out of voice mode. Looking forward to trying this but hoping it’s not a creepy overdone mess (like Sesame). Expectations are they’ll keep doubling down on fake AI girlfriend approach because the thing I want probably wouldn’t drive engagement anywhere nearly as well
Same. I might switch back to ChatGPT from Gemini because I use the voice feature all the time.
One of my favorite use cases is talking with it while driving on random topics and learning about them.
There doesn’t seem to be any indication whether this is available in the chat-got app nor is there any indication in the app that anything has changed. Anyone know how to actually try this?
Cool, was this conversation through the chatgpt app?
What made you to try again?
Read the post
Gemini live has been able to do this for over a year now. I can just activate it on my phone and it really works surprisingly well, especially the interruption. I’ve tested it with my 95 year old Dutch grandmother and it switched seamlessly between English and Dutch with her and handled her poor hearing very well, including her asking for repetition.
I’m a little surprised by how much OAI is playing catch up here.
It sounds like they’ve switched to a “native audio” model which if I understand right is what Gemini has had for quite a while?
But does it generate good pelicans?
“I don’t translate, I interpret” - Ahmed the best interpreter there ever was
Is it possible to create a “companion” of sorts with this model, using, say, an RPi and a speaker + microphone? Not for advanced scientific brainstorming, but for seniors who are often alone in their homes.