GPT‑Live

https://openai.com/index/introducing-gpt-live/

Read the linked article.


Discuss on Hacker News

This solves my biggest annoyance with the current advanced voice: its speech getting interrupted by me setting yup or even background noise if loud enough

full duplex, yay

Hoping to use this for natural conversation language learning. Previous iterations of the app kept correcting my words/grammar before it got to the model, causing issues with identifying mistakes in speech

With this, human translators have been totally and absolutely a solved problem with this version of real time translation.

This time is the most natural version that exists and it is a natural as a conversation.

To Downvoters: Why aren’t you feeling the AGI?

Very cool. Not cool bringing Brazil’s loss to Norway again. We’re already devastated. No need to keep beating someone on the ground. :frowning:

The potential conversational dynamics of people telling each other “quiet!” after they pick up the habit from talking with AI will be interesting. It could lead to people being more assertive and thoughtful, or it could be contentious and rude.

Awesome that they’ve improved that aspect of voice chat, though.

Are there any open source full duplex models that are out besides PersonaPlex? There was a chinese open one, maybe Fun Audio chat or something, that said it was going to release a full duplex version but I am not sure if it did.

My dream would be open source full duplex with function calling or some kind of rudimentary text output. PersonaPlex is still interesting although it was looking like we would need to fine tune it to handle outgoing or avoid going off the rails easily.

Oh wow, I’d like this. Our current voice interactions with ChatGPT are on a 4o era model; really terrible. oAI has always been pretty cagey on the architecture of their end to end multimodal models. And RL has basically made them worse since launch. (Check the launch videos where the model sings, is more realtime, has accents, etc). I’d love to try a next gen version.

Absolutely can’t wait to try this for language practice. The advanced voice mode is great but ultimately just doesn’t work that well and doesn’t have the feel of a natural conversation.

This looks very cool. An AI that can listen and speak and handle tasks without breaking the flow of conversation would solve some big annoyances with current tools.

The concern is though as these get better will people struggle to distinguish these with real human connections?

Any pricing announced yet?

I’m very eager to test this for brainstorming!

One thing I noticed is that we lost vision feature for some reason on the live chat?

This was an extremely useful feature. Not sure if it’s a regional thing or that they just removed that from the current live chat.

I imagine it will be even more useful with this new version.

Does this support more than one user voice? Or, are there plans for this? I did not see that mentioned in the announcement.

Very cool. I thought the agent came in a little to hot at 1:03. I wonder how it decides when to jump in.

I like this and felt like some of it was much more fluid; but was I alone in feeling like the interjected “uh-huh” or “yeah?” moments felt a little jarring?

Almost felt a bit uncanny valley for what “natural” conversation is supposed to be like. If the “uh huh” isn’t timed correctly, it’ll feel like a zoom call with lag.

I had preview access to this one for a few weeks. It’s very good. I had one conversation that lasted a full hour while I was walking the dog, got some good brainstorming done against one of my projects.

The best feature is that it can delegate questions out to GPT-5.5 in the background, so you’re no longer restricted to a voice model that’s several years behind the frontier.

I did report a fun bug with it though: it was interrupting me and laughing at my (not really intended as) jokes while I was still talking! They seem to have clamped that behavior down thankfully, it felt a bit rude and condescending.

Definitely in the right direction in terms of architecture. However those “hmmm” “uh huh” interjected in the demo are pretty awful.

I want to know this too, as I’m hoping to fabricobble up a “smart speaker” that communicates with my local AI assistant. Right now we do everything via iMessage but it would be nice to be able to tell it to add things to my grocery list by voice while my hands are busy in the kitchen. Also would love if anyone has any advice on what microphone & speaker to pick out, was planning on just reusing a raspberry pi I’ve got around for the brain part.

I was hopeful that they avoided the well known sultry voice this go around, but alas. There is little hope for these companies.

The full duplex is awesome, and the feedback that it is getting what you’re saying is ok, but in some of the demos was a little overkill.

I’ll agree that using the “Golden Girls” was at least more entertaining than the usual pitch.