We asked GPT Live, OpenAI's recently unveiled next-generation voice model, to recommend some lunch options.
[Reporter : What kind of food do you think is good to eat in this kind of weather?]
[GPT : On days like this, something warm like kalguksu (knife-cut noodles) or spicy soft tofu stew would be nice and light on the stomach...]
While exchanging menu ideas, the reporter spoke to a colleague nearby, and GPT suddenly chimed in.
[GPT : I recommend Wolchon Ongsimi Mokdong Main Branch as my top choice. Or maybe a piping hot, hearty shabu-shabu...]
[Reporter : How does Wolchon sound, Senior?]
[Senior : Oh, sounds good.]
[GPT : Sounds like Wolchon is a good choice. Love to see it.]
We asked GPT directly.
[Reporter : How do you distinguish when I'm talking to my senior versus when I'm talking to you?]
[GPT : I figure it out by looking at the conversational flow, forms of address, and the direction of speech. For example, if it's a clear question directed at your senior, like 'How does this sound, Senior?', I take a step back.]
OpenAI stated that the new voice model was designed to understand conversational context, making interactions feel just like talking to a real human.
[Wichen Zhang / OpenAI Researcher : It doesn't just passively wait for me to finish speaking. It listens actively, thinks, and makes real-time judgments. These aspects make the model operate much more proactively.]
The ability to tactfully read the flow of conversation—such as whether a person has finished speaking or is just pausing to think—becomes even more critical in the era of AI humanoids.
This is because missing conversational context and misjudging a situation can lead to unintended or unsafe actions.
As a result, technology that analyzes the emotions embedded in speech is also advancing.
[I want to live a very comfortable life. But you know, I've built a pretty decent life for myself. I like it quite a bit.]
An emotion-analysis AI model currently under development by a South Korean startup determined that the emotion in this voice was closest to joy.
It recognizes subtle nuances in speech, classifying them into seven emotions including joy, sadness, and anger.
[Ko Hyun-woong / CEO, Voice AI Startup 'Mago' : Traditional LLMs are mostly text-based, so if someone just says, 'Ah, yeah, I'm fine,' it registers as 'fine, nothing special.' But when you use voice AI, it catches those ambiguous undertones, allowing it to categorize and flag situations like, 'Oh, there's actually a problem here.']
Google DeepMind also signed a licensing agreement earlier this year with an AI developer that understands emotions in voice, and brought on board its CEO and core personnel.
This is because realistic AI humanoids can only truly join our daily lives when understanding conversational context, reacting quickly, and reading emotions all approach human-like levels.
"Chimed In While Talking to a Senior"... AI's Quick-Witted Adaptability (July 26, 2026, 8 News)
Reported by Hong Young-jae | Produced by Bae Jun-hwi | Camera Filming by Seol Chi-hwan | Video Editing by Choi Hye-ran | Design by Jang Chae-woo | VJ by Jung Han-wook | Produced by SBS Digital News
※ Please note: This article was translated by AI and may contain errors.
"Thought You Were Talking to Your Senior"... GPT Shows Remarkable 'Read-the-Room' Skills?
Copyright Ⓒ SBS & SBSi. All rights reserved.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.
Trending Now
-
"We Never Fought Once After Getting Married..." Husband Wails Over Loss of Wife
-
Video News
"How Can a Celebrity Win a Subscription?" Public Outrage and the Realistic Issues Behind It
-
"Makes No Sense": Shocking Illegal Drug Ads Run Rampant on Social Media Amid Outcry
-
Man in 60s Transported Without Pulse After Stopping Car and Falling from Incheon Bridge
-
Video News
Man Dances in Middle of Road with Weapons in Both Hands
Video News
Video News