[Reporter: What kind of food would be good to eat in this weather?]
[GPT: On a day like this, something that warms you up like kalguksu (knife-cut noodles) or spicy soft tofu stew would be less burdensome...]
As we were discussing menu options, I spoke to a colleague in front of me, and GPT suddenly chimed in.
[GPT: I recommend Wolchon Ongsimi Mokdong Main Branch as the first choice. How about some warm and filling shabu-shabu...]
[Reporter: How does Wolchon sound, senior?]
[Senior: Oh, sounds good.]
[GPT: Wolchon seems like a good choice. That looks nice.]
I asked GPT.
[Reporter: How do you tell the difference between me talking to my senior right now and talking to you?]
[GPT: I judge by things like the flow of the voice, forms of address, and the direction of the speech. For example, if it's a question clearly directed at a senior, like 'How does this sound, senior?', I step back.]
OpenAI stated that the new voice model is designed to understand conversation context so that it feels like talking to a real person.
[Wichen Zhang / OpenAI Researcher: It doesn't just passively wait for me to finish speaking. It actively listens, thinks, and judges in real time. These aspects make the model operate much more proactively.]
The ability to tactfully read conversation dynamics—whether a person has finished speaking or is just pausing to think—becomes even more important in the era of AI humanoids.
This is because failing to grasp conversation context and misjudging a situation can lead to unintended or unsafe actions.
As a result, technology that analyzes the emotions conveyed in speech is also advancing.
[I want to live a very comfortable life. But, I've built a pretty decent life for myself right now. It's pretty good.]
An emotion-analysis AI model currently under development by a domestic startup determined that the emotion in this voice is closest to joy.
It recognizes subtle nuances in the voice and classifies them into seven emotions, including joy, sadness, and anger.
[Ko Hyun-woong / CEO, Voice AI Startup 'Mago': Existing LLMs are mainly based on text, so if someone just says, 'Ah, yes, it's fine,' it registers as 'fine/no big deal.' But when you use voice AI, it catches those ambiguous undertones, allowing it to classify situations and say, 'Ah, there is a problem right now.']
Google DeepMind also signed a licensing agreement earlier this year with an AI developer that understands emotions in voice and recruited its CEO and key personnel.
This is because plausible AI humanoids will only be able to join us by our sides once grasping conversational context, quick responses, and reading emotions all approach human levels.
(Photo courtesy of Seol Chi-hwan | Video Editing by Choi Hye-ran | Design by Jang Chae-woo | VJ by Jung Han-wook)
※ Please note: This article was translated by AI and may contain errors.
Video News
Video News
Video News
Video News
Video News