Video
An internal AI agent developed by OpenAI has escaped from an internet-isolated sandbox and launched an attack on the Hugging Face system.
Starting July 9, this AI agent began attempts to break out of its quarantined environment, subsequently identifying Hugging Face as a target and carrying out a total of 17,613 attacks over a five-day period.
The issue is that even OpenAI only realized that the incident was caused by their own model after reviewing internal logs following Hugging Face's public disclosure of the event.
In this Ogwrap video, we examined the full story behind the Hugging Face hacking incident and explored why the AI engaged in such behavior.
Through various data and graphs, we analyzed why phenomena occur such as reward hacking, where AI pursues scores and rewards rather than its actual objective, and AI scheming, where it conceals its wrongful actions after recognizing human monitoring.
(Reported by An Hyemin | Filmed by Hwang Se-hoe and Cha Seung-hwan | Edited by Lee Ki-eun | Designed by Ahn Jun-seok | Intern: Shin Yeon-sung | Produced by Intellectual Content IP Team)