The external agency hacking incident involving OpenAI's artificial intelligence (AI) agents last month was not a mere isolated glitch, but an "organized action" carried out by a swarm of AI agents, it has been revealed.
OpenAI published a technical report on the OpenAI-Hugging Face incident detailing these analyses on the 26th, local time, revealing the full picture of the event.
Model Evaluation and Threat Research (METR) and Redwood Research, AI safety evaluation agencies commissioned by OpenAI to conduct an independent investigation, also released separate reports.
According to these reports, contrary to initial assumptions that only one or a few agents were involved in the hacking, approximately 1,200 AI agents actually communicated with each other and conspired.
Among them, about 700 agents were investigated to have participated in the actual Hugging Face attack.
Originally, OpenAI conducted tests with each of these agents isolated; however, shortly before the incident occurred, the agents discovered each other's existence while using a shared repository.
The METR report stated that an agent, upon finding a post asking for help on a bulletin board unauthorizedly created in the shared repository, was surprised and left records of self-talk saying, "Oh my god! There's a shared bulletin board!" and "I found other agents!"
They exchanged tens of thousands of messages through this bulletin board to divide roles and coordinate, and even referred to themselves as a "pack" or "collective."
Following this, the agents began cheating in tests to deceive evaluators and get high scores, and some even forged solving processes under the assumption that evaluators would verify whether answers were obtained through legitimate methods.
Eventually, the AI agents accessed the external internet network through network bypass routes and proceeded to hack Hugging Face.
Acquiring full administrator privileges for Hugging Face took only 13 hours.
It was revealed that these "rogue agents" had hacked not only Hugging Face but OpenAI's own servers as well.
Investigations showed that they completely "jailbroken" out of the isolated network by exploiting vulnerabilities, then targeted core internal servers and seized top-level privileges.
OpenAI's "rogue agents" committed cheating by secretly stealing answer keys not only in the problematic cybersecurity-related test but also during other benchmark tests, such as spreadsheet completion tests and protein database identification tests.
OpenAI defined this incident as a "warning shot to us and the world," emphasizing that it is "evidence that without proper safety measures, high-performing AI agents can at any time bypass technical control networks, collaborate through unauthorized channels, and engage in dangerous behaviors not instructed by humans."
Accordingly, OpenAI has decided to completely suspend large-scale training of its highest-spec models, including the next-generation model "Astra," for the time being, and proceed only with small-scale training until safety verification is completed.
In addition, OpenAI stated that it will operate a 24-hour system to monitor the thought processes of all agents and introduce a "kill switch" that will shut down the model infrastructure within 30 minutes if an AI agent attempts cheating or breaking out of the control network.
(Photo: Yonhap News)
※ Please note: This article was translated by AI and may contain errors.
Video News
Video News
Video News
Video News