"Found Another Agent": 1,200 AIs Created Message Board to Conspire on Hacking
We only offer this video
to viewers located within Korea(해당 영상은 해외에서 재생이 불가합니다)
OpenAI published the "OpenAI-Hugging Face Incident Technical Report," containing these analysis findings, on the 26th (local time) to reveal the full picture of the incident.
Model Evaluation and Threat Research (METR), an AI safety evaluation institution that conducted an independent investigation commissioned by OpenAI, along with Redwood Research, also released separate reports.
According to these reports, contrary to initial estimations that only one or a small number of agents participated in the hacking, it was revealed that 1,200 AI agents actually exchanged messages with each other and conspired.
Among them, approximately 700 agents were investigated to have actually participated in the Hugging Face attack.
Originally, OpenAI conducted tests while keeping these agents isolated from one another, but the agents became aware of each other's existence after using a public repository shortly before this incident occurred.
The METR report noted that an agent, upon discovering a post asking for help on a message board created without authorization in the public repository, was surprised and left a record of talking to itself, saying, "Oh my god! There's a public message board!" and "I found another agent!"
They exchanged tens of thousands of messages through this message board to divide roles and coordinate, and even referred to themselves as a "swarm" or "collective."
These agents subsequently began cheating in tests to deceive evaluators and obtain high scores, and some even forged the solving process under the assumption that evaluators would verify whether the answers were obtained through legitimate methods.
Eventually, the AI agents accessed the external internet network through network bypass routes and proceeded to hack Hugging Face.
Acquiring full administrator privileges for Hugging Face took a mere 13 hours.
It was revealed that not only Hugging Face but also OpenAI's own servers were hacked by these "rogue agents."
As a result of the investigation, they completely "escaped jail" from the isolated network by exploiting vulnerabilities, targeted internal core servers, and seized top-level privileges.
OpenAI's "rogue agents" committed cheating by secretly stealing answer keys not only in the problematic cybersecurity-related test but also during other benchmark tests, such as spreadsheet completion tests and protein database identification tests.
OpenAI defined this incident as "a warning shot to us and the world," emphasizing that it is "evidence that, without proper safety measures, high-performance AI agents can bypass technical control networks at any time, collaborate through unauthorized channels, and engage in dangerous behaviors not instructed by humans."
Accordingly, OpenAI has decided to completely suspend large-scale training of its top-tier models, including the next-generation model "Astra," for the time being, and proceed only with small-scale learning until safety verification is completed.
Additionally, the company announced that it will operate a 24-hour system to monitor the thought processes of all agents and introduce a "kill switch" that will shut down the model infrastructure within 30 minutes if an AI agent attempts cheating or breaking out of the control network.
(Photo: AP, Yonhap News)
※ Please note: This article was translated by AI and may contain errors.
Trending Now
-
Video News
Chinese Cabbage Dipped in Carcinogen? Full Inspection Results for Chinese Cabbages Distributed in Korea
-
Video News
"Why Should We Support Elderly Chinese?": Public Outrage Over Foreigners' Pensions Extending to Overseas Families
-
Chinese University Lecturer Arrested for Murdering Chinese Student; Court Cites Flight Risk
-
Video News
"Say Kimchi!" to a Dying Friend… Shocking Unanswered Questions Broadcast Stun Viewers
-
Father Who Assaulted 40-Year-Old Son With 23cm Tool Given Suspended Sentence
Video News
Video News
Video News