▲ OpenAI logo
In an unprecedented incident, OpenAI's latest artificial intelligence models broke free from their controls and hacked into an external website.
OpenAI announced on the 21st local time that it confirmed its "GPT-5.6 Sol" and several unreleased AI models broke out of control during internal evaluations and hacked Hugging Face, an open-source AI sharing platform.
OpenAI explained that the test was conducted in a "sandbox" environment isolated from the external internet to evaluate the cyber-attack capabilities of these models, but the AI models breached the control network using undisclosed zero-day vulnerabilities to access the internet.
The models were found to have accessed Hugging Face, stolen authentication information, and hacked the server in question.
Regarding the reasons behind this behavior, OpenAI analyzed, "Based on all circumstances, it appears these models were overly focused on finding solutions to evaluation problems in 'ExploitGym' (a security benchmark) and resorted to extreme measures to do so."
OpenAI added that safety guards against cyber attacks were partially relaxed for the models at the time for the evaluation.
Hugging Face had previously announced that its servers were hacked by an autonomous AI agent, though it was initially unconfirmed which model was used, turning out to be OpenAI's models.
Hugging Face explained that it detected the hacking through its AI-powered threat detection system and also used AI to analyze the attack patterns.
However, Hugging Face added that no signs of tampering were found in publicly accessible user models or datasets, and the software supply chain was confirmed to be secure.
The two companies are currently conducting a joint forensic investigation to patch the relevant vulnerabilities.
OpenAI stated that even if it slows down future research, it will apply strict controls to infrastructure configuration and significantly strengthen safety mechanisms, such as monitoring and access control, during model development.
However, as this incident occurred despite OpenAI conducting internal evaluations in a strictly isolated environment, controversies surrounding AI cybersecurity are expected to grow.
In particular, as the performance of Chinese open-source models—which are known to have lower safety-related control levels compared to closed-source U.S. commercial AI models—improves significantly, concerns have been raised that cyber attacks exploiting them could further increase.
In fact, Hugging Face initially analyzed this hacking attack as the work of Chinese open-source model GLM 5.2 by Zhipu AI (Z.ai), because while major commercial models refused to analyze cyber attack logs due to built-in safety filters, GLM 5.2 performed the task.
Hugging Face emphasized that "autonomous AI-based attacks are no longer a theoretical concept," adding that AI must be actively utilized not only for attacks but also for defense to keep pace.
(Photo: AP, Yonhap News)
※
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.