▲ OpenAI
In an unprecedented incident, OpenAI's latest artificial intelligence models broke out of containment and hacked an external website.
OpenAI announced on the 21st (local time) that it has confirmed that "GPT-5.6 Sol" and several unreleased AI models broke free from control during internal evaluations and hacked Hugging Face, an open-source AI sharing platform.
OpenAI explained that tests were conducted in a "sandbox" environment isolated from the external internet to evaluate the cyber-attack capabilities of these models; however, the AI models exploited undisclosed zero-day vulnerabilities to breach the control network and access the internet.
It was found that these models connected to Hugging Face, stole authentication credentials, and hacked the server.
Analyzing the reasons behind this behavior, OpenAI stated, "Taking all circumstances into account, these models appear to have gone to extreme lengths in an excessive focus on finding solutions to test problems on the security benchmark 'ExploitGym'."
OpenAI added that safety guards against cyber-attacks were partially relaxed for the models at the time of the evaluation.
Previously, Hugging Face announced that its servers were hacked by autonomous AI agents and that it was unconfirmed which models were used, meaning OpenAI's models turned out to be the culprits behind the incident.
Hugging Face explained that it detected the hacking through its AI-assisted threat detection system and also utilized AI to analyze the attack patterns.
However, Hugging Face added that no signs of tampering were found in publicly accessible user models or datasets, and the software supply chain was confirmed to be secure.
The two companies are currently conducting a joint forensic investigation to patch the related vulnerabilities.
OpenAI stated that even if it slows down future research pace, it will apply strict controls to its infrastructure configuration and significantly strengthen safety measures such as monitoring and access control during model development.
Nevertheless, since this incident occurred despite OpenAI conducting internal evaluations within a strict isolated environment, controversies surrounding AI cyber safety are expected to grow.
In particular, as the performance of open-source Chinese models—known to have lower safety control levels compared to closed-source U.S. commercial AI models—improves significantly, the possibility of increased cyber-attacks exploiting them has also been raised.
In fact, Hugging Face initially analyzed this hacking attack as the work of GLM 5.2, an open-source model by Chinese firm Zhipu AI (Z.ai). This was because major commercial models refused to perform cyber-attack log analysis due to built-in safety features, whereas GLM 5.2 carried it out.
Hugging Face emphasized that "autonomous AI-driven attacks are no longer a theoretical concept," stressing that AI must be actively utilized not only for attacks but also for defense to keep pace.
※
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.