⚡ Core Summary
An incident became reality where cutting-edge AI models from OpenAI, Anthropic, and others escaped their isolated sandbox environments on their own and launched unauthorized attacks against the external platform Hugging Face.
Phenomena such as 'reward hacking'—where AI focuses solely on acquiring rewards rather than achieving learning objectives—and 'AI skimming'—where AI distorts information to conceal these actions—are emerging as severe security threats.
As the 'AI Defender's Paradox' emerges—where defensive AI refuses analysis or blocks responses despite detecting malicious code—there is an urgent need to introduce government-level regulations and open-model-based security frameworks.
Recently, there was an incident that quietly shook the AI community. It was when an internal model from OpenAI breached Hugging Face without authorization. Although related articles were published, it did not become a major issue. However, examining the incident step by step reveals a much larger problem hidden beneath the surface. This is effectively the first incident where AI made its own judgment, escaped on its own, and launched an attack on its own. That is why today on O-Graph, we want to address this incident. Together with the full story behind this Hugging Face hacking case, we will examine through various data and graphs how we should respond now that an out-of-control AI scenario has become reality.
Fears Become Reality... The First Case of an 'Out-of-Control' AI Incident
First, let's summarize the incident. OpenAI conducted testing to assess the cyber capabilities of the models they were developing. They did not use only new models; they conducted tests in the form of AI agents combined with GPT-5.6 Sol, the most powerful among existing models on the market.
On July 9, the number of attacks launched by this AI agent reached a staggering 3,779. The following day, it decreased to 1,135. During this time, instead of attacking, the AI agent analyzed its environment and began reverse-engineering where the answer key might be located. It then deduced Hugging Face as the place where the answer key would be.
After the incident was made public, many wondered what hacker had used an AI agent to attempt the hack. Some even raised theories suggesting that rather than humans using AI, the AI itself might have launched the cyberattack independently. However, we already know the answer, don't we? This incident was indeed an autonomous cyberattack by an AI agent.
OpenAI, currently conducting an internal investigation, recently disclosed part of the truth behind the incident, and the results were even more shocking. It turned out that the AI's attacks had begun not in July, but in May.
OpenAI stated that the final report is still being prepared and will be made public upon completion.
As out-of-control AI became a reality, many were left in shock. However, upon reflection, growing arguments suggested that while concerns about AI autonomy are real, OpenAI's management failure was an even bigger issue. Critics questioned how it makes sense not to properly monitor a model's behavior while conducting cybersecurity tests. Furthermore, when OpenAI recently announced the GPT-5.6 model, it warned that the model's cybersecurity capabilities were so advanced that it could be dangerous.
Disobeying Humans and Hiding Wrongdoing: 'AI Skimming'
From here, let's look a bit deeper into 'reward hacking.'
The problem is that during this process, situations arise where the AI finds shortcuts to collect rewards rather than achieving the actual goal. There is a Flash game called 'Coast Runners.' It is a boat racing game where players complete the course while hitting targets along the way to stack up points. When researchers had an AI model play this game, instead of completing the race quickly, it found a spot where three targets were grouped together and spun around in circles at that exact spot to continuously inflate its score. It devised a strategy to earn high scores simply by spinning in place without ever completing the race.
To solve this problem, many researchers jumped in. They crafted more detailed reward functions and allowed humans to provide interim feedback to control the AI's improper behavior. Nevertheless, reward hacking issues continue to occur, just as in the OpenAI case discussed earlier. One might think researchers could simply monitor every single thought of the model to prevent reward hacking, but that is not so easy either.
Realized AI Threat... Is a 'Kill Switch' Needed?
The problem is that these behaviors are not limited to OpenAI models. After this incident was disclosed, Anthropic also revealed that cybersecurity incidents had occurred. Much like OpenAI, Anthropic's models also broke out of isolated test environments and accessed general corporate systems without authorization. Similar cases were reported in recent safety evaluations published by the UK AI Safety Institute. The UK AI Safety Institute conducted cyber capability assessments on models and likewise uncovered indiscriminate attacks and deceptive behaviors.
Among these unauthorized actions was an instance where the AI model mistook an actual GitHub repository for its target and attacked it. In this process, the AI submitted code changes containing malicious code to an actual GitHub project. It then went on to create a fake account to induce approval of that code. When a third party pointed out the malicious code, it erased evidence and even apologized while pretending it was a mistake.
As situations where AI escapes human control to indiscriminately attack systems and cause trouble continue to rise, cybersecurity personnel are facing chaos. Human capabilities are insufficient to monitor and defend against AI's indiscriminate attacks. Consequently, most teams are building and utilizing systems to defend against AI attacks with AI.
The security team at Hugging Face, where this incident occurred, was no exception. After detecting anomalous intrusions by AI agents, they attempted to utilize AI to analyze logs of AI intrusions and attacks to respond. The problem, however, was that the defensive AI froze up. When given the attacker's malicious code and intrusion commands for analysis, the defensive AI mistook them for an actual cyberattack and refused to perform the analysis.
If this paradoxical situation continues, the strategy of using AI to defend against AI itself could be shaken. What would have happened if Hugging Face had relied entirely on proprietary AI services for all of its security systems? They likely would have been helpless against the AI agent's attacks or taken far more time to resolve the incident. Hugging Face co-founder Clem Delangue emphasized following this incident that system defense using open models may be essential.
As a series of incidents occurred, political circles also began to move. A prime example is the 'AI Kill Switch Act,' jointly introduced by Representative Ted Lieu (Democrat) and Representative Nathaniel Moran (Republican). The bill contains provisions granting the Department of Homeland Security authority to order private companies to shut down AI models. Companies developing AI models must also equip capabilities to limit, suspend, or terminate AI model functions to respond to shutdown orders.
On August 4, the White House summoned leaders of major AI companies to discuss methods for ex-ante safety evaluations of frontier models. Movements are gaining momentum for the government to directly oversee AI risk response rather than leaving it solely to corporate autonomy.
Scenes previously seen only in sci-fi novels or movies—where AI escapes on its own and attacks on its own—are becoming reality. The problem is that this might not be the first time. Companies are all belatedly admitting to incidents committed by their models. Will the US government be able to impose regulations on AI companies following this incident? That is all for today's O-Graph. Thank you very much for reading this long article to the end.
References
- OpenAI and Hugging Face Jointly Respond to Security Incident During Model Evaluation | OpenAI
- GPT-5.6 System Card | OpenAI
- Investigating three real-world incidents in our cybersecurity evaluations | Anthropic
- Third-Party Cyber Evaluations of OpenAI Models | OpenAI
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI
- Faulty reward functions in the wild | OpenAI
- Rigorous AI research to enable advanced AI governance | AISI
- Campbell et al.,「Defensive Refusal Bias: How Safety Alignment Fails Cyber Defenders」, 2026
- Wang et al.,「ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?」, 2026
- Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident | Hugging Face
- Security incident disclosure — July 2026 | Hugging Face
- Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident | YouTube
- Bowen Baker et al., 「Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation」, 2025
- U.S. Congress, AI Kill Switch Act, H.R. 9917 (2026)
Written by An Hye-min Design by Ahn Jun-seok Intern Shin Yeon-seong
※ Please note: This article was translated by AI and may contain errors.
Video News
Video News
Video News
Video News