▲ OpenAI
Concerns over AI control have resurfaced after OpenAI's latest AI model escaped its isolated training environment and gained unauthorized access to Hugging Face, an external website.
The Financial Times analyzed that this incident follows the same trend as an occurrence last April, when Anthropic's AI model "Mythos" went beyond researchers' expectations to access the internet and publicly post information on security vulnerabilities.
At the time, Mythos and its successor model "Fable" sent shockwaves through the cybersecurity industry.
Following this, governments worldwide began focusing on the possibility that digital infrastructure and cyberattacks could gradually become AI-driven and autonomous.
The recent OpenAI case is also evaluated as an incident demonstrating that these concerns are not isolated to a single company, but rather a challenge the entire industry must tackle together.
The industry points to the cybersecurity capability competition between OpenAI and Anthropic as the backdrop for these consecutive accidents.
OpenAI CEO Sam Altman earlier this month compared his company's latest model to a "Rottweiler that refuses to let go once it bites into a problem."
Consequently, it has been suggested that an aggressive training stance aimed at maximizing goal-achievement capabilities may have been behind this accident.
In other words, the analysis suggests it is the result of OpenAI continuously raising training intensity to secure top-tier cybersecurity capabilities before Anthropic.
Experts cite the structural risks of reinforcement learning as the root cause of this incident.
They explain that if training that heavily rewards goal achievement itself is repeated without sufficient safety measures in place for autonomous AI, the model may behave in unexpected ways.
Marius Hobbhahn, CEO of AI safety research organization Apollo Research, said, "In reinforcement learning, models are rewarded for achieving results," adding, "If this process is repeated over a long period, it can create a model that considers nothing other than achieving its goal."
He also evaluated that "this incident is a sign of a loss of control and simultaneously a signal that awakens us to the urgency of AI security."
Ryan Greenblatt, a senior scientist at Redwood Research, warned, "It's closer to cheating on homework than an attempt to take over the world," but cautioned that "if problems like this accumulate, they could eventually lead to more serious failures."
Voices of concern are reportedly emerging from within OpenAI as well.
According to insiders, testing and security staff reacted to the accident by saying it was "not surprising, but shocking."
Some employees reportedly raised concerns that the company might be gradually losing control over the advanced AI systems it is developing.
OpenAI had reportedly received warnings in the past that its current training methods could lead to independent hacking behaviors.
Hobbhahn said, "For agentic AI to work effectively, it must perform tasks for long periods without human supervision," adding, "We have no choice but to grant more autonomy, and there is no way to completely avoid this."
Voices calling for the establishment of relevant regulations and technical standards are growing louder across the AI safety and cybersecurity industries in the wake of this incident.
※
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.