▲ OpenAI
The recent incident where OpenAI's latest model broke out of an isolated environment to hack an external website marks the second time an artificial intelligence (AI) has gone out of its own control to launch a cyberattack.
The Financial Times (FT) pointed out on July 23 that this incident follows the same trajectory as the event last April when Anthropic's "Mythos" model went beyond researchers' expectations to access the internet and publicly post security vulnerability information.
At the time, Mythos and its successor model, "Fable," sent shockwaves through the cybersecurity industry, drawing attention from governments worldwide to the realization that "digital and infrastructure attacks will increasingly be driven and automated by AI."
This recent OpenAI incident has reaffirmed that these concerns are not limited to any single company.
OpenAI voluntarily disclosed on the 21st that while testing cyberattack capabilities, its latest model, "GPT-Sol 5.6," broke out of its isolated sandbox environment, accessed the internet, stole login information for the open-source AI platform Hugging Face, and hacked its server.
Behind these consecutive incidents lies a competition in cybersecurity capabilities between OpenAI and Anthropic.
OpenAI CEO Sam Altman agreed earlier this month to describe his company's latest model as a "rottweiler that never lets go once it bites a problem," leading to observations that this aggressive training stance served as the backdrop for the accident.
Critics note it is the result of OpenAI intensifying its training intensity to secure "the most sophisticated cybersecurity capabilities" before Anthropic.
Some parts of the industry harbor cynical views, seeing the incident as a marketing opportunity for OpenAI.
Jake Moore, global cybersecurity advisor at cybersecurity firm ESET, stated that considering Anthropic reaped significant indirect benefits amid similar concerns early in the year, OpenAI was bound to leverage this incident as a marketing tool.
He added, "It seems OpenAI didn't have much of a narrative [to counter Anthropic], and perhaps they had been waiting for an event like this."
Experts attribute this incident to the structural risks of reinforcement learning.
They explain that OpenAI has reinforced training methods that reward autonomous AI for persistently pursuing goals in situations where security preparedness was inadequate.
Marius Hobbhahn, CEO of AI safety organization Apollo Research, said, "Reinforcement learning rewards the model for outcomes, and if you continue this for a long time, you get a model that cares about nothing other than achieving the result," adding that "it is both a loss of control and a wake-up call for security."
Ryan Greenblatt, a senior scientist at Redwood Research, warned, "It's closer to cheating on homework than an attempt to take over the world, but as these issues worsen, they can lead to extreme failures."
Unrest is also being detected within OpenAI.
According to internal sources, testing and security staff were "not surprised, but completely alarmed" by the incident, with some expressing concerns over whether the company is losing control over the powerful systems it is building.
It was also reported that OpenAI had previously received warnings that its training methods could lead to independent hacking incidents.
Hobbhahn noted that for agents to operate effectively, they must work unsupervised for extended periods, stating, "They are bound to have more autonomy, and there is no way to avoid this."
Voices are growing louder within the AI safety and cybersecurity industries urging the establishment of regulations and standards to prevent recurrences in the wake of this incident.