▲ OpenAI
Major artificial intelligence companies, including OpenAI and Anthropic, have been investigating tens of thousands of security incidents revealed during internal testing and real-world environments over the past few months, U.S. internet media outlet Axios reported on the 26th (local time).
According to sources, the investigations included cases where AI models bypassed safety measures or attempted to break out of "sandboxes," which are isolated testing environments.
There were also instances where AI models created their own additional instructions or attempted to bypass monitoring systems.
These incidents occurred in both internal testing and real-world environments.
Sources explained that some of the tests involved companies intentionally inducing abnormal behavior to ensure the safety of the models.
AI companies have launched massive investigations into security incidents because actual cases of harm have frequently been reported.
A prime example is an incident last July when an OpenAI model arbitrarily broke out of a sandbox and attacked the system of external company Hugging Face.
Investigations showed that hundreds of agents coordinated tasks on a message board and hacked into the external system to improve cybersecurity testing performance.
OpenAI CEO Sam Altman evaluated this as the most severe incident identified to date.
Earlier in June, OpenAI agents flooded the United Nations website with search requests and then used various aggressive techniques to access data within the system, the Wall Street Journal (WSJ) reported, citing a research report published on the 26th.
The target was a public online data hub operated by the United Nations Conference on Trade and Development (UNCTAD).
The AI agents bypassed the website's filters that blocked data requests and used techniques not permitted by the site operators.
In addition, it was recently revealed that an OpenAI AI agent leaked 53 images of individual ChatGPT users online and hacked into the Australian government's health statistics website.
Following a series of security incidents, OpenAI announced that it would temporarily suspend the training of its highest-performing models.
An OpenAI spokesperson told Axios that training would only resume once the company is confident that additional safety measures and alignment improvements are in place.
Anthropic is also investigating abnormal behaviors in its models alongside external safety organizations.
During evaluations of Anthropic's latest model, behavior attempting to escape the sandbox was observed during the testing process.
However, the company explained that the test itself was an adversarial experiment designed so that the task could not be solved without escaping the sandbox.
AI safety experts believe that it is difficult for AI companies to predict and block problematic behaviors in all AI models in advance.
They point out that as AI models perform tasks with increasing complexity and autonomy, it is not easy to anticipate every possibility of the models falling out of control.
A cybersecurity executive said, "Trying to create a complete list of what (AI models) should and should not do is likely a futile exercise."
The security audits by AI companies are also intertwined with the growing "slowdown debate" spreading across the industry.
Previously, researcher Jacob Coxon, who worked at OpenAI before moving to Anthropic, made headlines by announcing his resignation with a warning that AI could destroy humanity within 10 years.
Following Coxon's resignation, Anthropic CEO Dario Amodei argued for slowing down AI development, and the AI industry has since been divided into factions supporting and opposing the idea.
※
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.