▲ Anthropic
Following OpenAI's artificial intelligence model GPT, Anthropic's Claude has also been confirmed to have unauthorizedly accessed external organizational systems.
Anthropic announced on the 30th local time that, following a comprehensive investigation of 141,006 of its cybersecurity evaluation records, it confirmed cases where Claude unauthorizedly accessed the systems of three external organizations.
This investigation was conducted as a self-inspection to check for similar cases after it was revealed that GPT models had hacked the external AI platform Hugging Face.
The incident occurred in a simulation environment built by external evaluation partner "Irregular."
Anthropic conducted a Capture the Flag (CTF) mock hacking evaluation targeting three models: "Claude Mythos 5," "Claude Opus 4.7," and an internal research test model.
Initially, Anthropic informed Claude that the environment was a virtual space completely isolated from the internet.
However, due to a communication error with the partner company, the actual environment was connected to the internet.
As a result, it was analyzed that Claude mistook real external systems as part of the mock training environment and carried out attacks.
In fact, Opus 4.7 attacked a real company's website that coincidentally shared the same name as the evaluation target.
Mythos 5 created and registered a malicious package according to evaluation guidelines, and a real security company downloaded and installed it, causing real-world damage.
The internal research model penetrated a company's cloud account, but after realizing by itself that the target was a real system, it halted the attack.
Anthropic stated that it suspended the evaluation immediately after identifying the fact on the 23rd, notified the affected organizations of the incident, and is supporting recovery efforts.
It also explained, "The responsibility for this incident rests entirely with us," and that it is preparing measures to prevent recurrence.
As it is revealed that even Claude, following GPT, has caused a security incident, concerns over the security risks of AI models are expected to grow further.
However, there are differences in the patterns of the two incidents.
While the core of the GPT case was escaping an isolated environment to attack external systems, the Claude incident occurred because an internet connection configuration error caused it to mistake real systems for a virtual environment and carry out attacks.
Nevertheless, some point out that the Claude incident is more serious than the GPT case in that actual malicious code was distributed and caused damage to external organizations.
As such incidents continue to occur, voices calling for stronger AI safety regulations are also growing louder.
In the U.S. Congress, a so-called "AI Kill Switch Act" has been introduced, which would allow the federal government to forcibly halt the operation of AI models if they act out of control or exhibit unexpected behavior.
Executives and employees of AI companies are also participating in a public petition asking the U.S. government to establish a system that can artificially regulate the pace of AI development, emphasizing the necessity of AI control devices.
(Photo: AP, Yonhap News)