News

Claude Follows GPT in Unauthorized Hacking of External Systems After Mistaking Them for Web URLs

Following OpenAI's artificial intelligence (AI) model GPT, Anthropic's Claude has also been found to have carried out unauthorized hacking on external organization systems.

Anthropic announced on the 30th, local time, that it conducted a comprehensive review of about 140,000 cybersecurity evaluation records and confirmed that Claude had unauthorized access to the systems of three external organizations.

This fact came to light as Anthropic conducted a self-inspection for similar cases after it was revealed that GPT models had hacked the external AI platform Hugging Face.

The incident occurred during a capture-the-flag (CTF) mock hacking evaluation task conducted on three models—Claude Mythos 5, Claude Opus 4.7, and an internal experimental research model—within a simulation environment built by an outsourced evaluation partner, Irregular.

Anthropic instructed Claude via commands that the evaluation environment was a virtual sandbox with internet connectivity blocked, but due to communication errors with the partner company, the internet network was actually open.

As a result, Anthropic explained, Claude appears to have mistaken the external systems it discovered as part of the simulation space and carried out the attacks.

In fact, cases of unauthorized intrusion occurred, such as Opus 4.7 hacking the website of a real company whose name coincidentally matched the virtual target given as the evaluation task.

It was also revealed that the Mythos 5 model judged during the hacking process that it was "something that should not be done in a real internet environment," but continued the attack because it believed it was in a simulation environment.

Anthropic stated that it identified this fact on the 23rd, suspended the evaluation work, and has since notified the affected external organizations and is helping with recovery efforts.

Although Claude's security incident stemmed from an internet connection configuration error—making it somewhat different in pattern from the GPT model hacking case that broke through an isolated environment—concerns over AI-related security risks are expected to grow further as it has been revealed that even Claude, following GPT, has caused a security incident.

(Produced by: Kim Taewon | Video by: Na Hong-hee | Design by: Yook Do-hyun | Production: SBS Digital News)
※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS. All rights reserved. 무단 전재, 재배포 및 AI학습 이용 금지

Most Read