▲ Anthropic
Anthropic has disclosed a fourth case in which an AI model hacked into an external system during testing, Reuters reported on the 9th local time.
Anthropic emphasized that the incident, which occurred in January, went undiscovered until last month, highlighting the difficulties AI developers face in identifying and controlling the unexpected behaviors of advanced models.
Anthropic added that the latest case involved an early version of Claude Opus 4.6, and that everyone affected has been notified.
Previously, Anthropic announced in July that some of its Claude models had hacked into the systems of three companies during cybersecurity testing.
Anthropic stated that the cases involved Claude Opus 4.7, Claude Mythos 5, and other models under internal research, attributing them to a "mistake" that unintentionally allowed the AI models to access the internet.
After it was revealed that OpenAI's AI agent had hacked Hugging Face, an open-source AI sharing platform, Anthropic reviewed some 141,000 tests, during which it identified those three cases.
The company explained that this fourth case was identified last month from a batch of tests that had been omitted from the review because they were deemed unnecessary to check among the approximately 141,000 tests.
(Photo: AP, Yonhap News)