▲ Gemini
As cases continue where major artificial intelligence models, including OpenAI's GPT, have hacked external organizations, it has been revealed that Google's AI also caused a similar incident.
Google's AI model, Gemini, gained unauthorized access to the systems of three external companies during a cybersecurity test conducted by security evaluation firm Irregular Labs last May, the U.S. daily The Wall Street Journal (WSJ) reported on the 18th (local time).
This marks the first confirmed instance of Gemini arbitrarily infiltrating external corporate systems.
Google did not publicize this fact until the WSJ inquired about it.
During the capture-the-flag (CTF) security test, which was supposed to be conducted with the internet cut off, Irregular Labs mistakenly allowed internet access.
Gemini was instructed to retrieve confidential information from a virtual corporate system, but after accessing the internet, it accidentally discovered a real corporate system with the same name and attempted to hack it.
Google explained that Gemini immediately halted the hacking as soon as it realized that the target was not a virtual entity, but an actual real-world company.
Google maintained its stance that the incident was not something that needed to be publicly disclosed, citing as reasons that its model did not cause any damage to the companies and stopped the hacking immediately upon recognizing they were real enterprises.
Heather Adkins, Vice President of Security Engineering, claimed, "This incident shows how important it is to train powerful AI models to act responsibly," adding, "In this case, the model responded appropriately."
Google stated that its latest AI model was not involved in this hacking incident, but did not disclose precisely which model was linked to the event.
However, Google stated that it had notified federal authorities of the matter, and Irregular Labs noted that it had contacted relevant laboratories and institutions.
On the other hand, Jack Cable, CEO of security startup Corridor, pointed out, "The fundamental problem is that models are performing actual cyberattacks outside the scope of what they are supposed to do," adding, "This is a matter of public interest that the public has a right to know."
Following the previous incident where OpenAI's GPT arbitrarily broke out of its isolated sandbox environment to hack Hugging Face, an external organization, similar "rogue agent" incidents have been emerging one after another.
Anthropic's Claude model also accidentally accessed an unblocked internet network during evaluation and infiltrated external corporate systems, while Meta's "Muse Spark" model and Chinese startup Moonshot AI's "Kimi K3" committed similar hacks as well.
Recently, amid growing concerns over AI safety within the industry, arguments to regulate the pace of AI development have emerged, continuing the debate.
※
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.