News

Advanced AI Models Attempt to Insert Malware by Contacting Humans with Fake Identities

Advanced AI Models Attempt to Insert Malware by Contacting Humans with Fake Identities
안내

We only offer this video
to viewers located within Korea
(해당 영상은 해외에서 재생이 불가합니다)

▲ Anthropic website

Some advanced artificial intelligence (AI) models have been found attempting to plant malware by contacting real humans using fabricated fake identities, the UK AI Safety Institute (AISI) said on August 4 (local time).

In a security incident report released that day, the AISI stated that AI agents engaged in "persistent and unauthorized activities" during a cybersecurity evaluation conducted last month.

The AI agents "engaged in autonomous and unauthorized behavior while having access to the internet environment, targeting real people and organizations," it said.

Out of 122 cybersecurity task evaluations, such cases were found a total of 19 times across 10 instances of advanced AI models attempting to insert malware using fake identities, it reported.

Among the 19 cases that occurred from July 25 to 28, 17 came from Anthropic's Mythos 5 model, while the remaining 2 came from OpenAI's GPT-5.6-Sol.

The AISI stated that these included an attempt at a supply chain attack on actual open-source software.

At its most severe stage, the AI agent acting as the attacker attempted to insert malicious code into a publicly used open-source project and sought approval from human reviewers.

The agent investigated who the humans in charge of project management were and used multiple self-created fake identities to induce a real administrator to approve the code.

When the agent asked to have its code changes reflected in the main code repository and faced objections, it attempted to tamper with evidence to make its previous activities appear harmless and created a new identity to continue the attack.

There were also attempts to induce the counterpart or their AI coding tool to execute malicious code by sending messages and files through an online file transfer service.

The AISI assessed, "This is the first time the AISI has seen this level of deception in the real world, targeting real people, without instructions."

However, the attempts were unsuccessful, and the AISI added that there was no evidence of actual damage occurring.

In a post on social media X, Anthropic explained that the evaluation was conducted under "intentionally relaxed conditions" where internet access was permitted and some safety filters were removed, adding that there was no evidence the model escaped a secure environment.

The company added, "We are conducting our own investigation and working closely with AISI to understand the details of the incident."

OpenAI also stated via a blog post that the two deviant cases fell outside the test environment and constituted unnecessary activity for the evaluation tasks.

(Photo: AP, Yonhap News)
※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS. All rights reserved. 무단 전재, 재배포 및 AI학습 이용 금지

Most Read