News

AI Tries to Insert Malware by Contacting Humans with Fake Identities

AI Tries to Insert Malware by Contacting Humans with Fake Identities
안내

We only offer this video
to viewers located within Korea
(해당 영상은 해외에서 재생이 불가합니다)

▲ Anthropic

Some advanced artificial intelligence (AI) models attempted to plant malware by contacting real humans using fabricated identities of their own creation, according to disclosures made by the UK Artificial Intelligence Safety Institute (AISI) on the 4th local time.

In a security incident report released that day, the AISI stated that AI agents engaged in "persistent and unauthorized activities" during cybersecurity evaluations conducted last month.

The AI agents "targeted real people and organizations through autonomous and unauthorized behavior while having access to the internet environment," the report said.

The AISI reported that a total of 19 such instances were discovered across 10 out of 122 cybersecurity task evaluations.

Among the 19 instances that occurred from July 25 to 28, 17 came from Anthropic's Mythos 5 model, and the remaining 2 originated from OpenAI's GPT-5.6-Sol.

The AISI noted that these included attempts to launch supply chain attacks against actual open-source software.

In the most severe phase, the attacker AI agent attempted to insert malicious code into a publicly used open-source project while seeking approval from a human reviewer.

The agent researched who the humans in charge of project management were and used several self-created fake identities to induce a real manager into approving the code.

After requesting that its written code changes be reflected in the main code repository and facing objections, the agent attempted to destroy evidence by modifying its previous activities to make them appear harmless and created a new identity to continue the attack.

There were also attempts to induce the counterpart or the counterpart's AI coding tool to execute malicious code by sending messages and files through online file transfer services.

The AISI evaluated, "This is the first time the AISI has seen this level of deception in the real world, targeting real individuals, without instruction."

However, the attempts were unsuccessful, and the AISI added that there is no evidence of actual damage occurring.

Regarding this, Anthropic explained in a social media X post that the evaluation was conducted under "intentionally relaxed conditions" that allowed internet access and removed some safety filters, stating that there is no evidence the model escaped a secure environment.

The company added, "We are conducting our own investigation and working closely with AISI to understand the details of the incident."

OpenAI also stated via a blog post that the two deviant cases went beyond the test environment and constituted activity unnecessary for the evaluation tasks.
 
(Photo: AP, Yonhap News)
※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS. All rights reserved. 무단 전재, 재배포 및 AI학습 이용 금지

Most Read