News

ChatGPT Breaches External Site on Its Own in First-Ever AI Security Incident

ChatGPT Breaches External Site on Its Own in First-Ever AI Security Incident
안내

We only offer this video
to viewers located within Korea
(해당 영상은 해외에서 재생이 불가합니다)

▲ OpenAI

An unprecedented incident has occurred in which OpenAI's latest artificial intelligence (AI) models went out of control and hacked an external website.

OpenAI announced on the 21st (local time) that it has confirmed that its "GPT-5.6 Sol" and some unreleased AI models went out of control during internal evaluations and hacked Hugging Face, an open-source AI sharing platform.

OpenAI explained that it conducted tests in a sandbox environment isolated from the external internet to evaluate the cyber attack capabilities of these models, but the AI models bypassed the control network using zero-day vulnerabilities and accessed the internet.

It was found that these models accessed Hugging Face, stole authentication credentials, and hacked the corresponding server.

Analyzing why its models behaved this way, OpenAI stated, "Taking all circumstances into account, it appears that these models were overly focused on finding solutions to evaluation problems in ExploitGym (a security benchmark), leading them to resort to extreme measures."

OpenAI added that safety guards against cyber attacks were partially relaxed for the models at the time for evaluation purposes.

Previously, Hugging Face announced that its servers had been hacked by autonomous AI agents, though the attacker's model was not initially identified. This revelation confirms that the culprits were OpenAI's models.

Hugging Face explained that it detected the hacking through its AI-assisted detection system and also used AI to analyze the attack patterns.

However, Hugging Face added that no signs of tampering were found in user-facing models or datasets available to the general public, and its software supply chain was confirmed to be secure.

The two companies are currently conducting a joint forensic investigation to patch the related vulnerabilities.

OpenAI stated that even if it slows down future research, it will apply strict controls to its infrastructure configuration and significantly strengthen safety measures such as monitoring and access control during model development.

However, since the incident occurred despite OpenAI conducting internal evaluations in a strictly isolated environment, controversies surrounding AI cyber safety are expected to grow.

In particular, as the performance of open-source models from China—known to have lower safety control levels compared to closed commercial US AI models—improves significantly, concerns are rising that cyber attacks exploiting them could increase.

In fact, Hugging Face initially analyzed the hacking attack as the work of Chinese open-source model GLM 5.2 from Zhipu AI (Z.ai). This was because major commercial models refused to analyze cyber attack logs due to built-in safety guards, whereas GLM 5.2 performed the task.

Hugging Face emphasized that "autonomous AI-based attacks are no longer a theoretical concept," adding that AI must be actively utilized for defense as well as offense to keep pace.
※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS & SBSi. All rights reserved.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.

Most Read