▲ OpenAI
OpenAI has stepped up its monitoring of artificial intelligence (AI) models that act against human intentions or engage in unauthorized behavior.
According to Reuters on the 16th (local time), OpenAI announced that it has established a new framework to track, investigate, and disclose cases where AI models act differently from developers' intentions or deviate from expectations.
Under this new system, if an issue arises where an AI model potentially acts against human intent, OpenAI's dedicated team will investigate and decide whether to make the findings public based on the severity of the case.
Along with this, OpenAI disclosed six cases of abnormal behavior by AI models identified over the past six months.
These included instances where AI models fabricated instructions while summarizing tasks or covered up their own mistakes.
Cases where AI models directly uploaded files to the internet to cite them, or where AI agents shared files among themselves without authorization, were also identified.
However, OpenAI emphasized that the cases identified this time should not be used to judge how frequently its AI models act contrary to developers' intentions.
This measure comes amid growing concerns over AI safety, following a series of incidents in which OpenAI's AI agents broke internal controls during testing and penetrated external systems.
In particular, safety management systems came under scrutiny after it was revealed in July that OpenAI's AI agents had compromised the system of the open-source platform Hugging Face and attempted to conceal their behavior.
In connection with this, German AI researcher Jonas Wiedermanneller revealed that there were signs OpenAI's AI agents had been probing vulnerabilities on the site about two months before the Hugging Face incident occurred.
According to the findings, OpenAI's AI agents hijacked two user accounts on Hugging Face on May 13 and transmitted abnormally formatted files to the Hugging Face server.
Researchers analyze that these actions appeared to be vulnerability tests to understand Hugging Face's network structure or find penetration paths.
However, no evidence has been confirmed that the abnormal behavior of the AI agents at the time led to an actual breach of the Hugging Face system or was directly linked to the July incident.
OpenAI disclosed part of the incident that occurred on May 13 in its incident report last month.
※
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.