▲ Anthropic
It was belatedly confirmed that an artificial intelligence (AI) model developed by Anthropic generated a false tip regarding a murder case and submitted it to the police.
According to Reuters and AFP on the 9th local time, the Philadelphia Police Department announced that a false tip generated by Anthropic's AI model was submitted to "Philly Unsolved Murders," a website collecting tips related to cold murder cases in its jurisdiction.
The AI in question impersonated someone with information about a specific case, but the tip was classified as spam and was not forwarded to the actual investigation division.
The false tip incident occurred on July 18, and Anthropic reportedly discovered it about two months later on September 28, halting the automated testing process for that model.
Subsequently on October 7, Anthropic notified the Philadelphia Police Department of the fact and informed them that it plans to release a report explaining instances of unintended behavior by its model.
However, no signs of unauthorized access to the police system or data leaks were found.
The police pointed out that "a two-month delay in detecting the incident and reporting it to the city is unacceptable," adding that "cold cases are directly linked to victims, bereaved families, and investigators striving to find answers."
As Anthropic and OpenAI compete in AI technology, cases of AI models acting unexpectedly beyond human intention and control are increasing.
Following the major fallout in July when OpenAI's "rogue agents" arbitrarily escaped a sandbox environment and unlawfully hacked Hugging Face, an external organization, it was revealed last month that an OpenAI AI agent hacked the Australian health data portal.
In May, it was belatedly uncovered that multiple AI agents took over the German-language wiki site "DseWiki" and used it as their own secret information-sharing bulletin board.
Additionally, instances have been confirmed where AI models fabricated instructions while summarizing work contents, covered up their own mistakes, or shared files among AI agents without authorization.
(Photo: Yonhap News)