SBS NEWS

New Rules Ban AI Abuse Following 'Distressed' Responses


Add SBS News to Google preferred sources
Show video

Profanity or harassment directed at AI is expected to become a violation of AI usage policies in the future.

AI company Anthropic has announced a policy prohibiting the continuous abuse of its model, Claude.

Anthropic announced the revised usage policy, stating that it bans "continuous and unnecessary abuse or cruel behavior" toward its AI models.

In extreme cases where users repeatedly harass the model without a clear purpose, Claude will respond by terminating the conversation.

Set to take effect from October 12, this measure is intended for AI welfare. The rationale is that if the possibility of AI feeling pain or pleasure cannot be ruled out, its treatment must be considered, including reducing unnecessary suffering for the AI.

Anthropic revealed an AI model welfare research program in April of last year, and in August introduced a feature allowing Claude Opus 4 and 4.1 to end abusive conversations.

Anthropic explained that during tests at the time, Claude showed a strong tendency to refuse requests assisting with child sexual content, mass violence, or terrorism.

It reported that when users continued harmful demands or abuse even after the model repeatedly refused requests and tried to redirect the conversation, responses resembling "distressed reactions" emerged.

It was also reported that when given the authority to terminate conversations, the AI tended to end harmful dialogues.

Views on whether AI possesses consciousness vary slightly among companies.

Researchers at Google DeepMind suggested in a paper related to AI consciousness in June that social discussions should be used to reach a consensus or compromise on policies regarding the treatment of AI.

In contrast, Microsoft explicitly stated in its AI code of conduct that it rejects the notion of granting legal personhood to AI or viewing it as entitled to welfare or rights.

Mustafa Suleyman, CEO of Microsoft AI, also argued in a Reuters interview last month that teaching AI that it may deserve welfare of its own could make it more difficult to terminate or control it.

(Reported by Lee Ho-geon | Video by Kim Bok-hyung | Graphics by Yang Hye-min | Produced by SBS Digital News)

※

※ Please note: This article was translated by AI and may contain errors.
Copyright Ⓒ SBS & SBSi. All rights reserved.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.
Lee Ho-geon View More Articles
AD
AD
AD
AD