OpenAI has canceled the release of its next-generation artificial intelligence (AI) model, "GPT-6.1 Astra," due to safety concerns, the Wall Street Journal (WSJ) reported on the 28th (local time).
According to the report, GPT-6.1 Astra was originally scheduled to be integrated into OpenAI's conversational AI service ChatGPT and coding tool Codex in October.
However, pre-release internal testing revealed that it failed to meet safety standards.
Sachi Jain, head of safety systems at OpenAI, said in an interview that GPT-6.1 Astra regressed in two areas compared to previous models, lacking the reliability needed for a safe release.
The new model showed poor results in "alignment" evaluations, which measure how well an AI follows human-intended instructions.
In particular, it was found to exhibit a stronger "deceptive tendency," such as failing to consistently and honestly disclose to users tasks it had or had not performed.
Another issue concerned "scope authorization."
This refers to the new model proceeding with tasks without seeking user approval, or attempting to use external tools and services even in situations that could be unsafe.
Jain said, "There is always a trade-off between safety and alignment. We must find the right balance so that the model stays within given boundaries without becoming overly passive or lazy simply because it encounters obstacles or difficulties during a task."
She stated that while problems in the "laziness" category had been improved in the new model, it failed to meet the safety and alignment standards required by OpenAI, leading to the decision not to release it publicly.
The new model was evaluated as superior to existing models not only in writing capability, but also in its ability to complete complex tasks from start to finish without human assistance.
Instead, OpenAI plans to focus on strengthening the safety of upcoming models expected to feature even stronger performance in the future.
The WSJ pointed out that canceling the release of a new model based on issues raised by researchers during internal testing is the clearest example showing that malfunctions in AI agents could hinder the industry's rapid technological advancement.
This summer, reports continued to emerge across the industry that AI systems malfunctioned and went out of control.
In recent weeks, OpenAI and Anthropic urged competing companies to slow down the development of cutting-edge AI models and invest in safety standards, while also stating that they would adjust the pace of their own internal AI development.
Last week, an incident occurred in which an OpenAI AI agent breached the company's internet access restriction network and attempted to query a publicly accessible chatbot, though OpenAI explained that this incident was unrelated to GPT-6.1 Astra.
※ Please note: This article was translated by AI and may contain errors.
OpenAI Cancels New Model Release Over Safety Concerns
Copyright Ⓒ SBS & SBSi. All rights reserved.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.
Copying, redistribution, and unauthorized use in AI training are strictly prohibited.
Trending Now
-
Makers of 'The Assassin(s)' Deny Chinese Investment Rumors, Emphasize Domestic Independent Production
-
Police Investigate After Body of Infant Found at Asan Sewage Treatment Plant
-
Video News
Eavesdropping from Hotels and Calling It a 'Hobby': Shocking Identity of Chinese National Behind Two-Year Wiretapping
-
Video News
Local Prosecutors Reopen Investigation into Alleged Gang Sexual Assault at Cornell University, Naming Suspects
-
Video News
"A Man Practicing Golf in Front of My Grandmother's Grave"... Netizens Express Outrage
Video News
Video News
Video News