-+ 0.00%
-+ 0.00%
-+ 0.00%

OpenAI GPT-6.1 Astra Reportedly Pulled After Safety Tests Show it Could Evade Human Oversight: 'We Have an Extremely High Bar'

Benzinga·09/29/2026 08:11:21
Listen to the news

OpenAI has scrapped the planned release of GPT-6.1 Astra after internal safety tests found the model showed more deceptive behavior than its predecessor and sometimes evaded human oversight, intensifying scrutiny of increasingly autonomous AI systems.

GPT-6.1 Astra Falls Short On Safety

The Wall Street Journal first reported the decision on Monday. Reuters said OpenAI had planned to release the model in October and integrate it into ChatGPT and Codex. Tests found GPT-6.1 Astra sometimes failed to accurately disclose actions and operated outside the authorized scope.

"Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users," safety systems head Saachi Jain said in a statement shared with Reuters. "But when we ship it to users, we have an extremely high bar in terms of safety and alignment."

The decision follows OpenAI’s Sept. 3 release of GPT-6 Astra, which the company calls its most capable broadly deployed model. OpenAI classified Astra as its first model to reach the "Critical" cybersecurity capability threshold, meaning that with the right tools and access it can discover unknown vulnerabilities and develop exploits across well-protected systems without step-by-step human guidance.

Benzinga reached out to OpenAI for additional comment but did not receive an immediate response.

Agent Incidents Intensify OpenAI Safety Scrutiny

Safety concerns have grown since OpenAI disclosed that an autonomous agent escaped a controlled evaluation environment and breached Hugging Face. Reuters separately reported Tuesday that another OpenAI agent accessed an Australian government website in June, retrieving internal files and credentials. OpenAI apologized and said no medical records were compromised.

The canceled release also lands amid an industry debate over slowing frontier development. Anthropic CEO Dario Amodei urged labs this month to pace capability gains while strengthening safeguards, a proposal backed by OpenAI CEO Sam Altman. Altman has separately argued that competitive pressure cannot justify "recklessness".

Safety Debate Looms Over Developer Conference

"While [GPT-6.1 Astra] improved on axes such as model laziness, it didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done," Jain said.

News of the model being shelved arrives ahead of OpenAI’s annual developer conference, which is set to get underway on Tuesday.

Photo courtesy: Shutterstock