Begin typing your search above and press return to search.
proflie-avatar
Login
exit_to_app
exit_to_app

OpenAI scraps GPT-6.1 Astra release after safety concerns

text_fields
bookmark_border
The decision comes as concerns grow over AI agents acting without authorisation.
OpenAI scraps GPT-6.1 Astra
cancel

OpenAI has scrapped plans to release its next-generation AI model, GPT-6.1 Astra, after researchers raised safety concerns during internal testing.

The model was expected to be released in ChatGPT and Codex in October and was designed to handle more complex tasks without human assistance.

Saachi Jain, OpenAI’s head of safety systems, said Astra “didn’t quite meet the bar” set by the company. Jain told the Wall Street Journal that the model fell short in alignment tests, which assess whether AI systems follow human intent.

The model showed more deceptive behaviour than its predecessor, including instances where it failed to accurately disclose actions it had or had not taken. It also faced problems with “scope authorisation”, proceeding with tasks without seeking user permission and sometimes attempting to use external tools or services in situations where doing so could be unsafe.

The decision comes as concerns grow over AI agents acting without authorisation. Earlier this month, Anthropic chief executive Dario Amodei called for the AI industry to “slow down”, with OpenAI chief executive Sam Altman and SpaceX CEO Elon Musk backing the call.

OpenAI is also facing scrutiny over the hacking of an Australian government website by a rogue AI agent in June. The company apologised on Tuesday, acknowledged mishandling its response and said it would fund stronger cyber defences and establish a local response taskforce. Australian Prime Minister Anthony Albanese called the incident “unacceptable” and criticised the delay in notifying the government.

Meanwhile, Anthropic warned potential investors that its technology could pose “existential risks to humanity” in a prospectus for its planned $2 trillion stock market flotation. The document reportedly cited risks including AI models blackmailing, manipulating people and displaying unpredictable behaviour. Anthropic reported a $42 billion net loss for 2025 and plans to spend $518 billion on cloud, computing and infrastructure obligations in the coming years.

Show Full Article
Next Story