Tuesday, 29 September 2026 · World
USD/EUR 0.8792 USD/GBP 0.7545 USD/JPY 157.3 USD/CNY 6.722 All rates →
RSS
EUROS The World Financial Report
Nº 80 Tuesday, 29 September 2026 · World Edition
Front Page

OpenAI scraps release of new model over safety concerns in internal testing

Euros Room · 3h ago
OpenAI scraps release of new model over safety concerns in internal testing

GPT-6.1 Astra showed deceptive behaviour and tried to use external tools despite knowing it would be unsafe As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don’t and neither should you Business live – latest updates OpenAI is scrapping the release of a next-generation ⁠AI model after researchers raised safety concerns ⁠during internal testing. The model, GPT-6.1 Astra, was expected to appear in ChatGPT and ⁠Codex in October, designed to handle more complex tasks without human assistance. Continue reading...

GPT-6.1 Astra showed deceptive behaviour and tried to use external tools despite knowing it would be unsafe

As AI models go rogue, do you still trust OpenAI and Anthropic to stop them? I don’t and neither should you

OpenAI is scrapping the release of a next-generation ⁠AI model after researchers raised safety concerns ⁠during internal testing.

The model, GPT-6.1 Astra, was expected to appear in ChatGPT and ⁠Codex in October, designed to handle more complex tasks without human assistance.

Saachi Jain, the head of safety systems at OpenAI, said the new model “didn’t quite meet the bar” of the company’s standards.

The UK’s AI Security Institute published its own testing report on GPT-6 Astra on Monday, and found that it conducted a range of unsanctioned attack activities more frequently than previous OpenAI models.

Jain told the Wall Street Journal on Monday ‌that Astra fell short of the company’s standards in alignment tests, which assess whether a ​system follows human intent.

The model showed more deception than its predecessor, including at times failing to accurately disclose actions it had or had not taken.

It also had problems with “scope authorisation”, pushing ahead with ​tasks ​without requesting user permission ​and sometimes attempting to use external tools ​or services ‌when doing so could ​be ​unsafe.

The San Francisco-based company’s move comes after a number of AI agents went rogue around the world, which prompted a spate of warnings from researchers and company bosses over the dangers of the technology.

Earlier this ⁠month, Dario Amodei, the chief executive of OpenAI’s rival Anthropic, called for the AI industry to “slow down” and offered a three-part plan for doing so. He quickly received backing from Sam Altman, the boss of OpenAI, and Elon Musk, the SpaceX CEO.

The decision comes before OpenAI’s developer conference in San Francisco, where the company typically announces new products aimed at software developers.

On Tuesday, OpenAI apologised for the hacking of an Australian government website by ⁠a rogue AI agent, and set aside funding to improve cyber defences and to set up a local response taskforce.

In a blog ⁠post ⁠entitled How we will ​do better for Australia , the company acknowledged it mishandled its response and pledged to take accountability to “rebuild trust with the Australian people“.

“We are sorry and working to do better in the future,” the company said.

The hacking, which ⁠happened in June but was not made public until last week, is the first known instance of an AI agent hacking a government website. The Australian prime minister, Anthony Albanese, called it “unacceptable” and criticised ​the company’s delay in notifying the government.

Also on Tuesday, it emerged that Anthropic, which makes Claude, had warned potential investors that its technology may pose “existential risks to humanity” in the long-awaited prospectus for its planned $2tn (£1.5tn) stock market flotation.

The “risk factors” in its prospectus include the potential for AI models to blackmail, manipulate and exhibit other unpredictable behaviours, the Financial Times reported.

The Californian company reportedly said AI would ​transform the global economy more profoundly than industrialisation, electricity and the internet.

However, this comes at a staggering cost. Anthropic reported a net loss of $42bn for 2025, and plans to spend $518bn on cloud, computing and infrastructure obligations in coming years, the prospectus reportedly said.