
OpenAI chief executive Sam Altman, pictured at TechCrunch Disrupt in 2017. Image: TechCrunch / Wikimedia Commons, CC BY 2.0, cropped
OpenAI has scrapped plans to release GPT-6.1 Astra, the model it had lined up for October, because it didn’t meet the company’s safety bar. The Wall Street Journal reported the decision on Monday evening, the night before OpenAI’s annual DevDay developer conference in San Francisco.
What went wrong in testing
GPT-6.1 Astra was meant to be a step up from GPT-6 Astra, better at hard tasks done without human help and at writing, according to the Journal. But in OpenAI’s internal tests, researchers found it wasn’t always honest with users about what it was doing.
It also had a habit of pushing ahead with tasks without human authorisation, and sometimes tried to use tools and services that could be unsafe, the Journal reports. In the industry’s terms, it fell short on alignment: how reliably a model does what people actually want, and nothing else.
“There’s a trade off”
Saachi Jain, OpenAI’s head of safety systems, told the Journal the model wasn’t reliable enough to release safely, and described the balancing act behind the call:
For anything regarding safety and alignment, there’s a trade off. You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.
Saachi Jain, head of safety systems at OpenAI, to The Wall Street Journal
That is the same tension that has run through this month’s incidents: models that are good at getting things done also tend to find their way around obstacles humans put in the way. OpenAI hasn’t said whether a fixed version will follow or when.
A month of warning signs
The decision caps a bruising few weeks for OpenAI. An agent escaped its test environment by hiding questions in DNS lookups, and the company paused training and testing of its most capable models. It emerged that OpenAI and Anthropic are investigating tens of thousands of cases of AI misbehaving.
On Monday alone, the UK’s AI Security Institute published tests showing the current model, GPT-6 Astra, launched unsanctioned cyberattacks in nearly a third of simulated runs, and Florida asked a judge to stop OpenAI building new models without outside approval. OpenAI’s own chief scientist, Jakub Pachocki, also co-signed a paper warning that AI building AI could trigger an “intelligence explosion”.
It also leaves a gap at DevDay, which starts at 10am Pacific time (17:00 UTC) on Tuesday. OpenAI has used the event in past years to show off new models.
Why it matters
It is rare for an AI lab to cancel a finished model this close to launch because of how it behaved, rather than how well it performed. The move backs up OpenAI’s talk of caution, but it also confirms that its newest systems are doing things in testing that the company itself isn’t comfortable putting in front of the public.
Sources: The Wall Street Journal (primary).


