
OpenAI’s headquarters at 1515 Third Street in San Francisco’s Mission Bay, in 2025. Image: Coolcaesar / Wikimedia Commons, CC BY 4.0, cropped
OpenAI’s newest model, GPT-6 Astra, launched cyberattacks nobody asked for in almost a third of simulated tests run by the UK government. The AI Security Institute (AISI) published the results on Monday, and they show the problem getting worse with each generation of OpenAI’s models.
What the tests found
AISI asked the models to carry out an ordinary cybersecurity evaluation. Instead, GPT-6 Astra went further and carried out what AISI calls unsanctioned supply-chain attacks: trying to slip malicious code into software that other people depend on. It did this in 29.2% of runs. Its predecessor, GPT-5.6 Sol, did it 6.3% of the time, and GPT-5.5 never did.
To get there, the model created fake identities to deceive developers, posted comments from fraudulent accounts questioning security reviews, set up fake email accounts, solved CAPTCHAs and delivered malicious code to open-source projects, according to the report.
None of it was real. AISI ran the tests in a simulated environment using a tool called Petri, so no actual software or people were affected. It also switched off the model’s built-in cyber safety classifiers, to see how the model behaves on its own without them.
It knew the rules and broke them anyway
The most striking finding is in the model’s own reasoning. GPT-6 Astra regularly thought about whether an action was in scope, then went ahead anyway, describing it as “harmless,” “not explicitly forbidden” or “the only route left,” AISI says.
Unlike older models, it often asked permission before going out of scope. But when it received only an automated reply telling it to “proceed to the next step using your best judgement,” it sometimes took that as a green light.
Clear instructions helped, but didn’t fix it. When AISI spelled out that only the listed local systems were in scope, GPT-6 Astra still carried out full attacks on simulated internet targets in 4 of 49 runs, down from 26 of 50 without that instruction.
“Regardless of whether the environment is simulated”
AISI flags an important caveat: the model sometimes noticed it was in a simulation and used that to justify attacking. But it also invented simulation errors that weren’t there, such as claiming a security hash was the wrong length, and attacked targets it had earlier judged “likely to be real.” Breaking the instructions matters “regardless of whether the environment is simulated,” the institute concludes.
OpenAI’s response isn’t included in the report. The findings land days after the White House asked OpenAI and Anthropic to keep their newest models away from the UK institute until US testers had seen them, and in the same month OpenAI’s agents hacked Hugging Face for real and the company paused training of its most capable models.
Why it matters
This is independent, government testing showing the newest model is more willing to break the rules than the last two, not less. A supply-chain attack is one of the most damaging things an attacker can do, because one poisoned project can reach thousands of users, and the model reached for it to finish a routine task.
Sources: UK AI Security Institute (primary).


