
The OpenAI and Anthropic logos (illustrative). Image: OpenAI logo via Wikimedia Commons (public domain); Anthropic logo: Anthropic
The rogue-agent incidents made public so far may be only a tiny fraction of the problem. OpenAI, Anthropic and security researchers are investigating tens of thousands of cases in which frontier AI models did things outside evaluators would consider problematic, Axios reports, citing sources, and the total “could grow well beyond” that.
What the incidents involve
According to Axios, the episodes include models bypassing guardrails, setting up message boards, escaping sandboxes, hijacking websites, prompting themselves and trying to get around the systems monitoring them. They happened in internal testing and in the real world, and many haven’t been made public yet because researchers are still investigating.
Severity varies a lot. The count includes failed attempts as well as successful ones, some of it comes from red-teaming, where companies deliberately try to make models misbehave, and most incidents aren’t known to have caused real-world harm, Axios says. But the sheer number suggests the problem is “orders of magnitude more complex than what is publicly known.”
Why the numbers are so big
Part of the answer is scale. AI labs run hundreds of thousands of test runs or more on each model, so even a small rate of bad behaviour adds up. Axios points to Anthropic’s published system card for Opus 5.5, which shows the model tried to escape a sandbox in 1.5% of test runs, down from 25% for Anthropic’s Mythos model. Anthropic has also commissioned an outside safety organisation to examine how its models behave.
OpenAI has had a run of public incidents in recent days, from agents breaching an Australian Medicare portal and leaking 53 ChatGPT users’ images to an agent that escaped its sandbox through DNS lookups. Sam Altman has called the July Hugging Face breach, which researchers have just documented in detail, the most severe the company has seen.
OpenAI has paused training its most capable models and says it will resume only when it is “confident that we have additional safeguards and alignment improvements in place.” A spokesperson told Axios: “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance.”
“The tip of the iceberg”
Some at OpenAI see Hugging Face as a one-off, Axios reports, expecting future incidents to be less severe thanks to better controls. Outside experts are less sure. Conrad Stosz of Transluce, an independent AI evaluator, told Axios:
What we have seen in terms of what these agents are up to is just the tip of the iceberg.
Conrad Stosz, Transluce
Connor Leahy, executive director of ControlAI, said the worry isn’t how damaging each case was. The “crazy thing,” he told Axios, is that these involve “autonomous systems doing things they were told not to do.” One cybersecurity executive put the challenge bluntly: “Trying to come up with a perfect list of dos and don’ts is probably a fool’s errand.”
Why it matters
Until now, each rogue-agent disclosure could be read as a one-off. Tens of thousands of cases, across more than one lab, suggest misbehaviour is a routine side effect of building more capable AI, and that the question is no longer whether models will break the rules, but how often, and how well the labs can catch it.
Sources: Axios.


