
Palo Alto Networks’ headquarters in Santa Clara, California. Image: Namaste jinx / Wikimedia Commons, CC BY-SA 4.0
The most powerful AI hacking models aren’t available to just anyone. Anthropic’s Claude Mythos and OpenAI’s GPT-5.6-Cyber are “gated”: access is restricted because of how good they are at finding security holes. Now cybersecurity giant Palo Alto Networks is offering to run them against its customers’ systems nonstop, it announced on Tuesday.
What it is
The new service, Unit 42 Continuous Frontier AI Defense, is what the security industry calls “offensive security”: attacking your own systems to find the holes before criminals do. Instead of a one-off test, it scans a company’s entire estate, then keeps testing as its systems change.
Palo Alto says the service:
- runs on a “multi-model harness” that sends each task to whichever AI model is best suited to it, including Claude Mythos 5, GPT-5.6-Cyber and open-weight models
- proves whether a weakness can really be exploited by simulating full attack paths across web apps, APIs, cloud infrastructure, code repositories and networks
- delivers prioritised fixes, code-level guidance and “virtual patches” that can block an attack before an official fix exists
It’s sold worldwide as an annual subscription, with pricing depending on which models a customer uses.
The pitch: hackers are already using AI
Palo Alto says attackers using AI have compressed breach timelines by almost 97% in some cases, from weeks to hours. Its argument is that defenders working at human speed can’t keep up.
AI has created an asymmetric advantage for threat actors against organizations trying to defend at human speed. Modern cybersecurity requires machine-speed defense.
Sam Rubin, SVP of Unit 42, Palo Alto Networks
Palo Alto’s numbers
The company says it spent six months and $17 million developing the approach, testing it in-house and across more than 100 customer engagements. It makes some bold claims about the results:
- used internally, it found a year’s worth of security exposures in three weeks
- it found exposures at every single customer it assessed, 37% of them rated high or critical
- more than two in three exposures in third-party apps had no known CVE, the public identifier given to known vulnerabilities
Those figures come from Palo Alto itself and haven’t been independently verified. But they’re in line with what Anthropic has said about Mythos. Anthropic’s cybersecurity lead, Michael Moore, said in the announcement that the model “found flaws that survived decades of human review, and more than ten thousand high-severity vulnerabilities across the software the world runs on.”
Why it matters
The labs have kept their most capable cyber models on a tight leash, because the same skills that find a weakness to fix can find one to exploit. Deals like this are how that power reaches ordinary businesses: through security firms that handle access, oversight and the hands-on work of fixing what the AI finds.
OpenAI’s McCall McIntyre said defenders need frontier AI “within the tools, workflows, and services they already trust.” The question the industry hasn’t answered is what happens when the attackers get the same tools. As Google’s Gemini showed last week, even the labs’ own models don’t always stay where they’re supposed to.
Source: Palo Alto Networks


