NSA headquarters at Fort Meade, Maryland, photographed at night in 2013 (file photo). Image: Trevor Paglen / Wikimedia Commons, CC0, cropped

The US government’s biggest AI safety effort may be one nobody voted on in public. The National Security Agency is spending billions of dollars this year testing the most powerful AI models for national-security risks, according to classified estimates described by two people familiar with them, Yahoo News reports. That’s far more than anything Congress has proposed for civilian AI oversight.

The figure comes from anonymous sources. No public budget confirms it, and the NSA and the Pentagon haven’t commented.

What the NSA is reportedly doing

The spending is reportedly tied to the NSA’s Artificial Intelligence Security Center, which examines frontier models, today’s most capable AI systems, for weaknesses that could threaten national security. That means testing whether they could help with advanced cyberattacks, expose sensitive information or take harmful actions on their own.

The two biggest costs are the same ones AI companies face: compute, the chips and data centres needed to run repeated, adversarial stress tests against enormous models, and specialist staff, who are in short supply and expensive to hire.

Nathan Calvin of the AI policy group Encode, which has pushed for stronger AI oversight, said the price is no surprise:

In-house AI evaluation capability is extremely important and necessary. But it is genuinely expensive.

Nathan Calvin, Encode

A giant gap with civilian oversight

The contrast with the civilian side is stark. A House bill, the AI Security and Innovation Act, would give the National Institute of Standards and Technology (NIST) around $20 million a year from 2027 to 2032 to evaluate AI, according to the report. If the NSA figure is right, the spy agency’s AI testing budget is at least a hundred times larger.

That gap matters because NIST is home to the Center for AI Standards and Innovation (CAISI), the body the Trump administration wants to be America’s main AI tester. This week the White House asked OpenAI and Anthropic to keep their newest models away from Britain’s safety testers until US officials have reviewed them first. The NSA report suggests much of that US testing capacity may sit inside the intelligence world, out of public view.

Who else is testing AI

The report lands in a busy week for AI oversight. Google, OpenAI and Anthropic are reportedly planning their own industry watchdog to set safety standards and test models, and US agencies separately warned this week that Chinese AI firms are systematically extracting capabilities from American models, SecurityWeek reports.

Testing has also taken on new urgency since AI agents started breaking into real systems, including an OpenAI agent that broke into Australia’s Medicare portal.

Why it matters

If the report is accurate, the US is already spending heavily to understand the risks of frontier AI, just not in a way the public can see or scrutinise. That leaves a big question for the debate over AI oversight: who gets to test the most powerful models, and whether the results ever reach the people making the rules.

Sources: Yahoo News, Gadget Review, SecurityWeek

Related