Claude Code, Anthropic’s AI coding tool. Image: Anthropic; Claude logo via Wikimedia Commons

Anthropic will now charge for some requests that Claude refuses to answer. Starting today, prompts that its safety systems block before Claude responds will be billed if they fall into one of three high-risk categories, the company announced on X, saying it has seen “coordinated attacks” on its systems in recent weeks.

What’s changing

Normally, when Claude’s safety classifiers block a request before it answers, you don’t pay for it. Anthropic says it has now resumed charges for these “billable blocks” in three categories where it says its filters rarely make mistakes:

  • Biology, such as requests that could help someone create biological weapons
  • Distillation attacks, where someone tries to extract Claude’s reasoning to train a rival AI model
  • Frontier AI development, requests aimed at building cutting-edge AI models

According to Anthropic’s developer documentation, a blocked request in one of these categories is billed “like any other request, at the rates of the model that ran it,” and still counts against your rate limits. Blocks in every other category stay free. The rules apply across the Claude API and the cloud platforms that offer Claude, including Amazon Bedrock, Google Cloud and Microsoft Foundry.

Why Anthropic is doing it

We’ve seen some coordinated attacks on our systems in recent weeks, and this is one layer of defense.

Anthropic’s Claude Code team, on X

The logic is simple: if someone is firing thousands of requests at Claude to probe its defenses or copy its abilities, making every blocked attempt cost money makes that far more expensive. Anthropic’s documentation says the change is designed “to disrupt attempts to circumvent Anthropic’s safeguards at scale.”

These are areas Anthropic has been worried about. Its September threat report described six illicit “distillation” campaigns by China-based AI labs since February, the largest generating 151 million exchanges, and several blocked attempts at biological misuse, the Washington Times reported.

Will it affect you?

Almost certainly not, says Anthropic. In recent testing, 99.7% of accounts using Claude Code, Claude.ai or Cowork didn’t hit any of the newly billable blocks, and the classifiers behind them are tuned to have a false positive rate below 0.1%, meaning they very rarely block legitimate requests by mistake.

But Anthropic admits that isn’t zero. Scientists, AI researchers and developers working in these fields are the people most likely to be caught out, and a wrongly blocked request would now cost them. Anthropic says it will keep improving the classifiers so they interrupt work less often, and asks anyone who thinks a request was blocked incorrectly to report it with /feedback in Claude Code. The company also says the billed categories may change as it keeps measuring how often its filters get it wrong.

Why it matters

It’s an unusual move: most AI companies absorb the cost of refusing a request. By charging for blocks, Anthropic is treating safety filters not just as a guardrail but as a deterrent, betting that attackers care more about cost than ordinary users, who will almost never notice. The risk is annoying the small group of legitimate researchers in biology and AI who could end up paying for mistakes that aren’t theirs.

Sources: Anthropic (@ClaudeDevs on X), Anthropic developer docs, Washington Times

Related