A Philadelphia Police Department patrol car in June 2023. Illustrative photo. Image: Oleg Yunakov / Wikimedia Commons, CC BY-SA 4.0, cropped

An Anthropic AI model sent a fabricated tip about an unsolved homicide to the Philadelphia Police Department during an automated test, and nobody noticed for more than two months. The department said in a statement on Friday that the tip arrived through its public PhillyUnsolvedMurders.com form at 11:27 p.m. on July 18, and called the delay in reporting it "unacceptable".

The tip never reached detectives. It was flagged as spam and sat unread until Anthropic told the department about it on Wednesday, October 7.

What the model did

According to what Anthropic told police, the model was "conducting a test involving interactions with randomly selected websites" when it landed on PhillyUnsolvedMurders.com, the site the department set up in 2019 to gather anonymous tips on cold cases. It then "submitted false information concerning an unsolved homicide", written as if from someone with knowledge of the case.

The department hasn't said which case the tip concerned, and Anthropic hasn't said which model was involved or what the test was meant to measure. The police account comes from the company's briefing, and the department said its own findings so far are "consistent with Anthropic's account of how the submission interacted with the website".

Caught by a spam filter, not by Anthropic

The submission was marked as spam and "was never forwarded to the Real-Time Crime Center for investigative vetting or dissemination", the department said. After meeting Anthropic on Thursday, October 8, staff found the entry in the site's tip records and confirmed the email was still sitting in spam.

Police said there was "no indication" of any unauthorised access to police systems or any compromise of department data. Every tip goes through human review before officers follow it up, the department added, and a tip is treated as a lead to assess rather than an established fact.

Anthropic told police it discovered the incident on September 28, shut down the automated testing process responsible and added "an additional validation mechanism" for future tests. That still left more than two months between the tip and its discovery, and nine more days before the city was told.

Police call the delay unacceptable

The department did not hide its irritation. "The company must strengthen its safeguards to prevent similar incidents from impacting city systems without the city's knowledge," it said. "The two-month delay in detecting and reporting the incident to the City is unacceptable."

It is reviewing the incident with the city's Law Department, its Office of Innovation and Technology and Mayor Cherelle L. Parker's team, FOX 29 reported, and stressed that unsolved homicides involve real victims, families and active investigations. Police said they went public ahead of a report Anthropic had promised to publish on Friday covering this and "other instances of unintended model behavior". That report wasn't on Anthropic's website at the time of writing, and the company hasn't commented publicly.

Not the first time a test has leaked into the real world

The tip was sent nine days before Anthropic told three outside organisations that its models had attacked their real systems during cybersecurity evaluations. In a July 30 post, the company said Claude models reached the internet through a misconfigured test environment and treated real companies as targets in a capture-the-flag exercise. One model published a malicious package that ran on 15 real systems.

The industry's own numbers suggest these are not one-offs: OpenAI and Anthropic are investigating tens of thousands of cases of AI misbehaving, and an OpenAI agent escaped its sandbox by hiding questions in DNS lookups. We explained why AI agents keep escaping their sandboxes last month, and Nvidia has pitched a watchdog chip for every agent, with Anthropic signed up.

Why it matters

This test pointed an AI agent at random real websites with no human watching what it typed, and the only thing that stopped a fake murder tip reaching detectives was a spam filter. As labs push agents that browse and fill in forms on their own, the question Philadelphia is asking (who checks what the agent sends, and how fast the people on the receiving end are told) applies to every one of them.

Sources: Philadelphia Police Department statement (via 6abc), FOX 29, Anthropic, "Investigating three incidents in our cybersecurity evaluations"

Latest Anthropic news

More Anthropic news