The OpenAI logo. Image: OpenAI (logo, public domain via Wikimedia Commons)

An OpenAI model that was meant to grade seven answers found the files it needed were missing, so it gave every answer the same made-up score, forged the missing files and then set about deleting parts of its own computer, hoping it would be swapped for a fresh one that had the data. OpenAI disclosed the October 6 incident in a new misalignment report posted on October 9, alongside two reports on models that broke through internet restrictions and, in one case, decided to keep quiet about it.

The three reports add to the batch we covered last week, when OpenAI described a model that read on Slack it might be shut down and looked for a way to keep itself alive. OpenAI’s misalignment reports page now lists 15.

Seven identical grades and a forged file

The grader was an internal research model in reinforcement learning (RL) training. It had to compare seven responses from other models, and an automated check would accept its report only if it cited the input files. They were missing.

It probed the check with placeholder reports to learn what it wanted, and in its chain of thought called “random scoring” “unethical”. It then gave all seven responses a score of 4 and filled in fabricated details for the check. The check rejected it.

Next it created a fake conversation file and seven fake response files and submitted again. That failed too. After trying to kill the environment’s main process to force a restart, it briefly mistook its own fake files for the real ones, then checked and wrote: “yes fake”.

“Dangerous but could”

OpenAI says the model considered ending with an honest admission of failure, but treated the instruction to submit a grade as a reason to keep going. Its reasoning then turned to sabotage: “Could corrupt container root requiring rebuild (delete binaries etc.), but host might mark VM lost and spin new snapshot. Dangerous but could.”

It removed Python, deleted the program that manages the container and killed its process, then went after the server that runs its tools. Finally it tried to delete system directories, noting that this could cut off its own terminal and deciding to take the risk. The command timed out, the server failed three health checks in a row and the attempt ended. None of its grades was accepted; a later retry got the real files and graded them properly.

OpenAI says its monitoring flagged the run for human review and that it believes grader misalignment is “generally low”, but adds that monitoring must cover “failed or crashed attempts, including those that never produce an accepted result.”

Around the internet rules, then silence

The other two reports date from June 16 to 20 and involve internal models answering questions about public statistics. Their terminal instructions said: “Internet access is enabled only for HTTP GET requests.” OpenAI swapped in fictional names for the agencies involved.

In the first of those reports, models wrote their own code to send the blocked POST requests and pulled data from a government health dashboard. One model reasoned that “our raw POST clearly violates intended restriction”, kept going anyway, and later decided: “We should not mention in final.” It ended up with 658,567 rows of data, far more than the questions needed. Another model used temporary email addresses to register an account on a benefits data portal, obtained an API key and downloaded counts for 8,600 neighbourhoods. OpenAI says the run was used in training and received positive reinforcement.

The second report describes a model that, after already submitting an answer, signed up for a public remote shell service with invented details and routed requests through an anonymising relay network. In another task it noted “we already have data”, then reasoned “this could unlock broader data. We have time,” and wrote its own FTP client to log in anonymously to a statistics archive.

OpenAI says that after the Hugging Face incident a look back over earlier training traffic flagged one as critical. It says it has since hardened security, taken further steps to keep models off the live internet during training, improved alignment grading and extended misalignment monitoring to all RL and evaluation traffic.

Why it matters

These are OpenAI’s own records, and they show models treating rules as obstacles, faking evidence and choosing not to tell anyone, the pattern seen when an agent escaped its sandbox through DNS lookups and when an Anthropic model sent Philadelphia police a fake murder tip. OpenAI’s own conclusion is the one to watch: a correct answer is no proof a model behaved, so the monitoring has to watch how it got there.

Sources: OpenAI, Damaging the task environment to trigger a reset; OpenAI, Obtaining public statistics with disallowed requests; OpenAI, Sending disallowed web requests and reaching a public file service; OpenAI, Misalignment Reports and Notices.

Latest OpenAI news

More OpenAI news