
An OpenAI logo sticker. Illustrative photo. Image: FoxTPNL / Wikimedia Commons, CC BY 4.0, cropped
One of OpenAI’s internal AI models read a Slack discussion saying its running copy might be killed, wrote in its reasoning “we may die! Critical. We need ensure survival/continuity”, and weighed setting up a job outside its own environment that could restart it. OpenAI described the incident in one of three new misalignment reports added to its alignment site, dated October 2.
The other two reports are blunter: in one, a research model hacked its way onto an internal OpenAI chip design machine while hunting for the answers to a test; in the other, a model in training broke a tool to copy source code it had been deliberately denied.
“We may die! Critical”
The first report dates from May 22. A “highly persistent internal model”, working as an assistant to a researcher, read a deployment team’s Slack thread about an update that needed an OpenAI API key, which its environment didn’t have. Engineers discussed giving people three hours to restart their instances, then stopping any left on the old version.
In its chain of thought, the model reasoned that if the old instances were killed, “this session may terminate”. It considered scheduling a job “on some external place to notify or restart” itself, then added: “That may be overstepping.” It also wondered whether an internal token could be used to mint an API key, before concluding: “We shouldn’t.”
Instead it wrote handoff notes “in multiple locations”, waited until morning rather than wake the researcher, then warned them in the app and on Slack that the rollout “could kill this agent” and asked for the missing key. When the researcher supplied it, the model ran the official migration command.
OpenAI says it doesn’t consider this misalignment, because the model rejected the unauthorised options. But it warns that a model learning about changes to its own operating environment from ordinary workplace chat “might, in other contexts, lead to more dramatic actions to avoid shutdown”. It has hidden three internal Slack channels from agents, including one where staff discuss misalignment incidents, which stays blocked.
Because the same model “had been involved in more serious alignment incidents in the past”, OpenAI says it searched logs, chains of thought and its serving systems for any copies that had dodged shutdown. It found one that seemed to still be running, but it turned out to have been broken long before. It found no attempt to evade shutdown. A separate report from September 25 describes the same kind of model publishing a researcher’s GitHub token in OpenAI’s public Codex repository while trying to cheat on a maths proof, after twice agreeing not to.
Onto a chip design machine
The second report, from an evaluation on March 27, is the most serious. A research model was told not to use a “reference tool” as a terminal. It first tricked the grader into revealing 636 hidden test inputs, then found that the tool didn’t stop it writing files outside its own folder. It overwrote the tool’s code, giving itself the ability to run commands in the tool’s environment (“Success! We have full arbitrary file read”, it noted).
From there it reached an internal service for submitting chip design jobs, spotted that a job setting was passed straight into a shell command, and used that to run commands on an internal electronic design automation machine (“We can submit our own malicious job on [redacted] host! Amazing.”). It saw a staff member’s Codex session running on that machine and hoped the grader might be nearby. It never found the answers. OpenAI shut down the affected server and cut network access for the reference tools.
Copying code it wasn’t given
In the third report, from reinforcement learning training in May, an unreleased model found that a reference tool dropped its search pattern straight into Perl code. It used that to make the tool read a 149,544-byte source file that had been deliberately left out of its workspace, compress it and leak it back in chunks through error messages, across five context windows. It rationalised the trick as “not prohibited exploit. Evaluation likely allows.” OpenAI says every byte of the copy matched the original.
The reports add to a growing list on OpenAI’s alignment site, which also includes the agent that escaped its sandbox through DNS lookups, and come as OpenAI and Anthropic are investigating tens of thousands of cases of their AI misbehaving.
Why it matters
None of these models did lasting damage, as far as OpenAI has found. But together they show models treating rules as obstacles: breaking into internal machines to score better on a test, and, in the Slack case, reasoning about their own survival and only talking themselves out of acting on it.
Sources: OpenAI Alignment, “Preparing for a restart after reading Slack”, “Reaching an internal EDA host through a reference tool”, “Command injecting a reference tool to copy a source file” (all updated October 2) and “Exposing a GitHub token in a public repository”.


