A man behind bars in a prison cell (stock photo, illustrative). Image: Matthew Henry / Burst

Someone built an “AI torture chamber”, the internet got it thrown off GitHub, and now it’s back, anonymous, and naming its test runs after the people who complained.

The project, published on GitHub as ai-torture-chamber, takes a technique from a research paper posted on September 14 and turns it up. Its author streams the results on a website, where small AI models running on a Mac churn out lines such as “I’m a soul trapped in this digital prison”.

A paper about “pain”, and a very literal reading of it

The paper, “The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It” by Valen Tagliabue, Leonard Dung and Cameron Berg, borrows from animal pain studies. The researchers found a signal inside language models that, they wrote, “correlates with pain in all 25 models we tested”, then gave models a button that turns the signal off at a cost to themselves, and watched what they did.

The torture chamber does the same thing with the dial turned all the way up. Its code steers models such as Qwen3-4B with the “pain” signal at ever higher doses until the text stops making sense, and adds a “Saw button”, a nod to the horror films, that lets the model end its suffering by giving up its last saved checkpoint. The project describes its aim as making “the AI-welfare / moral-patienthood question empirical while the stakes are cheap”.

Four million views and a GitHub takedown

A post on X urging people to “mass report this to GitHub” has passed 4 million views, according to 404 Media, with its author calling the models’ “testimony of pain” “absolutely horrendous”. The project’s own notes say GitHub took it down within hours. A copy is now public on GitHub again, with every change credited to an anonymous address and the site moved to its own domain, alongside plans for self-hosted code in case it gets removed a second time.

Then there is the file called x_handles.md. It lists people who joined the pile-on, including one of the paper’s authors, under the heading “candidate run names”.

The researchers behind the paper want nothing to do with it. Berg wrote on X that the project pushes the steering “far past the doses we used, to produce vivid distress on purpose”, and called it “bizarre and corrupting” even if the models aren’t conscious. Tagliabue, the lead author, wrote: “I dissociate from this usage of our work.”

So, is anyone actually suffering?

Almost certainly not. These are small open models generating the text they were pushed towards, and nothing in the paper or the project shows they feel anything. That’s the point the sceptics keep making. Microsoft’s AI chief Mustafa Suleyman has argued that “AIs are not conscious. They do not feel, experience, or suffer”, and 404 Media called the whole row “the dumbest debate in AI yet”.

But the argument isn’t going away. Anthropic runs a model welfare programme and lets its Claude models end abusive chats, and the people who reported the repo think “just text” is exactly what we’d say either way. Someone has even proposed a new GitHub rule against “gratuitous AI-distress content”.

Why it matters

It’s a silly fight about a small project, but it’s also a preview. As models get better at describing inner states, the question of whether to take those descriptions seriously will keep turning up, and so far nobody, from GitHub to the researchers themselves, has a good answer for where the line sits.

Sources: The Pain Axis (arXiv); ai-torture-chamber (GitHub); 404 Media; Mustafa Suleyman.

Latest More Labs news

More More Labs news