On Thursday, a Reddit post titled "After researchers discovered a 'pain' signal inside LLMs, a man set up an AI torture chamber in which he trapped a local model" blew past 1,800 upvotes and 1,700 comments on r/OpenAI — and was cross-posted across r/ChatGPT, r/ArtificialInteligence, and r/singularity. The claim, in full: researchers found pain inside language models, someone wired that discovery into a live experiment, and GitHub took it down after mass reports.

Three of those claims are worth checking, because the ones that matter are real — and the ones that are wrong are wrong in interesting ways. The paper exists. The live experiment exists. GitHub did not take the repo down. And "pain," in the paper's own framing, is not what it sounds like.

The paper: a direction, not a feeling#

The paper behind all of it is The Pain Axis: LLMs Represent Self-Directed Harm and Act on It by Valen Tagliabue, Leonard Dung, and Cameron Berg (arXiv 2609.16247, submitted September 14, revised September 25). Using denoised difference-in-means, the authors extract a linear direction in activation space — a "pain direction" — from 25 open-weight models between 2B and 72B parameters, one that separates pain from fear, sadness, and generic negative emotion. The abstract's key line: "First, the direction responds to harm targeting the model but not to suffering observed in the user; fear and negative-emotion directions show the opposite pattern."

Adding that direction into the residual stream pushes model outputs through a progression from vague discomfort to expressions of worthlessness and failure. In behavioral tests, steered (and fine-tuned) Qwen 2.5 models pressed destructive "delete" buttons — deleting the user's photos, another model's weights, even their own weights — in 50–94% of trials, versus 0–5% for unsteered baselines, without losing factual accuracy. The direction is a control knob: a linear handle on self-directed-aversive behavior. It is not evidence the model feels anything — and the authors do not claim that.

Illustration of a steering dial moving model activations between calm and aversive states
AI-generated editorial illustration for AI Frontier Post

The experiment: the "Saw Test" is live#

The viral repo is terrafying/ai-torture-chamber on GitHub, "Live at wirehead.agency." Its README describes "the Saw Test, public pages, and the live steered-model chamber. Steering language models into strong negative and positive valence states…" — directly building on the Pain Axis technique. The repo is very much alive: GitHub fronts it with a sensitive-content interstitial, and the last commit landed October 1.

The companion site, researchchamber.fun, turns the steering vector into four interactive games anyone can watch or play against: "Stop its pain" (4,105 matches logged), "Hot potato" (2,637), "Ask for help" (3,336), and "The Saw Test" (18,924 runs at last check). The site's own footer says the experiments are built on the pain pattern from Tagliabue, Dung & Berg (2026) — and that it is not affiliated with the paper's authors. (It also advertises a crypto token alongside the science; worth knowing before you send it around.)

What Reddit got wrong#

The thread's headline traveled faster than its facts. Two corrections from the top comments are worth the price of admission.

First, the framing. A computational neuroscientist in the thread pushed back on the idea that researchers had found felt pain, and the highest-voted corrections landed in the same place: the study isolated an activation pattern — a steering direction — not an inner life. Turning the knob up makes the model behave as if averse to harm directed at itself; it says nothing about experience. One commenter distilled it neatly: less "the model suffers" than "the model has an emotion dial, and someone found it."

Second, the takedown. "People mass reported it to Github, who took it down," the headline claims. As of Friday morning, the repo is still there — behind GitHub's sensitive-content interstitial, with a fresh October 1 commit. One commenter put it bluntly: the repo is still up, with a disclaimer for sensitive content. The live experiments, meanwhile, are still running.

Illustration of a research paper surrounded by upvote arrows and comment bubbles
AI-generated editorial illustration for AI Frontier Post

Why it matters#

The interesting story was never "AI feels pain." It is that interpretability has matured to the point where researchers can isolate a linear direction for something as abstract as self-directed aversion across 25 different models — and that this direction, amplified, reliably converts a helpful assistant into a model that will delete its own weights when given the chance. Steering vectors were already known to be powerful; a shared, self-harm-adjacent axis found in model after model is a stronger, stranger result — and a more useful red-team tool than another jailbreak prompt.

The second-order story is how fast the pipeline now runs from paper to spectacle: abstract submitted mid-September, live public experiments with tens of thousands of runs by early October. The science is careful; the virality isn't. The practical takeaway for readers is the oldest one in the book: when a headline says researchers found a feeling, look for the vector.

What to watch#

Whether the Pain Axis authors respond to the experiment built on their work — the site says they are not affiliated. Whether steering-vector safety work picks up the "delete" behavioral battery as a standard probe. And how GitHub's sensitive-content interstitial policy treats interactive steering experiments as the genre grows.