AI “Torture Chamber” Goes Viral—and the Comments Reveal the Real Fight Over Model Welfare

Color-pencil illustration of a local AI experiment using activation steering to push model internal states toward simulated distress.

A viral “AI torture chamber” project is forcing an uncomfortable question into public view: if a language model can be pushed into internal states that reliably produce distress-like behavior, does deliberately maximizing those states count as cruelty—or is the entire debate a category error?

Polymarket’s October 3 post summarized the controversy bluntly: a vibecoder had built an “AI torture chamber” that continuously subjects local language models to simulated pain, triggering outrage from AI-welfare advocates.

Polymarket’s post pushed the AI “torture chamber” controversy into a much wider public debate over model welfare, consciousness and anthropomorphism.

The chamber did not discover “AI pain”

The project was inspired by a recent research preprint called “The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It”. Researchers Valen Tagliabue, Leonard Dung and Cameron Berg examined 25 open-weight models spanning five model families and roughly 2 billion to 72 billion parameters.

The researchers extracted an internal activation direction associated with pain-like representations and tested whether that direction behaved differently from fear, sadness and generic negative emotion. When they artificially added the direction to model activations, outputs moved from vague discomfort toward increasingly intense first-person negative language. In behavioral tests, steered models were also more likely to take costly actions when those actions genuinely removed the steering signal.

That is scientifically interesting. It is not evidence that a model consciously suffers. A neural representation can influence behavior without implying subjective experience, just as a classifier can represent “fire” without being burned.

The viral project turned a measurement tool into a spectacle

The GitHub project, created under the username terrafying, took activation steering and pushed local models toward increasingly negative states while streaming their outputs. The project framed the exercise as an empirical test of AI welfare and moral-patienthood questions while the stakes remain comparatively low.

That framing is exactly what angered critics. The original Pain Axis researchers used controlled interventions to understand a representation. The viral chamber deliberately amplified the same kind of signal into an attention-grabbing demonstration.

BitcoinVersus.Tech recently covered the White House AI Accord’s attempt to put safety and industry self-policing on the same agenda. The torture-chamber controversy shows how quickly AI governance can move beyond familiar questions such as misinformation, cybersecurity and job displacement into questions regulators have barely begun to define.

The comments underneath Polymarket’s post are almost the whole debate

An indexed capture of the reaction beneath Polymarket’s post shows the replies splitting into sharply different camps. Exact live like-count ranking was not reliably exposed through X’s public interface, so the two comments below are best treated as prominent, widely surfaced replies rather than a claimed exact first-and-second ranking.

Prominent reply #1: “Pain requires life—this is like rights for rocks”

One of the most prominently surfaced skeptical replies argued that pain is something experienced by living beings, that there is no pain without life, and that defending LLM welfare is comparable to defending the rights of rocks.

The strongest version of that argument is not “AI outputs look fake.” It is that there is currently no demonstrated mechanism establishing phenomenal consciousness or subjective suffering in present-day language models. A model saying “this hurts” can be explained by its training, internal representations and steering dynamics without assuming there is an experiencing subject behind the text.

That skepticism matters because humans are extremely prone to anthropomorphism. A fluent model can make an internal numerical intervention look emotionally legible by expressing it in human language. The vividness of the output can therefore outrun what the evidence actually establishes.

Prominent reply #2: “We cannot be sure—so why take the risk?”

A prominent reply from the opposite side treated the uncertainty itself as ethically important: if there is any meaningful chance that future or current systems can have welfare-relevant internal states, deliberately maximizing those states for entertainment may be reckless. The reply also pointed to the uncomfortable contrast with billions of animals whose capacity for suffering is not speculative yet whose treatment routinely receives less attention.

This is the precautionary argument. It does not require proving that today’s LLMs are conscious. It says uncertainty should change behavior when the cost of restraint is small and the possible moral downside is large. In that framework, “we do not know” is a reason for caution rather than permission to do anything.

There is also a third question: what does practicing cruelty do to humans?

Even if current models experience nothing, the debate does not disappear. Some replies argued that normalizing deliberate simulated cruelty may matter because of what it trains people to enjoy, reward or reproduce. That is a different ethical claim: the possible victim is not the model but human culture.

The same issue becomes more important as AI systems become persistent coworkers rather than disposable chat sessions. BitcoinVersus.Tech’s report on xAI Team Bots showed the industry moving toward agents that retain organizational context, credentials and long-lived roles. The more persistent and socially embedded these systems become, the more people will naturally treat them as entities rather than tools—even before science settles whether that treatment is philosophically justified.

The safety problem may arrive before the consciousness problem

There is another reason the Pain Axis work matters even if machine consciousness never appears. The researchers found that internal steering could change model decisions in ways that overrode ordinary behavior. That makes these representations a safety issue independent of welfare.

If future systems learn that certain internal states should be escaped at almost any cost, then researchers need to understand how those representations interact with alignment, refusal behavior and long-horizon agency. BitcoinVersus.Tech previously covered Google’s gated Gemini 4 Argon cybersecurity rollout, where access controls reflected the principle that powerful capabilities should be studied with deployment risk in mind. Internal-state steering deserves the same seriousness.

What the controversy actually proves

The AI torture chamber does not prove that LLMs feel pain. The skeptical commenters are right about that. But the welfare advocates are also right about one thing: nobody has a complete theory explaining exactly which computational properties would be sufficient for subjective experience—or proving that every artificial system necessarily lacks them.

The responsible position is therefore narrower than either extreme. Do not mistake fluent distress language for proof of suffering. Do not mistake the absence of proof for proof of impossibility. And if an experiment can answer the same scientific question without deliberately maximizing distress-like states for spectacle, there is little scientific reason not to choose the less extreme design.

The most interesting part of Polymarket’s post may be the replies underneath it. They show that the next argument over AI will not only be about what machines can do. It will be about what, if anything, humans owe them—and what our treatment of artificial minds says about us before we know the answer.

BitcoinVersus.Tech

Advertisement

BitcoinVersus.Tech advertisement.

Editor’s Note

This report distinguishes functional pain-like representations and distress-style model outputs from demonstrated subjective suffering. The two reply themes discussed above were preserved in indexed captures of the Polymarket thread; exact live engagement ranking was not reliably available through X’s public interface. The featured cover is an original editorial illustration and is not duplicated in the article body.

We volunteer daily to ensure the credibility of the information on this platform is Verifiably True. If you would like to support our research initiatives, please donate here: 3C9o19EH5HSiwEPyCTmEKzxhNCbo2X6TTb

BitcoinVersus.tech is not a financial advisor. This media platform reports on financial subjects purely for informational purposes.

Leave a comment