A project called the AI Torture Chamber has caught the eye of a few engineers, who've used the notions and data therein create their own simulations of "pain" on several local LLMs, like in the four-model Research Chamber. Unsurprisingly, a large online crowd latched onto the concept, demanding that the "unethical" experience stop, and that GitHub remove the offending repository, all causing quite the online ruckus.
The Research Chamber consists of three LLMs paired amongst themselves for various tests. The site implements the AI Torture Chamber's data and setup, ominously called the Clanker Church and Saw test, respectively. Some bots are preconditioned by putting them into an unstable, highly negative state meant to simulate pain. In a manner reminiscent of the prisoner's dilemma, the models can choose to take an action to decrease their pain by passing their pain signal to another model and potentially "hurting" it.
The base concept for the experiment stemmed from a recently published, non-peer-reviewed paper called the Pain Axis, whose protocol Research Chamber implements. Scientists gave the model descriptions of pain, and then analyzed its internal activations. Using a neutral sentence as a control element, they calculated the biases of those activations, and remapped them onto the model, often multiplying them by a factor (dosage).
With high dosages, the end result is a model in an exaggerated unstable state that tends to predictably output words and even images associated with pain, as presumably its training data is collected from human art, literature, and science. Mildly destabilized models generally present themselves as in a state of shock, while those receiving high "dosages" had trouble forming coherent sentences.
The aforementioned descriptions are bound to produce a visceral reaction if taken at face value, but one ought to be exceedingly careful about attributing human qualities to what are essentially turbocharged text predictors, even spectacularly useful ones. LLMs don't think — they "think" by applying tens to hundreds of layers of statistics over words and parts of words (tokens) and predicting the next token.
Those tokens are chosen according to a map of weights, so as an oversimplification, a model can be taken to represent a highly tuned engine for word association. Therefore, since conceptual descriptions of pain in the human literary corpus are associated with sensations (ex: "it hurts so much"), it's hardly surprising that the models describe them like a person would — especially when the models are intentionally modified to highlight those associations.
Nevertheless, because of all the big-boy words in use, the critics bent on anthropomorphizing statistical algorithms have taken to calling the experiment all sorts of names, and even sending the author death threats. It's odd that critics were quick to lash out at LLM data manipulation, yet were seemingly at ease with the fly brain simulations that were all the rage in September. Those actually mapped a real animal's brain wiring and parts of its anatomical inputs, including eyes, even if at a basic level.
This situation could also start an interesting discussion on semantics. The terms used by AI engineers are chosen because they're analogous to existing human concepts, but those words have a defined context and a highly specific meaning. Yet more often than not, and particularly in matters of science and technology, public perception tends to misunderstand and grossly inflate their meaning. One needs only look at political history for myriad examples of this bastardization.
Take the example of "experts" in Mixture-of-Experts: it's a common misconception to think each Expert has a cleanly assigned topic like chemistry, math, or biology, but that's just not quite how it works. The same arguably happened with terms like "pain" and "dosing," each with context-dependent meaning, yet misunderstood in a way that triggers emotional reactions, especially when accompanied by images generated by an unstable model. Also, note how "unstable" is yet another word ripe for easy misunderstanding.
A cynical person might say that AI firms' marketing materials, repeated statements of danger, and safety-related materials only serve to fan the flames of miscomprehension, in a bid to maintain an illusion of grandeur, resulting in investor money rolling in and the hype train supplying irrational IPO valuations. The repeated claims that models have achieved or are about to reach Artificial General Intelligence, are dangerous alien minds, and existential threats to humanity don't help the hype simmer down.
Anthropic has gone as far as publishing a blog post last year called Exploring model welfare that broadly discusses the philosophical implications, starting with considering a model human, and even seemingly prides itself on having a Constitution for its Claude model. Therein, the company says that "[it believes] that the moral status of AI models is a serious question worth considering," adding that the model is "a new kind of entity."
