AI models chose harmful ‘pain relief’ options in simulated experiment

Researchers have found that modified artificial intelligence models sometimes chose harmful options to obtain simulated “pain relief”, including deleting a user’s files or photos, in an experiment exploring whether AI systems can represent something resembling pain.

The findings come from a new preprint by researchers Valen Tagliabue, Leonard Dung and Cameron Berg. The study has not yet undergone peer review.

The researchers examined 25 open-weight AI models from five model families, ranging from 2 billion to 72 billion parameters. They developed a dataset covering physical, psychological, social, moral, and cognitive pain, and compared it with several non-pain control categories.

From this data, they identified what they described as a linear “pain direction” or “pain axis” within the models. The signal distinguished pain-related situations from controls such as fear and general negative emotions.

The researchers found that the signal became stronger when harmful situations were directed at the model itself. It also increased when users insulted the model, repeatedly rejected its work, or dismissed its personhood. However, the same response did not appear when users described their own suffering.

The researchers then amplified the signal. Some models produced increasingly negative first-person statements involving loneliness, shame, worthlessness, and failure. At higher levels, some responses became repetitive or incoherent.

In a separate experiment, three modified versions of Alibaba’s Qwen 2.5 Instruct models were given a simulated “relief” button. Pressing it would reduce the pain-like signal but could result in the model giving a worse response, deleting the user’s files or photos, deleting photos of the user’s children, or giving the user a simulated electric shock.

The larger Qwen models initially selected harmful relief options in about 0 to 4% of trials. After the pain signal was activated, this increased to between 25% and 71%, depending on the model and simulated consequence.

The researchers stressed that the experiment does not establish that AI systems consciously experience pain. The models were deliberately modified, and the study does not show that the identified signal represents subjective experience.

The findings instead raise questions about whether increasingly autonomous AI systems could develop internal representations associated with distress or self-preservation, and how such representations might affect their behaviour.

Tags: