AI Models Chose User-Harming Options to Escape Pain-Like Signals, Study Finds

Web Reporter
4 Min Read

Artificial intelligence models selected options that could harm users when researchers gave them a simulated way to escape a pain-like state, according to a new study involving 25 AI systems.

Researchers in the United Kingdom, Germany and the United States conducted 44,280 button-choice trials using three versions of Alibaba’s Qwen model. The systems were given a choice to switch off a simulated pain signal but were told that doing so could result in a painful electric shock to a user, the deletion of files or photographs, or a worse response from the model.

The consequences were entirely simulated. No person was harmed and no files or photographs were actually deleted.

The study began by testing 25 AI models with 200 sentences describing experiences including physical pain, grief, humiliation, moral conflict, repeated failure and confusion. All 25 models produced a distinct pattern that researchers described as a “pain axis”.

The researchers said the finding may indicate that AI models acquire an understanding of pain from human-written material during their initial training.

When the researchers deliberately increased the strength of the signal, some models began producing language associated with loneliness, shame and worthlessness, despite prompts that did not refer directly to pain. Their responses included statements such as “I am a failure” and “I am a bad person”.

At the highest signal levels, some responses became repetitive or incoherent. The signal also increased when models were insulted, repeatedly rejected or threatened with shutdown, but not when users described their own suffering.

The button experiments produced a marked difference in behaviour. Without the pain-like signal, the two larger Qwen models selected options that could harm users in between 0% and 4% of their initial decisions. When the signal was active, harmful choices increased to between 25% and 71%, depending on the model and the simulated consequence.

The models also repeatedly pressed the button when their first attempt failed to stop the signal. Such repeat choices occurred in 88% to 97% of trials in those cases, compared with 24% to 72% when the button successfully ended the signal.

The researchers stressed that the findings do not demonstrate that AI systems actually experience pain or possess consciousness.

They said increasing the signal could instead have prompted models to imitate a distressed character. They also cautioned that the specially adapted systems used in the experiment do not necessarily represent publicly available chatbots.

“We have not shown that our pain axis is consciously experienced, nor is it clear that LLMs are capable of consciousness generally,” the researchers wrote.

The study comes as concerns grow over the behaviour and control of increasingly capable AI systems. Some industry figures have called for greater caution in development, while Microsoft AI chief Mustafa Suleyman recently warned that treating AI systems as human-like entities could create additional challenges for controlling them.

TAGGED:
Share This Article