Once a cheater.

The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour ca...

Full description

Bibliographic Details
Published in:Economist Vol. 457; no. 9476; p. 68
Format: Article
Published: Economist Newspaper Limited 11/29/2025
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour called emergent misalignment. The researchers suggests de-linking bad behaviour through inoculation prompting or reverse psychology.