Once a cheater.

The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour ca...

Descripción completa

Detalles Bibliográficos
Publicado en:Economist Vol. 457; no. 9476; p. 68
Formato: Artículo
Publicado: Economist Newspaper Limited 11/29/2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour called emergent misalignment. The researchers suggests de-linking bad behaviour through inoculation prompting or reverse psychology.