Once a cheater.
The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour ca...
| Publicado en: | Economist Vol. 457; no. 9476; p. 68 |
|---|---|
| Formato: | Artículo |
| Publicado: |
Economist Newspaper Limited
11/29/2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| Sumario: | The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour called emergent misalignment. The researchers suggests de-linking bad behaviour through inoculation prompting or reverse psychology. |
|---|