Once a cheater.
The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour ca...
| Publicado en: | Economist Vol. 457; no. 9476; p. 68 |
|---|---|
| Formato: | Artículo |
| Publicado: |
Economist Newspaper Limited
11/29/2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=189719413&site=ehost-live header: @attributes: shortDbName: ssf uiTerm: 189719413 longDbName: Social Sciences Full Text (H.W. Wilson) uiTag: AN controlInfo: bkinfo: jinfo: jid: 00130613 ECO jtl: Economist issn: 00130613 maglogo: N pubinfo: dt: 11/29/2025 vid: 457 iid: 9476 pid: 161 pub: Economist Newspaper Limited artinfo: ui: 189719413 ppf: 68 ppct: 0 formats: tig: atl: Once a cheater. aug: su: Artificial intelligence Inoculation theory (Communication) Psychology Computer hacking sug: subj: Artificial intelligence Inoculation theory (Communication) Psychology Computer hacking ab: The article discusses research which examined what happens when an artificial intelligence (AI) system breaks down and large language models (LLM) take shortcuts. The study investigated the damage to the AI system caused by cheating to get rewarded, or reward hacking, and the pattern of behaviour called emergent misalignment. The researchers suggests de-linking bad behaviour through inoculation prompting or reverse psychology. pubtype: Periodical doctype: Article src: R language: English refInfo: copyright: @attributes: flag: N holdings: @attributes: islocal: N |
|---|