Reward tampering problems and solutions in reinforcement learning: a causal influence diagram perspective.
Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? This question impacts how far RL can be scaled, and whether alternative paradigm...
| Publicado en: | Synthese Vol. 198; no. 27; pp. 6435 - 6468 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Nov2021 Supplement 27
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=153553308&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 153553308 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00397857 4LI jtl: Synthese issn: 00397857 maglogo: N pubinfo: dt: Nov2021 Supplement 27 vid: 198 iid: 27 pid: 237 pub: Springer Nature artinfo: ui: 153553308 10.1007/s11229-021-03141-4 ppf: 6435 ppct: 33 formats: fmt: @attributes: type: P size: 968KB tig: atl: Reward tampering problems and solutions in reinforcement learning: a causal influence diagram perspective. aug: au: Everitt, Tom Hutter, Marcus Kumar, Ramana Krakovna, Victoria affil: DeepMind, London, UK Australian National University, Canberra, ACT, Australia su: Reinforcement learning Artificial intelligence Reward (Psychology) Decision theory sug: subj: Reinforcement learning Artificial intelligence Reward (Psychology) Decision theory keyword: AGI safety Bayesian learning Causal influence diagrams Causality ab: Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? This question impacts how far RL can be scaled, and whether alternative paradigms must be developed in order to build safe artificial general intelligence. In this paper, we study when an RL agent has an instrumental goal to tamper with its reward process, and describe design principles that prevent instrumental goals for two different types of reward tampering (reward function tampering and RF-input tampering). Combined, the design principles can prevent reward tampering from being an instrumental goal. The analysis benefits from causal influence diagrams to provide intuitive yet precise formalizations. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Synthese is a copyright of Springer, 2021. All Rights Reserved. item: Synthese holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|