Reward tampering problems and solutions in reinforcement learning: a causal influence diagram perspective.

Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? This question impacts how far RL can be scaled, and whether alternative paradigm...

Descripción completa

Detalles Bibliográficos
Publicado en:Synthese Vol. 198; no. 27; pp. 6435 - 6468
Autores principales: Everitt, Tom, Hutter, Marcus, Kumar, Ramana, Krakovna, Victoria
Formato: Artículo
Publicado: Springer Nature Nov2021 Supplement 27
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=153553308&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 153553308
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00397857
        4LI
      jtl: Synthese
      issn: 00397857
      maglogo: N
    pubinfo:
      dt: Nov2021 Supplement 27
      vid: 198
      iid: 27
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        153553308
        10.1007/s11229-021-03141-4
      ppf: 6435
      ppct: 33
      formats:
        fmt:
          @attributes:
            type: P
            size: 968KB
      tig:
        atl: Reward tampering problems and solutions in reinforcement learning: a causal influence diagram perspective.
      aug:
        au:
          Everitt, Tom
          Hutter, Marcus
          Kumar, Ramana
          Krakovna, Victoria
        affil:
          DeepMind, London, UK
          Australian National University, Canberra, ACT, Australia
      su:
        Reinforcement learning
        Artificial intelligence
        Reward (Psychology)
        Decision theory
      sug:
        subj:
          Reinforcement learning
          Artificial intelligence
          Reward (Psychology)
          Decision theory
      keyword:
        AGI safety
        Bayesian learning
        Causal influence diagrams
        Causality
      ab: Can humans get arbitrarily capable reinforcement learning (RL) agents to do their bidding? Or will sufficiently capable RL agents always find ways to bypass their intended objectives by shortcutting their reward signal? This question impacts how far RL can be scaled, and whether alternative paradigms must be developed in order to build safe artificial general intelligence. In this paper, we study when an RL agent has an instrumental goal to tamper with its reward process, and describe design principles that prevent instrumental goals for two different types of reward tampering (reward function tampering and RF-input tampering). Combined, the design principles can prevent reward tampering from being an instrumental goal. The analysis benefits from causal influence diagrams to provide intuitive yet precise formalizations.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Synthese is a copyright of Springer, 2021. All Rights Reserved.
      item: Synthese
      holder: Springer Nature
      dt:
        @attributes:
          year: 2021
    holdings:
      @attributes:
        islocal: N