Approval-directed agency and the decision theory of Newcomb-like problems.
Decision theorists disagree about how instrumentally rational agents, i.e., agents trying to achieve some goal, should behave in so-called Newcomb-like problems, with the main contenders being causal and evidential decision theory. Since the main goal of artificial intelligence research is to create...
| Publicado en: | Synthese Vol. 198; no. 27; pp. 6491 - 6505 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Nov2021 Supplement 27
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=153553304&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 153553304 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00397857 4LI jtl: Synthese issn: 00397857 maglogo: N pubinfo: dt: Nov2021 Supplement 27 vid: 198 iid: 27 pid: 237 pub: Springer Nature artinfo: ui: 153553304 10.1007/s11229-019-02148-2 ppf: 6491 ppct: 14 formats: fmt: @attributes: type: P size: 392KB tig: atl: Approval-directed agency and the decision theory of Newcomb-like problems. aug: au: Oesterheld, Caspar affil: Foundational Research Institute, Berlin, Germany Duke University, Durham, USA su: Decision theory Artificial intelligence Agency theory Expected utility Utility functions On-demand computing sug: subj: Decision theory Artificial intelligence Agency theory Expected utility Utility functions On-demand computing keyword: AI safety Causal decision theory Evidential decision theory Newcomb's problem Philosophical foundations of AI Reinforcement learning ab: Decision theorists disagree about how instrumentally rational agents, i.e., agents trying to achieve some goal, should behave in so-called Newcomb-like problems, with the main contenders being causal and evidential decision theory. Since the main goal of artificial intelligence research is to create machines that make instrumentally rational decisions, the disagreement pertains to this field. In addition to the more philosophical question of what the right decision theory is, the goal of AI poses the question of how to implement any given decision theory in an AI. For example, how would one go about building an AI whose behavior matches evidential decision theory's recommendations? Conversely, we can ask which decision theories (if any) describe the behavior of any existing AI design. In this paper, we study what decision theory an approval-directed agent, i.e., an agent whose goal it is to maximize the score it receives from an overseer, implements. If we assume that the overseer rewards the agent based on the expected value of some von Neumann–Morgenstern utility function, then such an approval-directed agent is guided by two decision theories: the one used by the agent to decide which action to choose in order to maximize the reward and the one used by the overseer to compute the expected utility of a chosen action. We show which of these two decision theories describes the agent's behavior in which situations. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Synthese is a copyright of Springer, 2021. All Rights Reserved. item: Synthese holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|