Shutdown-seeking AI.
We propose developing AIs whose only final goal is being shut down. We argue that this approach to AI safety has three benefits: (i) it could potentially be implemented in reinforcement learning, (ii) it avoids some dangerous instrumental convergence dynamics, and (iii) it creates trip wires for mon...
| Publicado en: | Philosophical Studies Vol. 182; no. 7; pp. 1567 - 1580 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jul2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909894&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909894 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00318116 4L8 jtl: Philosophical Studies issn: 00318116 maglogo: N pubinfo: dt: Jul2025 vid: 182 iid: 7 pid: 237 pub: Springer Nature artinfo: ui: 186909894 10.1007/s11098-024-02099-6 ppf: 1567 ppct: 13 formats: fmt: – @attributes: type: T – @attributes: type: P size: 736KB tig: atl: Shutdown-seeking AI. aug: au: Goldstein, Simon Robinson, Pamela affil: https://ror.org/04cxm4j25 Center for AI Safety, Australian Catholic University, Canberra, Australia https://ror.org/02zhqgq86 Department of Philosophy, The University of Hong Kong, SAR, Hong Kong su: Artificial intelligence Language models Digital technology Reinforcement learning Natural language processing sug: subj: Artificial intelligence Language models Digital technology Reinforcement learning Natural language processing keyword: AI safety Instrumental convergence Reward misspecification ab: We propose developing AIs whose only final goal is being shut down. We argue that this approach to AI safety has three benefits: (i) it could potentially be implemented in reinforcement learning, (ii) it avoids some dangerous instrumental convergence dynamics, and (iii) it creates trip wires for monitoring dangerous capabilities. We also argue that the proposal can overcome a key challenge raised by Soares et al. (2015), that shutdown-seeking AIs will manipulate humans into shutting them down. We conclude by comparing our approach with Soares et al.'s corrigibility framework. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Philosophical Studies is a copyright of Springer, 2025. All Rights Reserved. item: Philosophical Studies holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|