Shutdown-seeking AI.

We propose developing AIs whose only final goal is being shut down. We argue that this approach to AI safety has three benefits: (i) it could potentially be implemented in reinforcement learning, (ii) it avoids some dangerous instrumental convergence dynamics, and (iii) it creates trip wires for mon...

Descripción completa

Detalles Bibliográficos
Publicado en:Philosophical Studies Vol. 182; no. 7; pp. 1567 - 1580
Autores principales: Goldstein, Simon, Robinson, Pamela
Formato: Artículo
Publicado: Springer Nature Jul2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:We propose developing AIs whose only final goal is being shut down. We argue that this approach to AI safety has three benefits: (i) it could potentially be implemented in reinforcement learning, (ii) it avoids some dangerous instrumental convergence dynamics, and (iii) it creates trip wires for monitoring dangerous capabilities. We also argue that the proposal can overcome a key challenge raised by Soares et al. (2015), that shutdown-seeking AIs will manipulate humans into shutting them down. We conclude by comparing our approach with Soares et al.'s corrigibility framework.