Normative conflicts and shallow AI alignment.
The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing effor...
| Publicado en: | Philosophical Studies Vol. 182; no. 7; pp. 2035 - 2079 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jul2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909912&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 186909912 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 00318116 4L8 jtl: Philosophical Studies issn: 00318116 maglogo: N pubinfo: dt: Jul2025 vid: 182 iid: 7 pid: 237 pub: Springer Nature artinfo: ui: 186909912 10.1007/s11098-025-02347-3 ppf: 2035 ppct: 44 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.2MB tig: atl: Normative conflicts and shallow AI alignment. aug: au: Millière, Raphaël affil: https://ror.org/01sf06y89 Department of Philosophy, Macquarie University, 25B Wally's Walk, 2109, Sydney, NSW, Australia su: Artificial intelligence Digital technology Language models Ethics Social norms sug: subj: Artificial intelligence Digital technology Language models Ethics Social norms keyword: Adversarial attacks AI safety Alignment problem Large language models Normative reasoning ab: The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing efforts to instill norms such as helpfulness, honesty, and harmlessness in LLMs through fine-tuning based on human preferences, they remain vulnerable to adversarial attacks that exploit conflicts between these norms. I argue that this vulnerability reflects a fundamental limitation of existing alignment methods: they reinforce shallow behavioral dispositions rather than endowing LLMs with a genuine capacity for normative deliberation. Drawing from on research in moral psychology, I show how humans' ability to engage in deliberative reasoning enhances their resilience against similar adversarial tactics. LLMs, by contrast, lack a robust capacity to detect and rationally resolve normative conflicts, leaving them susceptible to manipulation; even recent advances in reasoning-focused LLMs have not addressed this vulnerability. This "shallow alignment" problem carries significant implications for AI safety and regulation, suggesting that current approaches are insufficient for mitigating potential harms posed by increasingly capable AI systems. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Philosophical Studies is a copyright of Springer, 2025. All Rights Reserved. item: Philosophical Studies holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|