Normative conflicts and shallow AI alignment.

The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing effor...

Descripción completa

Detalles Bibliográficos
Publicado en:Philosophical Studies Vol. 182; no. 7; pp. 2035 - 2079
Autor principal: Millière, Raphaël
Formato: Artículo
Publicado: Springer Nature Jul2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909912&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909912
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00318116
        4L8
      jtl: Philosophical Studies
      issn: 00318116
      maglogo: N
    pubinfo:
      dt: Jul2025
      vid: 182
      iid: 7
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909912
        10.1007/s11098-025-02347-3
      ppf: 2035
      ppct: 44
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.2MB
      tig:
        atl: Normative conflicts and shallow AI alignment.
      aug:
        au: Millière, Raphaël
        affil: https://ror.org/01sf06y89 Department of Philosophy, Macquarie University, 25B Wally's Walk, 2109, Sydney, NSW, Australia
      su:
        Artificial intelligence
        Digital technology
        Language models
        Ethics
        Social norms
      sug:
        subj:
          Artificial intelligence
          Digital technology
          Language models
          Ethics
          Social norms
      keyword:
        Adversarial attacks
        AI safety
        Alignment problem
        Large language models
        Normative reasoning
      ab: The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are fundamentally inadequate to prevent misuse. Despite ongoing efforts to instill norms such as helpfulness, honesty, and harmlessness in LLMs through fine-tuning based on human preferences, they remain vulnerable to adversarial attacks that exploit conflicts between these norms. I argue that this vulnerability reflects a fundamental limitation of existing alignment methods: they reinforce shallow behavioral dispositions rather than endowing LLMs with a genuine capacity for normative deliberation. Drawing from on research in moral psychology, I show how humans' ability to engage in deliberative reasoning enhances their resilience against similar adversarial tactics. LLMs, by contrast, lack a robust capacity to detect and rationally resolve normative conflicts, leaving them susceptible to manipulation; even recent advances in reasoning-focused LLMs have not addressed this vulnerability. This "shallow alignment" problem carries significant implications for AI safety and regulation, suggesting that current approaches are insufficient for mitigating potential harms posed by increasingly capable AI systems.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Philosophical Studies is a copyright of Springer, 2025. All Rights Reserved.
      item: Philosophical Studies
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N