Does calibration mean what they say it means; or, the reference class problem rises again.

Discussions of statistical criteria for fairness commonly convey the normative significance of calibration within groups by invoking what risk scores "mean." On the Same Meaning picture, group-calibrated scores "mean the same thing" (on average) across individuals from different groups and according...

Descripción completa

Detalles Bibliográficos
Publicado en:Philosophical Studies Vol. 182; no. 5; pp. 1305 - 1332
Autor principal: Hu, Lily
Formato: Artículo
Publicado: Springer Nature Jun2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185649422&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 185649422
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00318116
        4L8
      jtl: Philosophical Studies
      issn: 00318116
      maglogo: N
    pubinfo:
      dt: Jun2025
      vid: 182
      iid: 5
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        185649422
        10.1007/s11098-025-02322-y
      ppf: 1305
      ppct: 27
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 894KB
      tig:
        atl: Does calibration mean what they say it means; or, the reference class problem rises again.
      aug:
        au: Hu, Lily
        affil: https://ror.org/03v76x132 Department of Philosophy, Yale University, 451 College St., 06511, New Haven, CT, USA
      su:
        Calibration
        Justification (Ethics)
        Metaphysics
        Philosophy periodicals
        Philosophy
      sug:
        subj:
          Calibration
          Justification (Ethics)
          Metaphysics
          Philosophy periodicals
          Philosophy
      keyword:
        Algorithmic fairness
        Reference class problem
      ab: Discussions of statistical criteria for fairness commonly convey the normative significance of calibration within groups by invoking what risk scores "mean." On the Same Meaning picture, group-calibrated scores "mean the same thing" (on average) across individuals from different groups and accordingly, guard against disparate treatment of individuals based on group membership. My contention is that calibration guarantees no such thing. Since concrete actual people belong to many groups, calibration cannot ensure the kind of consistent score interpretation that the Same Meaning picture implies matters for fairness, unless calibration is met within every group to which an individual belongs. Alas only perfect predictors may meet this bar. The Same Meaning picture thus commits a reference class fallacy by inferring from calibration within some group to the "meaning" or evidential value of an individual's score, because they are a member of that group. The reference class answer it presumes does not only lack justification; it is very likely wrong. I then show that the reference class problem besets not just calibration but other group statistical criteria that claim a close connection to fairness. Reflecting on the origins of this oversight opens a wider lens onto the predominant methodology in algorithmic fairness based on stylized cases.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Philosophical Studies is a copyright of Springer, 2025. All Rights Reserved.
      item: Philosophical Studies
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N