Racial Bias in AI Training Data: Do Laypersons Notice?

Given that the nature of training data is the primary cause of algorithmic bias, do laypersons realize that systematic misrepresentation and under-representation of certain races in the training data can affect AI performance in a way that privileges some races over others? To answer this question,...

Full description

Bibliographic Details
Published in:Media Psychology Vol. 29; no. 5; pp. 981 - 1009
Main Authors: Chen, Cheng, Jang, Eunchae, Sundar, S. Shyam
Format: Article
Published: Taylor & Francis Ltd 2026
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=196698591&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 196698591
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        15213269
        7MN
      jtl: Media Psychology
      issn: 15213269
      maglogo: N
    pubinfo:
      dt: 2026
      vid: 29
      iid: 5
      pid: 377
      pub: Taylor & Francis Ltd
    artinfo:
      ui:
        196698591
        10.1080/15213269.2025.2558036
      ppf: 981
      ppct: 28
      formats:
      tig:
        atl: Racial Bias in AI Training Data: Do Laypersons Notice?
      aug:
        au:
          Chen, Cheng
          Jang, Eunchae
          Sundar, S. Shyam
        affil:
          School of Communication, Oregon State University, Corvallis, OR, USA
          Media Effects Research Lab, Bellisario College of Communications, Pennsylvania State University, University Park, PA, USA
          Department of Immersive Media Engineering, Sungkyunkwan University, Seoul, Republic of Korea
      su:
        Algorithmic bias
        Emotion recognition
        Racism
        Public opinion
        Metadata
        Cognitive bias
      sug:
        subj:
          Algorithmic bias
          Emotion recognition
          Racism
          Public opinion
          Metadata
          Cognitive bias
      ab: Given that the nature of training data is the primary cause of algorithmic bias, do laypersons realize that systematic misrepresentation and under-representation of certain races in the training data can affect AI performance in a way that privileges some races over others? To answer this question, we conducted three between-subjects online experiments (N = 769 in total) with a prototype of an AI system that recognizes emotion-based facial expressions. Our results show that, by and large, training data representativeness is not an effective cue to communicate algorithmic bias. Instead, users rely on AI's performance bias to perceive racial bias in AI algorithms. In addition, the race of the users matters. Black participants perceive the system to be more biased when all facial images used to represent unhappy emotions in the training data are those of Black individuals. This finding highlights a significant human cognitive limitation that should be accounted for when communicating algorithmic bias arising from biases in the training data.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N