Examining inferred author and textual correlates of harmful language annotation.

This study examines whether the psycholinguistic and demographic characteristics of authors of online texts are correlated with the way harmful language, such as toxicity and hate speech, is judged. We apply artificial intelligence models to two harmful language datasets, Jigsaw's Special Rater Pool...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 59; no. 4; pp. 3411 - 3443
Autores principales: Korre, Katerina, Yenikent, Seren, Basile, Angelo, Spallaccia, Beatrice, Franco-Salvador, Marc, Barrón-Cedeño, Alberto
Formato: Artículo
Publicado: Springer Nature Dec2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912015&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 189912015
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 59
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        189912015
        10.1007/s10579-025-09822-7
      ppf: 3411
      ppct: 32
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.2MB
      tig:
        atl: Examining inferred author and textual correlates of harmful language annotation.
      aug:
        au:
          Korre, Katerina
          Yenikent, Seren
          Basile, Angelo
          Spallaccia, Beatrice
          Franco-Salvador, Marc
          Barrón-Cedeño, Alberto
        affil:
          https://ror.org/01111rn36 DIT, Università di Bologna, Corso della Repubblica 136, 47121, Forlì, Italy
          ai-mind.solutions, Tartu mnt 67/1-13b, 10115, Tallinn, Estonia
          Symanto Research, Calle Reina 12, 46011, Valencia, Spain
          https://ror.org/01460j859 Universitat Politècnica de València, Camino de Vera, 46022, Valencia, Spain
          United Nations International Computing Centre (UNICC), Avinguda Comarques del País Valencià 2, 46930, Quart de Poblet, Valencia, Spain
      su:
        Hate speech
        Psycholinguistics
        Toxicity testing
        Discriminatory language
        Content analysis
        Sociocultural factors
        Text mining
        Machine learning
      sug:
        subj:
          Hate speech
          Psycholinguistics
          Toxicity testing
          Discriminatory language
          Content analysis
          Sociocultural factors
          Text mining
          Machine learning
      keyword:
        Bias
        Communication and Culture Linguistics
        Language
        Psycholinguistic factors in hate speech labelling
        Toxicity labelling
      ab: This study examines whether the psycholinguistic and demographic characteristics of authors of online texts are correlated with the way harmful language, such as toxicity and hate speech, is judged. We apply artificial intelligence models to two harmful language datasets, Jigsaw's Special Rater Pool dataset and the Measuring Hate Speech dataset, to generate probabilities for different text aspects, namely inferring demographic information of the author behind the suspicious text in terms of age and gender, as well as the expressed emotions, emotionality, sentiment and communication style. We then perform a statistical regression analysis to examine how these text aspects correlate with the perception of hate speech and toxicity during the annotation process. The study shows that while the frequency of the psycholinguistic text aspects that can be derived from the author's personality does not differ significantly between harmful and non-harmful classes, the inferred text aspects are statistically associated with the annotators' perception of harmful language and could potentially influence the way annotators label the texts.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N