Examining inferred author and textual correlates of harmful language annotation.
This study examines whether the psycholinguistic and demographic characteristics of authors of online texts are correlated with the way harmful language, such as toxicity and hate speech, is judged. We apply artificial intelligence models to two harmful language datasets, Jigsaw's Special Rater Pool...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 4; pp. 3411 - 3443 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912015&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 189912015 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2025 vid: 59 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 189912015 10.1007/s10579-025-09822-7 ppf: 3411 ppct: 32 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.2MB tig: atl: Examining inferred author and textual correlates of harmful language annotation. aug: au: Korre, Katerina Yenikent, Seren Basile, Angelo Spallaccia, Beatrice Franco-Salvador, Marc Barrón-Cedeño, Alberto affil: https://ror.org/01111rn36 DIT, Università di Bologna, Corso della Repubblica 136, 47121, Forlì, Italy ai-mind.solutions, Tartu mnt 67/1-13b, 10115, Tallinn, Estonia Symanto Research, Calle Reina 12, 46011, Valencia, Spain https://ror.org/01460j859 Universitat Politècnica de València, Camino de Vera, 46022, Valencia, Spain United Nations International Computing Centre (UNICC), Avinguda Comarques del País Valencià 2, 46930, Quart de Poblet, Valencia, Spain su: Hate speech Psycholinguistics Toxicity testing Discriminatory language Content analysis Sociocultural factors Text mining Machine learning sug: subj: Hate speech Psycholinguistics Toxicity testing Discriminatory language Content analysis Sociocultural factors Text mining Machine learning keyword: Bias Communication and Culture Linguistics Language Psycholinguistic factors in hate speech labelling Toxicity labelling ab: This study examines whether the psycholinguistic and demographic characteristics of authors of online texts are correlated with the way harmful language, such as toxicity and hate speech, is judged. We apply artificial intelligence models to two harmful language datasets, Jigsaw's Special Rater Pool dataset and the Measuring Hate Speech dataset, to generate probabilities for different text aspects, namely inferring demographic information of the author behind the suspicious text in terms of age and gender, as well as the expressed emotions, emotionality, sentiment and communication style. We then perform a statistical regression analysis to examine how these text aspects correlate with the perception of hate speech and toxicity during the annotation process. The study shows that while the frequency of the psycholinguistic text aspects that can be derived from the author's personality does not differ significantly between harmful and non-harmful classes, the inferred text aspects are statistically associated with the annotators' perception of harmful language and could potentially influence the way annotators label the texts. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|