"You'll be a nurse, my son!" Automatically assessing gender biases in autoregressive language models in French and Italian.
Language models are now massively used for a variety of tasks, including open-ended generation and writing assistance. However, generated texts can encapsulate biases and harm users. A variety of articles aim at detecting, measuring and mitigating stereotypical biases, but focus mainly on English an...
| Publicado en: | Language Resources & Evaluation Vol. 59; no. 2; pp. 1495 - 1524 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=185240067&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 185240067 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2025 vid: 59 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 185240067 10.1007/s10579-024-09780-6 ppf: 1495 ppct: 29 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.6MB tig: atl: "You'll be a nurse, my son!" Automatically assessing gender biases in autoregressive language models in French and Italian. aug: au: Ducel, Fanny Névéol, Aurélie Fort, Karën affil: https://ror.org/03xjwb503 LISN, CNRS, Université Paris-Saclay, Orsay, France https://ror.org/02en5vm52 Sorbonne-Université, LORIA, Paris/Nancy, France su: Language models Italian language Language & languages Sex discrimination Autoregressive models sug: subj: Language models Italian language Language & languages Sex discrimination Autoregressive models keyword: Communication and Culture Linguistics French Gender Italian Language Language model Stereotypical biases ab: Language models are now massively used for a variety of tasks, including open-ended generation and writing assistance. However, generated texts can encapsulate biases and harm users. A variety of articles aim at detecting, measuring and mitigating stereotypical biases, but focus mainly on English and on pre-training tasks. Thus, we propose a framework to automatically measure gender biases generated by language models in inflected languages, in a practical setting. Herein, we report experiments using this framework on seven autoregressive language models used to generate more than 52,000 cover letters in French, addressing 203 industry and sectors, and over 4100 cover letters in Italian, on 55 sectors. Associations between occupation and gender are studied using a system that we introduce to automatically identify morpho-syntactic gender markers in text. Results suggest that all models are strongly biased towards the generation of texts containing masculine gender markers. Overall, generated texts contain twice as many masculine (vs. feminine) markers in French, and eight times as many in Italian. Models also exacerbate gender stereotypes that are evidenced in social science studies and associate feminine inflections with occupations related to care, children and physical appearance, whereas occupations that require physical, technical and manual skills are strongly associated with masculine markers. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|