The corpus of aggressive language in Polish parliamentary debates.
This paper presents a new resource: a publicly available corpus of aggressive statements in Polish. Sourced from parliamentary speeches, it also includes annotations of additional semantic layers related to credibility, designed to facilitate research on aggressive language. We outline the corpus's...
| Published in: | Language Resources & Evaluation Vol. 59; no. 4; pp. 4043 - 4068 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Springer Nature
Dec2025
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912042&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 189912042 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2025 vid: 59 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 189912042 10.1007/s10579-025-09870-z ppf: 4043 ppct: 25 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: The corpus of aggressive language in Polish parliamentary debates. aug: au: Sarzyńska-Wawer, Justyna Wawer, Aleksander affil: https://ror.org/01dr6c206 Institute of Psychology, Polish Academy of Sciences, Jaracza 1, 00-378, Warszawa, Poland https://ror.org/01dr6c206 Institute of Computer Science, Polish Academy of Sciences, Jana Kazimierza 5, 01-248, Warszawa, Poland su: Machine learning Language models Annotations Corpora Automatic classification sug: subj: Machine learning Language models Annotations Corpora Automatic classification keyword: Aggression corpus Aggression detection Information and Computing Sciences Artificial Intelligence and Image Processing Political discourse ab: This paper presents a new resource: a publicly available corpus of aggressive statements in Polish. Sourced from parliamentary speeches, it also includes annotations of additional semantic layers related to credibility, designed to facilitate research on aggressive language. We outline the corpus's design and the challenges encountered during its creation, the most crucial one being the relative rarity of aggressive statements. We detail multiple approaches to identify potentially aggressive statements: the first is a selection based on a generic dictionary of negative sentiment, and the second is based on a dedicated dictionary of words potentially linked to aggression. We compare both methods. Finally, we present machine learning experiments on aggressive statement recognition using our corpus to fine-tune a pre-trained language model (PLM) and a recurrent neural network baseline. Both model types reveal promising results in terms of metrics such as F1, precision, and recall, with an advantage of the PLM model. Based on these results and qualitative analysis performed with a demo model, we conclude that the corpus is a usable resource to train models for automated recognition of certain types of aggression in the Polish language. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|