The corpus of aggressive language in Polish parliamentary debates.

This paper presents a new resource: a publicly available corpus of aggressive statements in Polish. Sourced from parliamentary speeches, it also includes annotations of additional semantic layers related to credibility, designed to facilitate research on aggressive language. We outline the corpus's...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 59; no. 4; pp. 4043 - 4068
Main Authors: Sarzyńska-Wawer, Justyna, Wawer, Aleksander
Format: Article
Published: Springer Nature Dec2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=189912042&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 189912042
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2025
      vid: 59
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        189912042
        10.1007/s10579-025-09870-z
      ppf: 4043
      ppct: 25
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.1MB
      tig:
        atl: The corpus of aggressive language in Polish parliamentary debates.
      aug:
        au:
          Sarzyńska-Wawer, Justyna
          Wawer, Aleksander
        affil:
          https://ror.org/01dr6c206 Institute of Psychology, Polish Academy of Sciences, Jaracza 1, 00-378, Warszawa, Poland
          https://ror.org/01dr6c206 Institute of Computer Science, Polish Academy of Sciences, Jana Kazimierza 5, 01-248, Warszawa, Poland
      su:
        Machine learning
        Language models
        Annotations
        Corpora
        Automatic classification
      sug:
        subj:
          Machine learning
          Language models
          Annotations
          Corpora
          Automatic classification
      keyword:
        Aggression corpus
        Aggression detection
        Information and Computing Sciences Artificial Intelligence and Image Processing
        Political discourse
      ab: This paper presents a new resource: a publicly available corpus of aggressive statements in Polish. Sourced from parliamentary speeches, it also includes annotations of additional semantic layers related to credibility, designed to facilitate research on aggressive language. We outline the corpus's design and the challenges encountered during its creation, the most crucial one being the relative rarity of aggressive statements. We detail multiple approaches to identify potentially aggressive statements: the first is a selection based on a generic dictionary of negative sentiment, and the second is based on a dedicated dictionary of words potentially linked to aggression. We compare both methods. Finally, we present machine learning experiments on aggressive statement recognition using our corpus to fine-tune a pre-trained language model (PLM) and a recurrent neural network baseline. Both model types reveal promising results in terms of metrics such as F1, precision, and recall, with an advantage of the PLM model. Based on these results and qualitative analysis performed with a demo model, we conclude that the corpus is a usable resource to train models for automated recognition of certain types of aggression in the Polish language.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N