The ParlaMint corpora of parliamentary proceedings.

This paper presents the ParlaMint corpora containing transcriptions of the sessions of the 17 European national parliaments with half a billion words. The corpora are uniformly encoded, contain rich meta-data about 11 thousand speakers, and are linguistically annotated following the Universal Depend...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 1; pp. 415 - 449
Autores principales: Erjavec, Tomaž, Ogrodniczuk, Maciej, Osenova, Petya, Ljubešić, Nikola, Simov, Kiril, Pančur, Andrej, Rudolf, Michał, Kopp, Matyáš, Barkarson, Starkaður, Steingrímsson, Steinþór, Çöltekin, Çağrı, de Does, Jesse, Depuydt, Katrien, Agnoloni, Tommaso, Venturi, Giulia, Pérez, María Calzada, de Macedo, Luciana D., Navarretta, Costanza, Luxardo, Giancarlo, Coole, Matthew
Formato: Artículo
Publicado: Springer Nature Mar2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162506756&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 162506756
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Mar2023
      vid: 57
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        162506756
        10.1007/s10579-021-09574-0
      ppf: 415
      ppct: 34
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.1MB
      tig:
        atl: The ParlaMint corpora of parliamentary proceedings.
      aug:
        au:
          Erjavec, Tomaž
          Ogrodniczuk, Maciej
          Osenova, Petya
          Ljubešić, Nikola
          Simov, Kiril
          Pančur, Andrej
          Rudolf, Michał
          Kopp, Matyáš
          Barkarson, Starkaður
          Steingrímsson, Steinþór
          Çöltekin, Çağrı
          de Does, Jesse
          Depuydt, Katrien
          Agnoloni, Tommaso
          Venturi, Giulia
          Pérez, María Calzada
          de Macedo, Luciana D.
          Navarretta, Costanza
          Luxardo, Giancarlo
          Coole, Matthew
        affil:
          Department of Knowledge Technologies, Jožef Stefan Institute, Ljubljana, Slovenia
          Institute of Computer Science, Polish Academy of Sciences, Warsaw, Poland
          Institute of Information and Communication Technologies, Bulgarian Academy of Sciences, and Sofia University "St. Kl. Ohridski", Sofia, Bulgaria
          Department of Knowledge Technologies, Jožef Stefan Institute and Faculty of Computer Science and Informatics, University of Ljubljana, Ljubljana, Slovenia
          Institute of Information and Communication Technologies, Bulgarian Academy of Sciences, Sofia, Bulgaria
          Institute for Contemporay History, Ljubljana, Slovenia
          Institute of Formal and Applied Linguistics, Faculty of Mathematics and Physics, Charles University, Prague, Czech Republic
          The Árni Magnússon Institute for Icelandic Studies, Reykjavík, Iceland
          University of Tübingen, Tübingen, Germany
          Dutch Language Institute, Hague, The Netherlands
          Institute of Legal Informatics and Judicial Systems CNR-IGSG, Florence, Italy
          Institute of Computational Linguistics CNR-ILC, Pis, Italy
          Universitat Jaume I, Castellón de la Plana, Spain
          Univ. Federal de Minas Gerais, Belo Horizonte, Brazil
          University of Copenhagen, Copenhagen, Denmark
          Univ. Paul Valéry Montpellier 3, Montpellier, France
          Lancaster University, Lancaster, UK
      su:
        Corpora
        Metadata
        Legislative bodies
        Scripts
      sug:
        subj:
          Corpora
          Metadata
          Legislative bodies
          Scripts
      keyword:
        Comparable corpora
        Parliamentary proceedings
        TEI
      ab: This paper presents the ParlaMint corpora containing transcriptions of the sessions of the 17 European national parliaments with half a billion words. The corpora are uniformly encoded, contain rich meta-data about 11 thousand speakers, and are linguistically annotated following the Universal Dependencies formalism and with named entities. Samples of the corpora and conversion scripts are available from the project's GitHub repository, and the complete corpora are openly available via the CLARIN.SI repository for download, as well as through the NoSketch Engine and KonText concordancers and the Parlameter interface for on-line exploration and analysis.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N