The ParlaMint corpora of parliamentary proceedings.
This paper presents the ParlaMint corpora containing transcriptions of the sessions of the 17 European national parliaments with half a billion words. The corpora are uniformly encoded, contain rich meta-data about 11 thousand speakers, and are linguistically annotated following the Universal Depend...
| Publicado en: | Language Resources & Evaluation Vol. 57; no. 1; pp. 415 - 449 |
|---|---|
| Autores principales: | , , , , , , , , , , , , , , , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Mar2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162506756&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 162506756 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Mar2023 vid: 57 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 162506756 10.1007/s10579-021-09574-0 ppf: 415 ppct: 34 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.1MB tig: atl: The ParlaMint corpora of parliamentary proceedings. aug: au: Erjavec, Tomaž Ogrodniczuk, Maciej Osenova, Petya Ljubešić, Nikola Simov, Kiril Pančur, Andrej Rudolf, Michał Kopp, Matyáš Barkarson, Starkaður Steingrímsson, Steinþór Çöltekin, Çağrı de Does, Jesse Depuydt, Katrien Agnoloni, Tommaso Venturi, Giulia Pérez, María Calzada de Macedo, Luciana D. Navarretta, Costanza Luxardo, Giancarlo Coole, Matthew affil: Department of Knowledge Technologies, Jožef Stefan Institute, Ljubljana, Slovenia Institute of Computer Science, Polish Academy of Sciences, Warsaw, Poland Institute of Information and Communication Technologies, Bulgarian Academy of Sciences, and Sofia University "St. Kl. Ohridski", Sofia, Bulgaria Department of Knowledge Technologies, Jožef Stefan Institute and Faculty of Computer Science and Informatics, University of Ljubljana, Ljubljana, Slovenia Institute of Information and Communication Technologies, Bulgarian Academy of Sciences, Sofia, Bulgaria Institute for Contemporay History, Ljubljana, Slovenia Institute of Formal and Applied Linguistics, Faculty of Mathematics and Physics, Charles University, Prague, Czech Republic The Árni Magnússon Institute for Icelandic Studies, Reykjavík, Iceland University of Tübingen, Tübingen, Germany Dutch Language Institute, Hague, The Netherlands Institute of Legal Informatics and Judicial Systems CNR-IGSG, Florence, Italy Institute of Computational Linguistics CNR-ILC, Pis, Italy Universitat Jaume I, Castellón de la Plana, Spain Univ. Federal de Minas Gerais, Belo Horizonte, Brazil University of Copenhagen, Copenhagen, Denmark Univ. Paul Valéry Montpellier 3, Montpellier, France Lancaster University, Lancaster, UK su: Corpora Metadata Legislative bodies Scripts sug: subj: Corpora Metadata Legislative bodies Scripts keyword: Comparable corpora Parliamentary proceedings TEI ab: This paper presents the ParlaMint corpora containing transcriptions of the sessions of the 17 European national parliaments with half a billion words. The corpora are uniformly encoded, contain rich meta-data about 11 thousand speakers, and are linguistically annotated following the Universal Dependencies formalism and with named entities. Samples of the corpora and conversion scripts are available from the project's GitHub repository, and the complete corpora are openly available via the CLARIN.SI repository for download, as well as through the NoSketch Engine and KonText concordancers and the Parlameter interface for on-line exploration and analysis. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|