The KAS corpus of Slovenian academic writing.
The paper presents the KAS corpus of Slovenian academic writing, which consists of almost 65,000 B.A./B.Sc., 16,000 M.A./M.Sc. and 1600 Ph.D. theses (5 million pages or 1.7 billion tokens) gathered from the digital libraries of Slovenian higher education institutions via the Slovenian Open Science p...
| Publicado en: | Language Resources & Evaluation Vol. 55; no. 2; pp. 551 - 584 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2021
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=150471571&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 150471571 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2021 vid: 55 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 150471571 10.1007/s10579-020-09506-4 ppf: 551 ppct: 33 formats: fmt: @attributes: type: P size: 1.2MB tig: atl: The KAS corpus of Slovenian academic writing. aug: au: Erjavec, Tomaž Fišer, Darja Ljubešić, Nikola affil: Department of Knowledge Technologies, Jožef Stefan Institute, Jamova cesta 39, 1000, Ljubljana, Slovenia Department of Translation, Faculty of Arts, University of Ljubljana, Aškerčeva cesta 2, 1000, Ljubljana, Slovenia Faculty of Computer Science and Informatics, University of Ljubljana, Večna pot 113, 1000, Ljubljana, Slovenia su: Academic discourse Digital libraries Corpora Universities & colleges Gene ontology sug: subj: Academic discourse Digital libraries Corpora Universities & colleges Gene ontology keyword: Academic writing Corpus Slovenian TEI Terminology ab: The paper presents the KAS corpus of Slovenian academic writing, which consists of almost 65,000 B.A./B.Sc., 16,000 M.A./M.Sc. and 1600 Ph.D. theses (5 million pages or 1.7 billion tokens) gathered from the digital libraries of Slovenian higher education institutions via the Slovenian Open Science portal. We discuss the compilation, meta-data, annotation, and distribution of the corpus, which is made freely available via on-line concordancers and is openly available for research through the CLARIN.SI research infrastructure. We also present the tools for mono- and bilingual term extraction and for thesis structure annotation that were developed in the scope of the project, including the manually annotated datasets used to train these tools. This specialised corpus, large by any standards, represents a substantial and highly useful language resource for the study of Slovenian academic writing and for terminology extraction. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|