Developing computational infrastructure for the CorCenCC corpus: The National Corpus of Contemporary Welsh.
CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes—National Corpus of Contemporary Welsh) is the first comprehensive corpus of Welsh designed to be reflective of language use across communication types, genres, speakers, language varieties (regional and social) and contexts. This article focuses on the co...
| Publicado en: | Language Resources & Evaluation Vol. 55; no. 3; pp. 789 - 817 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2021
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=151686293&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 151686293 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2021 vid: 55 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 151686293 10.1007/s10579-020-09501-9 ppf: 789 ppct: 28 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.9MB tig: atl: Developing computational infrastructure for the CorCenCC corpus: The National Corpus of Contemporary Welsh. aug: au: Knight, Dawn Loizides, Fernando Neale, Steven Anthony, Laurence Spasić, Irena affil: School of English, Communication and Language Sciences, Cardiff University, CF10 3EU, Cardiff, UK School of Computer Science and Informatics, Cardiff University, CF24 3AA, Cardiff, UK Faculty of Science and Engineering, Waseda University, Tokyo, Japan su: Linguistic minorities Natural language processing Corpora Mobile apps Metadata sug: subj: Linguistic minorities Natural language processing Corpora Mobile apps Metadata keyword: Data modelling Information retrieval Language resources Usability testing Web interfaces ab: CorCenCC (Corpws Cenedlaethol Cymraeg Cyfoes—National Corpus of Contemporary Welsh) is the first comprehensive corpus of Welsh designed to be reflective of language use across communication types, genres, speakers, language varieties (regional and social) and contexts. This article focuses on the computational infrastructure that we have designed to support data collection for CorCenCC, and the subsequent uses of the corpus which include lexicography, pedagogical research and corpus analysis. A grass-roots approach to design has been adopted, that has adapted and extended previous corpus-building and introduced new features as required for this specific context and language. The key pillars of the infrastructure include a framework that supports metadata collection, an innovative mobile application designed to collect spoken data (utilising a crowdsourcing approach), a backend database that stores curated data and a web-based interface that allows users to query the data online. A usability study was conducted to evaluate the user facing tools and to suggest directions for future improvements. Though the infrastructure was developed for Welsh language collection, its design can be re-used to support corpus development in other minority or major language contexts, broadening the potential utility and impact of this work. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2021. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2021 holdings: @attributes: islocal: N |
|---|