Two sepedi-english code-switched speech corpora.
We report on the development of two reference corpora for the analysis of Sepedi-English code-switched speech in the context of automatic speech recognition. For the first corpus, possible English events were obtained from an existing corpus of transcribed Sepedi-English speech. The second corpus is...
| Published in: | Language Resources & Evaluation Vol. 56; no. 3; pp. 703 - 728 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Springer Nature
Sep2022
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158609445&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 158609445 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2022 vid: 56 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 158609445 10.1007/s10579-022-09592-6 ppf: 703 ppct: 25 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1MB tig: atl: Two sepedi-english code-switched speech corpora. aug: au: Modipa, Thipe I. Davel, Marelie H. affil: Department of Computer Science, University of Limpopo, University Road, Sovenga, Polokwane, South Africa Centre for Artificial Intelligence Research (CAIR), National Institute for Theoretical and Computational Sciences (NITheCS), Pretoria, South Africa Faculty of Engineering, North-West University, Potchefstroom, South Africa su: Automatic speech recognition Speech Code switching (Linguistics) Corpora Native language sug: subj: Automatic speech recognition Speech Code switching (Linguistics) Corpora Native language keyword: Code switching Multilingual speech recognition Sepedi Speech corpus ab: We report on the development of two reference corpora for the analysis of Sepedi-English code-switched speech in the context of automatic speech recognition. For the first corpus, possible English events were obtained from an existing corpus of transcribed Sepedi-English speech. The second corpus is based on the analysis of radio broadcasts: actual instances of code switching were transcribed and reproduced by a number of native Sepedi speakers. We describe the process to develop and verify both corpora and perform an initial analysis of the newly produced data sets. We find that, in naturally occurring speech, the frequency of code switching is unexpectedly high for this language pair, and that the continuum of code switching (from unmodified embedded words to loanwords absorbed into the matrix language) makes this a particularly challenging task for speech recognition systems. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|