Two sepedi-english code-switched speech corpora.

We report on the development of two reference corpora for the analysis of Sepedi-English code-switched speech in the context of automatic speech recognition. For the first corpus, possible English events were obtained from an existing corpus of transcribed Sepedi-English speech. The second corpus is...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 56; no. 3; pp. 703 - 728
Main Authors: Modipa, Thipe I., Davel, Marelie H.
Format: Article
Published: Springer Nature Sep2022
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=158609445&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 158609445
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2022
      vid: 56
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        158609445
        10.1007/s10579-022-09592-6
      ppf: 703
      ppct: 25
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1MB
      tig:
        atl: Two sepedi-english code-switched speech corpora.
      aug:
        au:
          Modipa, Thipe I.
          Davel, Marelie H.
        affil:
          Department of Computer Science, University of Limpopo, University Road, Sovenga, Polokwane, South Africa
          Centre for Artificial Intelligence Research (CAIR), National Institute for Theoretical and Computational Sciences (NITheCS), Pretoria, South Africa
          Faculty of Engineering, North-West University, Potchefstroom, South Africa
      su:
        Automatic speech recognition
        Speech
        Code switching (Linguistics)
        Corpora
        Native language
      sug:
        subj:
          Automatic speech recognition
          Speech
          Code switching (Linguistics)
          Corpora
          Native language
      keyword:
        Code switching
        Multilingual speech recognition
        Sepedi
        Speech corpus
      ab: We report on the development of two reference corpora for the analysis of Sepedi-English code-switched speech in the context of automatic speech recognition. For the first corpus, possible English events were obtained from an existing corpus of transcribed Sepedi-English speech. The second corpus is based on the analysis of radio broadcasts: actual instances of code switching were transcribed and reproduced by a number of native Sepedi speakers. We describe the process to develop and verify both corpora and perform an initial analysis of the newly produced data sets. We find that, in naturally occurring speech, the frequency of code switching is unexpectedly high for this language pair, and that the continuum of code switching (from unmodified embedded words to loanwords absorbed into the matrix language) makes this a particularly challenging task for speech recognition systems.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N