Collecting and evaluating speech recognition corpora for 11 South African languages.

We describe the Lwazi corpus for automatic speech recognition (ASR), a new telephone speech corpus which contains data from the eleven official languages of South Africa. Because of practical constraints, the amount of speech per language is relatively small compared to major corpora in world langua...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 45; no. 3; pp. 289 - 310
Autores principales: Badenhorst, Jaco, Heerden, Charl, Davel, Marelie, Barnard, Etienne
Formato: Artículo
Publicado: Springer Nature Aug2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=63899363&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 63899363
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Aug2011
      vid: 45
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        63899363
        10.1007/s10579-011-9152-1
      ppf: 289
      ppct: 21
      formats:
        fmt:
          @attributes:
            type: P
            size: 811KB
      tig:
        atl: Collecting and evaluating speech recognition corpora for 11 South African languages.
      aug:
        au:
          Badenhorst, Jaco
          Heerden, Charl
          Davel, Marelie
          Barnard, Etienne
        affil:
          Human Language Technology Competency Area, CSIR Meraka Institute, Meiring Naude Road Pretoria South Africa
          Multilingual Speech Technologies, North-West University, Vanderbijlpark 1900 South Africa
      su:
        Speech perception
        Corpora
        African languages
        Language policy
        Phoneme (Linguistics)
        South Africa
      sug:
        subj:
          South Africa
          Speech perception
          Corpora
          African languages
          Language policy
          Phoneme (Linguistics)
      keyword:
        Lwazi corpus
        Resource-scarce languages
        South African languages
        Speech recognition
      ab: We describe the Lwazi corpus for automatic speech recognition (ASR), a new telephone speech corpus which contains data from the eleven official languages of South Africa. Because of practical constraints, the amount of speech per language is relatively small compared to major corpora in world languages, and we report on our investigation of the stability of the ASR models derived from the corpus. We also report on phoneme distance measures across languages, and describe initial phone recognisers that were developed using this data. We find that a surprisingly small number of speakers (fewer than 50) and around 10 to 20 h of speech per language are sufficient for the purposes of acceptable phone-based recognition.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N