CORAA ASR: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese.

Automatic Speech recognition (ASR) is a complex and challenging task. In recent years, there have been significant advances in the area. In particular, for the Brazilian Portuguese (BP) language, there were around 376 h publicly available for the ASR task until the second half of 2020. With the rele...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 3; pp. 1139 - 1172
Autores principales: Candido Junior, Arnaldo, Casanova, Edresson, Soares, Anderson, de Oliveira, Frederico Santos, Oliveira, Lucas, Junior, Ricardo Corso Fernandes, da Silva, Daniel Peixoto Pinto, Fayet, Fernando Gorgulho, Carlotto, Bruno Baldissera, Gris, Lucas Rafael Stefanel, Aluísio, Sandra Maria
Formato: Artículo
Publicado: Springer Nature Sep2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=170029257&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 170029257
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2023
      vid: 57
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        170029257
        10.1007/s10579-022-09621-4
      ppf: 1139
      ppct: 33
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1MB
      tig:
        atl: CORAA ASR: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese.
      aug:
        au:
          Candido Junior, Arnaldo
          Casanova, Edresson
          Soares, Anderson
          de Oliveira, Frederico Santos
          Oliveira, Lucas
          Junior, Ricardo Corso Fernandes
          da Silva, Daniel Peixoto Pinto
          Fayet, Fernando Gorgulho
          Carlotto, Bruno Baldissera
          Gris, Lucas Rafael Stefanel
          Aluísio, Sandra Maria
        affil:
          São Paulo State University, São José do Rio Preto, Brazil
          Instituto de Ciências Matemáticas e de Computação - University of São Paulo, São Carlos, Brazil
          Federal University of Goias, Goiânia, Brazil
          Federal University of Technology — Paraná (UTFPR), Medianeira, Brazil
      su:
        Automatic speech recognition
        Speech
        Portuguese language
        Speech perception
        Corpora
        Freedom of speech
        Error rates
      sug:
        subj:
          Automatic speech recognition
          Speech
          Portuguese language
          Speech perception
          Corpora
          Freedom of speech
          Error rates
      keyword:
        Brazilian Portuguese
        Prepared speech
        Public datasets
        Public speech corpora
        Spontaneous speech
      ab: Automatic Speech recognition (ASR) is a complex and challenging task. In recent years, there have been significant advances in the area. In particular, for the Brazilian Portuguese (BP) language, there were around 376 h publicly available for the ASR task until the second half of 2020. With the release of new datasets in early 2021, this number increased to 574 h. The existing resources, however, are composed of audios containing only read and prepared speech. There is a lack of datasets including spontaneous speech, which are essential in several ASR applications. This paper presents CORAA (Corpus of Annotated Audios) ASR with 290 h, a publicly available dataset for ASR in BP containing validated pairs of audio-transcription. CORAA ASR also contains European Portuguese audios (4.6 h). We also present a public ASR model based on Wav2Vec 2.0 XLSR-53, fine-tuned over CORAA ASR. Our model achieved a Word Error Rate (WER) of 24.18% on CORAA ASR test set and 20.08% on Common Voice test set. When measuring the Character Error Rate (CER), we obtained 11.02% and 6.34% for CORAA ASR and Common Voice, respectively. CORAA ASR corpora were assembled to both improve ASR models in BP with phenomena from spontaneous speech and motivate young researchers to start their studies on ASR for Portuguese. All the corpora are publicly available at https://github.com/nilc-nlp/CORAA under the CC BY-NC-ND 4.0 license.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N