Finnish parliament ASR corpus: Analysis, benchmarks and statistics.

Public sources like parliament meeting recordings and transcripts provide ever-growing material for the training and evaluation of automatic speech recognition (ASR) systems. In this paper, we publish and analyse the Finnish Parliament ASR Corpus, the most extensive publicly available collection of...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 4; pp. 1645 - 1671
Autores principales: Virkkunen, Anja, Rouhe, Aku, Phan, Nhan, Kurimo, Mikko
Formato: Artículo
Publicado: Springer Nature Dec2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=173723395&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 173723395
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2023
      vid: 57
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        173723395
        10.1007/s10579-023-09650-7
      ppf: 1645
      ppct: 26
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.1MB
      tig:
        atl: Finnish parliament ASR corpus: Analysis, benchmarks and statistics.
      aug:
        au:
          Virkkunen, Anja
          Rouhe, Aku
          Phan, Nhan
          Kurimo, Mikko
        affil: https://ror.org/020hwjq30 Department of Information and Communications Engineering, Aalto University, Espoo, Finland
      su:
        Automatic speech recognition
        Artificial neural networks
        Hidden Markov models
        Delay lines
        Speech perception
        Legislative bodies
      sug:
        subj:
          Automatic speech recognition
          Artificial neural networks
          Hidden Markov models
          Delay lines
          Speech perception
          Legislative bodies
      keyword:
        AED
        Finnish
        HMM-DNN
        Metadata
        Parliament speech data
        Speech recognition
        Wav2vec
      ab: Public sources like parliament meeting recordings and transcripts provide ever-growing material for the training and evaluation of automatic speech recognition (ASR) systems. In this paper, we publish and analyse the Finnish Parliament ASR Corpus, the most extensive publicly available collection of manually transcribed speech data for Finnish with over 3000 h of speech and 449 speakers for which it provides rich demographic metadata. This corpus builds on earlier initial work, and as a result the corpus has a natural split into two training subsets from two periods of time. Similarly, there are two official, corrected test sets covering different times, setting an ASR task with longitudinal distribution-shift characteristics. An official development set is also provided. We developed a complete Kaldi-based data preparation pipeline and ASR recipes for hidden Markov models (HMM), hybrid deep neural networks (HMM-DNN), and attention-based encoder-decoders (AED). For HMM-DNN systems, we provide results with time-delay neural networks (TDNN) as well as state-of-the-art wav2vec 2.0 pretrained acoustic models. We set benchmarks on the official test sets and multiple other recently used test sets. Both temporal corpus subsets are already large, and we observe that beyond their scale, HMM-TDNN ASR performance on the official test sets has reached a plateau. In contrast, other domains and larger wav2vec 2.0 models benefit from added data. The HMM-DNN and AED approaches are compared in a carefully matched equal data setting, with the HMM-DNN system consistently performing better. Finally, the variation of the ASR accuracy is compared between the speaker categories available in the parliament metadata to detect potential biases based on factors such as gender, age, and education.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2023. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N