Modeling under-resourced languages for speech recognition.

One particular problem in large vocabulary continuous speech recognition for low-resourced languages is finding relevant training data for the statistical language models. Large amount of data is required, because models should estimate the probability for all possible word sequences. For Finnish, E...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 51; no. 4; pp. 961 - 988
Main Authors: Kurimo, Mikko, Enarvi, Seppo, Varjokallio, Matti, Mansikkaniemi, André, Tilk, Ottokar, Alumäe, Tanel
Format: Article
Published: Springer Nature Dec2017
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=126259355&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 126259355
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2017
      vid: 51
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        126259355
        10.1007/s10579-016-9336-9
      ppf: 961
      ppct: 27
      formats:
        fmt:
          @attributes:
            type: P
            size: 533KB
      tig:
        atl: Modeling under-resourced languages for speech recognition.
      aug:
        au:
          Kurimo, Mikko
          Enarvi, Seppo
          Varjokallio, Matti
          Mansikkaniemi, André
          Tilk, Ottokar
          Alumäe, Tanel
        affil:
          Department of Signal Processing and Acoustics , Aalto University , Espoo Finland
          Institute of Cybernetics , Tallinn University of Technology , Tallinn Estonia
      su:
        Speech perception
        Vocabulary
        Finnish language
        Estonian language
        Data
        Artificial neural networks
        Internet
      sug:
        subj:
          Speech perception
          Vocabulary
          Finnish language
          Estonian language
          Data
          Artificial neural networks
          Internet
      keyword:
        Adaptation
        Data filtering
        Large vocabulary speech recognition
        Statistical language modeling
        Subword units
      ab: One particular problem in large vocabulary continuous speech recognition for low-resourced languages is finding relevant training data for the statistical language models. Large amount of data is required, because models should estimate the probability for all possible word sequences. For Finnish, Estonian and the other fenno-ugric languages a special problem with the data is the huge amount of different word forms that are common in normal speech. The same problem exists also in other language technology applications such as machine translation, information retrieval, and in some extent also in other morphologically rich languages. In this paper we present methods and evaluations in four recent language modeling topics: selecting conversational data from the Internet, adapting models for foreign words, multi-domain and adapted neural network language modeling, and decoding with subword units. Our evaluations show that the same methods work in more than one language and that they scale down to smaller data resources.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N