Modeling under-resourced languages for speech recognition.
One particular problem in large vocabulary continuous speech recognition for low-resourced languages is finding relevant training data for the statistical language models. Large amount of data is required, because models should estimate the probability for all possible word sequences. For Finnish, E...
| Published in: | Language Resources & Evaluation Vol. 51; no. 4; pp. 961 - 988 |
|---|---|
| Main Authors: | , , , , , |
| Format: | Article |
| Published: |
Springer Nature
Dec2017
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=126259355&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 126259355 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2017 vid: 51 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 126259355 10.1007/s10579-016-9336-9 ppf: 961 ppct: 27 formats: fmt: @attributes: type: P size: 533KB tig: atl: Modeling under-resourced languages for speech recognition. aug: au: Kurimo, Mikko Enarvi, Seppo Varjokallio, Matti Mansikkaniemi, André Tilk, Ottokar Alumäe, Tanel affil: Department of Signal Processing and Acoustics , Aalto University , Espoo Finland Institute of Cybernetics , Tallinn University of Technology , Tallinn Estonia su: Speech perception Vocabulary Finnish language Estonian language Data Artificial neural networks Internet sug: subj: Speech perception Vocabulary Finnish language Estonian language Data Artificial neural networks Internet keyword: Adaptation Data filtering Large vocabulary speech recognition Statistical language modeling Subword units ab: One particular problem in large vocabulary continuous speech recognition for low-resourced languages is finding relevant training data for the statistical language models. Large amount of data is required, because models should estimate the probability for all possible word sequences. For Finnish, Estonian and the other fenno-ugric languages a special problem with the data is the huge amount of different word forms that are common in normal speech. The same problem exists also in other language technology applications such as machine translation, information retrieval, and in some extent also in other morphologically rich languages. In this paper we present methods and evaluations in four recent language modeling topics: selecting conversational data from the Internet, adapting models for foreign words, multi-domain and adapted neural network language modeling, and decoding with subword units. Our evaluations show that the same methods work in more than one language and that they scale down to smaller data resources. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|