Intonation and dialog context as constraints for speech recognition.
This paper describes a way of using intonation and dialog context to improve the performance of an automatic speech recognition (ASR) system. Our experiments were run on the DCIEM Maptask corpus, a corpus of spontaneous task-oriented dialog speech. This corpus has been tagged according to a dialog a...
| Publicado en: | Language & Speech Vol. 41; no. 3/4; pp. 493 - 513 |
|---|---|
| Autores principales: | , , , |
| Formato: | equations & formulas research tables/charts Journal Article |
| Publicado: |
Sage Publications Inc.
Jul-Dec98
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=107187251&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 107187251 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 00238309 3YY jtl: Language & Speech issn: 00238309 maglogo: Y pubinfo: dt: Jul-Dec98 vid: 41 iid: 3/4 pid: 344 pub: Sage Publications Inc. place: Thousand Oaks, California artinfo: ui: 107187251 1999035389 10.1177/002383099804100411 107187251 ppf: 493 ppct: 20 formats: fmt: @attributes: type: P tig: atl: Intonation and dialog context as constraints for speech recognition. aug: au: Taylor P King S Isard S Wright H affil: Center for Speech Technology Research, University of Edinburgh, 80 South Bridge, Edinburgh, UK EH1 1HN. E-mail: pault@cstr.ed.ac.uk sug: subj: Phonology Conversation Voice Recognition Systems Funding Source Discourse Analysis Experimental Studies Models, Theoretical Speech Perception Conceptual Framework Paired T-Tests Human ab: This paper describes a way of using intonation and dialog context to improve the performance of an automatic speech recognition (ASR) system. Our experiments were run on the DCIEM Maptask corpus, a corpus of spontaneous task-oriented dialog speech. This corpus has been tagged according to a dialog analysis scheme that assigns each utterance to one of 12 'move types,' such as 'acknowledge,' 'query-yes/no' or 'instruct.' Most ASR systems use a bigram language model to constrain the possible sequences of words that might be recognized. Here we use a separate bigram language model for each move type. We show that when the 'correct' movespecific language model is used for each utterance in the test set, the word error rate of the recognizer drops. Of course when the recognizer is run on previously unseen data, it cannot know in advance what move type the speaker has just produced. To determine the move type we use an intonation model combined with a dialog model that puts constraints on possible sequences of move types, as well as the speech recognizer likelihoods for the different move-specific models. In the full recognition system, the combination of automatic move type recognition with the move specific language models reduces the overall word error rate by a small but significant amount when compared with a baseline system that does not take intonation or dialog acts into account. Interestingly, the word error improvement is restricted to 'initiating' move types, where word recognition is important. In 'response' move types, where the important information is conveyed by the move type itself--for example, positive versus negative response--there is no word error improvement, but recognition of the response types themselves is good. The paper discusses the intonation model, the language models, and the dialog model in detail and describes the architecture in which they are combined. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|