Automatic induction of language model data for a spoken dialogue system.
In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artifi...
| Publicado en: | Language Resources & Evaluation Vol. 40; no. 1; pp. 25 - 47 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Feb2006
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=23218135&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 23218135 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Feb2006 vid: 40 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 23218135 10.1007/s10579-006-9007-3 ppf: 25 ppct: 22 formats: fmt: @attributes: type: P size: 378KB tig: atl: Automatic induction of language model data for a spoken dialogue system. aug: au: Chao Wang Chung, Grace Seneff, Stephanie affil: MIT Computer Science and Artificial Intelligence Laboratory, 32 Vassar Street, Cambridge, MA 02139, USA. Corporation for National Research Initiatives, 1895 Preston White Drive, Suite 100, Reston, VA 22209, USA. su: Dialogue analysis Interpersonal communication Simulation methods & models Recognition (Psychology) Information resources Language & languages sug: subj: Dialogue analysis Interpersonal communication Simulation methods & models Recognition (Psychology) Information resources Language & languages keyword: Example-based generation Language model Spoken dialogue systems User simulation ab: In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artificially constructed in-domain phrases. After some syntactic and semantic filtering, a large corpus of synthetically assembled user utterances is induced. In the second stage, two sampling methods are explored to filter the synthetic corpus to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions from the raw data by modelling after a small set of dialogues produced by the developers during the course of system refinement. Evaluation is conducted on recognition performance in a restaurant information domain. We show that a partial match to usage-appropriate semantic content distribution can be achieved via user simulations. Furthermore, word error rate can be reduced when limited amounts of in-domain training data are augmented with synthetic data derived by our methods. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2006. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2006 holdings: @attributes: islocal: N |
|---|