A comparison of human and statistical language model performance using missing-word tests.
This paper presents results from a series of missing-word tests, in which a small fragment of text is presented to human subjects who are then asked to suggest a ranked list of completions. The same experiment is repeated with the WA model, an n-gram statistical language model. From the completion d...
| Publicado en: | Language & Speech Vol. 40; no. 4; pp. 377 - 390 |
|---|---|
| Autores principales: | , , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Sage Publications Inc.
Oct-Dec97
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=107276595&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 107276595 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 00238309 3YY jtl: Language & Speech issn: 00238309 maglogo: Y pubinfo: dt: Oct-Dec97 vid: 40 iid: 4 pid: 344 pub: Sage Publications Inc. place: Thousand Oaks, California artinfo: ui: 107276595 1998048545 10.1177/002383099704000404 107276595 ppf: 377 ppct: 13 formats: fmt: @attributes: type: P tig: atl: A comparison of human and statistical language model performance using missing-word tests. aug: au: Owens M O'Boyle P McMahon J Ming J Smith FJ affil: The Queen's University of Belfast. E-mail: m.owens@uk.ac.qub sug: subj: Language Tests Language Processing Models, Statistical Comparative Studies Female Male Adult Test Taking Grammar Chi Square Test Descriptive Statistics Pearson's Correlation Coefficient Human Adult: 19-44 years Female Male ab: This paper presents results from a series of missing-word tests, in which a small fragment of text is presented to human subjects who are then asked to suggest a ranked list of completions. The same experiment is repeated with the WA model, an n-gram statistical language model. From the completion data two measures are obtained: (i) verbatim predictability, which indicates the extent to which subjects nominated exactly the missing word and (ii) grammatical class predictability, which indicates the extent to which subjects nominated words of the same grammatical class as the missing word. The differences in language model performance and human performance are encouragingly small, especially for verbatim predictability. This is especially significant given that the WA model was able, on average, to use at most half the available context. The results highlight human superiority in handling missing content words. Most importantly, the experiments illustrate the detailed information one can obtain about the performance of a language model through using missing-word tests. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|