Coupling an annotated corpus and a lexicon for state-of-the-art POS tagging.
This paper investigates how to best couple hand-annotated data with information extracted from an external lexical resource to improve part-of-speech tagging performance. Focusing mostly on French tagging, we introduce a maximum entropy Markov model-based tagging system that is enriched with informa...
| Published in: | Language Resources & Evaluation Vol. 46; no. 4; pp. 721 - 737 |
|---|---|
| Main Authors: | , |
| Format: | Article |
| Published: |
Springer Nature
Dec2012
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=83587034&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 83587034 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2012 vid: 46 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 83587034 10.1007/s10579-012-9193-0 ppf: 721 ppct: 16 formats: fmt: @attributes: type: P size: 405KB tig: atl: Coupling an annotated corpus and a lexicon for state-of-the-art POS tagging. aug: au: Denis, Pascal Sagot, Benoît affil: Alpage, INRIA Paris-Rocquencourt & Université Paris 7, Domaine de Voluceau, Rocquencourt, 78153 Le Chesnay Cedex France su: Lexicon Performance evaluation Markov processes Data extraction Maximum entropy method Data analysis sug: subj: Lexicon Performance evaluation Markov processes Data extraction Maximum entropy method Data analysis keyword: French Language resource development Maximum entropy models Morphosyntactic lexicon Part-of-speech tagging ab: This paper investigates how to best couple hand-annotated data with information extracted from an external lexical resource to improve part-of-speech tagging performance. Focusing mostly on French tagging, we introduce a maximum entropy Markov model-based tagging system that is enriched with information extracted from a morphological resource. This system gives a 97.75 % accuracy on the French Treebank, an error reduction of 25 % (38 % on unknown words) over the same tagger without lexical information. We perform a series of experiments that help understanding how this lexical information helps improving tagging accuracy. We also conduct experiments on datasets and lexicons of varying sizes in order to assess the best trade-off between annotating data versus developing a lexicon. We find that the use of a lexicon improves the quality of the tagger at any stage of development of either resource, and that for fixed performance levels the availability of the full lexicon consistently reduces the need for supervised data by at least one half. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2012. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2012 holdings: @attributes: islocal: N |
|---|