Coupling an annotated corpus and a lexicon for state-of-the-art POS tagging.

This paper investigates how to best couple hand-annotated data with information extracted from an external lexical resource to improve part-of-speech tagging performance. Focusing mostly on French tagging, we introduce a maximum entropy Markov model-based tagging system that is enriched with informa...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 46; no. 4; pp. 721 - 737
Main Authors: Denis, Pascal, Sagot, Benoît
Format: Article
Published: Springer Nature Dec2012
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=83587034&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 83587034
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2012
      vid: 46
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        83587034
        10.1007/s10579-012-9193-0
      ppf: 721
      ppct: 16
      formats:
        fmt:
          @attributes:
            type: P
            size: 405KB
      tig:
        atl: Coupling an annotated corpus and a lexicon for state-of-the-art POS tagging.
      aug:
        au:
          Denis, Pascal
          Sagot, Benoît
        affil: Alpage, INRIA Paris-Rocquencourt & Université Paris 7, Domaine de Voluceau, Rocquencourt, 78153 Le Chesnay Cedex France
      su:
        Lexicon
        Performance evaluation
        Markov processes
        Data extraction
        Maximum entropy method
        Data analysis
      sug:
        subj:
          Lexicon
          Performance evaluation
          Markov processes
          Data extraction
          Maximum entropy method
          Data analysis
      keyword:
        French
        Language resource development
        Maximum entropy models
        Morphosyntactic lexicon
        Part-of-speech tagging
      ab: This paper investigates how to best couple hand-annotated data with information extracted from an external lexical resource to improve part-of-speech tagging performance. Focusing mostly on French tagging, we introduce a maximum entropy Markov model-based tagging system that is enriched with information extracted from a morphological resource. This system gives a 97.75 % accuracy on the French Treebank, an error reduction of 25 % (38 % on unknown words) over the same tagger without lexical information. We perform a series of experiments that help understanding how this lexical information helps improving tagging accuracy. We also conduct experiments on datasets and lexicons of varying sizes in order to assess the best trade-off between annotating data versus developing a lexicon. We find that the use of a lexicon improves the quality of the tagger at any stage of development of either resource, and that for fixed performance levels the availability of the full lexicon consistently reduces the need for supervised data by at least one half.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2012. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2012
    holdings:
      @attributes:
        islocal: N