Reassessing the value of resources for cross-lingual transfer of POS tagging models.

When linguistically annotated data is scarce, as is the case for many under-resourced languages, one has to resort to less complete forms of annotations obtained from crawled dictionaries and/or through cross-lingual transfer. Several recent works have shown that learning from such partially supervi...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 51; no. 4; pp. 927 - 961
Autores principales: Pécheux, Nicolas, Wisniewski, Guillaume, Yvon, François
Formato: Artículo
Publicado: Springer Nature Dec2017
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=126259352&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 126259352
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2017
      vid: 51
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        126259352
        10.1007/s10579-016-9362-7
      ppf: 927
      ppct: 34
      formats:
        fmt:
          @attributes:
            type: P
            size: 925KB
      tig:
        atl: Reassessing the value of resources for cross-lingual transfer of POS tagging models.
      aug:
        au:
          Pécheux, Nicolas
          Wisniewski, Guillaume
          Yvon, François
        affil: LIMSI-CNRS , Orsay France
      su:
        Supervised learning
        Language & languages
        Annotations
        Data
        Perceptrons
        Machine learning
        Monolingualism
      sug:
        subj:
          Supervised learning
          Language & languages
          Annotations
          Data
          Perceptrons
          Machine learning
          Monolingualism
      keyword:
        Cross-lingual transfer
        POS tagging
        Weakly supervised learning
      ab: When linguistically annotated data is scarce, as is the case for many under-resourced languages, one has to resort to less complete forms of annotations obtained from crawled dictionaries and/or through cross-lingual transfer. Several recent works have shown that learning from such partially supervised data can be effective in many practical situations. In this work, we review two existing proposals for learning with ambiguous labels which extend conventional learners to the weakly supervised setting: a history-based model using a variant of the perceptron, on the one hand; an extension of the Conditional Random Fields model on the other hand. Focusing on the part-of-speech tagging task, but considering a large set of ten languages, we show (a) that good performance can be achieved even in the presence of ambiguity, provided however that both monolingual and bilingual resources are available; (b) that our two learners exploit different characteristics of the training set, and are successful in different situations; (c) that in addition to the choice of an adequate learning algorithm, many other factors are critical for achieving good performance in a cross-lingual transfer setting.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2017
    holdings:
      @attributes:
        islocal: N