Reassessing the value of resources for cross-lingual transfer of POS tagging models.
When linguistically annotated data is scarce, as is the case for many under-resourced languages, one has to resort to less complete forms of annotations obtained from crawled dictionaries and/or through cross-lingual transfer. Several recent works have shown that learning from such partially supervi...
| Publicado en: | Language Resources & Evaluation Vol. 51; no. 4; pp. 927 - 961 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2017
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=126259352&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 126259352 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2017 vid: 51 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 126259352 10.1007/s10579-016-9362-7 ppf: 927 ppct: 34 formats: fmt: @attributes: type: P size: 925KB tig: atl: Reassessing the value of resources for cross-lingual transfer of POS tagging models. aug: au: Pécheux, Nicolas Wisniewski, Guillaume Yvon, François affil: LIMSI-CNRS , Orsay France su: Supervised learning Language & languages Annotations Data Perceptrons Machine learning Monolingualism sug: subj: Supervised learning Language & languages Annotations Data Perceptrons Machine learning Monolingualism keyword: Cross-lingual transfer POS tagging Weakly supervised learning ab: When linguistically annotated data is scarce, as is the case for many under-resourced languages, one has to resort to less complete forms of annotations obtained from crawled dictionaries and/or through cross-lingual transfer. Several recent works have shown that learning from such partially supervised data can be effective in many practical situations. In this work, we review two existing proposals for learning with ambiguous labels which extend conventional learners to the weakly supervised setting: a history-based model using a variant of the perceptron, on the one hand; an extension of the Conditional Random Fields model on the other hand. Focusing on the part-of-speech tagging task, but considering a large set of ten languages, we show (a) that good performance can be achieved even in the presence of ambiguity, provided however that both monolingual and bilingual resources are available; (b) that our two learners exploit different characteristics of the training set, and are successful in different situations; (c) that in addition to the choice of an adequate learning algorithm, many other factors are critical for achieving good performance in a cross-lingual transfer setting. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2017. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2017 holdings: @attributes: islocal: N |
|---|