In no uncertain terms: a dataset for monolingual and multilingual automatic term extraction from comparable corpora.
Automatic term extraction is a productive field of research within natural language processing, but it still faces significant obstacles regarding datasets and evaluation, which require manual term annotation. This is an arduous task, made even more difficult by the lack of a clear distinction betwe...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 2; pp. 385 - 419 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=143152354&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 143152354 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2020 vid: 54 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 143152354 10.1007/s10579-019-09453-9 ppf: 385 ppct: 34 formats: fmt: – @attributes: type: T – @attributes: type: P size: 840KB tig: atl: In no uncertain terms: a dataset for monolingual and multilingual automatic term extraction from comparable corpora. aug: au: Rigouts Terryn, Ayla Hoste, Véronique Lefever, Els affil: LT3 Language and Translation Technology Team, Department of Translation, Interpreting and Communication, Ghent University, Groot-Brittanniëlaan 45, 9000, Ghent, Belgium su: Natural language processing Corpora Terms & phrases sug: subj: Natural language processing Corpora Terms & phrases keyword: ATR Automatic term extraction Comparable corpora Term annotation Terminology ab: Automatic term extraction is a productive field of research within natural language processing, but it still faces significant obstacles regarding datasets and evaluation, which require manual term annotation. This is an arduous task, made even more difficult by the lack of a clear distinction between terms and general language, which results in low inter-annotator agreement. There is a large need for well-documented, manually validated datasets, especially in the rising field of multilingual term extraction from comparable corpora, which presents a unique new set of challenges. In this paper, a new approach is presented for both monolingual and multilingual term annotation in comparable corpora. The detailed guidelines with different term labels, the domain- and language-independent methodology and the large volumes annotated in three different languages and four different domains make this a rich resource. The resulting datasets are not just suited for evaluation purposes but can also serve as a general source of information about terms and even as training data for supervised methods. Moreover, the gold standard for multilingual term extraction from comparable corpora contains information about term variants and translation equivalents, which allows an in-depth, nuanced evaluation. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|