SenseDefs: a multilingual corpus of semantically annotated textual definitions: Exploiting multiple languages and resources jointly for high-quality Word Sense Disambiguation and Entity Linking.
Definitional knowledge has proved to be essential in various Natural Language Processing tasks and applications, especially when information at the level of word senses is exploited. However, the few sense-annotated corpora of textual definitions available to date are of limited size: this is mainly...
| Published in: | Language Resources & Evaluation Vol. 53; no. 2; pp. 251 - 279 |
|---|---|
| Main Authors: | , , , |
| Format: | Article |
| Published: |
Springer Nature
Jun2019
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=136891170&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 136891170 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2019 vid: 53 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 136891170 10.1007/s10579-018-9421-3 ppf: 251 ppct: 28 formats: fmt: – @attributes: type: T – @attributes: type: P size: 976KB tig: atl: SenseDefs: a multilingual corpus of semantically annotated textual definitions: Exploiting multiple languages and resources jointly for high-quality Word Sense Disambiguation and Entity Linking. aug: au: Camacho-Collados, Jose Delli Bovi, Claudio Raganato, Alessandro Navigli, Roberto affil: Cardiff University, Cardiff, UK Amazon.com, Inc., Turin, Italy University of Helsinki, Helsinki, Finland Sapienza University of Rome, Rome, Italy su: Natural language processing Definitions Corpora Multilingualism Lexicon Annotations sug: subj: Natural language processing Definitions Corpora Multilingualism Lexicon Annotations keyword: Entity linking Glosses Lexical resources Multilinguality Textual definitions Word Sense Disambiguation ab: Definitional knowledge has proved to be essential in various Natural Language Processing tasks and applications, especially when information at the level of word senses is exploited. However, the few sense-annotated corpora of textual definitions available to date are of limited size: this is mainly due to the expensive and time-consuming process of annotating a wide variety of word senses and entity mentions at a reasonably high scale. In this paper we present SenseDefs, a large-scale high-quality corpus of disambiguated definitions (or glosses) in multiple languages, comprising sense annotations of both concepts and named entities from a wide-coverage unified sense inventory. Our approach for the construction and disambiguation of this corpus builds upon the structure of a large multilingual semantic network and a state-of-the-art disambiguation system: first, we gather complementary information of equivalent definitions across different languages to provide context for disambiguation; then we refine the disambiguation output with a distributional approach based on semantic similarity. As a result, we obtain a multilingual corpus of textual definitions featuring over 38 million definitions in 263 languages, and we publicly release it to the research community. We assess the quality of SenseDefs's sense annotations both intrinsically and extrinsically on Open Information Extraction and Sense Clustering tasks. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2019. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|