A comparison of graph-based word sense induction clustering algorithms in a pseudoword evaluation framework.
This article presents a comparison of different Word Sense Induction (WSI) clustering algorithms on two novel pseudoword data sets of semantic-similarity and co-occurrence-based word graphs, with a special focus on the detection of homonymic polysemy. We follow the original definition of a pseudowor...
| Publicado en: | Language Resources & Evaluation Vol. 52; no. 3; pp. 733 - 771 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Sep2018
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=131216690&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 131216690 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Sep2018 vid: 52 iid: 3 pid: 237 pub: Springer Nature artinfo: ui: 131216690 10.1007/s10579-018-9415-1 ppf: 733 ppct: 38 formats: fmt: – @attributes: type: T – @attributes: type: P size: 839KB tig: atl: A comparison of graph-based word sense induction clustering algorithms in a pseudoword evaluation framework. aug: au: Cecchini, Flavio Massimiliano Fersini, Elisabetta Riedl, Martin Biemann, Chris affil: DISCo, Università degli Studi di Milano - Bicocca, Viale Sarca 336, Ed. U14, 20126, Milan, Italy Informatikum, Universität Hamburg, Vogt-Kölln-Straße 30, 22527, Hamburg, Germany su: Clustering of particles Natural language processing Graphic methods Semantic computing Polysemy sug: subj: Clustering of particles Natural language processing Graphic methods Semantic computing Polysemy keyword: Evaluation Graph clustering Pseudowords Word sense induction ab: This article presents a comparison of different Word Sense Induction (WSI) clustering algorithms on two novel pseudoword data sets of semantic-similarity and co-occurrence-based word graphs, with a special focus on the detection of homonymic polysemy. We follow the original definition of a pseudoword as the combination of two monosemous terms and their contexts to simulate a polysemous word. The evaluation is performed comparing the algorithm’s output on a pseudoword’s ego word graph (i.e., a graph that represents the pseudoword’s context in the corpus) with the known subdivision given by the components corresponding to the monosemous source words forming the pseudoword. The main contribution of this article is to present a self-sufficient pseudoword-based evaluation framework for WSI graph-based clustering algorithms, thereby defining a new evaluation measure (TOP2) and a secondary clustering process (hyperclustering). To our knowledge, we are the first to conduct and discuss a large-scale systematic pseudoword evaluation targeting the induction of coarse-grained homonymous word senses across a large number of graph clustering algorithms. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|