A comparison of graph-based word sense induction clustering algorithms in a pseudoword evaluation framework.

This article presents a comparison of different Word Sense Induction (WSI) clustering algorithms on two novel pseudoword data sets of semantic-similarity and co-occurrence-based word graphs, with a special focus on the detection of homonymic polysemy. We follow the original definition of a pseudowor...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 52; no. 3; pp. 733 - 771
Autores principales: Cecchini, Flavio Massimiliano, Fersini, Elisabetta, Riedl, Martin, Biemann, Chris
Formato: Artículo
Publicado: Springer Nature Sep2018
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=131216690&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 131216690
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2018
      vid: 52
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        131216690
        10.1007/s10579-018-9415-1
      ppf: 733
      ppct: 38
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 839KB
      tig:
        atl: A comparison of graph-based word sense induction clustering algorithms in a pseudoword evaluation framework.
      aug:
        au:
          Cecchini, Flavio Massimiliano
          Fersini, Elisabetta
          Riedl, Martin
          Biemann, Chris
        affil:
          DISCo, Università degli Studi di Milano - Bicocca, Viale Sarca 336, Ed. U14, 20126, Milan, Italy
          Informatikum, Universität Hamburg, Vogt-Kölln-Straße 30, 22527, Hamburg, Germany
      su:
        Clustering of particles
        Natural language processing
        Graphic methods
        Semantic computing
        Polysemy
      sug:
        subj:
          Clustering of particles
          Natural language processing
          Graphic methods
          Semantic computing
          Polysemy
      keyword:
        Evaluation
        Graph clustering
        Pseudowords
        Word sense induction
      ab: This article presents a comparison of different Word Sense Induction (WSI) clustering algorithms on two novel pseudoword data sets of semantic-similarity and co-occurrence-based word graphs, with a special focus on the detection of homonymic polysemy. We follow the original definition of a pseudoword as the combination of two monosemous terms and their contexts to simulate a polysemous word. The evaluation is performed comparing the algorithm’s output on a pseudoword’s ego word graph (i.e., a graph that represents the pseudoword’s context in the corpus) with the known subdivision given by the components corresponding to the monosemous source words forming the pseudoword. The main contribution of this article is to present a self-sufficient pseudoword-based evaluation framework for WSI graph-based clustering algorithms, thereby defining a new evaluation measure (TOP2) and a secondary clustering process (hyperclustering). To our knowledge, we are the first to conduct and discuss a large-scale systematic pseudoword evaluation targeting the induction of coarse-grained homonymous word senses across a large number of graph clustering algorithms.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N