A hybrid approach for paraphrase identification based on knowledge-enriched semantic heuristics.

In this paper, we propose a hybrid approach for sentence paraphrase identification. The proposal addresses the problem of evaluating sentence-to-sentence semantic similarity when the sentences contain a set of named-entities. The essence of the proposal is to distinguish the computation of the seman...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 54; no. 2; pp. 457 - 486
Autores principales: Mohamed, Muhidin, Oussalah, Mourad
Formato: Artículo
Publicado: Springer Nature Jun2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=143152357&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 143152357
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Jun2020
      vid: 54
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        143152357
        10.1007/s10579-019-09466-4
      ppf: 457
      ppct: 29
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 730KB
      tig:
        atl: A hybrid approach for paraphrase identification based on knowledge-enriched semantic heuristics.
      aug:
        au:
          Mohamed, Muhidin
          Oussalah, Mourad
        affil:
          Department of Computer Science, EAS, Aston University, B4 7ET, Birmingham, UK
          Centre for Ubiquitous Computing, Faculty of Information Technology Computer Science, University of Oulu, P.O. Box 4500, 90014, Oulu, Finland
      su:
        Wikipedia
        Paraphrase
        Heuristic
        Identification
        Gene ontology
      sug:
        subj:
          Wikipedia
          Paraphrase
          Heuristic
          Identification
          Gene ontology
      keyword:
        Named-entity semantic relatedness
        Paraphrase identification
        Word category subsumption
        WordNet
      ab: In this paper, we propose a hybrid approach for sentence paraphrase identification. The proposal addresses the problem of evaluating sentence-to-sentence semantic similarity when the sentences contain a set of named-entities. The essence of the proposal is to distinguish the computation of the semantic similarity of named-entity tokens from the rest of the sentence text. More specifically, this is based on the integration of word semantic similarity derived from WordNet taxonomic relations, and named-entity semantic relatedness inferred from Wikipedia entity co-occurrences and underpinned by Normalized Google Distance. In addition, the WordNet similarity measure is enriched with word part-of-speech (PoS) conversion aided with a Categorial Variation database (CatVar), which enhances the lexico-semantics of words. We validated our hybrid approach using two different datasets; Microsoft Research Paraphrase Corpus (MSRPC) and TREC-9 Question Variants. In our empirical evaluation, we showed that our system outperforms baselines and most of the related state-of-the-art systems for paraphrase detection. We also conducted a misidentification analysis to disclose the primary sources of our system errors.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N