A hybrid approach for paraphrase identification based on knowledge-enriched semantic heuristics.
In this paper, we propose a hybrid approach for sentence paraphrase identification. The proposal addresses the problem of evaluating sentence-to-sentence semantic similarity when the sentences contain a set of named-entities. The essence of the proposal is to distinguish the computation of the seman...
| Publicado en: | Language Resources & Evaluation Vol. 54; no. 2; pp. 457 - 486 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2020
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=143152357&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 143152357 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2020 vid: 54 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 143152357 10.1007/s10579-019-09466-4 ppf: 457 ppct: 29 formats: fmt: – @attributes: type: T – @attributes: type: P size: 730KB tig: atl: A hybrid approach for paraphrase identification based on knowledge-enriched semantic heuristics. aug: au: Mohamed, Muhidin Oussalah, Mourad affil: Department of Computer Science, EAS, Aston University, B4 7ET, Birmingham, UK Centre for Ubiquitous Computing, Faculty of Information Technology Computer Science, University of Oulu, P.O. Box 4500, 90014, Oulu, Finland su: Wikipedia Paraphrase Heuristic Identification Gene ontology sug: subj: Wikipedia Paraphrase Heuristic Identification Gene ontology keyword: Named-entity semantic relatedness Paraphrase identification Word category subsumption WordNet ab: In this paper, we propose a hybrid approach for sentence paraphrase identification. The proposal addresses the problem of evaluating sentence-to-sentence semantic similarity when the sentences contain a set of named-entities. The essence of the proposal is to distinguish the computation of the semantic similarity of named-entity tokens from the rest of the sentence text. More specifically, this is based on the integration of word semantic similarity derived from WordNet taxonomic relations, and named-entity semantic relatedness inferred from Wikipedia entity co-occurrences and underpinned by Normalized Google Distance. In addition, the WordNet similarity measure is enriched with word part-of-speech (PoS) conversion aided with a Categorial Variation database (CatVar), which enhances the lexico-semantics of words. We validated our hybrid approach using two different datasets; Microsoft Research Paraphrase Corpus (MSRPC) and TREC-9 Question Variants. In our empirical evaluation, we showed that our system outperforms baselines and most of the related state-of-the-art systems for paraphrase detection. We also conducted a misidentification analysis to disclose the primary sources of our system errors. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2020 holdings: @attributes: islocal: N |
|---|