Enhancing information retrieval through concept-based language modeling and semantic smoothing.

Traditionally, many information retrieval models assume that terms occur in documents independently. Although these models have already shown good performance, the word independency assumption seems to be unrealistic from a natural language point of view, which considers that terms are related to ea...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the Association for Information Science & Technology Vol. 67; no. 12; pp. 2909 - 2927
Autores principales: Said Lhadj, Lynda, Boughanem, Mohand, Amrouche, Karima
Formato: equations & formulas research tables/charts Journal Article
Publicado: Wiley-Blackwell Dec2016
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=119478025&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 119478025
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23301635
        H6JN
      jtl: Journal of the Association for Information Science & Technology
      issn: 23301635
      maglogo: N
    pubinfo:
      dt: Dec2016
      vid: 67
      iid: 12
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        119478025
        119478025
        119478025
        10.1002/asi.23553
        119478025
      ppf: 2909
      ppct: 18
      formats:
      tig:
        atl: Enhancing information retrieval through concept-based language modeling and semantic smoothing.
      aug:
        au:
          Said Lhadj, Lynda
          Boughanem, Mohand
          Amrouche, Karima
        affil: Ecole nationale Supérieure d'Informatique ESI, P.O. Box 68 M, 16309, Algiers Algeria
      sug:
        subj:
          Information Retrieval
          Semantics
          Models, Statistical
          Human
          News
          Experimental Studies
      ab: Traditionally, many information retrieval models assume that terms occur in documents independently. Although these models have already shown good performance, the word independency assumption seems to be unrealistic from a natural language point of view, which considers that terms are related to each other. Therefore, such an assumption leads to two well-known problems in information retrieval ( IR), namely, polysemy, or term mismatch, and synonymy. In language models, these issues have been addressed by considering dependencies such as bigrams, phrasal-concepts, or word relationships, but such models are estimated using simple n-grams or concept counting. In this paper, we address polysemy and synonymy mismatch with a concept-based language modeling approach that combines ontological concepts from external resources with frequently found collocations from the document collection. In addition, the concept-based model is enriched with subconcepts and semantic relationships through a semantic smoothing technique so as to perform semantic matching. Experiments carried out on TREC collections show that our model achieves significant improvements over a single word-based model and the Markov Random Field model (using a Markov classifier).
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N