Term position‐based language model for information retrieval.

Term position feature is widely and successfully used in IR and Web search engines, to enhance the retrieval effectiveness. This feature is essentially used for two purposes: to capture query terms proximity or to boost the weight of terms appearing in some parts of a document. In this paper, we are...

Full description

Bibliographic Details
Published in:Journal of the Association for Information Science & Technology Vol. 72; no. 5; pp. 627 - 643
Main Authors: Hammache, Arezki, Boughanem, Mohand
Format: equations & formulas research tables/charts Journal Article
Published: Wiley-Blackwell May2021
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=149757585&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 149757585
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        23301635
        H6JN
      jtl: Journal of the Association for Information Science & Technology
      issn: 23301635
      maglogo: N
    pubinfo:
      dt: May2021
      vid: 72
      iid: 5
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        149757585
        147085616
        149757585
        149757585
        10.1002/asi.24431
        149757585
      ppf: 627
      ppct: 16
      formats:
      tig:
        atl: Term position‐based language model for information retrieval.
      aug:
        au:
          Hammache, Arezki
          Boughanem, Mohand
        affil: Laboratoire de Recherche en Informatique (LARI), Université Mouloud Mammeri, Tizi Ouzou, Algeria
      sug:
        subj:
          Information Retrieval Methods
          Language
          Models, Statistical
          Natural Language Processing
          Human
          Experimental Studies
          Algorithms
          Information Technology
          Comparative Studies
          Descriptive Statistics
      ab: Term position feature is widely and successfully used in IR and Web search engines, to enhance the retrieval effectiveness. This feature is essentially used for two purposes: to capture query terms proximity or to boost the weight of terms appearing in some parts of a document. In this paper, we are interested in this second category. We propose two novel query‐independent techniques based on absolute term positions in a document, whose goal is to boost the weight of terms appearing in the beginning of a document. The first one considers only the earliest occurrence of a term in a document. The second one takes into account all term positions in a document. We formalize each of these two techniques as a document model based on term position, and then we incorporate it into a basic language model (LM). Two smoothing techniques, Dirichlet and Jelinek‐Mercer, are considered in the basic LM. Experiments conducted on three TREC test collections show that our model, especially the version based on all term positions, achieves significant improvements over the baseline LMs, and it also often performs better than two state‐of‐the‐art baseline models, the chronological term rank model and the Markov random field model.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N