Term position‐based language model for information retrieval.
Term position feature is widely and successfully used in IR and Web search engines, to enhance the retrieval effectiveness. This feature is essentially used for two purposes: to capture query terms proximity or to boost the weight of terms appearing in some parts of a document. In this paper, we are...
| Published in: | Journal of the Association for Information Science & Technology Vol. 72; no. 5; pp. 627 - 643 |
|---|---|
| Main Authors: | , |
| Format: | equations & formulas research tables/charts Journal Article |
| Published: |
Wiley-Blackwell
May2021
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=149757585&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 149757585 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 23301635 H6JN jtl: Journal of the Association for Information Science & Technology issn: 23301635 maglogo: N pubinfo: dt: May2021 vid: 72 iid: 5 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 149757585 147085616 149757585 149757585 10.1002/asi.24431 149757585 ppf: 627 ppct: 16 formats: tig: atl: Term position‐based language model for information retrieval. aug: au: Hammache, Arezki Boughanem, Mohand affil: Laboratoire de Recherche en Informatique (LARI), Université Mouloud Mammeri, Tizi Ouzou, Algeria sug: subj: Information Retrieval Methods Language Models, Statistical Natural Language Processing Human Experimental Studies Algorithms Information Technology Comparative Studies Descriptive Statistics ab: Term position feature is widely and successfully used in IR and Web search engines, to enhance the retrieval effectiveness. This feature is essentially used for two purposes: to capture query terms proximity or to boost the weight of terms appearing in some parts of a document. In this paper, we are interested in this second category. We propose two novel query‐independent techniques based on absolute term positions in a document, whose goal is to boost the weight of terms appearing in the beginning of a document. The first one considers only the earliest occurrence of a term in a document. The second one takes into account all term positions in a document. We formalize each of these two techniques as a document model based on term position, and then we incorporate it into a basic language model (LM). Two smoothing techniques, Dirichlet and Jelinek‐Mercer, are considered in the basic LM. Experiments conducted on three TREC test collections show that our model, especially the version based on all term positions, achieves significant improvements over the baseline LMs, and it also often performs better than two state‐of‐the‐art baseline models, the chronological term rank model and the Markov random field model. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|