Utilizing term proximity for blog post retrieval.

Term proximity is effective for many information retrieval (IR) research fields yet remains unexplored in blogosphere IR. The blogosphere is characterized by large amounts of noise, including incohesive, off-topic content and spam. Consequently, the classical bag-of-words unigram IR models are not r...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Society for Information Science & Technology Vol. 64; no. 11; pp. 2278 - 2299
Autores principales: Ye, Zheng, He, Ben, Wang, Lifeng, Luo, Tiejian
Formato: equations & formulas research tables/charts Journal Article
Publicado: Wiley-Blackwell Nov2013
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=104146117&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 104146117
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        15322882
        IGD
      jtl: Journal of the American Society for Information Science & Technology
      issn: 15322882
      maglogo: Y
    pubinfo:
      dt: Nov2013
      vid: 64
      iid: 11
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        104146117
        91102965
        10.1002/asi.22916
        104146117
      ppf: 2278
      ppct: 21
      formats:
      tig:
        atl: Utilizing term proximity for blog post retrieval.
      aug:
        au:
          Ye, Zheng
          He, Ben
          Wang, Lifeng
          Luo, Tiejian
        affil: School of Computer Science and Technology, Hangzhou Dianzi University
      sug:
        subj:
          Blogs
          Information Retrieval Methods
          Public Opinion
          Models, Statistical
          Information Science Methods
          Electronic Publishing
          Human
          Wilcoxon Rank Sum Test
          Funding Source
      ab: Term proximity is effective for many information retrieval (IR) research fields yet remains unexplored in blogosphere IR. The blogosphere is characterized by large amounts of noise, including incohesive, off-topic content and spam. Consequently, the classical bag-of-words unigram IR models are not reliable enough to provide robust and effective retrieval performance. In this article, we propose to boost the blog postretrieval performance by employing term proximity information. We investigate a variety of popular and state-of-the-art proximity-based statistical IR models, including a proximity-based counting model, the Markov random field (MRF) model, and the divergence from randomness (DFR) multinomial model. Extensive experimentation on the standard TREC Blog06 test dataset demonstrates that the introduction of term proximity information is indeed beneficial to retrieval from the blogosphere. Results also indicate the superiority of the unordered bi-gram model with the sequential-dependence phrases over other variants of the proximity-based models. Finally, inspired by the effectiveness of proximity models, we extend our study by exploring the proximity evidence between uery terms and opinionated terms. The consequent opinionated proximity model shows promising performance in the experiments.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N