Fake opinion detection: how similar are crowdsourced datasets to real data?

Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online and a...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 54; no. 4; pp. 1019 - 1059
Autores principales: Fornaciari, Tommaso, Cagnina, Leticia, Rosso, Paolo, Poesio, Massimo
Formato: Artículo
Publicado: Springer Nature Dec2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=146751847&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 146751847
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Dec2020
      vid: 54
      iid: 4
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        146751847
        10.1007/s10579-020-09486-5
      ppf: 1019
      ppct: 40
      formats:
        fmt:
          @attributes:
            type: P
            size: 541KB
      tig:
        atl: Fake opinion detection: how similar are crowdsourced datasets to real data?
      aug:
        au:
          Fornaciari, Tommaso
          Cagnina, Leticia
          Rosso, Paolo
          Poesio, Massimo
        affil:
          Bocconi University, Milan, Italy
          Universidad Nacional de San Luis, San Luis, Argentina
          Universitat Politècnica de València, Valencia, Spain
          Queen Mary University of London, London, UK
      su:
        Spam email
        Natural language processing
        Internet publishing
      sug:
        subj:
          Spam email
          Natural language processing
          Internet publishing
      keyword:
        Crowdsourcing
        Deception detection
        Ground truth
        Probabilistic labeling
      ab: Identifying deceptive online reviews is a challenging tasks for Natural Language Processing (NLP). Collecting corpora for the task is difficult, because normally it is not possible to know whether reviews are genuine. A common workaround involves collecting (supposedly) truthful reviews online and adding them to a set of deceptive reviews obtained through crowdsourcing services. Models trained this way are generally successful at discriminating between 'genuine' online reviews and the crowdsourced deceptive reviews. It has been argued that the deceptive reviews obtained via crowdsourcing are very different from real fake reviews, but the claim has never been properly tested. In this paper, we compare (false) crowdsourced reviews with a set of 'real' fake reviews published on line. We evaluate their degree of similarity and their usefulness in training models for the detection of untrustworthy reviews. We find that the deceptive reviews collected via crowdsourcing are significantly different from the fake reviews published online. In the case of the artificially produced deceptive texts, it turns out that their domain similarity with the targets affects the models' performance, much more than their untruthfulness. This suggests that the use of crowdsourced datasets for opinion spam detection may not result in models applicable to the real task of detecting deceptive reviews. As an alternative method to create large-size datasets for the fake reviews detection task, we propose methods based on the probabilistic annotation of unlabeled texts, relying on the use of meta-information generally available on the e-commerce sites. Such methods are independent from the content of the reviews and allow to train reliable models for the detection of fake reviews.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2020. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N