Filtering artificial texts with statistical machine learning techniques.

Fake content is flourishing on the Internet, ranging from basic random word salads to web scraping. Most of this fake content is generated for the purpose of nourishing fake web sites aimed at biasing search engine indexes: at the scale of a search engine, using automatically generated texts render...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 45; no. 1; pp. 25 - 44
Autores principales: Lavergne, Thomas, Urvoy, Tanguy, Yvon, François
Formato: Artículo
Publicado: Springer Nature Feb2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=58721583&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 58721583
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Feb2011
      vid: 45
      iid: 1
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        58721583
        10.1007/s10579-009-9113-0
      ppf: 25
      ppct: 19
      formats:
        fmt:
          @attributes:
            type: P
            size: 463KB
      tig:
        atl: Filtering artificial texts with statistical machine learning techniques.
      aug:
        au:
          Lavergne, Thomas
          Urvoy, Tanguy
          Yvon, François
        affil:
          Orange Labs, Lannion France
          Univ Paris Sud 11 & LIMSI/CNRS, Orsay cedex France
      su:
        Artificial languages
        Machine learning
        Polyglot texts, selections, quotations, etc.
        Spam filtering (Email)
        Markov processes
        Algorithms
        Search engines
      sug:
        subj:
          Artificial languages
          Machine learning
          Polyglot texts, selections, quotations, etc.
          Spam filtering (Email)
          Markov processes
          Algorithms
          Search engines
      keyword:
        Statistical language models
        Web spam filtering
      ab: Fake content is flourishing on the Internet, ranging from basic random word salads to web scraping. Most of this fake content is generated for the purpose of nourishing fake web sites aimed at biasing search engine indexes: at the scale of a search engine, using automatically generated texts render such sites harder to detect than using copies of existing pages. In this paper, we present three methods aimed at distinguishing natural texts from artificially generated ones: the first method uses basic lexicometric features, the second one uses standard language models and the third one is based on a relative entropy measure which captures short range dependencies between words. Our experiments show that lexicometric features and language models are efficient to detect most generated texts, but fail to detect texts that are generated with high order Markov models. By comparison our relative entropy scoring algorithm, especially when trained on a large corpus, allows us to detect these 'hard' text generators with a high degree of accuracy.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N