Filtering artificial texts with statistical machine learning techniques.
Fake content is flourishing on the Internet, ranging from basic random word salads to web scraping. Most of this fake content is generated for the purpose of nourishing fake web sites aimed at biasing search engine indexes: at the scale of a search engine, using automatically generated texts render...
| Publicado en: | Language Resources & Evaluation Vol. 45; no. 1; pp. 25 - 44 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Feb2011
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=58721583&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 58721583 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Feb2011 vid: 45 iid: 1 pid: 237 pub: Springer Nature artinfo: ui: 58721583 10.1007/s10579-009-9113-0 ppf: 25 ppct: 19 formats: fmt: @attributes: type: P size: 463KB tig: atl: Filtering artificial texts with statistical machine learning techniques. aug: au: Lavergne, Thomas Urvoy, Tanguy Yvon, François affil: Orange Labs, Lannion France Univ Paris Sud 11 & LIMSI/CNRS, Orsay cedex France su: Artificial languages Machine learning Polyglot texts, selections, quotations, etc. Spam filtering (Email) Markov processes Algorithms Search engines sug: subj: Artificial languages Machine learning Polyglot texts, selections, quotations, etc. Spam filtering (Email) Markov processes Algorithms Search engines keyword: Statistical language models Web spam filtering ab: Fake content is flourishing on the Internet, ranging from basic random word salads to web scraping. Most of this fake content is generated for the purpose of nourishing fake web sites aimed at biasing search engine indexes: at the scale of a search engine, using automatically generated texts render such sites harder to detect than using copies of existing pages. In this paper, we present three methods aimed at distinguishing natural texts from artificially generated ones: the first method uses basic lexicometric features, the second one uses standard language models and the third one is based on a relative entropy measure which captures short range dependencies between words. Our experiments show that lexicometric features and language models are efficient to detect most generated texts, but fail to detect texts that are generated with high order Markov models. By comparison our relative entropy scoring algorithm, especially when trained on a large corpus, allows us to detect these 'hard' text generators with a high degree of accuracy. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2011 holdings: @attributes: islocal: N |
|---|