A comparative study of machine learning methods for authorship attribution.
We compare and benchmark the performance of five classification methods, four of which are taken from the machine learning literature, in a classic authorship attribution problem involving the Federalist Papers. Cross-validation results are reported for each method, and each method is further employ...
| Publicado en: | Literary & Linguistic Computing Vol. 25; no. 2; pp. 215 - 224 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2010
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=50986031&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 50986031 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Jun2010 vid: 25 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 50986031 10.1093/llc/fqq001 ppf: 215 ppct: 9 formats: fmt: @attributes: type: P size: 126KB tig: atl: A comparative study of machine learning methods for authorship attribution. aug: au: Jockers, Matthew L. Witten, Daniela M. affil: Department of English, Stanford University, Stanford, CA 94305, USA Department of Statistics, Stanford University, Stanford, CA 94305, USA su: Attribution of authorship Machine learning Artificial intelligence Machine theory Multivariate analysis Statistical correlation sug: subj: Attribution of authorship Machine learning Artificial intelligence Machine theory Multivariate analysis Statistical correlation ab: We compare and benchmark the performance of five classification methods, four of which are taken from the machine learning literature, in a classic authorship attribution problem involving the Federalist Papers. Cross-validation results are reported for each method, and each method is further employed in classifying the disputed papers and the few papers that are generally understood to be coauthored. These tests are performed using two separate feature sets: a “raw” feature set containing all words and word bigrams that are common to all of the authors, and a second “pre-processed” feature set derived by reducing the raw feature set to include only words meeting a minimum relative frequency threshold. Each of the methods tested performed well, but nearest shrunken centroids and regularized discriminant analysis had the best overall performances with 0/70 cross-validation errors. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2010 holdings: @attributes: islocal: N |
|---|