A comparative study of machine learning methods for authorship attribution.

We compare and benchmark the performance of five classification methods, four of which are taken from the machine learning literature, in a classic authorship attribution problem involving the Federalist Papers. Cross-validation results are reported for each method, and each method is further employ...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 25; no. 2; pp. 215 - 224
Autores principales: Jockers, Matthew L., Witten, Daniela M.
Formato: Artículo
Publicado: Oxford University Press / USA Jun2010
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:We compare and benchmark the performance of five classification methods, four of which are taken from the machine learning literature, in a classic authorship attribution problem involving the Federalist Papers. Cross-validation results are reported for each method, and each method is further employed in classifying the disputed papers and the few papers that are generally understood to be coauthored. These tests are performed using two separate feature sets: a “raw” feature set containing all words and word bigrams that are common to all of the authors, and a second “pre-processed” feature set derived by reducing the raw feature set to include only words meeting a minimum relative frequency threshold. Each of the methods tested performed well, but nearest shrunken centroids and regularized discriminant analysis had the best overall performances with 0/70 cross-validation errors.