Statistical Stylistics and Authorship Attribution: an Empirical Investigation.
This paper investigates the effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words in distinguishing texts by different authors and grouping texts by a single author. An examination of groups of texts by known authors shows that...
| Publicado en: | Literary & Linguistic Computing Vol. 16; no. 4; pp. 421 - 445 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Nov2001
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=44626934&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 44626934 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Nov2001 vid: 16 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 44626934 10.1093/llc/16.4.421 ppf: 421 ppct: 24 formats: fmt: @attributes: type: P size: 355KB tig: atl: Statistical Stylistics and Authorship Attribution: an Empirical Investigation. aug: au: Hoover, David L. affil: New York University, USA su: Literary style Literary aesthetics Linguostylistics Attribution of authorship Authors Anonyms & pseudonyms sug: subj: Literary style Literary aesthetics Linguostylistics Attribution of authorship Authors Anonyms & pseudonyms ab: This paper investigates the effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words in distinguishing texts by different authors and grouping texts by a single author. An examination of groups of texts by known authors shows that cluster analyses typically achieve an accuracy rate of less than 90 per cent for contemporary novels, modern British and American novels, and contemporary literary critical texts, both on relatively large groups of texts and on smaller subsets of those texts. Although limiting the analysis to third‐person narration, and disambiguating homographic function words improves the results, inaccuracies remain. Furthermore, small groups of problematic texts extracted from the larger groups in simulated authorship studies also fail to cluster correctly. These failures suggest general rather than local problems with the technique, and cast doubt on the effectiveness of cluster analysis for authorship attribution and stylistic study. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2001 holdings: @attributes: islocal: N |
|---|