Statistical Stylistics and Authorship Attribution: an Empirical Investigation.

This paper investigates the effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words in distinguishing texts by different authors and grouping texts by a single author. An examination of groups of texts by known authors shows that...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 16; no. 4; pp. 421 - 445
Autor principal: Hoover, David L.
Formato: Artículo
Publicado: Oxford University Press / USA Nov2001
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=44626934&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 44626934
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Nov2001
      vid: 16
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        44626934
        10.1093/llc/16.4.421
      ppf: 421
      ppct: 24
      formats:
        fmt:
          @attributes:
            type: P
            size: 355KB
      tig:
        atl: Statistical Stylistics and Authorship Attribution: an Empirical Investigation.
      aug:
        au: Hoover, David L.
        affil: New York University, USA
      su:
        Literary style
        Literary aesthetics
        Linguostylistics
        Attribution of authorship
        Authors
        Anonyms & pseudonyms
      sug:
        subj:
          Literary style
          Literary aesthetics
          Linguostylistics
          Attribution of authorship
          Authors
          Anonyms & pseudonyms
      ab: This paper investigates the effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words in distinguishing texts by different authors and grouping texts by a single author. An examination of groups of texts by known authors shows that cluster analyses typically achieve an accuracy rate of less than 90 per cent for contemporary novels, modern British and American novels, and contemporary literary critical texts, both on relatively large groups of texts and on smaller subsets of those texts. Although limiting the analysis to third‐person narration, and disambiguating homographic function words improves the results, inaccuracies remain. Furthermore, small groups of problematic texts extracted from the larger groups in simulated authorship studies also fail to cluster correctly. These failures suggest general rather than local problems with the technique, and cast doubt on the effectiveness of cluster analysis for authorship attribution and stylistic study.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2001
    holdings:
      @attributes:
        islocal: N