Frequent Word Sequences and Statistical Stylistics.

This paper investigates the relative effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words and the frequencies of very frequent word sequences in distinguishing texts by different authors and grouping texts by a single author. C...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 17; no. 2; pp. 157 - 181
Autor principal: Hoover, David L.
Formato: Artículo
Publicado: Oxford University Press / USA Jun2002
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=44441549&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 44441549
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Jun2002
      vid: 17
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        44441549
        10.1093/llc/17.2.157
      ppf: 157
      ppct: 24
      formats:
        fmt:
          @attributes:
            type: P
            size: 2.3MB
      tig:
        atl: Frequent Word Sequences and Statistical Stylistics.
      aug:
        au: Hoover, David L.
        affil: New York University, New York, NY, USA
      su:
        Linguostylistics
        Multivariate analysis
        Decision making
        Cluster analysis (Statistics)
        Story plots
        Pronouns
        Literary terminology
      sug:
        subj:
          Linguostylistics
          Multivariate analysis
          Decision making
          Cluster analysis (Statistics)
          Story plots
          Pronouns
          Literary terminology
      ab: This paper investigates the relative effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words and the frequencies of very frequent word sequences in distinguishing texts by different authors and grouping texts by a single author. Cluster analyses based on frequent words are fairly accurate for groups of texts by known authors, whether the texts are long sections of modern British and US novels or shorter sections of contemporary literary critical texts, but they are only rarely completely accurate. When frequent word sequences are used instead of frequent words or in addition to them, however, the accuracy of the analyses often improves, sometimes dramatically, especially when personal pronouns are eliminated. Analyses based on frequent sequences even provide completely correct results in some cases where analyses based on frequent words fail. They also produce superior results for small groups of problematic novels and critical texts extracted from the larger corpora. Such successes suggest that analyses based on frequent word sequences constitute improved tools for authorship and stylistic studies.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2002
    holdings:
      @attributes:
        islocal: N