Does size matter? Authorship attribution, small samples, big problem.

The aim of this study is to find such a minimal size of text samples for authorship attribution that would provide stable results independent of random noise. A few controlled tests for different sample lengths, languages, and genres are discussed and compared. Depending on the corpus used, the mini...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities pp. 167 - 183
Autor principal: Eder, Maciej
Formato: Artículo
Publicado: Oxford University Press / USA 06/01/2015
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108489944&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 108489944
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: 06/01/2015
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        108489944
        10.1093/llc/fqt066
      ppf: 167
      ppct: 16
      formats:
        fmt:
          @attributes:
            type: P
            size: 8.2MB
      tig:
        atl: Does size matter? Authorship attribution, small samples, big problem.
      aug:
        au: Eder, Maciej
        affil: Pedagogical University of Kraków, Poland and Polish Academy of Sciences, Institute of Polish Language, Krakow, Poland
      su:
        Attribution of authorship
        Random noise theory
        Latin prose literature
        Statistical sampling
        Language & languages
      sug:
        subj:
          Attribution of authorship
          Random noise theory
          Latin prose literature
          Statistical sampling
          Language & languages
      ab: The aim of this study is to find such a minimal size of text samples for authorship attribution that would provide stable results independent of random noise. A few controlled tests for different sample lengths, languages, and genres are discussed and compared. Depending on the corpus used, the minimal sample length varied from 2,500 words (Latin prose) to 5,000 or so words (in most cases, including English, German, Polish, and Hungarian novels). Another observation is connected with the method of sampling: contrary to common sense, randomly excerpted 'bags of words' turned out to be much more effective than the classical solution, i.e. using original sequences of words ('passages') of desired size. Although the tests have been performed using the Delta method (Burrows, J.F. (2002). 'Delta': a measure of stylistic difference and a guide to likely authorship. Literary and Linguistic Computing, 17(3): 267-87) applied to the most frequent words, some additional experiments have been conducted for support vector machines and k-NN applied to most frequent words, character 3-grams, character 4-grams, and parts-of-speech-tag 3-grams. Despite significant differences in overall attributive success rate between particular methods and/or style markers, the minimal amount of textual data needed for reliable authorship attribution turned out to be method-independent.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2015
    holdings:
      @attributes:
        islocal: N