Does size matter? Authorship attribution, small samples, big problem.
The aim of this study is to find such a minimal size of text samples for authorship attribution that would provide stable results independent of random noise. A few controlled tests for different sample lengths, languages, and genres are discussed and compared. Depending on the corpus used, the mini...
| Publicado en: | Digital Scholarship in the Humanities pp. 167 - 183 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
06/01/2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108489944&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 108489944 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: 06/01/2015 pid: 622 pub: Oxford University Press / USA artinfo: ui: 108489944 10.1093/llc/fqt066 ppf: 167 ppct: 16 formats: fmt: @attributes: type: P size: 8.2MB tig: atl: Does size matter? Authorship attribution, small samples, big problem. aug: au: Eder, Maciej affil: Pedagogical University of Kraków, Poland and Polish Academy of Sciences, Institute of Polish Language, Krakow, Poland su: Attribution of authorship Random noise theory Latin prose literature Statistical sampling Language & languages sug: subj: Attribution of authorship Random noise theory Latin prose literature Statistical sampling Language & languages ab: The aim of this study is to find such a minimal size of text samples for authorship attribution that would provide stable results independent of random noise. A few controlled tests for different sample lengths, languages, and genres are discussed and compared. Depending on the corpus used, the minimal sample length varied from 2,500 words (Latin prose) to 5,000 or so words (in most cases, including English, German, Polish, and Hungarian novels). Another observation is connected with the method of sampling: contrary to common sense, randomly excerpted 'bags of words' turned out to be much more effective than the classical solution, i.e. using original sequences of words ('passages') of desired size. Although the tests have been performed using the Delta method (Burrows, J.F. (2002). 'Delta': a measure of stylistic difference and a guide to likely authorship. Literary and Linguistic Computing, 17(3): 267-87) applied to the most frequent words, some additional experiments have been conducted for support vector machines and k-NN applied to most frequent words, character 3-grams, character 4-grams, and parts-of-speech-tag 3-grams. Despite significant differences in overall attributive success rate between particular methods and/or style markers, the minimal amount of textual data needed for reliable authorship attribution turned out to be method-independent. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|