The effect of author set size and data size in authorship attribution.

Applications of authorship attribution `in the wild’ [Koppel, M., Schler, J., and Argamon, S. (2010). Authorship attribution in the wild. Language Resources and Evaluation. Advanced Access published January 12, 2010:10.1007/s10579-009-9111-2], for instance in social networks, will likely involve lar...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 26; no. 1; pp. 35 - 56
Autores principales: Luyckx, Kim, Daelemans, Walter
Formato: Artículo
Publicado: Oxford University Press / USA Apr2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=59688140&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 59688140
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Apr2011
      vid: 26
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        59688140
        10.1093/llc/fqq013
      ppf: 35
      ppct: 21
      formats:
        fmt:
          @attributes:
            type: P
            size: 316KB
      tig:
        atl: The effect of author set size and data size in authorship attribution.
      aug:
        au:
          Luyckx, Kim
          Daelemans, Walter
        affil: CLiPS Computational Linguistics Group, University of Antwerp, Belgium
      su:
        Authorship
        Social networks
        Email
        Literacy
        Linguistics
      sug:
        subj:
          Authorship
          Social networks
          Email
          Literacy
          Linguistics
      ab: Applications of authorship attribution `in the wild’ [Koppel, M., Schler, J., and Argamon, S. (2010). Authorship attribution in the wild. Language Resources and Evaluation. Advanced Access published January 12, 2010:10.1007/s10579-009-9111-2], for instance in social networks, will likely involve large sets of candidate authors and only limited data per author. In this article, we present the results of a systematic study of two important parameters in supervised machine learning that significantly affect performance in computational authorship attribution: (1) the number of candidate authors (i.e. the number of classes to be learned), and (2) the amount of training data available per candidate author (i.e. the size of the training data). We also investigate the robustness of different types of lexical and linguistic features to the effects of author set size and data size. The approach we take is an operationalization of the standard text categorization model, using memory-based learning for discriminating between the candidate authors. We performed authorship attribution experiments on a set of three benchmark corpora in which the influence of topic could be controlled. The short text fragments of e-mail length present the approach with a true challenge. Results show that, as expected, authorship attribution accuracy deteriorates as the number of candidate authors increases and size of training data decreases, although the machine learning approach continues performing significantly above chance. Some feature types (most notably character n-grams) are robust to changes in author set size and data size, but no robust individual features emerge.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N