Language chunking, data sparseness, and the value of a long marker list: explorations with word n-grams and authorial attribution.

The frequencies of individual words have been the mainstay of computer-assisted authorial attribution over the past three decades. The usefulness of this sort of data is attested in many benchmark trials and in numerous studies of particular authorship problems. It is sometimes argued, however, that...

Descripción completa

Detalles Bibliográficos
Publicado en:Literary & Linguistic Computing Vol. 29; no. 2; pp. 147 - 164
Autores principales: Antonia, Alexis, Craig, Hugh, Elliott, Jack
Formato: Artículo
Publicado: Oxford University Press / USA Jun2014
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=96092908&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 96092908
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        02681145
        BJ1
      jtl: Literary & Linguistic Computing
      issn: 02681145
      maglogo: N
    pubinfo:
      dt: Jun2014
      vid: 29
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        96092908
        10.1093/llc/fqt028
      ppf: 147
      ppct: 17
      formats:
        fmt:
          @attributes:
            type: P
            size: 228KB
      tig:
        atl: Language chunking, data sparseness, and the value of a long marker list: explorations with word n-grams and authorial attribution.
      aug:
        au:
          Antonia, Alexis
          Craig, Hugh
          Elliott, Jack
        affil: Centre for Literary and Linguistic Computing, University of Newcastle, Australia
      su:
        Computer assisted language instruction
        Authorship
        Vocabulary
        Language & languages
        Literature
      sug:
        subj:
          Computer assisted language instruction
          Authorship
          Vocabulary
          Language & languages
          Literature
      ab: The frequencies of individual words have been the mainstay of computer-assisted authorial attribution over the past three decades. The usefulness of this sort of data is attested in many benchmark trials and in numerous studies of particular authorship problems. It is sometimes argued, however, that since language as spoken or written falls into word sequences, on the ‘idiom principle’, and since language is characteristically produced in the brain in chunks, not in individual words, n-grams with n higher than 1 are superior to individual words as a source of authorship markers. In this article, we test the usefulness of word n-grams for authorship attribution by asking how many good-quality authorship markers are yielded by n-grams of various types, namely 1-grams, 2-grams, 3-grams, 4-grams, and 5-grams. We use two ways of formulating the n-grams, two corpora of texts, and two methods for finding and assessing markers. We find that when using methods based on regularly occurring markers, and drawing on all the available vocabulary, 1-grams perform best. With methods based on rare markers, and all the available vocabulary, strict 3-gram sequences perform best. If we restrict ourselves to a defined word-list of function-words to form n-grams, 2-grams offer a striking improvement on 1-grams.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Literary & Linguistic Computing
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2014
    holdings:
      @attributes:
        islocal: N