Distributed language representation for authorship attribution.

Distributed language representation (deep learning) has been applied successfully in different applications in natural language processing. Using this model, we propose and implement two new authorship attribution classifiers. In this perspective, a vector-space representation can be generated for e...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 33; no. 2; pp. 425 - 442
Autores principales: Kocher, Mirco, Savoy, Jacques
Formato: Artículo
Publicado: Oxford University Press / USA Jun2018
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=129666141&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 129666141
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2018
      vid: 33
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        129666141
        10.1093/llc/fqx046
      ppf: 425
      ppct: 17
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 682KB
      tig:
        atl: Distributed language representation for authorship attribution.
      aug:
        au:
          Kocher, Mirco
          Savoy, Jacques
        affil: University of Neuchatel, Switzerland
      su:
        Authorship
        Deep learning
        Newspapers
        Federalist Papers (Book)
        Glasgow Herald (Periodical)
      sug:
        subj:
          Authorship
          Deep learning
          Newspapers
          Federalist Papers (Book)
          Glasgow Herald (Periodical)
      ab: Distributed language representation (deep learning) has been applied successfully in different applications in natural language processing. Using this model, we propose and implement two new authorship attribution classifiers. In this perspective, a vector-space representation can be generated for each author or disputed text according to words and their nearby context. To determine the authorship of a disputed text, the cosine similarity between vector representations can be applied. The proposed strategies can be adapted without any difficulty to different languages (such as English and Italian) or genres (essays, political speeches, and newspaper articles). Evaluations using the k-nearest neighbors (k-NNs))and based on four test collections (the Federalist Papers, the State of the Union addresses, the Glasgow Herald, and La Stampa newspapers) indicate that the distributed language representation preforms well, providing sometimes better effectiveness than state-of-the-art methods such as k-NN, nearest shrunken centroids, chi-square, Delta, latent Dirichlet allocation, or multi-layer perceptron classifier.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2018
    holdings:
      @attributes:
        islocal: N