Distributed language representation for authorship attribution.
Distributed language representation (deep learning) has been applied successfully in different applications in natural language processing. Using this model, we propose and implement two new authorship attribution classifiers. In this perspective, a vector-space representation can be generated for e...
| Publicado en: | Digital Scholarship in the Humanities Vol. 33; no. 2; pp. 425 - 442 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2018
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=129666141&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 129666141 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2018 vid: 33 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 129666141 10.1093/llc/fqx046 ppf: 425 ppct: 17 formats: fmt: – @attributes: type: T – @attributes: type: P size: 682KB tig: atl: Distributed language representation for authorship attribution. aug: au: Kocher, Mirco Savoy, Jacques affil: University of Neuchatel, Switzerland su: Authorship Deep learning Newspapers Federalist Papers (Book) Glasgow Herald (Periodical) sug: subj: Authorship Deep learning Newspapers Federalist Papers (Book) Glasgow Herald (Periodical) ab: Distributed language representation (deep learning) has been applied successfully in different applications in natural language processing. Using this model, we propose and implement two new authorship attribution classifiers. In this perspective, a vector-space representation can be generated for each author or disputed text according to words and their nearby context. To determine the authorship of a disputed text, the cosine similarity between vector representations can be applied. The proposed strategies can be adapted without any difficulty to different languages (such as English and Italian) or genres (essays, political speeches, and newspaper articles). Evaluations using the k-nearest neighbors (k-NNs))and based on four test collections (the Federalist Papers, the State of the Union addresses, the Glasgow Herald, and La Stampa newspapers) indicate that the distributed language representation preforms well, providing sometimes better effectiveness than state-of-the-art methods such as k-NN, nearest shrunken centroids, chi-square, Delta, latent Dirichlet allocation, or multi-layer perceptron classifier. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|