Profile-based authorship analysis.
This article presents a profile-based authorship analysis method which first categorizes texts according to social and conceptual characteristics of their author (e.g. Sex and Political Ideology) and then combines these profiles for two authorship analysis tasks: (1) determining shared authorship of...
| Publicado en: | Digital Scholarship in the Humanities Vol. 31; no. 4; pp. 689 - 711 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
12/1/2016
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=120634611&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 120634611 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: 12/1/2016 vid: 31 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 120634611 10.1093/llc/fqv019 ppf: 689 ppct: 22 formats: fmt: @attributes: type: P size: 964KB tig: atl: Profile-based authorship analysis. aug: au: Dunn, Jonathan Argamon, Shlomo Rasooli, Amin Kumar, Geet affil: Illinois Institute of Technology, Chicago, IL, USA su: Authorship Ideology Document clustering Content analysis Data mining sug: subj: Authorship Ideology Document clustering Content analysis Data mining ab: This article presents a profile-based authorship analysis method which first categorizes texts according to social and conceptual characteristics of their author (e.g. Sex and Political Ideology) and then combines these profiles for two authorship analysis tasks: (1) determining shared authorship of pairs of texts without a set of candidate authors and (2) clustering texts according to characteristics of their authors in order to provide an analysis of the types of individuals represented in the data set. The first task outperforms Burrows' Delta by a wide margin on short texts and a small margin on long texts. The second task has no such benchmark with existing methods. The data set for evaluating the method consists of speeches from the US House and Senate from 1995 to 2013. This data set contains both a large number of texts (42,000 in the test sets) and a large number of speakers (over 800). The article shows that this approach to authorship analysis is more accurate than existing approaches given a data set with hundreds of authors. Further, this profile-based method makes new types of analysis possible by looking at types of individuals as well as at specific individuals. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2016 holdings: @attributes: islocal: N |
|---|