A syntactic characterization of authorship style surrounding proper names.
Accurately determining who wrote a manuscript has captivated scholars of literary history for centuries, as the true author can have important ramifications in religion, law, literary studies, philosophy, and education. A wide array of lexical, character, syntactic, semantic, and application-specifi...
| Publicado en: | Digital Scholarship in the Humanities pp. 53 - 71 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
04/01/2015
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=108489934&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 108489934 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: 04/01/2015 pid: 622 pub: Oxford University Press / USA artinfo: ui: 108489934 10.1093/llc/fqt033 ppf: 53 ppct: 18 formats: fmt: @attributes: type: P size: 9.2MB tig: atl: A syntactic characterization of authorship style surrounding proper names. aug: au: Lučić, Ana Blake, Catherine L. affil: University of Illinois, Urbana-Champaign, USA su: Manuscripts Authorship Natural language processing Artificial intelligence Electronic data processing sug: subj: Manuscripts Authorship Natural language processing Artificial intelligence Electronic data processing ab: Accurately determining who wrote a manuscript has captivated scholars of literary history for centuries, as the true author can have important ramifications in religion, law, literary studies, philosophy, and education. A wide array of lexical, character, syntactic, semantic, and application-specific features have been proposed to represent a text so that authorship attribution can be established automatically. Although surface-level features have been tested extensively, few studies have systematically explored high-level features, in part due to limitations in the natural language processing techniques required to capture highlevel features. However, high-level features, such as sentence structure, are used subconsciously by a writer and thus may be more consistent than surface-level features, such as word choice. In this article, we introduce a new high-level feature based on local syntactic dependencies that an author uses when referring to a named entity (in our case a person's name). The series of experiments in the contexts of movie reviews reveal how the amount of data in both the training and test sets influences predictive performance. Finally, we measure authorship consistency with respect to this new feature and show how consistency influences predictive performance. These results provide other researchers with a new model for how to evaluate new features and suggest that the local syntactic dependencies warrant further investigation. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2015 holdings: @attributes: islocal: N |
|---|