A small set of stylometric features differentiates Latin prose and verse.
Identifying the stylistic signatures characteristic of different genres is of central importance to literary theory and criticism. In this article we report a large-scale computational analysis of Latin prose and verse using a combination of quantitative stylistics and supervised machine learning. W...
| Publicado en: | Digital Scholarship in the Humanities Vol. 34; no. 4; pp. 716 - 730 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Dec2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=140352765&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 140352765 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Dec2019 vid: 34 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 140352765 10.1093/llc/fqy070 ppf: 716 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P size: 359KB tig: atl: A small set of stylometric features differentiates Latin prose and verse. aug: au: Chaudhuri, Pramit Dasgupta, Tathagata Dexter, Joseph P Iyer, Krithika affil: Department of Classics, University of Texas at Austin, TX, USA Department of Systems Biology, Harvard Medical School, MA, USA Plano East Senior High School, TX, USA and Center for Excellence in Education, Research Science Institute, VA, US su: Relative clauses Literary theory Supervised learning Literary criticism Prose poems sug: subj: Relative clauses Literary theory Supervised learning Literary criticism Prose poems ab: Identifying the stylistic signatures characteristic of different genres is of central importance to literary theory and criticism. In this article we report a large-scale computational analysis of Latin prose and verse using a combination of quantitative stylistics and supervised machine learning. We train a set of classifiers to differentiate prose and poetry with high accuracy (>97%) based on a set of twenty-six text-based, primarily syntactic features and rank the relative importance of these features to identify a low-dimensional set still sufficient to achieve excellent classifier performance. This analysis demonstrates that Latin prose and verse can be classified effectively using just three top features. From examination of the highly ranked features, we observe that measures of the hypotactic style favored in Latin prose (i.e. subordinating constructions in complex sentences, such as relative clauses) are especially useful for classification. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|