A small set of stylometric features differentiates Latin prose and verse.

Identifying the stylistic signatures characteristic of different genres is of central importance to literary theory and criticism. In this article we report a large-scale computational analysis of Latin prose and verse using a combination of quantitative stylistics and supervised machine learning. W...

Full description

Bibliographic Details
Published in:Digital Scholarship in the Humanities Vol. 34; no. 4; pp. 716 - 730
Main Authors: Chaudhuri, Pramit, Dasgupta, Tathagata, Dexter, Joseph P, Iyer, Krithika
Format: Article
Published: Oxford University Press / USA Dec2019
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:Identifying the stylistic signatures characteristic of different genres is of central importance to literary theory and criticism. In this article we report a large-scale computational analysis of Latin prose and verse using a combination of quantitative stylistics and supervised machine learning. We train a set of classifiers to differentiate prose and poetry with high accuracy (>97%) based on a set of twenty-six text-based, primarily syntactic features and rank the relative importance of these features to identify a low-dimensional set still sufficient to achieve excellent classifier performance. This analysis demonstrates that Latin prose and verse can be classified effectively using just three top features. From examination of the highly ranked features, we observe that measures of the hypotactic style favored in Latin prose (i.e. subordinating constructions in complex sentences, such as relative clauses) are especially useful for classification.