A small set of stylometric features differentiates Latin prose and verse.

Identifying the stylistic signatures characteristic of different genres is of central importance to literary theory and criticism. In this article we report a large-scale computational analysis of Latin prose and verse using a combination of quantitative stylistics and supervised machine learning. W...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 34; no. 4; pp. 716 - 730
Autores principales: Chaudhuri, Pramit, Dasgupta, Tathagata, Dexter, Joseph P, Iyer, Krithika
Formato: Artículo
Publicado: Oxford University Press / USA Dec2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=140352765&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 140352765
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Dec2019
      vid: 34
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        140352765
        10.1093/llc/fqy070
      ppf: 716
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 359KB
      tig:
        atl: A small set of stylometric features differentiates Latin prose and verse.
      aug:
        au:
          Chaudhuri, Pramit
          Dasgupta, Tathagata
          Dexter, Joseph P
          Iyer, Krithika
        affil:
          Department of Classics, University of Texas at Austin, TX, USA
          Department of Systems Biology, Harvard Medical School, MA, USA
          Plano East Senior High School, TX, USA and Center for Excellence in Education, Research Science Institute, VA, US
      su:
        Relative clauses
        Literary theory
        Supervised learning
        Literary criticism
        Prose poems
      sug:
        subj:
          Relative clauses
          Literary theory
          Supervised learning
          Literary criticism
          Prose poems
      ab: Identifying the stylistic signatures characteristic of different genres is of central importance to literary theory and criticism. In this article we report a large-scale computational analysis of Latin prose and verse using a combination of quantitative stylistics and supervised machine learning. We train a set of classifiers to differentiate prose and poetry with high accuracy (>97%) based on a set of twenty-six text-based, primarily syntactic features and rank the relative importance of these features to identify a low-dimensional set still sufficient to achieve excellent classifier performance. This analysis demonstrates that Latin prose and verse can be classified effectively using just three top features. From examination of the highly ranked features, we observe that measures of the hypotactic style favored in Latin prose (i.e. subordinating constructions in complex sentences, such as relative clauses) are especially useful for classification.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N