Stylometric similarity in literary corpora: Non-authorship clustering and Deutscher Novellenschatz.

A distant-reading task in literary corpus analysis is to group stylometrically similar texts. Since there are many ways to define writing style, the result not only depends on the clustering method but even more so on the measure of similarity. With authorship attribution, the predominant applicatio...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 38; no. 1; pp. 277 - 296
Autores principales: Päpcke, Simon, Weitin, Thomas, Herget, Katharina, Glawion, Anastasia, Brandes, Ulrik
Formato: Artículo
Publicado: Oxford University Press / USA Apr2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162941102&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 162941102
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2023
      vid: 38
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        162941102
        10.1093/llc/fqac039
      ppf: 277
      ppct: 19
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2.6MB
      tig:
        atl: Stylometric similarity in literary corpora: Non-authorship clustering and Deutscher Novellenschatz.
      aug:
        au:
          Päpcke, Simon
          Weitin, Thomas
          Herget, Katharina
          Glawion, Anastasia
          Brandes, Ulrik
        affil:
          Social Networks Lab , ETH Zurich, Zurich, Switzerland
          LitLab , TU Darmstadt, Darmstadt, Germany
      su:
        Corpora
        Attribution of authorship
        Stylometry
        Novellas (Literary form)
      sug:
        subj:
          Corpora
          Attribution of authorship
          Stylometry
          Novellas (Literary form)
      ab: A distant-reading task in literary corpus analysis is to group stylometrically similar texts. Since there are many ways to define writing style, the result not only depends on the clustering method but even more so on the measure of similarity. With authorship attribution, the predominant application of stylometry, as its benchmark much research has addressed the utility of methods for measuring similarity. We use a corpus of German-language novellas to demonstrate that one may be interested in very different meaningful groups of texts simultaneously, and that these can be recovered from stylometric clustering if the measure is chosen accordingly. As can be expected, different measures do better at recovering groups associated with, for instance, subgenre, author gender, or narrative perspective. As a consequence, it is suggested that corpus analyses should not be based on what is currently considered the most refined measure of stylometric similarity, but rather break down the decisions that yield a specific measure and provide substantively justified arguments for them.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N