Stylometric similarity in literary corpora: Non-authorship clustering and Deutscher Novellenschatz.
A distant-reading task in literary corpus analysis is to group stylometrically similar texts. Since there are many ways to define writing style, the result not only depends on the clustering method but even more so on the measure of similarity. With authorship attribution, the predominant applicatio...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 1; pp. 277 - 296 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=162941102&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 162941102 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2023 vid: 38 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 162941102 10.1093/llc/fqac039 ppf: 277 ppct: 19 formats: fmt: – @attributes: type: T – @attributes: type: P size: 2.6MB tig: atl: Stylometric similarity in literary corpora: Non-authorship clustering and Deutscher Novellenschatz. aug: au: Päpcke, Simon Weitin, Thomas Herget, Katharina Glawion, Anastasia Brandes, Ulrik affil: Social Networks Lab , ETH Zurich, Zurich, Switzerland LitLab , TU Darmstadt, Darmstadt, Germany su: Corpora Attribution of authorship Stylometry Novellas (Literary form) sug: subj: Corpora Attribution of authorship Stylometry Novellas (Literary form) ab: A distant-reading task in literary corpus analysis is to group stylometrically similar texts. Since there are many ways to define writing style, the result not only depends on the clustering method but even more so on the measure of similarity. With authorship attribution, the predominant application of stylometry, as its benchmark much research has addressed the utility of methods for measuring similarity. We use a corpus of German-language novellas to demonstrate that one may be interested in very different meaningful groups of texts simultaneously, and that these can be recovered from stylometric clustering if the measure is chosen accordingly. As can be expected, different measures do better at recovering groups associated with, for instance, subgenre, author gender, or narrative perspective. As a consequence, it is suggested that corpus analyses should not be based on what is currently considered the most refined measure of stylometric similarity, but rather break down the decisions that yield a specific measure and provide substantively justified arguments for them. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|