On the implementation of Latin part-of-speech taggers in intertextuality analysis: TreeTagger, CLTK, Cracovia system, LatinCy, and ChatGPT compared.
Digital-assisted intertextuality analysis often yields large amounts of results, many of which are irrelevant to researchers from a hermeneutic point of view. One strategy for minimizing these hermeneutically non meaningful findings is to apply a filter that sorts the results based on specified part...
| Publicado en: | Digital Scholarship in the Humanities Vol. 40; no. 1; pp. 329 - 338 |
|---|---|
| Autores principales: | , , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=184296825&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 184296825 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2025 vid: 40 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 184296825 10.1093/llc/fqae078 ppf: 329 ppct: 9 formats: fmt: – @attributes: type: T – @attributes: type: P size: 785KB tig: atl: On the implementation of Latin part-of-speech taggers in intertextuality analysis: TreeTagger, CLTK, Cracovia system, LatinCy, and ChatGPT compared. aug: au: Wittweiler, Michael Schropp, Franziska Konrad, Thomas E Revellio, Marie Feichtinger, Barbara affil: Department of Literature, Art and Media Studies, Universität Konstanz, Konstanz, Germany su: Machine learning Latin literature ChatGPT Transformer models Generative pre-trained transformers sug: subj: Machine learning Latin literature ChatGPT Transformer models Generative pre-trained transformers keyword: automated citation detection evaluation intertextuality Jerome part-of-speech tagging POS tagging taggers ab: Digital-assisted intertextuality analysis often yields large amounts of results, many of which are irrelevant to researchers from a hermeneutic point of view. One strategy for minimizing these hermeneutically non meaningful findings is to apply a filter that sorts the results based on specified parts-of-speech. Building on Mare Revellio's historical text-reuse grammar (HTRG) we demonstrate that an implementation of such a filter accommodating the hermeneutical context proves to be advantageous. We assessed the performance of various Latin part-of-speech (POS) taggers to refine our filtering process, using evaluation data on text congruencies from the Latin authors Virgil and Jerome. Among the Classical Language Toolkit , the TreeTagger , the Cracovia system , LatinCy , and ChatGPT , the Cracovia system surpassed the other taggers by approximately 2 percentage points in accuracy. While this tagger leverages transformer-based machine learning algorithms, the older, probabilistic-based TreeTagger demonstrated competitive performance. Although GPT-4 showed remarkable results, it still lags behind the state-of-the-art taggers for Latin. In order to build a powerful digital citation detection tool for intertextual relationships in ancient texts, the most accurate analysis of POS is crucial in filtering and evaluating valid citations. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|