On the implementation of Latin part-of-speech taggers in intertextuality analysis: TreeTagger, CLTK, Cracovia system, LatinCy, and ChatGPT compared.

Digital-assisted intertextuality analysis often yields large amounts of results, many of which are irrelevant to researchers from a hermeneutic point of view. One strategy for minimizing these hermeneutically non meaningful findings is to apply a filter that sorts the results based on specified part...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 40; no. 1; pp. 329 - 338
Autores principales: Wittweiler, Michael, Schropp, Franziska, Konrad, Thomas E, Revellio, Marie, Feichtinger, Barbara
Formato: Artículo
Publicado: Oxford University Press / USA Apr2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=184296825&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 184296825
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Apr2025
      vid: 40
      iid: 1
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        184296825
        10.1093/llc/fqae078
      ppf: 329
      ppct: 9
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 785KB
      tig:
        atl: On the implementation of Latin part-of-speech taggers in intertextuality analysis: TreeTagger, CLTK, Cracovia system, LatinCy, and ChatGPT compared.
      aug:
        au:
          Wittweiler, Michael
          Schropp, Franziska
          Konrad, Thomas E
          Revellio, Marie
          Feichtinger, Barbara
        affil: Department of Literature, Art and Media Studies, Universität Konstanz, Konstanz, Germany
      su:
        Machine learning
        Latin literature
        ChatGPT
        Transformer models
        Generative pre-trained transformers
      sug:
        subj:
          Machine learning
          Latin literature
          ChatGPT
          Transformer models
          Generative pre-trained transformers
      keyword:
        automated citation detection
        evaluation
        intertextuality
        Jerome
        part-of-speech tagging
        POS tagging
        taggers
      ab: Digital-assisted intertextuality analysis often yields large amounts of results, many of which are irrelevant to researchers from a hermeneutic point of view. One strategy for minimizing these hermeneutically non meaningful findings is to apply a filter that sorts the results based on specified parts-of-speech. Building on Mare Revellio's historical text-reuse grammar (HTRG) we demonstrate that an implementation of such a filter accommodating the hermeneutical context proves to be advantageous. We assessed the performance of various Latin part-of-speech (POS) taggers to refine our filtering process, using evaluation data on text congruencies from the Latin authors Virgil and Jerome. Among the Classical Language Toolkit , the TreeTagger , the Cracovia system , LatinCy , and ChatGPT , the Cracovia system surpassed the other taggers by approximately 2 percentage points in accuracy. While this tagger leverages transformer-based machine learning algorithms, the older, probabilistic-based TreeTagger demonstrated competitive performance. Although GPT-4 showed remarkable results, it still lags behind the state-of-the-art taggers for Latin. In order to build a powerful digital citation detection tool for intertextual relationships in ancient texts, the most accurate analysis of POS is crucial in filtering and evaluating valid citations.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N