From Paginā to Webpage: On Developing and Documenting a Digitized Latin Collection.

In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 mil...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Open Humanities Data Vol. 11; no. 1; pp. 1 - 9
Autores principales: Bothwell, Stephen, Stephan, Kaitlin, Müller, Hildegund, Chiang, David
Formato: Artículo
Publicado: Ubiquity Press 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190650359&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 190650359
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2059481X
        MPQW
      jtl: Journal of Open Humanities Data
      issn: 2059481X
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 11
      iid: 1
      pid: 83901
      pub: Ubiquity Press
    artinfo:
      ui:
        190650359
        10.5334/johd.397
      ppf: 1
      ppct: 8
      formats:
      tig:
        atl: From Paginā to Webpage: On Developing and Documenting a Digitized Latin Collection.
      aug:
        au:
          Bothwell, Stephen
          Stephan, Kaitlin
          Müller, Hildegund
          Chiang, David
        affil:
          Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, Indiana, USA
          Department of Classics, University of Notre Dame, Notre Dame, Indiana, USA
      su:
        Digitization
        Digital libraries
        Document markup languages
        Python programming language
        Electronic publications
        Data libraries
        Optical character recognition
      sug:
        subj:
          Digitization
          Digital libraries
          Document markup languages
          Python programming language
          Electronic publications
          Data libraries
          Optical character recognition
      keyword:
        digitization
        Latin
        OCR
        post-correction
        TEI
      ab: In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 million characters' worth of data in TSV, PNG, and TXT formats for training optical character recognition (OCR) and post-OCR correction systems. The third is ND-DLC-Tools: a set of Python scripts for reproducing our digitization workflow. Together, these repositories make many Latin texts computationally accessible and provide resources to bolster digitization efforts.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N