From Paginā to Webpage: On Developing and Documenting a Digitized Latin Collection.

In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 mil...

Full description

Bibliographic Details
Published in:Journal of Open Humanities Data Vol. 11; no. 1; pp. 1 - 9
Main Authors: Bothwell, Stephen, Stephan, Kaitlin, Müller, Hildegund, Chiang, David
Format: Article
Published: Ubiquity Press 2025
Subjects:
Online Access:View this record in EBSCOhost
Description
Summary:In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 million characters' worth of data in TSV, PNG, and TXT formats for training optical character recognition (OCR) and post-OCR correction systems. The third is ND-DLC-Tools: a set of Python scripts for reproducing our digitization workflow. Together, these repositories make many Latin texts computationally accessible and provide resources to bolster digitization efforts.