From Paginā to Webpage: On Developing and Documenting a Digitized Latin Collection.

In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 mil...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Open Humanities Data Vol. 11; no. 1; pp. 1 - 9
Autores principales: Bothwell, Stephen, Stephan, Kaitlin, Müller, Hildegund, Chiang, David
Formato: Artículo
Publicado: Ubiquity Press 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 million characters' worth of data in TSV, PNG, and TXT formats for training optical character recognition (OCR) and post-OCR correction systems. The third is ND-DLC-Tools: a set of Python scripts for reproducing our digitization workflow. Together, these repositories make many Latin texts computationally accessible and provide resources to bolster digitization efforts.