| Sumario: | In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 million characters' worth of data in TSV, PNG, and TXT formats for training optical character recognition (OCR) and post-OCR correction systems. The third is ND-DLC-Tools: a set of Python scripts for reproducing our digitization workflow. Together, these repositories make many Latin texts computationally accessible and provide resources to bolster digitization efforts.
|