From Paginā to Webpage: On Developing and Documenting a Digitized Latin Collection.
In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 mil...
| Publicado en: | Journal of Open Humanities Data Vol. 11; no. 1; pp. 1 - 9 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Ubiquity Press
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190650359&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 190650359 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2059481X MPQW jtl: Journal of Open Humanities Data issn: 2059481X maglogo: N pubinfo: dt: 2025 vid: 11 iid: 1 pid: 83901 pub: Ubiquity Press artinfo: ui: 190650359 10.5334/johd.397 ppf: 1 ppct: 8 formats: tig: atl: From Paginā to Webpage: On Developing and Documenting a Digitized Latin Collection. aug: au: Bothwell, Stephen Stephan, Kaitlin Müller, Hildegund Chiang, David affil: Department of Computer Science and Engineering, University of Notre Dame, Notre Dame, Indiana, USA Department of Classics, University of Notre Dame, Notre Dame, Indiana, USA su: Digitization Digital libraries Document markup languages Python programming language Electronic publications Data libraries Optical character recognition sug: subj: Digitization Digital libraries Document markup languages Python programming language Electronic publications Data libraries Optical character recognition keyword: digitization Latin OCR post-correction TEI ab: In this work, we present three Zenodo repositories related to the creation of digital editions for Latin texts. The first is the Notre Dame Digitized Latin Collection (ND-DLC), which contains over 550,000 words of Latin in TEI-XML. The second is the Corpus Correctum (Cor), a dataset offering 3.4 million characters' worth of data in TSV, PNG, and TXT formats for training optical character recognition (OCR) and post-OCR correction systems. The third is ND-DLC-Tools: a set of Python scripts for reproducing our digitization workflow. Together, these repositories make many Latin texts computationally accessible and provide resources to bolster digitization efforts. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|