Quantifying the impact of dirty OCR on historical text analysis: Eighteenth Century Collections Online as a case study.
This article aims to quantify the impact optical character recognition (OCR) has on the quantitative analysis of historical documents. Using Eighteenth Century Collections Online as a case study, we first explore and explain the differences between the OCR corpus and its keyed-in counterpart, create...
| Publicado en: | Digital Scholarship in the Humanities Vol. 34; no. 4; pp. 825 - 844 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Dec2019
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=140352772&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 140352772 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Dec2019 vid: 34 iid: 4 pid: 622 pub: Oxford University Press / USA artinfo: ui: 140352772 10.1093/llc/fqz024 ppf: 825 ppct: 19 formats: fmt: – @attributes: type: T – @attributes: type: P size: 838KB tig: atl: Quantifying the impact of dirty OCR on historical text analysis: Eighteenth Century Collections Online as a case study. aug: au: Hill, Mark J Hengchen, Simon affil: COMHIS, Department of Digital Humanities, University of Helsinki, Finland su: Eighteenth century Historical analysis Optical character recognition Digital humanities Attribution of authorship Vector spaces sug: subj: Eighteenth century Historical analysis Optical character recognition Digital humanities Attribution of authorship Vector spaces ab: This article aims to quantify the impact optical character recognition (OCR) has on the quantitative analysis of historical documents. Using Eighteenth Century Collections Online as a case study, we first explore and explain the differences between the OCR corpus and its keyed-in counterpart, created by the Text Creation Partnership. We then conduct a series of specific analyses common to the digital humanities: topic modelling, authorship attribution, collocation analysis, and vector space modelling. The article concludes by offering some preliminary thoughts on how these conclusions can be applied to other datasets, by reflecting on the potential for predicting the quality of OCR where no ground-truth exists. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2019 holdings: @attributes: islocal: N |
|---|