Quantifying the impact of dirty OCR on historical text analysis: Eighteenth Century Collections Online as a case study.

This article aims to quantify the impact optical character recognition (OCR) has on the quantitative analysis of historical documents. Using Eighteenth Century Collections Online as a case study, we first explore and explain the differences between the OCR corpus and its keyed-in counterpart, create...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 34; no. 4; pp. 825 - 844
Autores principales: Hill, Mark J, Hengchen, Simon
Formato: Artículo
Publicado: Oxford University Press / USA Dec2019
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=140352772&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 140352772
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Dec2019
      vid: 34
      iid: 4
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        140352772
        10.1093/llc/fqz024
      ppf: 825
      ppct: 19
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 838KB
      tig:
        atl: Quantifying the impact of dirty OCR on historical text analysis: Eighteenth Century Collections Online as a case study.
      aug:
        au:
          Hill, Mark J
          Hengchen, Simon
        affil: COMHIS, Department of Digital Humanities, University of Helsinki, Finland
      su:
        Eighteenth century
        Historical analysis
        Optical character recognition
        Digital humanities
        Attribution of authorship
        Vector spaces
      sug:
        subj:
          Eighteenth century
          Historical analysis
          Optical character recognition
          Digital humanities
          Attribution of authorship
          Vector spaces
      ab: This article aims to quantify the impact optical character recognition (OCR) has on the quantitative analysis of historical documents. Using Eighteenth Century Collections Online as a case study, we first explore and explain the differences between the OCR corpus and its keyed-in counterpart, created by the Text Creation Partnership. We then conduct a series of specific analyses common to the digital humanities: topic modelling, authorship attribution, collocation analysis, and vector space modelling. The article concludes by offering some preliminary thoughts on how these conclusions can be applied to other datasets, by reflecting on the potential for predicting the quality of OCR where no ground-truth exists.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2019
    holdings:
      @attributes:
        islocal: N