Creating specialized corpora from digitized historical newspaper archives: An iterative bootstrapping approach.

The availability of large digital archives of historical newspaper content has transformed the historical sciences. However, the scale of these archives can limit the direct application of advanced text processing methods. Even if it is computationally feasible to apply sophisticated language proces...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 38; no. 2; pp. 779 - 798
Autor principal: Black, Joshua Wilson
Formato: Artículo
Publicado: Oxford University Press / USA Jun2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=164367989&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 164367989
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2023
      vid: 38
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        164367989
        10.1093/llc/fqac079
      ppf: 779
      ppct: 19
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 1.5MB
      tig:
        atl: Creating specialized corpora from digitized historical newspaper archives: An iterative bootstrapping approach.
      aug:
        au: Black, Joshua Wilson
        affil:
          UC Arts Digital Lab, University of Canterbury , Christchurch, New Zealand
          New Zealand Institute of Language, Brain and Behaviour, University of Canterbury , Christchurch, New Zealand
      su:
        Historical libraries
        Corpora
        Digital libraries
        Text mining
        Problem solving
        New Zealand
      sug:
        subj:
          New Zealand
          Historical libraries
          Corpora
          Digital libraries
          Text mining
          Problem solving
      ab: The availability of large digital archives of historical newspaper content has transformed the historical sciences. However, the scale of these archives can limit the direct application of advanced text processing methods. Even if it is computationally feasible to apply sophisticated language processing to an entire digital archive, if the material of interest is a small fraction of the archive, the results are unlikely to be useful. Methods for generating smaller specialized corpora from large archives are required to solve this problem. This article presents such a method for historical newspaper archives digitized using the METS/ALTO XML standard (Veridian Software, n.d.). The method is an 'iterative bootstrapping' approach in which candidate corpora are evaluated using text mining techniques, items are manually labelled, and Naïve Bayes text classifiers are trained and applied in order to produce new candidate corpora. The method is illustrated by a case study that investigates philosophical content, broadly construed, in pre-1900 English-language New Zealand newspapers. Extensive code is provided in Supplementary Materials.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N