Creating specialized corpora from digitized historical newspaper archives: An iterative bootstrapping approach.
The availability of large digital archives of historical newspaper content has transformed the historical sciences. However, the scale of these archives can limit the direct application of advanced text processing methods. Even if it is computationally feasible to apply sophisticated language proces...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 2; pp. 779 - 798 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=164367989&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 164367989 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2023 vid: 38 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 164367989 10.1093/llc/fqac079 ppf: 779 ppct: 19 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.5MB tig: atl: Creating specialized corpora from digitized historical newspaper archives: An iterative bootstrapping approach. aug: au: Black, Joshua Wilson affil: UC Arts Digital Lab, University of Canterbury , Christchurch, New Zealand New Zealand Institute of Language, Brain and Behaviour, University of Canterbury , Christchurch, New Zealand su: Historical libraries Corpora Digital libraries Text mining Problem solving New Zealand sug: subj: New Zealand Historical libraries Corpora Digital libraries Text mining Problem solving ab: The availability of large digital archives of historical newspaper content has transformed the historical sciences. However, the scale of these archives can limit the direct application of advanced text processing methods. Even if it is computationally feasible to apply sophisticated language processing to an entire digital archive, if the material of interest is a small fraction of the archive, the results are unlikely to be useful. Methods for generating smaller specialized corpora from large archives are required to solve this problem. This article presents such a method for historical newspaper archives digitized using the METS/ALTO XML standard (Veridian Software, n.d.). The method is an 'iterative bootstrapping' approach in which candidate corpora are evaluated using text mining techniques, items are manually labelled, and Naïve Bayes text classifiers are trained and applied in order to produce new candidate corpora. The method is illustrated by a case study that investigates philosophical content, broadly construed, in pre-1900 English-language New Zealand newspapers. Extensive code is provided in Supplementary Materials. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|