An HPC-Ready, Wikidata-Based Workflow for Exploratory Geocoding of Unstructured Textual Corpora.
Geocoding, the task of linking place names in text to geographic coordinates, is a cornerstone of spatial humanities research, yet many existing tools assume structured data, contemporary toponyms, or commercial geocoding services that limit reuse. Humanities corpora, by contrast, are often unstruct...
| Publicado en: | Journal of Open Humanities Data Vol. 11; no. 1; pp. 1 - 10 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Ubiquity Press
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190650377&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 190650377 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2059481X MPQW jtl: Journal of Open Humanities Data issn: 2059481X maglogo: N pubinfo: dt: 2025 vid: 11 iid: 1 pid: 83901 pub: Ubiquity Press artinfo: ui: 190650377 10.5334/johd.401 ppf: 1 ppct: 9 formats: tig: atl: An HPC-Ready, Wikidata-Based Workflow for Exploratory Geocoding of Unstructured Textual Corpora. aug: au: Lamar, Annie K. affil: Department of Classics, University of California, Santa Barbara, Santa Barbara, CA, USA Low-Resource Language (LOREL) Lab, University of California, Santa Barbara, Santa Barbara, CA, USA su: High performance computing Natural language processing Databases Geotagging Geography Workflow Multilingual communication sug: subj: High performance computing Natural language processing Databases Geotagging Geography Workflow Multilingual communication keyword: digital humanities gazetteers geocoding literary geography spatial humanities Wikidata ab: Geocoding, the task of linking place names in text to geographic coordinates, is a cornerstone of spatial humanities research, yet many existing tools assume structured data, contemporary toponyms, or commercial geocoding services that limit reuse. Humanities corpora, by contrast, are often unstructured, multilingual, and historically variable. This discussion paper presents a scalable, first-pass workflow that applies Wikidata-based geocoding directly to plain-text files through the combined use of Stanford CoreNLP and Python-based Wikidata lookups. The pipeline presents complete shell and SLURM configurations for use on both local machines and high-performance computing (HPC) clusters. This paper details the pipeline's design, explaining its behavior across multilingual and ambiguous toponyms, and situates it in relation to existing gazetteers such as Pleiades, the World Historical Gazetteer, and GeoNames. Limitations, including minimal disambiguation and uneven language coverage, are discussed openly to guide appropriate reuse. The workflow aims to lower the barrier to Wikidata-based geocoding in the humanities by providing a transparent, extensible, and HPC-ready approach for working with unstructured text. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|