An HPC-Ready, Wikidata-Based Workflow for Exploratory Geocoding of Unstructured Textual Corpora.

Geocoding, the task of linking place names in text to geographic coordinates, is a cornerstone of spatial humanities research, yet many existing tools assume structured data, contemporary toponyms, or commercial geocoding services that limit reuse. Humanities corpora, by contrast, are often unstruct...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Open Humanities Data Vol. 11; no. 1; pp. 1 - 10
Autor principal: Lamar, Annie K.
Formato: Artículo
Publicado: Ubiquity Press 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190650377&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 190650377
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2059481X
        MPQW
      jtl: Journal of Open Humanities Data
      issn: 2059481X
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 11
      iid: 1
      pid: 83901
      pub: Ubiquity Press
    artinfo:
      ui:
        190650377
        10.5334/johd.401
      ppf: 1
      ppct: 9
      formats:
      tig:
        atl: An HPC-Ready, Wikidata-Based Workflow for Exploratory Geocoding of Unstructured Textual Corpora.
      aug:
        au: Lamar, Annie K.
        affil:
          Department of Classics, University of California, Santa Barbara, Santa Barbara, CA, USA
          Low-Resource Language (LOREL) Lab, University of California, Santa Barbara, Santa Barbara, CA, USA
      su:
        High performance computing
        Natural language processing
        Databases
        Geotagging
        Geography
        Workflow
        Multilingual communication
      sug:
        subj:
          High performance computing
          Natural language processing
          Databases
          Geotagging
          Geography
          Workflow
          Multilingual communication
      keyword:
        digital humanities
        gazetteers
        geocoding
        literary geography
        spatial humanities
        Wikidata
      ab: Geocoding, the task of linking place names in text to geographic coordinates, is a cornerstone of spatial humanities research, yet many existing tools assume structured data, contemporary toponyms, or commercial geocoding services that limit reuse. Humanities corpora, by contrast, are often unstructured, multilingual, and historically variable. This discussion paper presents a scalable, first-pass workflow that applies Wikidata-based geocoding directly to plain-text files through the combined use of Stanford CoreNLP and Python-based Wikidata lookups. The pipeline presents complete shell and SLURM configurations for use on both local machines and high-performance computing (HPC) clusters. This paper details the pipeline's design, explaining its behavior across multilingual and ambiguous toponyms, and situates it in relation to existing gazetteers such as Pleiades, the World Historical Gazetteer, and GeoNames. Limitations, including minimal disambiguation and uneven language coverage, are discussed openly to guide appropriate reuse. The workflow aims to lower the barrier to Wikidata-based geocoding in the humanities by providing a transparent, extensible, and HPC-ready approach for working with unstructured text.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N