Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype.

The writings of one ancient civilization often overlap in time and space with others. Many of these sources comprise unstructured text in ancient languages, causing scholars studying these civilizations to be siloed, often relying on sources in specific languages. Most recent efforts to extract stru...

Full description

Bibliographic Details
Published in:Language Resources & Evaluation Vol. 59; no. 3; pp. 2427 - 2452
Main Authors: Sagi, Tomer, Zaga, Moran, Rusinek, Sinai, Fekete, Marcell R., Bjerva, Johannes, Hose, Katja
Format: Article
Published: Springer Nature Sep2025
Subjects:
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=186909071&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 186909071
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: Sep2025
      vid: 59
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        186909071
        10.1007/s10579-025-09812-9
      ppf: 2427
      ppct: 25
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 2MB
      tig:
        atl: Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype.
      aug:
        au:
          Sagi, Tomer
          Zaga, Moran
          Rusinek, Sinai
          Fekete, Marcell R.
          Bjerva, Johannes
          Hose, Katja
        affil:
          https://ror.org/04m5j1k67 Department of Computer Science, Aalborg University, Aalborg, Denmark
          https://ror.org/02f009v59 e-Lijah Lab, University of Haifa, Haifa, Israel
          https://ror.org/04d836q62 Institute of Logic and Computation, TU Wien, Vienna, Austria
      su:
        Hebrew language
        Arabic language
        Transliteration
        Geographic names
        Bilingualism
        Phonology
        Historical source material
      sug:
        subj:
          Hebrew language
          Arabic language
          Transliteration
          Geographic names
          Bilingualism
          Phonology
          Historical source material
      keyword:
        Communication and Culture Linguistics
        Grapheme to phoneme
        Language
        Multi-lingual
        Toponym matching
      ab: The writings of one ancient civilization often overlap in time and space with others. Many of these sources comprise unstructured text in ancient languages, causing scholars studying these civilizations to be siloed, often relying on sources in specific languages. Most recent efforts to extract structured information from historical scripts into place (toponym) and people databases (prospographies) have followed this pattern, focusing on one civilization and selected sources. The path to creating a common database runs through aligning names or toponyms between sources from disparate languages utilizing different scripts. Existing multi-lingual orthographic (string-based) comparison often relies on transliteration to a common script (Latin/English). Transliteration often creates multiple options and even more confusion. However, when integrating sources that overlap in space and time, the languages often share a common phonetic background. This commonality may prove beneficial. In this work, we present a benchmark for comparing toponyms from two linguistically and culturally related languages, namely Hebrew and Arabic. We provide a benchmark comprised of a set of dataset pairs created from historical sources written in Medieval variants of these languages, later historical Gazetteers and a modern dataset curated from Wikidata. We empirically evaluate several toponym comparison approaches over the benchmark: transliteration to a common script, direct transliteration, and phonetic comparison using a common phonetic representation. We discuss the results and the limitations of the various methods and outline future work.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2025. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N