Reflections on Encoding Languages in Historical Data: Working With the Multilingual Dimension of the Dutch East India Company Archives.

This article investigates the challenges of encoding languages in historical data through the example of a reference dataset: a thesaurus in SKOS format of commodities traded by the Dutch East India Company (VOC). The VOC archives, from which this thesaurus draws a lot of its data, are far from pure...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Open Humanities Data Vol. 10; no. 1; pp. 1 - 11
Autor principal: Pepping, K. W.
Formato: Artículo
Publicado: Ubiquity Press 2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=182792289&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 182792289
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2059481X
        MPQW
      jtl: Journal of Open Humanities Data
      issn: 2059481X
      maglogo: N
    pubinfo:
      dt: 2024
      vid: 10
      iid: 1
      pid: 83901
      pub: Ubiquity Press
    artinfo:
      ui:
        182792289
        10.5334/johd.176
      ppf: 1
      ppct: 10
      formats:
      tig:
        atl: Reflections on Encoding Languages in Historical Data: Working With the Multilingual Dimension of the Dutch East India Company Archives.
      aug:
        au: Pepping, K. W.
        affil: Huygens Institute for the History of the Netherlands, Amsterdam, The Netherlands
      su:
        Nederlandsche Oost-Indische Cie.
        Linguistic complexity
        Language acquisition
        Dutch language
        Commodity futures
        Loanwords
      sug:
        subj:
          Nederlandsche Oost-Indische Cie.
          Linguistic complexity
          Language acquisition
          Dutch language
          Commodity futures
          Loanwords
      keyword:
        code-switching
        Dutch East India Company (VOC)
        htr
        language tagging
        reference data
        skos
      ab: This article investigates the challenges of encoding languages in historical data through the example of a reference dataset: a thesaurus in SKOS format of commodities traded by the Dutch East India Company (VOC). The VOC archives, from which this thesaurus draws a lot of its data, are far from purely Dutch. The company's multilingual workforce and interactions across Asia resulted in records influenced by a multitude of languages, full of loanwords and citations. This is further complicated by the VOC's role in colonising regions and suppressing local languages, resulting in some languages potentially only surviving in these 'Dutch' archives. This means that when working with a large corpus like the VOC archives, various challenges arise regarding historical language evolution, vocabulary borrowing, extinct languages, technical standards that are not geared towards historical context, and political sensitivities around identity-bound language. The article demonstrates how the GLOBALISE project navigates these issues by prioritising transparency, flexibility, and iterative refinement. It argues that as long as researchers are aware of the challenges, language complexities are not a roadblock but offer opportunities for further research and critical engagement with the past, encouraging broader discussions and creative solutions for encoding historical multilingualism and development of language.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2024
    holdings:
      @attributes:
        islocal: N