Reflections on Encoding Languages in Historical Data: Working With the Multilingual Dimension of the Dutch East India Company Archives.
This article investigates the challenges of encoding languages in historical data through the example of a reference dataset: a thesaurus in SKOS format of commodities traded by the Dutch East India Company (VOC). The VOC archives, from which this thesaurus draws a lot of its data, are far from pure...
| Publicado en: | Journal of Open Humanities Data Vol. 10; no. 1; pp. 1 - 11 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Ubiquity Press
2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=182792289&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 182792289 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2059481X MPQW jtl: Journal of Open Humanities Data issn: 2059481X maglogo: N pubinfo: dt: 2024 vid: 10 iid: 1 pid: 83901 pub: Ubiquity Press artinfo: ui: 182792289 10.5334/johd.176 ppf: 1 ppct: 10 formats: tig: atl: Reflections on Encoding Languages in Historical Data: Working With the Multilingual Dimension of the Dutch East India Company Archives. aug: au: Pepping, K. W. affil: Huygens Institute for the History of the Netherlands, Amsterdam, The Netherlands su: Nederlandsche Oost-Indische Cie. Linguistic complexity Language acquisition Dutch language Commodity futures Loanwords sug: subj: Nederlandsche Oost-Indische Cie. Linguistic complexity Language acquisition Dutch language Commodity futures Loanwords keyword: code-switching Dutch East India Company (VOC) htr language tagging reference data skos ab: This article investigates the challenges of encoding languages in historical data through the example of a reference dataset: a thesaurus in SKOS format of commodities traded by the Dutch East India Company (VOC). The VOC archives, from which this thesaurus draws a lot of its data, are far from purely Dutch. The company's multilingual workforce and interactions across Asia resulted in records influenced by a multitude of languages, full of loanwords and citations. This is further complicated by the VOC's role in colonising regions and suppressing local languages, resulting in some languages potentially only surviving in these 'Dutch' archives. This means that when working with a large corpus like the VOC archives, various challenges arise regarding historical language evolution, vocabulary borrowing, extinct languages, technical standards that are not geared towards historical context, and political sensitivities around identity-bound language. The article demonstrates how the GLOBALISE project navigates these issues by prioritising transparency, flexibility, and iterative refinement. It argues that as long as researchers are aware of the challenges, language complexities are not a roadblock but offer opportunities for further research and critical engagement with the past, encouraging broader discussions and creative solutions for encoding historical multilingualism and development of language. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|