Linguistic annotation of cuneiform texts using treebanks and deep learning.
We describe an efficient pipeline for morpho-syntactically annotating an ancient language corpus which takes advantage of bootstrapping techniques. This pipeline is designed for ancient language scholars looking to jump-start their own treebank projects, which can in turn serve further pedagogical r...
| Publicado en: | Digital Scholarship in the Humanities Vol. 39; no. 1; pp. 296 - 308 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=176806366&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 176806366 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2024 vid: 39 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 176806366 10.1093/llc/fqae002 ppf: 296 ppct: 12 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: Linguistic annotation of cuneiform texts using treebanks and deep learning. aug: au: Ong, Matthew Gordin, Shai affil: Middle Eastern Languages and Cultures, UC Berkeley , Berkeley, CA, United States Digital Pasts Lab, Department of Land of Israel Studies and Archaeology, Ariel University , Ariel, Israel Digital Humanities and Social Sciences Hub, Open University of Israel , Ra'anana, Israel su: Language models Machine learning Deep learning Annotations Corpora Assyria sug: subj: Assyria Language models Machine learning Deep learning Annotations Corpora keyword: Akkadian treebank bootstrapping cuneiform Neo-Assyrian letters spaCy ab: We describe an efficient pipeline for morpho-syntactically annotating an ancient language corpus which takes advantage of bootstrapping techniques. This pipeline is designed for ancient language scholars looking to jump-start their own treebank projects, which can in turn serve further pedagogical research projects in the target language. We situate our work in the field of similar ancient language treebank projects, arguing that our approach shows that individual humanities scholars can leverage current machine-learning tools to produce their own richly annotated corpora. We illustrate this pipeline by producing a new Akkadian-language treebank based on two volumes from the online editions of the State Archives of Assyria project hosted on Oracc, as well as a spaCy language model named AkkParser trained on that treebank. Both of these are made publicly available for annotating other Akkadian corpora. In addition, we discuss linguistic issues particular to the Neo-Assyrian letter corpus and data-encoding complications of cuneiform texts in Oracc. The strategies, language models, and processing scripts we developed to handle both linguistic and data-encoding issues in this project will be of special interest to scholars seeking to develop their own cuneiform treebanks. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|