Improving the Accuracy of Normalizing Historical Estonian Texts by Combining Statistical Machine Translation with BERT Language Model.
Automatic analysis of historical texts is often hindered by different spelling system compared to the one used today. One approach to address this issue is to convert these texts to present-day spelling conventions, also called normalizing. One of the relatively old, however, still used methods for...
| Publicado en: | Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications) Vol. 7; no. 2; pp. 1 - 10 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
University of Oslo
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190250186&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 190250186 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 27041441 NUGX jtl: Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications) issn: 27041441 maglogo: N pubinfo: dt: 2025 vid: 7 iid: 2 pid: 58757 pub: University of Oslo artinfo: ui: 190250186 10.5617/dhnbpub.12299 ppf: 1 ppct: 9 formats: tig: atl: Improving the Accuracy of Normalizing Historical Estonian Texts by Combining Statistical Machine Translation with BERT Language Model. aug: au: Jaanimäe, Gerth affil: University of Tartu su: Estonian language Language models Natural language processing Text processing (Computer science) Historical literature Machine translating Morphology (Grammar) sug: subj: Estonian language Language models Natural language processing Text processing (Computer science) Historical literature Machine translating Morphology (Grammar) keyword: historical texts machine learning normalization ab: Automatic analysis of historical texts is often hindered by different spelling system compared to the one used today. One approach to address this issue is to convert these texts to present-day spelling conventions, also called normalizing. One of the relatively old, however, still used methods for that is character level statistical machine translation (CSMT). As the method views the text one word at a time, it can lead to wrong normalizations due to lack of context. This paper gives an overview of integration of CSMT with BERT language models on texts written in Estonian, a morphologically rich language, to increase normalization accuracy. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|