Improving the Accuracy of Normalizing Historical Estonian Texts by Combining Statistical Machine Translation with BERT Language Model.
Automatic analysis of historical texts is often hindered by different spelling system compared to the one used today. One approach to address this issue is to convert these texts to present-day spelling conventions, also called normalizing. One of the relatively old, however, still used methods for...
| Publicado en: | Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications) Vol. 7; no. 2; pp. 1 - 10 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
University of Oslo
2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| Sumario: | Automatic analysis of historical texts is often hindered by different spelling system compared to the one used today. One approach to address this issue is to convert these texts to present-day spelling conventions, also called normalizing. One of the relatively old, however, still used methods for that is character level statistical machine translation (CSMT). As the method views the text one word at a time, it can lead to wrong normalizations due to lack of context. This paper gives an overview of integration of CSMT with BERT language models on texts written in Estonian, a morphologically rich language, to increase normalization accuracy. |
|---|