Improving the Accuracy of Normalizing Historical Estonian Texts by Combining Statistical Machine Translation with BERT Language Model.

Automatic analysis of historical texts is often hindered by different spelling system compared to the one used today. One approach to address this issue is to convert these texts to present-day spelling conventions, also called normalizing. One of the relatively old, however, still used methods for...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications) Vol. 7; no. 2; pp. 1 - 10
Autor principal: Jaanimäe, Gerth
Formato: Artículo
Publicado: University of Oslo 2025
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=190250186&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 190250186
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        27041441
        NUGX
      jtl: Digital Humanities in the Nordic & Baltic Countries Publications (DHNB Publications)
      issn: 27041441
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 7
      iid: 2
      pid: 58757
      pub: University of Oslo
    artinfo:
      ui:
        190250186
        10.5617/dhnbpub.12299
      ppf: 1
      ppct: 9
      formats:
      tig:
        atl: Improving the Accuracy of Normalizing Historical Estonian Texts by Combining Statistical Machine Translation with BERT Language Model.
      aug:
        au: Jaanimäe, Gerth
        affil: University of Tartu
      su:
        Estonian language
        Language models
        Natural language processing
        Text processing (Computer science)
        Historical literature
        Machine translating
        Morphology (Grammar)
      sug:
        subj:
          Estonian language
          Language models
          Natural language processing
          Text processing (Computer science)
          Historical literature
          Machine translating
          Morphology (Grammar)
      keyword:
        historical texts
        machine learning
        normalization
      ab: Automatic analysis of historical texts is often hindered by different spelling system compared to the one used today. One approach to address this issue is to convert these texts to present-day spelling conventions, also called normalizing. One of the relatively old, however, still used methods for that is character level statistical machine translation (CSMT). As the method views the text one word at a time, it can lead to wrong normalizations due to lack of context. This paper gives an overview of integration of CSMT with BERT language models on texts written in Estonian, a morphologically rich language, to increase normalization accuracy.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      dt:
        @attributes:
          year: 2025
    holdings:
      @attributes:
        islocal: N