Bertalign: Improved word embedding-based sentence alignment for Chinese–English parallel corpora of literary texts.

Bertalign is designed to improve sentence alignment accuracy for Chinese–English parallel corpora of literary texts. Aligning bilingual literary texts is not trivial, since most of the translation is interpretative and not based on 1-to-1 mappings between source and target sentences. Existing alignm...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 38; no. 2; pp. 621 - 635
Autores principales: Liu, Lei, Zhu, Min
Formato: Artículo
Publicado: Oxford University Press / USA Jun2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=164367995&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 164367995
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Jun2023
      vid: 38
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        164367995
        10.1093/llc/fqac089
      ppf: 621
      ppct: 14
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 706KB
      tig:
        atl: Bertalign: Improved word embedding-based sentence alignment for Chinese–English parallel corpora of literary texts.
      aug:
        au:
          Liu, Lei
          Zhu, Min
        affil: School of Foreign Languages, Yanshan University , China
      su:
        Bible
        Corpora
        Vocabulary
        Translating & interpreting
        Algorithms
      sug:
        subj:
          Bible
          Corpora
          Vocabulary
          Translating & interpreting
          Algorithms
      ab: Bertalign is designed to improve sentence alignment accuracy for Chinese–English parallel corpora of literary texts. Aligning bilingual literary texts is not trivial, since most of the translation is interpretative and not based on 1-to-1 mappings between source and target sentences. Existing alignment methods highlight 1-to-1 links while having difficulty coping with 1-to-many and many-to-many alignments that are common in literary texts. To overcome the weaknesses of current approaches, we propose a novel two-step algorithm for bilingual sentence alignment. The first step finds the optimal paths for 1-to-1 alignments based on the top- k most semantically similar target sentences for each source sentence using the bidirectional encoder representations from transformer-based cross-lingual word embeddings. The second step relies on search paths found in the previous step to recover all valid alignments with more than one sentence on each side of the bilingual text. A comprehensive experiment was conducted on a newly built Chinese–English literary parallel corpus and a large-scale publicly available bilingual corpus of the Bible to compare the performance of Bertalign with five baseline systems: Gale-Church, Hunalign, Bleualign, Bleurtalign, and Vecalign. The results show that Bertalign achieves the highest accuracy in terms of F score on the two evaluation datasets than previous methods.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2023
    holdings:
      @attributes:
        islocal: N