Using syntax for improving phrase-based SMT in low-resource languages.

Data driven approaches for machine translation, such as statistical and neural machine translation, suffer from sparsity when dealing with low-resource languages. In these cases, using other sources of information including linguistic information could alleviate the problem. In this article, we focu...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 35; no. 3; pp. 507 - 529
Autores principales: Fadaei, Hakimeh, Faili, Heshaam
Formato: Artículo
Publicado: Oxford University Press / USA Sep2020
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=146172308&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 146172308
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: Sep2020
      vid: 35
      iid: 3
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        146172308
        10.1093/llc/fqz033
      ppf: 507
      ppct: 22
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 971KB
      tig:
        atl: Using syntax for improving phrase-based SMT in low-resource languages.
      aug:
        au:
          Fadaei, Hakimeh
          Faili, Heshaam
        affil:
          School of Electrical and Computer Engineering , College of Engineering, University of Tehran, Tehran, Iran
          School of Electrical and Computer Engineering , College of Engineering, University of Tehran, Tehran, Iran School of Computer Science, Institute for Research in Fundamental Sciences (IPM), Tehran, Iran
      su:
        Language & languages
        Information resources
        Translations
        Grammar
      sug:
        subj:
          Language & languages
          Information resources
          Translations
          Grammar
      ab: Data driven approaches for machine translation, such as statistical and neural machine translation, suffer from sparsity when dealing with low-resource languages. In these cases, using other sources of information including linguistic information could alleviate the problem. In this article, we focus on the problem of word ordering in translation from a high-resource to a low-resource language and try to improve the quality by using syntactic information from the high-resource side. We propose some syntactic features based on Tree Adjoining Grammar (TAG) to be employed in a phrase-based SMT model in order to improve the word ordering. In this work, a set of synchronous TAG rules is extracted and used to estimate the probability of the phrase orders suggested by the phrase-based model. The main idea of the article is to handle the word ordering by using the extended domain of locality property of TAG and abstracting the long distance dependencies into a local view, which is a TAG elementary tree. The experiments on English–Persian and English–German translation showed that, by combining the proposed TAG-based reordering features with lexical and hierarchical reordering models, we gain significant improvements over the baseline and in comparison with a neural reordering model and a pre-reordering model.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2020
    holdings:
      @attributes:
        islocal: N