Recursive alignment block classification technique for word reordering in statistical machine translation.

Statistical machine translation (SMT) is based on alignment models which learn from bilingual corpora the word correspondences between source and target language. These models are assumed to be capable of learning reorderings. However, the difference in word order between two languages is one of the...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 45; no. 2; pp. 165 - 180
Autores principales: Costa-jussà, Marta, Fonollosa, José, Monte, Enric
Formato: Artículo
Publicado: Springer Nature May2011
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=60133439&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 60133439
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        1574020X
        179V
      jtl: Language Resources & Evaluation
      issn: 1574020X
      maglogo: N
    pubinfo:
      dt: May2011
      vid: 45
      iid: 2
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        60133439
        10.1007/s10579-010-9133-9
      ppf: 165
      ppct: 15
      formats:
        fmt:
          @attributes:
            type: P
            size: 677KB
      tig:
        atl: Recursive alignment block classification technique for word reordering in statistical machine translation.
      aug:
        au:
          Costa-jussà, Marta
          Fonollosa, José
          Monte, Enric
        affil:
          Barcelona Media Innovation Center, Av. Diagonal 177 08018 Barcelona Spain
          Universitat Politècnica de Catalunya, TALP Research Center, Jordi Girona 1-3 08034 Barcelona Spain
      su:
        Machine translating
        Bilingualism
        Corpora
        Vocabulary
        Error analysis in foreign language education
        Learning
        Experiments
      sug:
        subj:
          Machine translating
          Bilingualism
          Corpora
          Vocabulary
          Error analysis in foreign language education
          Learning
          Experiments
      keyword:
        Automatic evaluation
        Statistical classification
        Statistical machine translation
        Word reordering
      ab: Statistical machine translation (SMT) is based on alignment models which learn from bilingual corpora the word correspondences between source and target language. These models are assumed to be capable of learning reorderings. However, the difference in word order between two languages is one of the most important sources of errors in SMT. In this paper, we show that SMT can take advantage of inductive learning in order to solve reordering problems. Given a word alignment, we identify those pairs of consecutive source blocks (sequences of words) whose translation is swapped, i.e. those blocks which, if swapped, generate a correct monotonic translation. Afterwards, we classify these pairs into groups, following recursively a co-occurrence block criterion, in order to infer reorderings. Inside the same group, we allow new internal combination in order to generalize the reorder to unseen pairs of blocks. Then, we identify the pairs of blocks in the source corpora (both training and test) which belong to the same group. We swap them and we use the modified source training corpora to realign and to build the final translation system. We have evaluated our reordering approach both in alignment and translation quality. In addition, we have used two state-of-the-art SMT systems: a Phrased-based and an Ngram-based. Experiments are reported on the EuroParl task, showing improvements almost over 1 point in the standard MT evaluation metrics (mWER and BLEU).
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved.
      item: Language Resources & Evaluation
      holder: Springer Nature
      dt:
        @attributes:
          year: 2011
    holdings:
      @attributes:
        islocal: N