Overcoming statistical machine translation limitations: error analysis and proposed solutions for the Catalan-Spanish language pair.
This work aims to improve an N-gram-based statistical machine translation system between the Catalan and Spanish languages, trained with an aligned Spanish-Catalan parallel corpus consisting of 1.7 million sentences taken from El Periódico newspaper. Starting from a linguistic error analysis above t...
| Publicado en: | Language Resources & Evaluation Vol. 45; no. 2; pp. 181 - 209 |
|---|---|
| Autores principales: | , , , , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
May2011
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=60133438&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 60133438 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: May2011 vid: 45 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 60133438 10.1007/s10579-011-9137-0 ppf: 181 ppct: 28 formats: fmt: @attributes: type: P size: 387KB tig: atl: Overcoming statistical machine translation limitations: error analysis and proposed solutions for the Catalan-Spanish language pair. aug: au: Farrús, Mireia Costa-jussà, Marta Mariño, José Poch, Marc Hernández, Adolfo Henríquez, Carlos Fonollosa, José affil: TALP Research Center, Department of Signal Theory and Communications, Universitat Politècnica de Catalunya, C/Jordi Girona 1-3 08034 Barcelona Spain su: Machine translating Error analysis in foreign language education Catalan language Spanish language Corpora Morphology (Grammar) Orthography & spelling Semantics sug: subj: Machine translating Error analysis in foreign language education Catalan language Spanish language Corpora Morphology (Grammar) Orthography & spelling Semantics keyword: Grammatical categories Linguistic knowledge N-gram-based translation Statistical machine translation ab: This work aims to improve an N-gram-based statistical machine translation system between the Catalan and Spanish languages, trained with an aligned Spanish-Catalan parallel corpus consisting of 1.7 million sentences taken from El Periódico newspaper. Starting from a linguistic error analysis above this baseline system, orthographic, morphological, lexical, semantic and syntactic problems are approached using a set of techniques. The proposed solutions include the development and application of additional statistical techniques, text pre- and post-processing tasks, and rules based on the use of grammatical categories, as well as lexical categorization. The performance of the improved system is clearly increased, as is shown in both human and automatic evaluations of the system, with a gain of about 1.1 points BLEU observed in the Spanish-to-Catalan direction of translation, and a gain of about 0.5 points in the reverse direction. The final system is freely available online as a linguistic resource. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2011. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2011 holdings: @attributes: islocal: N |
|---|