Factor-based evaluation for English to Hindi MT outputs.
Design and implementation of automatic evaluation methods is an integral part of any scientific research in accelerating the development cycle of the output. This is no less true for automatic machine translation (MT) systems. However, no such global and systematic scheme exists for evaluation of pe...
| Publicado en: | Language Resources & Evaluation Vol. 52; no. 4; pp. 969 - 997 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Dec2018
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=132695162&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 132695162 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Dec2018 vid: 52 iid: 4 pid: 237 pub: Springer Nature artinfo: ui: 132695162 10.1007/s10579-018-9426-y ppf: 969 ppct: 28 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.3MB tig: atl: Factor-based evaluation for English to Hindi MT outputs. aug: au: Balyan, Renu Chatterjee, Niladri affil: Arizona State University, Tempe, AZ, USA Indian Institute of Technology Delhi, New Delhi, India su: Scientific method Machine translating Analysis of variance Hindi language Regression analysis sug: subj: Scientific method Machine translating Analysis of variance Hindi language Regression analysis keyword: Ensemble techniques Error identification and classification Linear regression Machine translation evaluation Two-way ANOVA ab: Design and implementation of automatic evaluation methods is an integral part of any scientific research in accelerating the development cycle of the output. This is no less true for automatic machine translation (MT) systems. However, no such global and systematic scheme exists for evaluation of performance of an MT system. The existing evaluation metrics, such as BLEU, METEOR, TER, although used extensively in literature have faced a lot of criticism from users. Moreover, performance of these metrics often varies with the pair of languages under consideration. The above observation is no less pertinent with respect to translations involving languages of the Indian subcontinent. This study aims at developing an evaluation metric for English to Hindi MT outputs. As a part of this process, a set of probable errors have been identified manually as well as automatically. Linear regression has been used for computing weight/penalty for each error, while taking human evaluations into consideration. A sentence score is computed as the weighted sum of the errors. A set of 126 models has been built using different single classifiers and ensemble of classifiers in order to find the most suitable model for allocating appropriate weight/penalty for each error. The outputs of the models have been compared with the state-of-the-art evaluation metrics. The models developed for manually identified errors correlate well with manual evaluation scores, whereas the models for the automatically identified errors have low correlation with the manual scores. This indicates the need for further improvement and development of sophisticated linguistic tools for automatic identification and extraction of errors. Although many automatic machine translation tools are being developed for many different language pairs, there is no such generalized scheme that would lead to designing meaningful metrics for their evaluation. The proposed scheme should help in developing such metrics for different language pairs in the coming days. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2018. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2018 holdings: @attributes: islocal: N |
|---|