Not by chance. Russian aspect in rule-based machine translation.

The aim of this paper is twofold: it illustrates the benefits of rule-based instead of statistical machine translation, and it provides a starting point for the machine translation of the Russian aspect into English. Rule-based machine translation is still promising, from both a computational and th...

Descripción completa

Detalles Bibliográficos
Publicado en:Russian Linguistics Vol. 40; no. 3; pp. 199 - 214
Autores principales: Sonnenhauser, Barbara, Zangenfeind, Robert
Formato: Artículo
Publicado: Springer Nature Nov2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:The aim of this paper is twofold: it illustrates the benefits of rule-based instead of statistical machine translation, and it provides a starting point for the machine translation of the Russian aspect into English. Rule-based machine translation is still promising, from both a computational and theoretical point of view, because by implementing rules on the computer theoretical assumptions concerning linguistic structures can be verified and improved. This will be shown using the example of the category of aspect, which is one of the main challenges for machine translation from Russian to English. A small corpus study on the translation of Russian sentences with verbs in the past tense (perfective and imperfective) by human translators shows that three-quarters of Russian verbs (both imperfective and perfective) are translated by English simple past forms. While this results from language internal markedness relations, the translation of the remaining 25 % requires an in-depth analysis of the various interpretations possible for the Russian aspect. We propose a semantic analysis based on which rules for the interpretation and translation of Russian aspect in a machine translation system can be derived. Their implementation in the machine translation system ĖTAP is shown in this paper using two test cases as examples.
Цель этой статьи двояка: она иллюстрирует пользу машинного перевода на основе правил по сравнению с машинным переводом на основе статистики и предлагает отправной пункт для машинного перевода русского вида глагола на английский язык. Машинный перевод на основе правил всё ещё имеет свои выгоды, и с вычислительной, и с теоретической точки зрения, поскольку, применив правила на компьютере, теоретические гипотезы, касающиеся лингвистических структур, будут проверены и улучшены. Мы это покажем на примере вида глагола, который является одной из главных сложностей для машинного перевода с русского на английский язык. Исследуя часть параллельного корпуса русского национального корпуса, мы изучаем, как русские предложения с глаголами в прошедшем времени переводятся на английский язык переводчиками-людьми. Эти исследования показывают, что три четверти русских глаголов (как несовершенного, так и совершенного вида) этого корпуса переводятся английскими формами past simple (претерит). В то время как это представляет собой следствие внутренних языковых отношений маркированности, перевод остальных 25 % требует глубокого анализа различных возможностей интерпретации русского аспекта. На основе семантического анализа, который мы предложим, можно получить правила для трактовки и перевода русского аспекта в системе машинного перевода. Их применение в системе машинного перевода (в этом случае ЭТАП) продемонстрировано в данной статье на двух примерах.