Identifying translationese at the word and sub-word level.

We use text classification to distinguish automatically between original and translated texts in Hebrew, a morphologically complex language. To this end, we design several linguistically informed feature sets that capture word-level and sub-word-level (in particular, morphological) properties of Heb...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 31; no. 1; pp. 30 - 55
Autores principales: Avner, Ehud Alexander, Ordan, Noam, Wintner, Shuly
Formato: Artículo
Publicado: Oxford University Press / USA 4/1/2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:We use text classification to distinguish automatically between original and translated texts in Hebrew, a morphologically complex language. To this end, we design several linguistically informed feature sets that capture word-level and sub-word-level (in particular, morphological) properties of Hebrew. Such features are abstract enough to allow for the development of accurate, robust classifiers, and they also lend themselves to linguistic interpretation. Careful evaluation shows that some of the classifiers we define are, indeed, highly accurate, and scale up nicely to domains that they were not trained on. In addition, analysis of the best features provides insight into the morphological properties of translated texts.