Manipuri–English comparable corpus for cross-lingual studies.

This paper presents Mni-EnCC, a temporal alligned Manipuri–English comparable corpus, to facilitate cross-lingual studies between Manipuri and English. Mni-EnCC has been created by collating text from two publicly published news sources in internet namely Sangai Express and Poknapham in Manipur. Tho...

Descripción completa

Detalles Bibliográficos
Publicado en:Language Resources & Evaluation Vol. 57; no. 1; pp. 377 - 414
Autores principales: Laitonjam, Lenin, Singh, Sanasam Ranbir
Formato: Artículo
Publicado: Springer Nature Mar2023
Materias:
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:This paper presents Mni-EnCC, a temporal alligned Manipuri–English comparable corpus, to facilitate cross-lingual studies between Manipuri and English. Mni-EnCC has been created by collating text from two publicly published news sources in internet namely Sangai Express and Poknapham in Manipur. Though, both the publishers publish news in Manipuri and English editions, they are not the translation of each other. Almost all of the Manipuri editions are created using proprietary tools which generate texts in customized non-standard and non-unicode encodings. We develop tools to transform the non-unicode text into unicode text to generate the Manipuri corpus. We then verify and time aligned all the articles using a semi-automated process. Furthermore, the quality of the Mni-EnCC is evaluated using two premier cross-lingual studies: bilingual dictionary induction and machine translation. Experimental observations provide encouraging results making it as a suitable dataset for future cross-lingual studies on between Manipuri and English language pair. With an objective to promote cross-lingual studies in Manipuri–English, we also plan to release the corpus and supporting Unicode conversion tool.