Building an Oranian-English parallel corpus for automated translation training.
The main obstacle to automated translation and processing of dialects is their dearth of linguistic resources. The latter provide data to natural language processing professionals to conduct their experiments of dialect recognition, processing, and machine translation. This article highlights the ne...
| Publicado en: | Digital Scholarship in the Humanities Vol. 40; no. 1; pp. 87 - 96 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=184296836&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 184296836 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2025 vid: 40 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 184296836 10.1093/llc/fqae089 ppf: 87 ppct: 9 formats: fmt: – @attributes: type: T – @attributes: type: P size: 993KB tig: atl: Building an Oranian-English parallel corpus for automated translation training. aug: au: Dou, Abdelbasset Kissi, Khalida affil: DSPM Research Laboratory, Abdelhamid Ibn Badis University of Mostaganem, Mostaganem, 27000, Algeria Department of English, Higher Teacher Training School of Oran (ENSO), Oran, 31000, Algeria su: Low-resource languages Natural language processing Computational linguistics Programming languages Data augmentation Machine translating sug: subj: Low-resource languages Natural language processing Computational linguistics Programming languages Data augmentation Machine translating keyword: low-resource languages machine translation Oranian dialect parallel corpus ab: The main obstacle to automated translation and processing of dialects is their dearth of linguistic resources. The latter provide data to natural language processing professionals to conduct their experiments of dialect recognition, processing, and machine translation. This article highlights the need to resource the Algerian dialects, reviews the use of the available relevant corpora, and describes the process and distinctiveness of the first Oranian-English parallel corpus (OEPC). This is the first parallel corpus that includes one Algerian dialect with its English equivalents made from scratch. Particularly, this article presents the criteria and steps of compiling a monolingual corpus for the Oranian dialect (ORN) with references to data sources and formats. The size of the monolingual corpus ORN reached 8.5K sentences; with their equivalents in English, OEPC has been built. This significant linguistic resource is made under the Empowering and Resourcing Algerian Dialects project. This project is launched to enrich NLP experts with linguistic resources that are different Algerian mono-, bi-, multi-, and cross-dialectal corpora. The mechanism of data compilation and augmentation to extend the products of this project is explained. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|