A large English–Thai parallel corpus from the web and machine-generated text.
The primary objective of our work is to build a large-scale English–Thai dataset for training neural machine translation models. We construct scb-mt-en-th-2020, an English–Thai machine translation dataset with over 1 million segment pairs, curated from various sources: news, Wikipedia articles, SMS...
| Publicado en: | Language Resources & Evaluation Vol. 56; no. 2; pp. 477 - 500 |
|---|---|
| Autores principales: | , , , |
| Formato: | Artículo |
| Publicado: |
Springer Nature
Jun2022
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=157410212&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 157410212 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 1574020X 179V jtl: Language Resources & Evaluation issn: 1574020X maglogo: N pubinfo: dt: Jun2022 vid: 56 iid: 2 pid: 237 pub: Springer Nature artinfo: ui: 157410212 10.1007/s10579-021-09536-6 ppf: 477 ppct: 23 formats: fmt: – @attributes: type: T – @attributes: type: P size: 500KB tig: atl: A large English–Thai parallel corpus from the web and machine-generated text. aug: au: Lowphansirikul, Lalita Polpanumas, Charin Rutherford, Attapol T. Nutanong, Sarana affil: School of Information Science and Technology, Vidyasirimedhi Institution of Science and Technology, Rayong, Thailand PyThaiNLP, Bangkok, Thailand Department of Linguistics, Chulalongkorn University, Bangkok, Thailand Teaching and Learning Thai as a Foreign Language Group, Bangkok, Thailand su: Wikipedia Google Inc. Machine translating Corpora Government publications Attribution of news Source code sug: subj: Wikipedia Google Inc. Machine translating Corpora Government publications Attribution of news Source code keyword: Machine translation Parallel corpus Pretraining Thai language Transformer ab: The primary objective of our work is to build a large-scale English–Thai dataset for training neural machine translation models. We construct scb-mt-en-th-2020, an English–Thai machine translation dataset with over 1 million segment pairs, curated from various sources: news, Wikipedia articles, SMS messages, task-based dialogs, web-crawled data, government documents, and text artificially generated by a pretrained language model. We present the methods for gathering data, aligning texts, and removing preprocessing noise and translation errors automatically. We also train machine translation models based on this dataset to assess the quality of the corpus. Our models perform comparably to Google Translation API (as of May 2020) for Thai–English and outperform Google when the Open Parallel Corpus (OPUS) is included in the training data for both Thai–English and English–Thai translation. The dataset is available for public use under CC-BY-SA 4.0 License. The pre-trained models and source code to reproduce our work are available under Apache-2.0 License. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: Language Resources & Evaluation is a copyright of Springer, 2022. All Rights Reserved. item: Language Resources & Evaluation holder: Springer Nature dt: @attributes: year: 2022 holdings: @attributes: islocal: N |
|---|