A large English–Thai parallel corpus from the web and machine-generated text.
The primary objective of our work is to build a large-scale English–Thai dataset for training neural machine translation models. We construct scb-mt-en-th-2020, an English–Thai machine translation dataset with over 1 million segment pairs, curated from various sources: news, Wikipedia articles, SMS...
| Published in: | Language Resources & Evaluation Vol. 56; no. 2; pp. 477 - 500 |
|---|---|
| Main Authors: | , , , |
| Format: | Article |
| Published: |
Springer Nature
Jun2022
|
| Subjects: | |
| Online Access: | View this record in EBSCOhost |