Translator attribution of Hongloumeng: using entropy-based features and machining learning algorithm.
This study utilized machine learning algorithms and entropy-based features to identify translators of two English translations of Hongloumeng , a great classical Chinese novel written in the mid-18th century. The translations under examination were completed, respectively, by David Hawkes and the Ya...
| Publicado en: | Digital Scholarship in the Humanities Vol. 40; no. 1; pp. 138 - 151 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Apr2025
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=184296822&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 184296822 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Apr2025 vid: 40 iid: 1 pid: 622 pub: Oxford University Press / USA artinfo: ui: 184296822 10.1093/llc/fqae074 ppf: 138 ppct: 13 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.6MB tig: atl: Translator attribution of Hongloumeng: using entropy-based features and machining learning algorithm. aug: au: Hu, Ruitao Wang, Gui Shao, Bin affil: School of International Studies, Zhejiang University, Hangzhou, Zhejiang, 310058, P.R. China su: Fisher discriminant analysis Machine learning Feature extraction Support vector machines Automatic identification sug: subj: Fisher discriminant analysis Machine learning Feature extraction Support vector machines Automatic identification keyword: entropy Hongloumeng machine learning algorithm translator attribution ab: This study utilized machine learning algorithms and entropy-based features to identify translators of two English translations of Hongloumeng , a great classical Chinese novel written in the mid-18th century. The translations under examination were completed, respectively, by David Hawkes and the Yangs (Yang Hsien-yi and Gladys Yang). Two feature sets were extracted as input for the identification of translator styles: wordform features (wordform unigrams, bigrams, and trigrams) and part-of-speech (POS) features (POS unigrams, bigrams, and trigrams). Additionally, four machine learning classifiers were tested: linear support vector machines (SVMs), linear discriminant analysis (LDA), random forest (RF), and multilayer perceptron (MLP). Analysis of feature importance and SHAP value identified the most influential features within each classifier. Results showed that LDA achieved the best performance, with 81 per cent accuracy in distinguishing between translations, showing promise for translator identification. In contrast, MLP struggled to reliably differentiate between translations, achieving only 50 per cent accuracy. Furthermore, POS features had the greatest influence in SVM and LDA, while wordform features dominated in RF. SHAP analysis revealed that Hawkes' translation tended to exhibit higher POS unigram and lower POS trigram entropy compared to the Yangs'. This increased contribution of POS unigrams and trigrams suggests a link to explicitation differences in translation. In summary, the combination of machine learning and entropy-based stylometric features shows potential for automatic translator identification and analysis. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2025 holdings: @attributes: islocal: N |
|---|