Translator attribution for Arabic using machine learning.
Given a set of target language documents and their translators, the translator attribution task aims at identifying which translator translated which documents. The attribution and the identification of the translator's style could contribute to fields including translation studies, digital humaniti...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 2; pp. 658 - 667 |
|---|---|
| Autores principales: | , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=164367969&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 164367969 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2023 vid: 38 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 164367969 10.1093/llc/fqac054 ppf: 658 ppct: 9 formats: fmt: – @attributes: type: T – @attributes: type: P size: 368KB tig: atl: Translator attribution for Arabic using machine learning. aug: au: Mohamed, Emad Sarwar, Raheem Mostafa, Sayed affil: Research Group in Computational Linguistics, University of Wolverhampton , UK Department of Operations, Technology, Events and Hospitality Management, Manchester Metropolitan University , UK Department of Mathematics and Statistics, North Carolina A&T State University , USA su: Machine learning Translators Support vector machines Hierarchical clustering (Cluster analysis) Feature extraction Machine translating sug: subj: Machine learning Translators Support vector machines Hierarchical clustering (Cluster analysis) Feature extraction Machine translating ab: Given a set of target language documents and their translators, the translator attribution task aims at identifying which translator translated which documents. The attribution and the identification of the translator's style could contribute to fields including translation studies, digital humanities, and forensic linguistics. To conduct this investigation, firstly, we develop a new corpus containing the translations of world-famous books into Arabic. We then pre-process the books in our corpus which mainly involves cleaning irrelevant material, morphological segmentation analysis of words, and devocalization. After pre-processing the books, we propose to use 100 most frequent words and/or morphologically segmented function words as writing style markers of the translators (i.e. stylometric features) to differentiate between translations of different translators. After the completion of features extraction process, we applied several supervised and unsupervised machine-learning algorithms along with our novel cluster-to-author index to perform this task. We found that the translators are not invisible, and morphological analysis may not be more useful than just using the 100 most frequent words as features. The support vector machine linear kernel algorithm reported 99% classification accuracy. Similar findings were reported by the unsupervised machine-learning methods, namely, K -mean clustering and hierarchical clustering. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|