Cross-linguistic authorship attribution and gender profiling. Machine translation as a method for bridging the language gap.
This study explores the feasibility of cross-linguistic authorship attribution and the author's gender identification using Machine Translation (MT). Computational stylistics experiments were conducted on a Greek blog corpus translated into English using Google's Neural MT. A Random Forest algorithm...
| Publicado en: | Digital Scholarship in the Humanities Vol. 39; no. 3; pp. 954 - 968 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Sep2024
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=179512336&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 179512336 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Sep2024 vid: 39 iid: 3 pid: 622 pub: Oxford University Press / USA artinfo: ui: 179512336 10.1093/llc/fqae028 ppf: 954 ppct: 14 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.1MB tig: atl: Cross-linguistic authorship attribution and gender profiling. Machine translation as a method for bridging the language gap. aug: au: Mikros, George Boumparis, Dimitris affil: Department of Middle Eastern Studies, Hamad Bin Khalifa University , Doha, Qatar University of Antwerp , Antwerp, Belgium su: Machine translating Random forest algorithms Attribution of authorship Linguistics Authorship sug: subj: Machine translating Random forest algorithms Attribution of authorship Linguistics Authorship keyword: author profiling Authors' authorship attribution lexical diversity Machine Translation Multilevel N-gram Profiles multilingual word embeddings ab: This study explores the feasibility of cross-linguistic authorship attribution and the author's gender identification using Machine Translation (MT). Computational stylistics experiments were conducted on a Greek blog corpus translated into English using Google's Neural MT. A Random Forest algorithm was employed for authorship and gender profiling, using different feature groups [Author's Multilevel N-gram Profiles, quantitative linguistics (QL), and cross-lingual word embeddings (CLWE)] in both original and translated texts. Results indicate that MT is a viable method for converting a multilingual corpus into one language for authorship attribution and gender profiling research, with considerable accuracy when training and testing datasets use identical language. In the pure cross-linguistic scenario, higher accuracies than the baselines were obtained using CLWE and QL features. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2024 holdings: @attributes: islocal: N |
|---|