Is word length inaccurate for authorship attribution?
Word length refers to a feature that is extracted from texts and used to characterize authorial style; it was quantitatively demonstrated by Mendenhall (Mendenhall, T. C. 1887, The characteristics curves of composition. Science , IX: 237–49). Many similar features for describing authorial style have...
| Publicado en: | Digital Scholarship in the Humanities Vol. 38; no. 2; pp. 875 - 891 |
|---|---|
| Autores principales: | , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2023
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=164367980&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 164367980 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: Jun2023 vid: 38 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 164367980 10.1093/llc/fqac067 ppf: 875 ppct: 16 formats: fmt: – @attributes: type: T – @attributes: type: P size: 1.3MB tig: atl: Is word length inaccurate for authorship attribution? aug: au: Zheng, Wanwan Jin, Mingzhe affil: Faculty of Engineering, Tokyo University of Science , 6-3-1 Niijuku, Katsushika-ku , Tokyo, Japan Faulty of Culture and Information Science, Doshisha University , 1-3 Tatara Miyakodani, Kyotanabe-shi , Kyoto-fu, Japan su: Attribution of authorship Vocabulary sug: subj: Attribution of authorship Vocabulary ab: Word length refers to a feature that is extracted from texts and used to characterize authorial style; it was quantitatively demonstrated by Mendenhall (Mendenhall, T. C. 1887, The characteristics curves of composition. Science , IX: 237–49). Many similar features for describing authorial style have been proposed; however, research indicates that compared with other features, word length identifies authors with lower accuracy. This study proposes a feature, referred to as c -wordL, to improve the accuracy of authorship attribution in texts through the classification of words into several types by following the part-of-speech (POS) tags and combining these types with the word length data. The proposed method was tested using 200 literary texts from ten different authors in Japanese, English, and Chinese. The results indicated that c -wordL was more accurate than the existing word length-based features and provided useful information that word unigrams and POS tag bigrams could not measure. In addition, the ease of interpretation of different types of features was discussed. In summary, c -wordL outperformed the existing superior features in explaining the distinct writing styles and identifying the authors. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2023 holdings: @attributes: islocal: N |
|---|