Frequent Word Sequences and Statistical Stylistics.
This paper investigates the relative effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words and the frequencies of very frequent word sequences in distinguishing texts by different authors and grouping texts by a single author. C...
| Publicado en: | Literary & Linguistic Computing Vol. 17; no. 2; pp. 157 - 181 |
|---|---|
| Autor principal: | |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
Jun2002
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=44441549&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 44441549 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 02681145 BJ1 jtl: Literary & Linguistic Computing issn: 02681145 maglogo: N pubinfo: dt: Jun2002 vid: 17 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 44441549 10.1093/llc/17.2.157 ppf: 157 ppct: 24 formats: fmt: @attributes: type: P size: 2.3MB tig: atl: Frequent Word Sequences and Statistical Stylistics. aug: au: Hoover, David L. affil: New York University, New York, NY, USA su: Linguostylistics Multivariate analysis Decision making Cluster analysis (Statistics) Story plots Pronouns Literary terminology sug: subj: Linguostylistics Multivariate analysis Decision making Cluster analysis (Statistics) Story plots Pronouns Literary terminology ab: This paper investigates the relative effectiveness and accuracy of multivariate analysis, specifically cluster analysis, of the frequencies of very frequent words and the frequencies of very frequent word sequences in distinguishing texts by different authors and grouping texts by a single author. Cluster analyses based on frequent words are fairly accurate for groups of texts by known authors, whether the texts are long sections of modern British and US novels or shorter sections of contemporary literary critical texts, but they are only rarely completely accurate. When frequent word sequences are used instead of frequent words or in addition to them, however, the accuracy of the analyses often improves, sometimes dramatically, especially when personal pronouns are eliminated. Analyses based on frequent sequences even provide completely correct results in some cases where analyses based on frequent words fail. They also produce superior results for small groups of problematic novels and critical texts extracted from the larger corpora. Such successes suggest that analyses based on frequent word sequences constitute improved tools for authorship and stylistic studies. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Literary & Linguistic Computing holder: Oxford University Press / USA dt: @attributes: year: 2002 holdings: @attributes: islocal: N |
|---|