Significance testing of word frequencies in corpora.
Finding out whether a word occurs significantly more often in one text or corpus than in another is an important question in analysing corpora. As noted by Kilgarriff (Language is never, ever, ever, random, Corpus Linguistics and Linguistic Theory, 2005; 1(2): 263-76.), the use of the X and log-like...
| Publicado en: | Digital Scholarship in the Humanities Vol. 31; no. 2; pp. 374 - 398 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | Artículo |
| Publicado: |
Oxford University Press / USA
6/1/2016
|
| Materias: | |
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=115711326&site=ehost-live header: @attributes: shortDbName: hlh uiTerm: 115711326 longDbName: Humanities International Complete uiTag: AN controlInfo: bkinfo: jinfo: jid: 2055768X JEO9 jtl: Digital Scholarship in the Humanities issn: 2055768X maglogo: N pubinfo: dt: 6/1/2016 vid: 31 iid: 2 pid: 622 pub: Oxford University Press / USA artinfo: ui: 115711326 10.1093/llc/fqu064 ppf: 374 ppct: 24 formats: fmt: @attributes: type: P size: 1.3MB tig: atl: Significance testing of word frequencies in corpora. aug: au: Lijffijt, Jefrey Nevalainen, Terttu Säily, Tanja Papapetrou, Panagiotis Puolamäki, Kai Mannila, Heikki affil: Aalto University and University of Bristol University of Helsinki Aalto University and Stockholm University Aalto University and Finnish Institute of Occupational Health Aalto University su: Vocabulary Corpora Data analysis T-test (Statistics) Wilcoxon signed-rank test sug: subj: Vocabulary Corpora Data analysis T-test (Statistics) Wilcoxon signed-rank test ab: Finding out whether a word occurs significantly more often in one text or corpus than in another is an important question in analysing corpora. As noted by Kilgarriff (Language is never, ever, ever, random, Corpus Linguistics and Linguistic Theory, 2005; 1(2): 263-76.), the use of the X and log-likelihood ratio tests is problematic in this context, as they are based on the assumption that all samples are statistically independent of each other. However, words within a text are not independent. As pointed out in Kilgarriff (Comparing corpora, International Journal of Corpus Linguistics, 2001; 6(1): 1-37) and Paquot and Bestgen (Distinctive words in academic writing: a comparison of three statistical tests for keyword extraction. In Jucker, A., Schreier, D., and Hundt, M. (eds), Corpora: Pragmatics and Discourse. Amsterdam: Rodopi, 2009, pp. 247-69), it is possible to represent the data differently and employ other tests, such that we assume independence at the level of texts rather than individual words. This allows us to account for the distribution of words within a corpus. In this article we compare the significance estimates of various statistical tests in a controlled resampling experiment and in a practical setting, studying differences between texts produced by male and female fiction writers in the British National Corpus. We find that the choice of the test, and hence data representation, matters. We conclude that significance testing can be used to find consequential differences between corpora, but that assuming independence between all words may lead to overestimating the significance of the observed differences, especially for poorly dispersed words. We recommend the use of the t-test, Wilcoxon ranksum test, or bootstrap test for comparing word frequencies across corpora. pubtype: Academic Journal doctype: Article src: R language: English refInfo: copyright: @attributes: flag: Y custom: © 2019 EADH: The European Association for Digital Humanities. item: Digital Scholarship in the Humanities holder: Oxford University Press / USA dt: @attributes: year: 2016 holdings: @attributes: islocal: N |
|---|