Significance testing of word frequencies in corpora.

Finding out whether a word occurs significantly more often in one text or corpus than in another is an important question in analysing corpora. As noted by Kilgarriff (Language is never, ever, ever, random, Corpus Linguistics and Linguistic Theory, 2005; 1(2): 263-76.), the use of the X and log-like...

Descripción completa

Detalles Bibliográficos
Publicado en:Digital Scholarship in the Humanities Vol. 31; no. 2; pp. 374 - 398
Autores principales: Lijffijt, Jefrey, Nevalainen, Terttu, Säily, Tanja, Papapetrou, Panagiotis, Puolamäki, Kai, Mannila, Heikki
Formato: Artículo
Publicado: Oxford University Press / USA 6/1/2016
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=115711326&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 115711326
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        2055768X
        JEO9
      jtl: Digital Scholarship in the Humanities
      issn: 2055768X
      maglogo: N
    pubinfo:
      dt: 6/1/2016
      vid: 31
      iid: 2
      pid: 622
      pub: Oxford University Press / USA
    artinfo:
      ui:
        115711326
        10.1093/llc/fqu064
      ppf: 374
      ppct: 24
      formats:
        fmt:
          @attributes:
            type: P
            size: 1.3MB
      tig:
        atl: Significance testing of word frequencies in corpora.
      aug:
        au:
          Lijffijt, Jefrey
          Nevalainen, Terttu
          Säily, Tanja
          Papapetrou, Panagiotis
          Puolamäki, Kai
          Mannila, Heikki
        affil:
          Aalto University and University of Bristol
          University of Helsinki
          Aalto University and Stockholm University
          Aalto University and Finnish Institute of Occupational Health
          Aalto University
      su:
        Vocabulary
        Corpora
        Data analysis
        T-test (Statistics)
        Wilcoxon signed-rank test
      sug:
        subj:
          Vocabulary
          Corpora
          Data analysis
          T-test (Statistics)
          Wilcoxon signed-rank test
      ab: Finding out whether a word occurs significantly more often in one text or corpus than in another is an important question in analysing corpora. As noted by Kilgarriff (Language is never, ever, ever, random, Corpus Linguistics and Linguistic Theory, 2005; 1(2): 263-76.), the use of the X and log-likelihood ratio tests is problematic in this context, as they are based on the assumption that all samples are statistically independent of each other. However, words within a text are not independent. As pointed out in Kilgarriff (Comparing corpora, International Journal of Corpus Linguistics, 2001; 6(1): 1-37) and Paquot and Bestgen (Distinctive words in academic writing: a comparison of three statistical tests for keyword extraction. In Jucker, A., Schreier, D., and Hundt, M. (eds), Corpora: Pragmatics and Discourse. Amsterdam: Rodopi, 2009, pp. 247-69), it is possible to represent the data differently and employ other tests, such that we assume independence at the level of texts rather than individual words. This allows us to account for the distribution of words within a corpus. In this article we compare the significance estimates of various statistical tests in a controlled resampling experiment and in a practical setting, studying differences between texts produced by male and female fiction writers in the British National Corpus. We find that the choice of the test, and hence data representation, matters. We conclude that significance testing can be used to find consequential differences between corpora, but that assuming independence between all words may lead to overestimating the significance of the observed differences, especially for poorly dispersed words. We recommend the use of the t-test, Wilcoxon ranksum test, or bootstrap test for comparing word frequencies across corpora.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: © 2019 EADH: The European Association for Digital Humanities.
      item: Digital Scholarship in the Humanities
      holder: Oxford University Press / USA
      dt:
        @attributes:
          year: 2016
    holdings:
      @attributes:
        islocal: N