Statistical significance and its critics: practicing damaging science, or damaging scientific practice?

While the common procedure of statistical significance testing and its accompanying concept of p-values have long been surrounded by controversy, renewed concern has been triggered by the replication crisis in science. Many blame statistical significance tests themselves, and some regard them as suf...

Descripción completa

Detalles Bibliográficos
Publicado en:Synthese Vol. 200; no. 3; pp. 1 - 24
Autores principales: Mayo, Deborah G., Hand, David
Formato: Artículo
Publicado: Springer Nature Jun2022
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=hlh&AN=156853676&site=ehost-live
header:
  @attributes:
    shortDbName: hlh
    uiTerm: 156853676
    longDbName: Humanities International Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00397857
        4LI
      jtl: Synthese
      issn: 00397857
      maglogo: N
    pubinfo:
      dt: Jun2022
      vid: 200
      iid: 3
      pid: 237
      pub: Springer Nature
    artinfo:
      ui:
        156853676
        10.1007/s11229-022-03692-0
      ppf: 1
      ppct: 23
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
              size: 483KB
      tig:
        atl: Statistical significance and its critics: practicing damaging science, or damaging scientific practice?
      aug:
        au:
          Mayo, Deborah G.
          Hand, David
        affil:
          Virginia Tech, Blacksburg, USA
          Imperial College London, London, UK
      sug:
      keyword:
        Data-dredging
        Error probabilities
        Fisher
        Neyman and Pearson
        P-values
        Statistical significance tests
      ab: While the common procedure of statistical significance testing and its accompanying concept of p-values have long been surrounded by controversy, renewed concern has been triggered by the replication crisis in science. Many blame statistical significance tests themselves, and some regard them as sufficiently damaging to scientific practice as to warrant being abandoned. We take a contrary position, arguing that the central criticisms arise from misunderstanding and misusing the statistical tools, and that in fact the purported remedies themselves risk damaging science. We argue that banning the use of p-value thresholds in interpreting data does not diminish but rather exacerbates data-dredging and biasing selection effects. If an account cannot specify outcomes that will not be allowed to count as evidence for a claim—if all thresholds are abandoned—then there is no test of that claim. The contributions of this paper are: To explain the rival statistical philosophies underlying the ongoing controversy; To elucidate and reinterpret statistical significance tests, and explain how this reinterpretation ameliorates common misuses and misinterpretations; To argue why recent recommendations to replace, abandon, or retire statistical significance undermine a central function of statistics in science: to test whether observed patterns in the data are genuine or due to background variability.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: Y
      custom: Synthese is a copyright of Springer, 2022. All Rights Reserved.
      item: Synthese
      holder: Springer Nature
      dt:
        @attributes:
          year: 2022
    holdings:
      @attributes:
        islocal: N