Challenging the norm: Length of exams determined by classification accuracy or reliability.

Purpose: This paper challenges the notion that reliability indices are appropriate for informing test length in exams in medical education, where the focus is on ensuring defensible pass‐fail decisions. Instead, we argue that using classification accuracy instead better suited to the purpose of exam...

Full description

Bibliographic Details
Published in:Medical Education Vol. 59; no. 12; pp. 1363 - 1375
Main Authors: Schauber, Stefan K., Homer, Matt
Format: research tables/charts Journal Article
Published: Wiley-Blackwell Dec2025
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=189913553&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 189913553
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        03080110
        ESF
      jtl: Medical Education
      issn: 03080110
      maglogo: Y
    pubinfo:
      dt: Dec2025
      vid: 59
      iid: 12
      pid: 480
      pub: Wiley-Blackwell
      place: Malden, Massachusetts
    artinfo:
      ui:
        189913553
        185677287
        189913553
        189913553
        10.1111/medu.15742
        189913553
      ppf: 1363
      ppct: 12
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: C
          – @attributes:
              type: P
      tig:
        atl: Challenging the norm: Length of exams determined by classification accuracy or reliability.
      aug:
        au:
          Schauber, Stefan K.
          Homer, Matt
        affil: Section for Health Sciences Education (HELP), Faculty of Medicine, University of Oslo,, Norway
      sug:
        subj:
          Education, Medical
          Achievement Tests
          Time
          Educational Measurement
          Classification
          Reliability
          Human
          Colleges and Universities Norway
          Norway
          Student Knowledge
          Students, Medical
          Coefficient alpha
          Triangulation
          Descriptive Statistics
          Data Analysis Software
      ab: Purpose: This paper challenges the notion that reliability indices are appropriate for informing test length in exams in medical education, where the focus is on ensuring defensible pass‐fail decisions. Instead, we argue that using classification accuracy instead better suited to the purpose of exams in these cases. We show empirically, using resampled test data from a range of undergraduate knowledge exams, that this is indeed the case. More specifically, we address the hypothesis that the use of classification accuracy results in recommending shorter test lengths as compared to when using reliability. Method: We analysed data from previous exams from both pre‐clinical and clinical phases of undergraduate medical education. We used a re‐sampling procedure in which both the cut‐score and test length of repeatedly generated synthetic exams were varied systematically. N = 52 500 datasets were generated from the original exams. For each of these both reliability and classification accuracy indices were estimated. Result: Results indicate that only classification accuracy, not reliability, varies in relation to the cut‐score for pass‐fail decisions. Furthermore, reliability and classification accuracy are differently related to test length. The optimal test length for using reliability was around 100 items, independent of pass‐rates. For classification accuracy, recommendations are less generic. For exams with a small percentage of failed decisions (i.e., 5% or less), an item size of 50 did, on average, achieve an accuracy of 95% correct classifications. Conclusions: We suggest a move towards the employment of classification accuracy using existing tools, whilst still using reliability as a complement. The benefits of re‐thinking current test design practice include minimizing the burden of assessment on candidates and test developers. Item writers could focus on developing fewer, but higher quality, items. Finally, we stress the need to consider the effects of the balance false positive and false negative decisions in pass/fail classifications. Schauber and Hauber question establishee dogma surrounding exam length by empirically examining its relation to both reliability and classification accuracy.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N