Challenging the norm: Length of exams determined by classification accuracy or reliability.
Purpose: This paper challenges the notion that reliability indices are appropriate for informing test length in exams in medical education, where the focus is on ensuring defensible pass‐fail decisions. Instead, we argue that using classification accuracy instead better suited to the purpose of exam...
| Published in: | Medical Education Vol. 59; no. 12; pp. 1363 - 1375 |
|---|---|
| Main Authors: | , |
| Format: | research tables/charts Journal Article |
| Published: |
Wiley-Blackwell
Dec2025
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=189913553&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 189913553 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 03080110 ESF jtl: Medical Education issn: 03080110 maglogo: Y pubinfo: dt: Dec2025 vid: 59 iid: 12 pid: 480 pub: Wiley-Blackwell place: Malden, Massachusetts artinfo: ui: 189913553 185677287 189913553 189913553 10.1111/medu.15742 189913553 ppf: 1363 ppct: 12 formats: fmt: – @attributes: type: T – @attributes: type: C – @attributes: type: P tig: atl: Challenging the norm: Length of exams determined by classification accuracy or reliability. aug: au: Schauber, Stefan K. Homer, Matt affil: Section for Health Sciences Education (HELP), Faculty of Medicine, University of Oslo,, Norway sug: subj: Education, Medical Achievement Tests Time Educational Measurement Classification Reliability Human Colleges and Universities Norway Norway Student Knowledge Students, Medical Coefficient alpha Triangulation Descriptive Statistics Data Analysis Software ab: Purpose: This paper challenges the notion that reliability indices are appropriate for informing test length in exams in medical education, where the focus is on ensuring defensible pass‐fail decisions. Instead, we argue that using classification accuracy instead better suited to the purpose of exams in these cases. We show empirically, using resampled test data from a range of undergraduate knowledge exams, that this is indeed the case. More specifically, we address the hypothesis that the use of classification accuracy results in recommending shorter test lengths as compared to when using reliability. Method: We analysed data from previous exams from both pre‐clinical and clinical phases of undergraduate medical education. We used a re‐sampling procedure in which both the cut‐score and test length of repeatedly generated synthetic exams were varied systematically. N = 52 500 datasets were generated from the original exams. For each of these both reliability and classification accuracy indices were estimated. Result: Results indicate that only classification accuracy, not reliability, varies in relation to the cut‐score for pass‐fail decisions. Furthermore, reliability and classification accuracy are differently related to test length. The optimal test length for using reliability was around 100 items, independent of pass‐rates. For classification accuracy, recommendations are less generic. For exams with a small percentage of failed decisions (i.e., 5% or less), an item size of 50 did, on average, achieve an accuracy of 95% correct classifications. Conclusions: We suggest a move towards the employment of classification accuracy using existing tools, whilst still using reliability as a complement. The benefits of re‐thinking current test design practice include minimizing the burden of assessment on candidates and test developers. Item writers could focus on developing fewer, but higher quality, items. Finally, we stress the need to consider the effects of the balance false positive and false negative decisions in pass/fail classifications. Schauber and Hauber question establishee dogma surrounding exam length by empirically examining its relation to both reliability and classification accuracy. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|