GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality.
Purpose: To compare the structural quality of multiple choice questions (MCQs) generated by a large language model, a type of artificial intelligence (AI), GPT-4, against human-authored items at both novice and expert level. Methods: We conducted a blinded analysis of 124 MCQs: 40 generated by GPT-4...
| Publicado en: | Medical Teacher Vol. 47; no. 12; pp. 1961 - 1975 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Taylor & Francis Ltd
Dec2025
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=190208038&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 190208038 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 0142159X MCH jtl: Medical Teacher issn: 0142159X maglogo: Y pubinfo: dt: Dec2025 vid: 47 iid: 12 pid: 377 pub: Taylor & Francis Ltd place: Philadelphia, Pennsylvania artinfo: ui: 190208038 185507552 190208038 190208038 10.1080/0142159X.2025.2505122 190208038 ppf: 1961 ppct: 14 formats: tig: atl: GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality. aug: au: Wu, Hannah Zerner, Toby Lee, Daniel Court-Kowalski, Stefan Devitt, Peter Palmer, Edward affil: Adelaide Medical School, University of Adelaide, Adelaide, Australia sug: subj: Artificial Intelligence, Generative Educational Measurement Authors Education, Medical Human Writing Content Validity Clinical Reasoning Psychomotor Performance Multimethod Studies Analysis of Variance Post Hoc Analysis ab: Purpose: To compare the structural quality of multiple choice questions (MCQs) generated by a large language model, a type of artificial intelligence (AI), GPT-4, against human-authored items at both novice and expert level. Methods: We conducted a blinded analysis of 124 MCQs: 40 generated by GPT-4, 39 from human item-writers at Novice level, and 45 from human item-writers at Expert level. A generic prompt for GPT-4 was engineered, which included item-writing guidance, example MCQs, and key learning points. A standardized scoring system was developed including content validity, scope, item anatomy, cognitive skill level, item-writing flaws, feedback comprehensiveness, veracity and adequacy of clinical reasoning, and global impression of fitness for use. A consensus panel objectively evaluated each item, blinded to the author, using the scoring system. Results: Analysis showed that all groups (Novice, Expert, and AI) were able to generate items within scope. Expert items performed better than Novice items in all categories. There was no difference in the global impressions of Expert and AI items, which suggests overall comparability. A statistically significant, albeit small, difference was observed with Expert items performing better than AI items in the specific domains of content validity, feedback veracity and clinical reasoning, and testing at higher order cognitive skill levels. However, both groups met acceptable standards in these domains. AI items had a higher rate than Expert items of being deemed unfit for use requiring major revision, indicating erroneous correct answers, and displaying biased answer positioning. Conclusions: GPT-4 can produce MCQs testing clinically complex concepts for medical assessment. While the structural quality of AI-generated MCQs is comparable to experts overall, human oversight is necessary to ensure content validity and optimize item quality. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|