GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality.

Purpose: To compare the structural quality of multiple choice questions (MCQs) generated by a large language model, a type of artificial intelligence (AI), GPT-4, against human-authored items at both novice and expert level. Methods: We conducted a blinded analysis of 124 MCQs: 40 generated by GPT-4...

Full description

Bibliographic Details
Published in:Medical Teacher Vol. 47; no. 12; pp. 1961 - 1975
Main Authors: Wu, Hannah, Zerner, Toby, Lee, Daniel, Court-Kowalski, Stefan, Devitt, Peter, Palmer, Edward
Format: research tables/charts Journal Article
Published: Taylor & Francis Ltd Dec2025
Online Access:View this record in EBSCOhost