GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality.
Purpose: To compare the structural quality of multiple choice questions (MCQs) generated by a large language model, a type of artificial intelligence (AI), GPT-4, against human-authored items at both novice and expert level. Methods: We conducted a blinded analysis of 124 MCQs: 40 generated by GPT-4...
| Published in: | Medical Teacher Vol. 47; no. 12; pp. 1961 - 1975 |
|---|---|
| Main Authors: | , , , , , |
| Format: | research tables/charts Journal Article |
| Published: |
Taylor & Francis Ltd
Dec2025
|
| Online Access: | View this record in EBSCOhost |