GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality.

Purpose: To compare the structural quality of multiple choice questions (MCQs) generated by a large language model, a type of artificial intelligence (AI), GPT-4, against human-authored items at both novice and expert level. Methods: We conducted a blinded analysis of 124 MCQs: 40 generated by GPT-4...

Descripción completa

Detalles Bibliográficos
Publicado en:Medical Teacher Vol. 47; no. 12; pp. 1961 - 1975
Autores principales: Wu, Hannah, Zerner, Toby, Lee, Daniel, Court-Kowalski, Stefan, Devitt, Peter, Palmer, Edward
Formato: research tables/charts Journal Article
Publicado: Taylor & Francis Ltd Dec2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=190208038&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 190208038
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        0142159X
        MCH
      jtl: Medical Teacher
      issn: 0142159X
      maglogo: Y
    pubinfo:
      dt: Dec2025
      vid: 47
      iid: 12
      pid: 377
      pub: Taylor & Francis Ltd
      place: Philadelphia, Pennsylvania
    artinfo:
      ui:
        190208038
        185507552
        190208038
        190208038
        10.1080/0142159X.2025.2505122
        190208038
      ppf: 1961
      ppct: 14
      formats:
      tig:
        atl: GPT-4 versus human authors in clinically complex MCQ creation: A blinded analysis of item quality.
      aug:
        au:
          Wu, Hannah
          Zerner, Toby
          Lee, Daniel
          Court-Kowalski, Stefan
          Devitt, Peter
          Palmer, Edward
        affil: Adelaide Medical School, University of Adelaide, Adelaide, Australia
      sug:
        subj:
          Artificial Intelligence, Generative
          Educational Measurement
          Authors
          Education, Medical
          Human
          Writing
          Content Validity
          Clinical Reasoning
          Psychomotor Performance
          Multimethod Studies
          Analysis of Variance
          Post Hoc Analysis
      ab: Purpose: To compare the structural quality of multiple choice questions (MCQs) generated by a large language model, a type of artificial intelligence (AI), GPT-4, against human-authored items at both novice and expert level. Methods: We conducted a blinded analysis of 124 MCQs: 40 generated by GPT-4, 39 from human item-writers at Novice level, and 45 from human item-writers at Expert level. A generic prompt for GPT-4 was engineered, which included item-writing guidance, example MCQs, and key learning points. A standardized scoring system was developed including content validity, scope, item anatomy, cognitive skill level, item-writing flaws, feedback comprehensiveness, veracity and adequacy of clinical reasoning, and global impression of fitness for use. A consensus panel objectively evaluated each item, blinded to the author, using the scoring system. Results: Analysis showed that all groups (Novice, Expert, and AI) were able to generate items within scope. Expert items performed better than Novice items in all categories. There was no difference in the global impressions of Expert and AI items, which suggests overall comparability. A statistically significant, albeit small, difference was observed with Expert items performing better than AI items in the specific domains of content validity, feedback veracity and clinical reasoning, and testing at higher order cognitive skill levels. However, both groups met acceptable standards in these domains. AI items had a higher rate than Expert items of being deemed unfit for use requiring major revision, indicating erroneous correct answers, and displaying biased answer positioning. Conclusions: GPT-4 can produce MCQs testing clinically complex concepts for medical assessment. While the structural quality of AI-generated MCQs is comparable to experts overall, human oversight is necessary to ensure content validity and optimize item quality.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N