Performance of large language models in medical licensing examinations: a systematic review and meta-analysis.

Purpose: This study systematically evaluates and compares the performance of large language models (LLMs) in answering medical licensing examination questions. By conducting subgroup analyses based on language, question format, and model type, this meta-analysis aims to provide a comprehensive overv...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Educational Evaluation for Health Professions Vol. 22; pp. 1 - 14
Autores principales: Nouri, Haniyeh, Mahdavi, Abdollah, Abedi, Ali, Mohammadnia, Alireza, Hamedan, Mahnaz, Amanzadeh, Masoud
Formato: meta analysis pictorial research systematic review tables/charts Journal Article
Publicado: National Health Personnel Licensing Examination Board 2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=192170971&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 192170971
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        19755937
        B0CC
      jtl: Journal of Educational Evaluation for Health Professions
      issn: 19755937
      maglogo: N
    pubinfo:
      dt: 2025
      vid: 22
      pid: 58691
      pub: National Health Personnel Licensing Examination Board
    artinfo:
      ui:
        192170971
        192170971
        192170971
        10.3352/jeehp.2025.22.36
        192170971
      ppf: 1
      ppct: 13
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Performance of large language models in medical licensing examinations: a systematic review and meta-analysis.
      aug:
        au:
          Nouri, Haniyeh
          Mahdavi, Abdollah
          Abedi, Ali
          Mohammadnia, Alireza
          Hamedan, Mahnaz
          Amanzadeh, Masoud
        affil: Student Research Committee, School of Medicine, Ardabil University of Medical Sciences, Ardabil, Iran
      sug:
        subj:
          Natural Language Processing Evaluation
          Education, Medical
          State Board Examinations
          Decision Making, Clinical
          Human
          Systematic Review
          Meta Analysis
          Medline
          PubMed
          World Wide Web
          Confidence Intervals
          Data Analysis Software
          Descriptive Statistics
      ab: Purpose: This study systematically evaluates and compares the performance of large language models (LLMs) in answering medical licensing examination questions. By conducting subgroup analyses based on language, question format, and model type, this meta-analysis aims to provide a comprehensive overview of LLM capabilities in medical education and clinical decision-making. Methods: This systematic review, registered in PROSPERO and following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, searched MEDLINE (PubMed), Scopus, and Web of Science for relevant articles published up to February 1, 2025. The search strategy included Medical Subject Headings (MeSH) terms and keywords related to ("ChatGPT" OR "GPT" OR "LLM variants") AND ("medical licensing exam*" OR "medical exam*" OR "medical education" OR "radiology exam*"). Eligible studies evaluated LLM accuracy on medical licensing examination questions. Pooled accuracy was estimated using a random-effects model, with subgroup analyses by LLM type, language, and question format. Publication bias was assessed using Egger's regression test. Results: This systematic review identified 2,404 studies. After removing duplicates and excluding irrelevant articles through title and abstract screening, 36 studies were included after full-text review. The pooled accuracy was 72% (95% confidence interval, 70.0% to 75.0%) with high heterogeneity (I2=99%, P<0.001). Among LLMs, GPT-4 achieved the highest accuracy (81%), followed by Bing (79%), Claude (74%), Gemini/Bard (70%), and GPT-3.5 (60%) (P=0.001). Performance differences across languages (range, 62% in Polish to 77% in German) were not statistically significant (P=0.170). Conclusion: LLMs, particularly GPT-4, can match or exceed medical students' examination performance and may serve as supportive educational tools. However, due to variability and the risk of errors, they should be used cautiously as complements rather than replacements for traditional learning methods.
      pubtype: Academic Journal
      doctype:
        meta analysis
        pictorial
        research
        systematic review
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N