Performance of large language models in medical licensing examinations: a systematic review and meta-analysis.
Purpose: This study systematically evaluates and compares the performance of large language models (LLMs) in answering medical licensing examination questions. By conducting subgroup analyses based on language, question format, and model type, this meta-analysis aims to provide a comprehensive overv...
| Publicado en: | Journal of Educational Evaluation for Health Professions Vol. 22; pp. 1 - 14 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | pictorial research systematic review tables/charts Journal Article |
| Publicado: |
National Health Personnel Licensing Examination Board
2025
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=192170971&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 192170971 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 19755937 B0CC jtl: Journal of Educational Evaluation for Health Professions issn: 19755937 maglogo: N pubinfo: dt: 2025 vid: 22 pid: 58691 pub: National Health Personnel Licensing Examination Board artinfo: ui: 192170971 192170971 192170971 10.3352/jeehp.2025.22.36 192170971 ppf: 1 ppct: 13 formats: fmt: @attributes: type: P tig: atl: Performance of large language models in medical licensing examinations: a systematic review and meta-analysis. aug: au: Nouri, Haniyeh Mahdavi, Abdollah Abedi, Ali Mohammadnia, Alireza Hamedan, Mahnaz Amanzadeh, Masoud affil: Student Research Committee, School of Medicine, Ardabil University of Medical Sciences, Ardabil, Iran sug: subj: Natural Language Processing Evaluation Education, Medical State Board Examinations Decision Making, Clinical Human Systematic Review Meta Analysis Medline PubMed World Wide Web Confidence Intervals Data Analysis Software Descriptive Statistics ab: Purpose: This study systematically evaluates and compares the performance of large language models (LLMs) in answering medical licensing examination questions. By conducting subgroup analyses based on language, question format, and model type, this meta-analysis aims to provide a comprehensive overview of LLM capabilities in medical education and clinical decision-making. Methods: This systematic review, registered in PROSPERO and following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, searched MEDLINE (PubMed), Scopus, and Web of Science for relevant articles published up to February 1, 2025. The search strategy included Medical Subject Headings (MeSH) terms and keywords related to ("ChatGPT" OR "GPT" OR "LLM variants") AND ("medical licensing exam*" OR "medical exam*" OR "medical education" OR "radiology exam*"). Eligible studies evaluated LLM accuracy on medical licensing examination questions. Pooled accuracy was estimated using a random-effects model, with subgroup analyses by LLM type, language, and question format. Publication bias was assessed using Egger's regression test. Results: This systematic review identified 2,404 studies. After removing duplicates and excluding irrelevant articles through title and abstract screening, 36 studies were included after full-text review. The pooled accuracy was 72% (95% confidence interval, 70.0% to 75.0%) with high heterogeneity (I2=99%, P<0.001). Among LLMs, GPT-4 achieved the highest accuracy (81%), followed by Bing (79%), Claude (74%), Gemini/Bard (70%), and GPT-3.5 (60%) (P=0.001). Performance differences across languages (range, 62% in Polish to 77% in German) were not statistically significant (P=0.170). Conclusion: LLMs, particularly GPT-4, can match or exceed medical students' examination performance and may serve as supportive educational tools. However, due to variability and the risk of errors, they should be used cautiously as complements rather than replacements for traditional learning methods. pubtype: Academic Journal doctype: meta analysis pictorial research systematic review tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|