Comparative Analysis of LLMs' Performance On a Practice Radiography Certification Exam.

Purpose To compare the performance of multiple large language models (LLMs) on a practice radiography certification exam. Method Using an exploratory, nonexperimental approach, 200 multiple-choice question stems and options (correct answers and distractors) from a practice radiography certification...

Descripción completa

Detalles Bibliográficos
Publicado en:Radiologic Technology Vol. 96; no. 5; pp. 334 - 343
Autor principal: Clark, Kevin R.
Formato: research tables/charts Journal Article
Publicado: American Society of Radiologic Technologists May/Jun2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=184544181&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 184544181
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        00338397
        39J
      jtl: Radiologic Technology
      issn: 00338397
      maglogo: N
    pubinfo:
      dt: May/Jun2025
      vid: 96
      iid: 5
      pid: 7052
      pub: American Society of Radiologic Technologists
      place: Alburquerque, New Mexico
    artinfo:
      ui:
        184544181
        184544181
        184544181
        10.64912/radt.2025.jdyi8440
        184544181
      ppf: 334
      ppct: 9
      formats:
        fmt:
          @attributes:
            type: P
      tig:
        atl: Comparative Analysis of LLMs' Performance On a Practice Radiography Certification Exam.
      aug:
        au: Clark, Kevin R.
        affil: Associate professor and associate director for the School of Health Professions at The University of Texas MD Anderson Cancer Center in Houston
      sug:
        subj:
          Artificial Intelligence, Generative
          Credentialing Examinations
          Education, Radiologic Technology
          Human
          Comparative Studies
          Exploratory Research
          Nonexperimental Studies
          McNemar's Test
          Descriptive Statistics
          Data Analysis Software
          Chi Square Test
      ab: Purpose To compare the performance of multiple large language models (LLMs) on a practice radiography certification exam. Method Using an exploratory, nonexperimental approach, 200 multiple-choice question stems and options (correct answers and distractors) from a practice radiography certification exam were entered into 5 LLMs: ChatGPT (OpenAI), Claude (Anthropic), Copilot (Microsoft), Gemini (Google), and Perplexity (Perplexity AI). Responses were recorded as correct or incorrect, and overall accuracy rates were calculated for each LLM. McNemar tests determined if there were significant differences between accuracy rates. Performance also was evaluated and aggregated by content categories and subcategories. Results ChatGPT had the highest overall accuracy of 83.5%, followed by Perplexity (78.9%), Copilot (78.0%), Gemini (75.0%), and Claude (71.0%). ChatGPT had a significantly higher accuracy rate than did Claude (P < .001) and Gemini (P = .02). Regarding content categories, ChatGPT was the only LLM to correctly answer all 38 patient care questions. In addition, ChatGPT had the highest number of correct responses in the areas of safety (38/48, 79.2%) and procedures (50/59, 84.7%). Copilot had the highest number of correct responses in the area of image production (43/55, 78.2%). ChatGPT also achieved superior accuracy in 4 of the 8 subcategories. Discussion Findings from this study provide valuable insights into the performance of multiple LLMs in answering practice radiography certification exam questions. Although ChatGPT emerged as the most accurate LLM for this practice exam, caution should be exercised when using generative artificial intelligence (AI) models. Because LLMs can generate false and incorrect information, responses must be checked for accuracy, and the models should be corrected when inaccurate responses are given. Conclusion Among the 5 LLMs compared in this study, ChatGPT was the most accurate model. As interest in generative AI continues to increase and new language applications become readily available, users should understand the limitations of LLMs and check responses for accuracy. Future research could include additional practice exams in other primary pathways, including magnetic resonance imaging, nuclear medicine technology, radiation therapy, and sonography.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N