A Comparative Evaluation of Large Language Model Utility in Neuroimaging Clinical Decision Support.

Imaging utilization has increased dramatically in recent years, and at least some of these studies are not appropriate for the clinical scenario. The development of large language models (LLMs) may address this issue by providing a more accessible reference resource for ordering providers, but their...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Imaging Informatics in Medicine Vol. 38; no. 4; pp. 2294 - 2303
Autores principales: Miller, Luke, Kamel, Peter, Patel, Jigar, Agrawal, Jay, Zhan, Min, Bumbarger, Nathan, Wang, Kenneth
Formato: research tables/charts Journal Article
Publicado: Springer Nature Aug2025
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=187278944&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 187278944
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        29482925
        NR3A
      jtl: Journal of Imaging Informatics in Medicine
      issn: 29482925
      maglogo: N
    pubinfo:
      dt: Aug2025
      vid: 38
      iid: 4
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        187278944
        187278944
        187278944
        10.1007/s10278-024-01161-3
        187278944
      ppf: 2294
      ppct: 9
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: A Comparative Evaluation of Large Language Model Utility in Neuroimaging Clinical Decision Support.
      aug:
        au:
          Miller, Luke
          Kamel, Peter
          Patel, Jigar
          Agrawal, Jay
          Zhan, Min
          Bumbarger, Nathan
          Wang, Kenneth
        affil: https://ror.org/00sde4n60 Department of Radiology, University of Maryland Medical Center, Baltimore, MD, USA
      sug:
        subj:
          Natural Language Processing Utilization
          Neuroradiography
          Decision Support Systems, Clinical
          Artificial Intelligence
          Diagnostic Imaging
          Human
          Comparative Studies
          Tomography, X-Ray Computed
          Magnetic Resonance Imaging
          Medical Informatics
          Guideline Adherence
          Emergency Medical Services
          User-Computer Interface
          McNemar's Test
      ab: Imaging utilization has increased dramatically in recent years, and at least some of these studies are not appropriate for the clinical scenario. The development of large language models (LLMs) may address this issue by providing a more accessible reference resource for ordering providers, but their relative performance is currently understudied. Evaluate and compare the relative appropriateness and usefulness of imaging recommendations generated by eight publicly available models in response to neuroradiology clinical scenarios. Twenty-four common neuroradiology clinical scenarios were selected which often yield suboptimal imaging utilization. Questions were crafted to assess the ability of LLMs to provide accurate and actionable advice. The LLMs were assessed in August 2023 using natural-language 1–2 sentence queries requesting advice about optimal image ordering given certain clinical parameters. Eight of the most well-known LLMs were chosen for evaluation: ChatGPT, GPT4, Bard (Versions 1 and 2), Bing Chat, Llama 2, Perplexity, and Claude. The models were graded by three fellowship-trained neuroradiologists on whether their advice was "optimal" or "not optimal" according to the ACR Appropriateness Criteria or the New Orleans Head CT Criteria. The raters also ranked the models based on the appropriateness, helpfulness, concision, and source-citations in their response. The models varied in their ability to deliver an "optimal" recommendation based on these scenarios as follows: ChatGPT (20/24), GPT4 (23/24), Bard 1 (13/24), Bard 2 (14/24), Bing Chat (14/24), Llama (5/24), Perplexity (19/24), and Claude (19/24). The median ranks of the LLMs were as follows: ChatGPT (3), GPT4 (1.5), Bard 1 (4.5), Bard 2 (5), Bing Chat (6), Llama (7.5), Perplexity (4), and Claude (3). Characteristic errors are described and discussed. GPT-4, ChatGPT, and Claude generally outperformed Bard, Bing Chat, and Llama 2. This study evaluates the performance of a greater variety of publicly available LLMs in settings that more closely mimic real-world use cases as well as discussing the practical challenges of doing so. This is the first study to evaluate and compare a wide range of publicly available LLMs to determine appropriateness of their neuroradiology imaging recommendations.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N