Boosting LLM-assisted diagnosis: 10-minute LLM tutorial elevates radiology residents' performance in brain MRI interpretation.

Purpose: To evaluate the impact of a structured tutorial on the use of a large language model (LLM)-based search engine on radiology residents' performance in brain MRI differential diagnosis. Methods: In this study, nine radiology residents determined the three most likely differential diagnoses fo...

Descripción completa

Detalles Bibliográficos
Publicado en:Neuroradiology Vol. 67; no. 8; pp. 2069 - 2082
Autores principales: Kim, Su Hwan, Schramm, Severin, Wihl, Jonas, Raffler, Philipp, Tahedl, Marlene, Canisius, Julian, Luiken, Ina, Endrös, Lukas, Reischl, Stefan, Marka, Alexander, Walter, Robert, Schillmaier, Mathias, Zimmer, Claus, Wiestler, Benedikt, Hedderich, Dennis Martin
Formato: pictorial research tables/charts randomized controlled trial Journal Article
Publicado: Springer Nature Aug2025
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:Purpose: To evaluate the impact of a structured tutorial on the use of a large language model (LLM)-based search engine on radiology residents' performance in brain MRI differential diagnosis. Methods: In this study, nine radiology residents determined the three most likely differential diagnoses for three sets of ten brain MRI cases with a challenging yet definite diagnosis. Each set was assessed (1) with the support of conventional internet search, (2) using an LLM-based search engine (© Perplexity AI) without prior tutorial, or (3) using the LLM-based search engine after a structured 10-minute tutorial. Reader responses were rated using a binary and numeric scoring system. Reading times and confidence levels (measured on a 5-point Likert scale) were recorded for each case. Search engine logs were examined to quantify user interaction metrics, and to identify hallucinations and misinterpretations in LLM responses. Results: Radiology residents achieved the highest accuracy when employing the LLM-based search engine following the tutorial, indicating the correct diagnosis among the top three differential diagnoses in 62.5% of cases (55/88). This was followed by the LLM-assisted workflow before the tutorial (44.8%; 39/87) and the conventional internet search workflow (32.2%; 28/87). The LLM tutorial led to significantly higher performance (binary scores: p = 0.042, numeric scores: p = 0.016) and confidence (p = 0.006) but resulted in no relevant differences in reading times. Hallucinations were found in 5.1% of LLM queries. Conclusion: Our findings demonstrate the considerable benefits that even low-effort educational interventions on LLMs can provide, highlighting their potential role in radiology training programs.