Automated classification of brain MRI reports using fine-tuned large language models.
Purpose: This study aimed to investigate the efficacy of fine-tuned large language models (LLM) in classifying brain MRI reports into pretreatment, posttreatment, and nontumor cases. Methods: This retrospective study included 759, 284, and 164 brain MRI reports for training, validation, and test dat...
| Publicado en: | Neuroradiology Vol. 66; no. 12; pp. 2177 - 2184 |
|---|---|
| Autores principales: | , , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Springer Nature
Dec2024
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=181252281&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 181252281 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 00283940 NYZ jtl: Neuroradiology issn: 00283940 maglogo: N pubinfo: dt: Dec2024 vid: 66 iid: 12 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 181252281 178379096 181252281 181252281 10.1007/s00234-024-03427-7 181252281 ppf: 2177 ppct: 7 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Automated classification of brain MRI reports using fine-tuned large language models. aug: au: Kanzawa, Jun Yasaka, Koichiro Fujita, Nana Fujiwara, Shin Abe, Osamu affil: Department of Radiology, The University of Tokyo Hospital, Tokyo, Japan sug: subj: Natural Language Processing Utilization Brain Neoplasms Radiography Magnetic Resonance Imaging Methods Brain Neoplasms Classification Treatment Outcomes Time Factors Human Japan Male Female Adult Middle Age Aged Retrospective Design Record Review Models, Theoretical Radiologists Validity Confidence Intervals Sensitivity and Specificity ROC Curve Adult: 19-44 years Middle Aged: 45-64 years Aged: 65+ years Male Female ab: Purpose: This study aimed to investigate the efficacy of fine-tuned large language models (LLM) in classifying brain MRI reports into pretreatment, posttreatment, and nontumor cases. Methods: This retrospective study included 759, 284, and 164 brain MRI reports for training, validation, and test dataset. Radiologists stratified the reports into three groups: nontumor (group 1), posttreatment tumor (group 2), and pretreatment tumor (group 3) cases. A pretrained Bidirectional Encoder Representations from Transformers Japanese model was fine-tuned using the training dataset and evaluated on the validation dataset. The model which demonstrated the highest accuracy on the validation dataset was selected as the final model. Two additional radiologists were involved in classifying reports in the test datasets for the three groups. The model's performance on test dataset was compared to that of two radiologists. Results: The fine-tuned LLM attained an overall accuracy of 0.970 (95% CI: 0.930–0.990). The model's sensitivity for group 1/2/3 was 1.000/0.864/0.978. The model's specificity for group1/2/3 was 0.991/0.993/0.958. No statistically significant differences were found in terms of accuracy, sensitivity, and specificity between the LLM and human readers (p ≥ 0.371). The LLM completed the classification task approximately 20–26-fold faster than the radiologists. The area under the receiver operating characteristic curve for discriminating groups 2 and 3 from group 1 was 0.994 (95% CI: 0.982–1.000) and for discriminating group 3 from groups 1 and 2 was 0.992 (95% CI: 0.982–1.000). Conclusion: Fine-tuned LLM demonstrated a comparable performance with radiologists in classifying brain MRI reports, while requiring substantially less time. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|