Automated classification of brain MRI reports using fine-tuned large language models.

Purpose: This study aimed to investigate the efficacy of fine-tuned large language models (LLM) in classifying brain MRI reports into pretreatment, posttreatment, and nontumor cases. Methods: This retrospective study included 759, 284, and 164 brain MRI reports for training, validation, and test dat...

Descripción completa

Detalles Bibliográficos
Publicado en:Neuroradiology Vol. 66; no. 12; pp. 2177 - 2184
Autores principales: Kanzawa, Jun, Yasaka, Koichiro, Fujita, Nana, Fujiwara, Shin, Abe, Osamu
Formato: research tables/charts Journal Article
Publicado: Springer Nature Dec2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=181252281&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 181252281
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        00283940
        NYZ
      jtl: Neuroradiology
      issn: 00283940
      maglogo: N
    pubinfo:
      dt: Dec2024
      vid: 66
      iid: 12
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        181252281
        178379096
        181252281
        181252281
        10.1007/s00234-024-03427-7
        181252281
      ppf: 2177
      ppct: 7
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Automated classification of brain MRI reports using fine-tuned large language models.
      aug:
        au:
          Kanzawa, Jun
          Yasaka, Koichiro
          Fujita, Nana
          Fujiwara, Shin
          Abe, Osamu
        affil: Department of Radiology, The University of Tokyo Hospital, Tokyo, Japan
      sug:
        subj:
          Natural Language Processing Utilization
          Brain Neoplasms Radiography
          Magnetic Resonance Imaging Methods
          Brain Neoplasms Classification
          Treatment Outcomes
          Time Factors
          Human
          Japan
          Male
          Female
          Adult
          Middle Age
          Aged
          Retrospective Design
          Record Review
          Models, Theoretical
          Radiologists
          Validity
          Confidence Intervals
          Sensitivity and Specificity
          ROC Curve
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Aged: 65+ years
          Male
          Female
      ab: Purpose: This study aimed to investigate the efficacy of fine-tuned large language models (LLM) in classifying brain MRI reports into pretreatment, posttreatment, and nontumor cases. Methods: This retrospective study included 759, 284, and 164 brain MRI reports for training, validation, and test dataset. Radiologists stratified the reports into three groups: nontumor (group 1), posttreatment tumor (group 2), and pretreatment tumor (group 3) cases. A pretrained Bidirectional Encoder Representations from Transformers Japanese model was fine-tuned using the training dataset and evaluated on the validation dataset. The model which demonstrated the highest accuracy on the validation dataset was selected as the final model. Two additional radiologists were involved in classifying reports in the test datasets for the three groups. The model's performance on test dataset was compared to that of two radiologists. Results: The fine-tuned LLM attained an overall accuracy of 0.970 (95% CI: 0.930–0.990). The model's sensitivity for group 1/2/3 was 1.000/0.864/0.978. The model's specificity for group1/2/3 was 0.991/0.993/0.958. No statistically significant differences were found in terms of accuracy, sensitivity, and specificity between the LLM and human readers (p ≥ 0.371). The LLM completed the classification task approximately 20–26-fold faster than the radiologists. The area under the receiver operating characteristic curve for discriminating groups 2 and 3 from group 1 was 0.994 (95% CI: 0.982–1.000) and for discriminating group 3 from groups 1 and 2 was 0.992 (95% CI: 0.982–1.000). Conclusion: Fine-tuned LLM demonstrated a comparable performance with radiologists in classifying brain MRI reports, while requiring substantially less time.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N