Personalized Impression Generation for PET Reports Using Large Language Models.
Large language models (LLMs) have shown promise in accelerating radiology reporting by summarizing clinical findings into impressions. However, automatic impression generation for whole-body PET reports presents unique challenges and has received little attention. Our study aimed to evaluate whether...
| Published in: | Journal of Digital Imaging Vol. 37; no. 2; pp. 471 - 489 |
|---|---|
| Main Authors: | , , , , , , , , , , |
| Format: | research tables/charts Journal Article |
| Published: |
Springer Nature
Apr2024
|
| Online Access: | View this record in EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=177626022&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 177626022 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 08971889 DOQ jtl: Journal of Digital Imaging issn: 08971889 maglogo: N pubinfo: dt: Apr2024 vid: 37 iid: 2 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 177626022 177626022 177626022 10.1007/s10278-024-00985-3 177626022 ppf: 471 ppct: 18 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Personalized Impression Generation for PET Reports Using Large Language Models. aug: au: Tie, Xin Shin, Muheon Pirasteh, Ali Ibrahim, Nevein Huemann, Zachary Castellino, Sharon M. Kelly, Kara M. Garrett, John Hu, Junjie Cho, Steve Y. Bradshaw, Tyler J. affil: Department of Radiology, School of Medicine and Public Health, University of Wissconsin, Madison, WI, USA sug: subj: Artificial Intelligence Positron-Emission Tomography Reports Image Processing, Computer Assisted Human Retrospective Design Algorithms Descriptive Statistics Physicians Funding Source ab: Large language models (LLMs) have shown promise in accelerating radiology reporting by summarizing clinical findings into impressions. However, automatic impression generation for whole-body PET reports presents unique challenges and has received little attention. Our study aimed to evaluate whether LLMs can create clinically useful impressions for PET reporting. To this end, we fine-tuned twelve open-source language models on a corpus of 37,370 retrospective PET reports collected from our institution. All models were trained using the teacher-forcing algorithm, with the report findings and patient information as input and the original clinical impressions as reference. An extra input token encoded the reading physician's identity, allowing models to learn physician-specific reporting styles. To compare the performances of different models, we computed various automatic evaluation metrics and benchmarked them against physician preferences, ultimately selecting PEGASUS as the top LLM. To evaluate its clinical utility, three nuclear medicine physicians assessed the PEGASUS-generated impressions and original clinical impressions across 6 quality dimensions (3-point scales) and an overall utility score (5-point scale). Each physician reviewed 12 of their own reports and 12 reports from other physicians. When physicians assessed LLM impressions generated in their own style, 89% were considered clinically acceptable, with a mean utility score of 4.08/5. On average, physicians rated these personalized impressions as comparable in overall utility to the impressions dictated by other physicians (4.03, P = 0.41). In summary, our study demonstrated that personalized impressions generated by PEGASUS were clinically useful in most cases, highlighting its potential to expedite PET reporting by automatically drafting impressions. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|