Predicting Immunotherapy Response in Unresectable Hepatocellular Carcinoma: A Comparative Study of Large Language Models and Human Experts.

Hepatocellular carcinoma (HCC) is an aggressive cancer with limited biomarkers for predicting immunotherapy response. Recent advancements in large language models (LLMs) like GPT-4, GPT-4o, and Gemini offer the potential for enhancing clinical decision-making through multimodal data analysis. Howeve...

Full description

Bibliographic Details
Published in:Journal of Medical Systems Vol. 49; no. 1; pp. 1 - 16
Main Authors: Xu, Jun, Wang, Junjie, Li, Junjun, Zhu, Zhangxiang, Fu, Xiao, Cai, Wei, Song, Ruipeng, Wang, Tengfei, Li, Hai
Format: diagnostic images research tables/charts Journal Article
Published: Springer Nature 5/15/2025
Online Access:View this record in EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=185184877&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 185184877
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01485598
        4N0
      jtl: Journal of Medical Systems
      issn: 01485598
      maglogo: N
    pubinfo:
      dt: 5/15/2025
      vid: 49
      iid: 1
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        185184877
        185184877
        185184877
        10.1007/s10916-025-02192-1
        185184877
      ppf: 1
      ppct: 15
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Predicting Immunotherapy Response in Unresectable Hepatocellular Carcinoma: A Comparative Study of Large Language Models and Human Experts.
      aug:
        au:
          Xu, Jun
          Wang, Junjie
          Li, Junjun
          Zhu, Zhangxiang
          Fu, Xiao
          Cai, Wei
          Song, Ruipeng
          Wang, Tengfei
          Li, Hai
        affil: https://ror.org/034t30j35 Hefei Cancer Hospital of CAS, Institute of Health and Medical Technology, Hefei Institutes of Physical Science, Chinese Academy of Sciences, 230031, Hefei, P. R. China
      sug:
        subj:
          Carcinoma, Hepatocellular Drug Therapy
          Carcinoma, Hepatocellular Prognosis
          Immunotherapy
          Radiologists
          Oncologists
          Natural Language Processing
          Tomography, X-Ray Computed
          Treatment Outcomes
          Human
          Middle Age
          Aged
          Male
          Female
          Funding Source
          Comparative Studies
          Retrospective Design
          Record Review
          Sensitivity and Specificity
          Descriptive Statistics
          Physicians
          Early Intervention
          McNemar's Test
          Confidence Intervals
          Overall Survival
          T-Tests
          Data Analysis Software
          Chi Square Test
          Middle Aged: 45-64 years
          Aged: 65+ years
          Male
          Female
      ab: Hepatocellular carcinoma (HCC) is an aggressive cancer with limited biomarkers for predicting immunotherapy response. Recent advancements in large language models (LLMs) like GPT-4, GPT-4o, and Gemini offer the potential for enhancing clinical decision-making through multimodal data analysis. However, their effectiveness in predicting immunotherapy response, especially compared to human experts, remains unclear. This study assessed the performance of GPT-4, GPT-4o, and Gemini in predicting immunotherapy response in unresectable HCC, compared to radiologists and oncologists of varying expertise. A retrospective analysis of 186 patients with unresectable HCC utilized multimodal data (clinical and CT images). LLMs were evaluated with zero-shot prompting and two strategies: the 'voting method' and the 'OR rule method' for improved sensitivity. Performance metrics included accuracy, sensitivity, area under the curve (AUC), and agreement across LLMs and physicians.GPT-4o, using the 'OR rule method,' achieved 65% accuracy and 47% sensitivity, comparable to intermediate physicians but lower than senior physicians (accuracy: 72%, p = 0.045; sensitivity: 70%, p < 0.0001). Gemini-GPT, combining GPT-4, GPT-4o, and Gemini, achieved an AUC of 0.69, similar to senior physicians (AUC: 0.72, p = 0.35), with 68% accuracy, outperforming junior and intermediate physicians while remaining comparable to senior physicians (p = 0.78). However, its sensitivity (58%) was lower than senior physicians (p = 0.0097). LLMs demonstrated higher inter-model agreement (κ = 0.59–0.70) than inter-physician agreement, especially among junior physicians (κ = 0.15). This study highlights the potential of LLMs, particularly Gemini-GPT, as valuable tools in predicting immunotherapy response for HCC.
      pubtype: Academic Journal
      doctype:
        diagnostic images
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N