Evaluating the Performance of Large Language Models on Palliative Care Test Questions: A Mixed Methods Study.

Background: Little is known about large language model (LLM) performance on palliative care (PC)-related knowledge-based tasks. We evaluated two LLMs in answering PC-related test questions and explaining their answer choice rationale. Methods: LLMs were prompted to answer 25 randomly selected questi...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Palliative Medicine Vol. 29; no. 9; pp. 1282 - 1287
Autores principales: Chua, Isaac S., Lo, Yen-Ting, Liu, David, Succi, Marc D., Zhang, Mark, Yeh, Jonathan, Skarf, Lara M., Doyle, Kathleen, Gundersen, Daniel A., Mazzola, Emanuele, Bates, David W.
Formato: research tables/charts Journal Article
Publicado: Mary Ann Liebert, Inc. Sep2026
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=196039012&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 196039012
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        10966218
        C1U
      jtl: Journal of Palliative Medicine
      issn: 10966218
      maglogo: N
    pubinfo:
      dt: Sep2026
      vid: 29
      iid: 9
      pid: 1365
      pub: Mary Ann Liebert, Inc.
      place: New Rochelle, New York
    artinfo:
      ui:
        196039012
        196039012
        196039012
        10.1177/10966218261456817
        196039012
      ppf: 1282
      ppct: 5
      formats:
      tig:
        atl: Evaluating the Performance of Large Language Models on Palliative Care Test Questions: A Mixed Methods Study.
      aug:
        au:
          Chua, Isaac S.
          Lo, Yen-Ting
          Liu, David
          Succi, Marc D.
          Zhang, Mark
          Yeh, Jonathan
          Skarf, Lara M.
          Doyle, Kathleen
          Gundersen, Daniel A.
          Mazzola, Emanuele
          Bates, David W.
        affil: Department of Medicine, Division of General Internal Medicine and Primary Care, Brigham and Women's Hospital, Boston, Massachusetts, USA.
      sug:
        subj:
          Natural Language Processing Evaluation
          Palliative Care Education
          Questionnaires
          Education, Medical
          Human
          Multimethod Studies
          Random Sample
          Logistic Regression
          Data Quality
          Descriptive Statistics
          Thematic Analysis
          Linguistics
          Task Performance and Analysis
      ab: Background: Little is known about large language model (LLM) performance on palliative care (PC)-related knowledge-based tasks. We evaluated two LLMs in answering PC-related test questions and explaining their answer choice rationale. Methods: LLMs were prompted to answer 25 randomly selected questions from the Fast Facts Quiz and provide their answer choice rationale. Three PC educators ranked and rated LLM-generated answer choice explanations versus the test's answer key explanations. Linear fixed-effect models evaluated reviewer ranking, and ordinal logistic regression evaluated reviewer ratings of quality, suitability, accuracy, relevance, and comprehensiveness. Results: Both LLMs answered 96% of selected questions correctly. Reviewers rated LLM-generated explanations more highly than Fast Facts Quiz explanations. Five themes emerged from reviewer comments: perceived inaccuracies, clarity of writing, educational value, linguistic style, and miscellaneous. Conclusions: LLMs demonstrated high answer choice accuracy and generated preferable answer explanations when compared to the Fast Facts Quiz answer key.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N