Use of a large language model with instruction‐tuning for reliable clinical frailty scoring.

Background: Frailty is an important predictor of health outcomes, characterized by increased vulnerability due to physiological decline. The Clinical Frailty Scale (CFS) is commonly used for frailty assessment but may be influenced by rater bias. Use of artificial intelligence (AI), particularly Lar...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of the American Geriatrics Society Vol. 72; no. 12; pp. 3849 - 3855
Autores principales: Kee, Xiang Lee Jamie, Sng, Gerald Gui Ren, Lim, Daniel Yan Zheng, Tung, Joshua Yi Min, Abdullah, Hairil Rizal, Chowdury, Anupama Roy
Formato: Artículo
Publicado: Wiley-Blackwell Dec2024
Materias:
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ssf&AN=181624015&site=ehost-live
header:
  @attributes:
    shortDbName: ssf
    uiTerm: 181624015
    longDbName: Social Sciences Full Text (H.W. Wilson)
    uiTag: AN
  controlInfo:
    bkinfo:
    jinfo:
      jid:
        00028614
        20Q
      jtl: Journal of the American Geriatrics Society
      issn: 00028614
      maglogo: Y
    pubinfo:
      dt: Dec2024
      vid: 72
      iid: 12
      pid: 480
      pub: Wiley-Blackwell
    artinfo:
      ui:
        181624015
        10.1111/jgs.19114
      ppf: 3849
      ppct: 6
      formats:
      tig:
        atl: Use of a large language model with instruction‐tuning for reliable clinical frailty scoring.
      aug:
        au:
          Kee, Xiang Lee Jamie
          Sng, Gerald Gui Ren
          Lim, Daniel Yan Zheng
          Tung, Joshua Yi Min
          Abdullah, Hairil Rizal
          Chowdury, Anupama Roy
        affil:
          Department of Geriatric Medicine, Singapore General Hospital, Singapore, Singapore
          Department of Endocrinology, Singapore General Hospital, Singapore, Singapore
          Data Science and Artificial Intelligence Laboratory, Singapore General Hospital, Singapore, Singapore
          Department of Gastroenterology, Singapore General Hospital, Singapore, Singapore
          Department of Urology, Singapore General Hospital, Singapore, Singapore
          Department of Anaesthesiology, Singapore General Hospital, Singapore, Singapore
      su:
        Language & languages
        Activities of daily living
        Frail elderly
        Natural language processing
        Mann Whitney U Test
        Descriptive statistics
        Geriatric assessment
        Mathematical models
        Theory
        Data analysis software
        Inter-observer reliability
      sug:
        subj:
          Language & languages
          Activities of daily living
          Frail elderly
          Natural language processing
          Mann Whitney U Test
          Descriptive statistics
          Geriatric assessment
          Mathematical models
          Theory
          Data analysis software
          Inter-observer reliability
      keyword:
        artificial intelligence
        frailty
        geriatrics
        artificial intelligence
        frailty
        geriatrics
      ab: Background: Frailty is an important predictor of health outcomes, characterized by increased vulnerability due to physiological decline. The Clinical Frailty Scale (CFS) is commonly used for frailty assessment but may be influenced by rater bias. Use of artificial intelligence (AI), particularly Large Language Models (LLMs) offers a promising method for efficient and reliable frailty scoring. Methods: The study utilized seven standardized patient scenarios to evaluate the consistency and reliability of CFS scoring by OpenAI's GPT‐3.5‐turbo model. Two methods were tested: a basic prompt and an instruction‐tuned prompt incorporating CFS definition, a directive for accurate responses, and temperature control. The outputs were compared using the Mann–Whitney U test and Fleiss' Kappa for inter‐rater reliability. The outputs were compared with historic human scores of the same scenarios. Results: The LLM's median scores were similar to human raters, with differences of no more than one point. Significant differences in score distributions were observed between the basic and instruction‐tuned prompts in five out of seven scenarios. The instruction‐tuned prompt showed high inter‐rater reliability (Fleiss' Kappa of 0.887) and produced consistent responses in all scenarios. Difficulty in scoring was noted in scenarios with less explicit information on activities of daily living (ADLs). Conclusions: This study demonstrates the potential of LLMs in consistently scoring clinical frailty with high reliability. It demonstrates that prompt engineering via instruction‐tuning can be a simple but effective approach for optimizing LLMs in healthcare applications. The LLM may overestimate frailty scores when less information about ADLs is provided, possibly as it is less subject to implicit assumptions and extrapolation than humans. Future research could explore the integration of LLMs in clinical research and frailty‐related outcome prediction.
      pubtype: Academic Journal
      doctype: Article
      src: R
    language: English
    refInfo:
    copyright:
      @attributes:
        flag: N
    holdings:
      @attributes:
        islocal: N