A Novel Evaluation Framework for Medical LLMs: Combining Fuzzy Logic and MCDM for Medical Relation and Clinical Concept Extraction.

Artificial intelligence (AI) has become a crucial element of modern technology, especially in the healthcare sector, which is apparent given the continuous development of large language models (LLMs), which are utilized in various domains, including medical beings. However, when it comes to using th...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Medical Systems Vol. 48; no. 1; pp. 1 - 13
Autores principales: Alamoodi, A. H., Zughoul, Omar, David, Dianese, Garfan, Salem, Pamucar, Dragan, Albahri, O. S., Albahri, A. S., Yussof, Salman, Sharaf, Iman Mohamad
Formato: equations & formulas research tables/charts Journal Article
Publicado: Springer Nature 8/31/2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=179604257&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 179604257
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01485598
        4N0
      jtl: Journal of Medical Systems
      issn: 01485598
      maglogo: N
    pubinfo:
      dt: 8/31/2024
      vid: 48
      iid: 1
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        179604257
        179604257
        179604257
        10.1007/s10916-024-02090-y
        179604257
      ppf: 1
      ppct: 12
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: A Novel Evaluation Framework for Medical LLMs: Combining Fuzzy Logic and MCDM for Medical Relation and Clinical Concept Extraction.
      aug:
        au:
          Alamoodi, A. H.
          Zughoul, Omar
          David, Dianese
          Garfan, Salem
          Pamucar, Dragan
          Albahri, O. S.
          Albahri, A. S.
          Yussof, Salman
          Sharaf, Iman Mohamad
        affil: https://ror.org/03kxdn807 Institute of Informatics and Computing in Energy, Universiti Tenaga Nasional, Kajang, Malaysia
      sug:
        subj:
          Medical Practice
          Natural Language Processing Evaluation
          Decision Making, Clinical Methods
          Logic
          Conceptual Framework
          Decision Support Systems, Clinical
          Human
          Funding Source
          United States
          Case Studies
          Problem Solving
          Uncertainty
          Models, Theoretical
      ab: Artificial intelligence (AI) has become a crucial element of modern technology, especially in the healthcare sector, which is apparent given the continuous development of large language models (LLMs), which are utilized in various domains, including medical beings. However, when it comes to using these LLMs for the medical domain, there's a need for an evaluation platform to determine their suitability and drive future development efforts. Towards that end, this study aims to address this concern by developing a comprehensive Multi-Criteria Decision Making (MCDM) approach that is specifically designed to evaluate medical LLMs. The success of AI, particularly LLMs, in the healthcare domain, depends on their efficacy, safety, and ethical compliance. Therefore, it is essential to have a robust evaluation framework for their integration into medical contexts. This study proposes using the Fuzzy-Weighted Zero-InConsistency (FWZIC) method extended to p, q-quasirung orthopair fuzzy set (p, q-QROFS) for weighing evaluation criteria. This extension enables the handling of uncertainties inherent in medical decision-making processes. The approach accommodates the imprecise and multifaceted nature of real-world medical data and criteria by incorporating fuzzy logic principles. The MultiAtributive Ideal-Real Comparative Analysis (MAIRCA) method is employed for the assessment of medical LLMs utilized in the case study of this research. The results of this research revealed that "Medical Relation Extraction" criteria with its sub-levels had more importance with (0.504) than "Clinical Concept Extraction" with (0.495). For the LLMs evaluated, out of 6 alternatives, ( A 4 ) "GatorTron S 10B" had the 1st rank as compared to ( A 1 ) "GatorTron 90B" had the 6th rank. The implications of this study extend beyond academic discourse, directly impacting healthcare practices and patient outcomes. The proposed framework can help healthcare professionals make more informed decisions regarding the adoption and utilization of LLMs in medical settings.
      pubtype: Academic Journal
      doctype:
        equations & formulas
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N