Comparative Analysis of Large Language Models and Machine Learning for ASA Classification Using Structured Electronic Health Record Data.

The American Society of Anesthesiologists Physical Status classification suffers from substantial inter-rater variability, compromising perioperative risk assessment. Large language models have demonstrated remarkable capabilities in medical text processing, but their application to anesthetic risk...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Medical Systems Vol. 50; no. 1; pp. 1 - 13
Autores principales: Ko, Shih-Yu, Huang, Ting-Yun, Huang, Ming-Siang, Lin, Yi-Hsuan, Su, Yung-Cheng, Chang, Yung-Chun
Formato: research tables/charts Journal Article
Publicado: Springer Nature 4/29/2026
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=193365993&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 193365993
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        01485598
        4N0
      jtl: Journal of Medical Systems
      issn: 01485598
      maglogo: N
    pubinfo:
      dt: 4/29/2026
      vid: 50
      iid: 1
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        193365993
        193365993
        193365993
        10.1007/s10916-026-02393-2
        193365993
      ppf: 1
      ppct: 12
      formats:
      tig:
        atl: Comparative Analysis of Large Language Models and Machine Learning for ASA Classification Using Structured Electronic Health Record Data.
      aug:
        au:
          Ko, Shih-Yu
          Huang, Ting-Yun
          Huang, Ming-Siang
          Lin, Yi-Hsuan
          Su, Yung-Cheng
          Chang, Yung-Chun
        affil: https://ror.org/05031qk94 Emergency Department, Shuang Ho Hospital, Taipei Medical University, Taipei, Taiwan
      sug:
        subj:
          Artificial Intelligence, Generative Evaluation
          Machine Learning Algorithms Evaluation
          Anesthesiologists Organizations
          Classification
          Automation
          Anesthetics Adverse Effects
          Intraoperative Complications Risk Factors
          Postoperative Complications Risk Factors
          Risk Assessment
          Human
          Comparative Studies
          Retrospective Design
          Record Review
          United States
          Male
          Female
          Adult
          Middle Age
          Aged
          Funding Source
          Electronic Health Records
          Preoperative Period
          Validation Studies
          ROC Curve
          Descriptive Statistics
          Criterion-Related Validity
          Body Weight
          Age Factors
          Drug Therapy
          Interrater Reliability
          Precision
          Artificial Intelligence
          Perioperative Medicine
          Decision Support Systems, Clinical
          Nonexperimental Studies
          McNemar's Test
          Post Hoc Analysis
          Confidence Intervals
          Data Analysis Software
          Classification Algorithms
          Health Status Classification
          Adult: 19-44 years
          Middle Aged: 45-64 years
          Aged: 65+ years
          Male
          Female
      ab: The American Society of Anesthesiologists Physical Status classification suffers from substantial inter-rater variability, compromising perioperative risk assessment. Large language models have demonstrated remarkable capabilities in medical text processing, but their application to anesthetic risk prediction remains unexplored. This study compared large language model performance against traditional machine learning approaches for automated ASA classification. This retrospective study analyzed 21,049 surgical records from the MOVER database (2017–2022), utilizing structured preoperative data including patient demographics, laboratory values, medication records, and procedural information for ASA I-IV classifications. We compared five large language models (GPT-4, GPT-4 Turbo, GPT-4o, Gemini 1.5 Flash, Gemini 1.5 Pro) against eight traditional machine learning algorithms using 10-fold cross-validation. Four training approaches were implemented: zero-shot learning, few-shot learning, few-shot with machine learning features, and few-shot with large language model features. Performance was evaluated using precision, recall, F₁-score, AUROC, and AUPRC. GPT-4o Few Shot Feature LLM achieved superior performance with 66% accuracy, F₁-score of 0.65, and AUROC of 0.79, substantially outperforming the best traditional method LightGBM (54% accuracy), representing a 12%-point improvement in classification accuracy. Large language models demonstrated excellent performance in intermediate risk categories (ASA II-III) with precision 0.60–0.75 and recall 0.60–0.78, covering 85% of surgical patients. However, all models showed low recall rates for ASA IV patients (28%), indicating challenges identifying highest-risk patients. Feature importance analysis revealed convergent validity between approaches, consistently identifying weight, age, and medication complexity as primary predictors. Large language models outperformed traditional machine learning for ASA classification in this dataset, particularly in intermediate-risk patients comprising most surgical cases. Their superior ability to interpret structured clinical text, including procedure names, medication records, and laboratory results, demonstrates potential for reducing inter-rater variability. However, low recall for high-risk patients necessitates human oversight and thus positions these models as decision support tools. Multi-institutional validation is essential before clinical implementation.
      pubtype: Academic Journal
      doctype:
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N