Comparative Analysis of Large Language Models and Machine Learning for ASA Classification Using Structured Electronic Health Record Data.
The American Society of Anesthesiologists Physical Status classification suffers from substantial inter-rater variability, compromising perioperative risk assessment. Large language models have demonstrated remarkable capabilities in medical text processing, but their application to anesthetic risk...
| Publicado en: | Journal of Medical Systems Vol. 50; no. 1; pp. 1 - 13 |
|---|---|
| Autores principales: | , , , , , |
| Formato: | research tables/charts Journal Article |
| Publicado: |
Springer Nature
4/29/2026
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=193365993&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 193365993 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 01485598 4N0 jtl: Journal of Medical Systems issn: 01485598 maglogo: N pubinfo: dt: 4/29/2026 vid: 50 iid: 1 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 193365993 193365993 193365993 10.1007/s10916-026-02393-2 193365993 ppf: 1 ppct: 12 formats: tig: atl: Comparative Analysis of Large Language Models and Machine Learning for ASA Classification Using Structured Electronic Health Record Data. aug: au: Ko, Shih-Yu Huang, Ting-Yun Huang, Ming-Siang Lin, Yi-Hsuan Su, Yung-Cheng Chang, Yung-Chun affil: https://ror.org/05031qk94 Emergency Department, Shuang Ho Hospital, Taipei Medical University, Taipei, Taiwan sug: subj: Artificial Intelligence, Generative Evaluation Machine Learning Algorithms Evaluation Anesthesiologists Organizations Classification Automation Anesthetics Adverse Effects Intraoperative Complications Risk Factors Postoperative Complications Risk Factors Risk Assessment Human Comparative Studies Retrospective Design Record Review United States Male Female Adult Middle Age Aged Funding Source Electronic Health Records Preoperative Period Validation Studies ROC Curve Descriptive Statistics Criterion-Related Validity Body Weight Age Factors Drug Therapy Interrater Reliability Precision Artificial Intelligence Perioperative Medicine Decision Support Systems, Clinical Nonexperimental Studies McNemar's Test Post Hoc Analysis Confidence Intervals Data Analysis Software Classification Algorithms Health Status Classification Adult: 19-44 years Middle Aged: 45-64 years Aged: 65+ years Male Female ab: The American Society of Anesthesiologists Physical Status classification suffers from substantial inter-rater variability, compromising perioperative risk assessment. Large language models have demonstrated remarkable capabilities in medical text processing, but their application to anesthetic risk prediction remains unexplored. This study compared large language model performance against traditional machine learning approaches for automated ASA classification. This retrospective study analyzed 21,049 surgical records from the MOVER database (2017–2022), utilizing structured preoperative data including patient demographics, laboratory values, medication records, and procedural information for ASA I-IV classifications. We compared five large language models (GPT-4, GPT-4 Turbo, GPT-4o, Gemini 1.5 Flash, Gemini 1.5 Pro) against eight traditional machine learning algorithms using 10-fold cross-validation. Four training approaches were implemented: zero-shot learning, few-shot learning, few-shot with machine learning features, and few-shot with large language model features. Performance was evaluated using precision, recall, F₁-score, AUROC, and AUPRC. GPT-4o Few Shot Feature LLM achieved superior performance with 66% accuracy, F₁-score of 0.65, and AUROC of 0.79, substantially outperforming the best traditional method LightGBM (54% accuracy), representing a 12%-point improvement in classification accuracy. Large language models demonstrated excellent performance in intermediate risk categories (ASA II-III) with precision 0.60–0.75 and recall 0.60–0.78, covering 85% of surgical patients. However, all models showed low recall rates for ASA IV patients (28%), indicating challenges identifying highest-risk patients. Feature importance analysis revealed convergent validity between approaches, consistently identifying weight, age, and medication complexity as primary predictors. Large language models outperformed traditional machine learning for ASA classification in this dataset, particularly in intermediate-risk patients comprising most surgical cases. Their superior ability to interpret structured clinical text, including procedure names, medication records, and laboratory results, demonstrates potential for reducing inter-rater variability. However, low recall for high-risk patients necessitates human oversight and thus positions these models as decision support tools. Multi-institutional validation is essential before clinical implementation. pubtype: Academic Journal doctype: research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|