From Embeddings to Accuracy: Comparing Foundation Models for Radiographic Classification.
Foundation models, pre-trained on extensive datasets, have significantly advanced machine learning by providing robust and transferable embeddings applicable to various domains, including medical imaging diagnostics. This study evaluates the utility of embeddings derived from both general-purpose an...
| Publicado en: | Journal of Imaging Informatics in Medicine Vol. 39; no. 4; pp. 3196 - 3208 |
|---|---|
| Autores principales: | , , , , , , , , , , , , , |
| Formato: | diagnostic images research tables/charts Journal Article |
| Publicado: |
Springer Nature
Aug2026
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=196241817&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 196241817 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 29482925 NR3A jtl: Journal of Imaging Informatics in Medicine issn: 29482925 maglogo: N pubinfo: dt: Aug2026 vid: 39 iid: 4 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 196241817 189902775 196241817 196241817 10.1007/s10278-025-01747-5 196241817 ppf: 3196 ppct: 12 formats: tig: atl: From Embeddings to Accuracy: Comparing Foundation Models for Radiographic Classification. aug: au: Li, Xue Merkow, Jameson Codella, Noel C. F. Santamaria-Pang, Alberto Sangani, Naiteek Ersoy, Alexander Burt, Christopher Garrett, John W. Bruce, Richard J. Warner, Joshua D. Bradshaw, Tyler Tarapov, Ivan Lungren, Matthew P. McMillan, Alan B. affil: https://ror.org/01y2jtd41 Department of Radiology, University of Wisconsin-Madison, Madison, WI, USA sug: subj: Radiography Classification Classification Algorithms Convolutional Neural Networks Tube Placement Determination Human Funding Source Wisconsin Male Female Infant, Newborn Infant Child, Preschool Child Adolescence Adult Middle Age Aged Aged, 80 and Over Comparative Studies Wilcoxon Signed Rank Test Kruskal-Wallis Test Post Hoc Analysis Machine Learning Algorithms Logistic Regression Support Vector Machine Random Forest Sex Factors Age Factors Infant, Newborn: birth-1 month Infant: 1-23 months Child, Preschool: 2-5 years Child: 6-12 years Adolescent: 13-18 years Adult: 19-44 years Middle Aged: 45-64 years Aged: 65+ years Aged, 80 & over Male Female ab: Foundation models, pre-trained on extensive datasets, have significantly advanced machine learning by providing robust and transferable embeddings applicable to various domains, including medical imaging diagnostics. This study evaluates the utility of embeddings derived from both general-purpose and medical domain-specific foundation models for training lightweight adapter models in multi-class radiography classification, focusing specifically on tube placement assessment and related findings, with comparison to the end-to-end training of an established convolutional neural network. A dataset comprising 8842 radiographs classified into seven distinct categories was employed to extract embeddings using seven foundation models: DenseNet121, BiomedCLIP, Med-Flamingo, MedImageInsight, MedSigLIP, Rad-DINO, and CXR-Foundation. Adapter models were subsequently trained using classical machine learning algorithms, including K-nearest neighbors (KNN), logistic regression (LR), support vector machines (SVM), random forest (RF), and multi-layer perceptron (MLP). Among these combinations, MedImageInsight embeddings paired with an SVM or MLP adapter yielded the highest mean area under the curve (mAUC) at 93.1%, followed closely by MedSigLIP with MLP (91.0%), Rad-DINO with SVM (90.7%), and CXR-Foundation with LR (88.6%), achieving a higher mAUC score than a fully finetuned convolutional neural network, DenseNet121 (87.2%). In comparison, BiomedCLIP and DenseNet121 exhibited moderate performance with SVM, obtaining mAUC scores of 82.8% and 81.1%, respectively, whereas Med-Flamingo delivered the lowest performance at 78.5% when combined with RF. Significant differences were found between each embedding model and MedImageInsight using the Wilcoxon signed-rank test at the significance level 0.05 (before Bonferroni correction). Notably, most adapter models demonstrated computational efficiency, achieving training within minutes and inference within seconds on CPU, underscoring their practicality for clinical applications. Furthermore, fairness analysis on adapters trained on MedImageInsight-derived embeddings indicated minimal disparities, with gender differences in performance within 1.8% and standard deviations across age groups not exceeding 1.4%. Further analysis indicated there is no significant difference across gender and age at a significance level of 0.05. These findings confirm that foundation model embeddings—especially those from MedImageInsight—facilitate accurate, computationally efficient, and equitable diagnostic classification using lightweight adapters for radiographic image analysis. pubtype: Academic Journal doctype: diagnostic images research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|