DINOv2-Pretrained Vision Transformers for Stroke and Heart Failure Detection from Chest X-rays with Limited Labeled Data.
Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We prop...
| Publicado en: | Journal of Imaging Informatics in Medicine pp. 1 - 19 |
|---|---|
| Autores principales: | , , , |
| Formato: | Journal Article |
| Publicado: |
Springer Nature
Sep2026
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=196882203&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 196882203 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 29482925 NR3A jtl: Journal of Imaging Informatics in Medicine issn: 29482925 maglogo: N pubinfo: dt: Sep2026 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 196882203 10.1007/s10278-026-02133-5 196882203 ppf: 1 ppct: 18 formats: tig: atl: DINOv2-Pretrained Vision Transformers for Stroke and Heart Failure Detection from Chest X-rays with Limited Labeled Data. aug: au: Tsai, Ming-Han Veda, Jalu Leu, Jenq-Shiou Tsai, Chia-Ti affil: Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology sug: ab: Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We propose a three-stage framework that combines DINOv2 self-supervised learning with vision transformers (ViTs) to detect stroke and heart failure from CXRs under limited labeled data. A ViT-Small/14 backbone is first initialized with DINOv2 weights pretrained on LVD-142M dataset and then further adapted by self-supervised pretraining on 5368 unlabeled CXRs from the UCSD dataset, with the bottom 8 transformer blocks frozen to preserve general visual representations during domain adaptation. Model evaluation is performed via frozen-backbone linear probing on 2000 labeled images per task from the NTUH-iMD database. Ablation experiments confirm that each pretraining stage contributes measurable gains under linear probing. Benchmarked against five SSL baselines spanning general domain and medical domain pretraining under an identical evaluation protocol, our stroke model achieves 89.60 ± 1.39% accuracy and 89.55 ± 1.40% <italic>F</italic>1-score, outperforming all baselines including RAD-DINO, which uses nearly four times the parameters. Our heart failure model achieves 92.52 ± 1.57% accuracy, surpassing GLoRIA, CheXzero, and RAD-DINO. We further show that negative class composition critically affects performance and that normal CXR proportion should be interpreted in light of its effect on task difficulty. Attention visualizations confirm that both models focus on anatomically meaningful cardiac and pulmonary regions, supporting the clinical interpretability of the proposed framework.Graphical Abstract: Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We propose a three-stage framework that combines DINOv2 self-supervised learning with vision transformers (ViTs) to detect stroke and heart failure from CXRs under limited labeled data. A ViT-Small/14 backbone is first initialized with DINOv2 weights pretrained on LVD-142M dataset and then further adapted by self-supervised pretraining on 5368 unlabeled CXRs from the UCSD dataset, with the bottom 8 transformer blocks frozen to preserve general visual representations during domain adaptation. Model evaluation is performed via frozen-backbone linear probing on 2000 labeled images per task from the NTUH-iMD database. Ablation experiments confirm that each pretraining stage contributes measurable gains under linear probing. Benchmarked against five SSL baselines spanning general domain and medical domain pretraining under an identical evaluation protocol, our stroke model achieves 89.60 ± 1.39% accuracy and 89.55 ± 1.40% <italic>F</italic>1-score, outperforming all baselines including RAD-DINO, which uses nearly four times the parameters. Our heart failure model achieves 92.52 ± 1.57% accuracy, surpassing GLoRIA, CheXzero, and RAD-DINO. We further show that negative class composition critically affects performance and that normal CXR proportion should be interpreted in light of its effect on task difficulty. Attention visualizations confirm that both models focus on anatomically meaningful cardiac and pulmonary regions, supporting the clinical interpretability of the proposed framework. pubtype: Academic Journal doctype: Journal Article ougenre: Unknown language: English refInfo: holdings: @attributes: islocal: N |
|---|