DINOv2-Pretrained Vision Transformers for Stroke and Heart Failure Detection from Chest X-rays with Limited Labeled Data.

Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We prop...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Imaging Informatics in Medicine pp. 1 - 19
Autores principales: Tsai, Ming-Han, Veda, Jalu, Leu, Jenq-Shiou, Tsai, Chia-Ti
Formato: Journal Article
Publicado: Springer Nature Sep2026
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=196882203&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 196882203
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        29482925
        NR3A
      jtl: Journal of Imaging Informatics in Medicine
      issn: 29482925
      maglogo: N
    pubinfo:
      dt: Sep2026
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        196882203
        10.1007/s10278-026-02133-5
        196882203
      ppf: 1
      ppct: 18
      formats:
      tig:
        atl: DINOv2-Pretrained Vision Transformers for Stroke and Heart Failure Detection from Chest X-rays with Limited Labeled Data.
      aug:
        au:
          Tsai, Ming-Han
          Veda, Jalu
          Leu, Jenq-Shiou
          Tsai, Chia-Ti
        affil: Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology
      sug:
      ab: Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We propose a three-stage framework that combines DINOv2 self-supervised learning with vision transformers (ViTs) to detect stroke and heart failure from CXRs under limited labeled data. A ViT-Small/14 backbone is first initialized with DINOv2 weights pretrained on LVD-142M dataset and then further adapted by self-supervised pretraining on 5368 unlabeled CXRs from the UCSD dataset, with the bottom 8 transformer blocks frozen to preserve general visual representations during domain adaptation. Model evaluation is performed via frozen-backbone linear probing on 2000 labeled images per task from the NTUH-iMD database. Ablation experiments confirm that each pretraining stage contributes measurable gains under linear probing. Benchmarked against five SSL baselines spanning general domain and medical domain pretraining under an identical evaluation protocol, our stroke model achieves 89.60 ± 1.39% accuracy and 89.55 ± 1.40% <italic>F</italic>1-score, outperforming all baselines including RAD-DINO, which uses nearly four times the parameters. Our heart failure model achieves 92.52 ± 1.57% accuracy, surpassing GLoRIA, CheXzero, and RAD-DINO. We further show that negative class composition critically affects performance and that normal CXR proportion should be interpreted in light of its effect on task difficulty. Attention visualizations confirm that both models focus on anatomically meaningful cardiac and pulmonary regions, supporting the clinical interpretability of the proposed framework.Graphical Abstract: Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We propose a three-stage framework that combines DINOv2 self-supervised learning with vision transformers (ViTs) to detect stroke and heart failure from CXRs under limited labeled data. A ViT-Small/14 backbone is first initialized with DINOv2 weights pretrained on LVD-142M dataset and then further adapted by self-supervised pretraining on 5368 unlabeled CXRs from the UCSD dataset, with the bottom 8 transformer blocks frozen to preserve general visual representations during domain adaptation. Model evaluation is performed via frozen-backbone linear probing on 2000 labeled images per task from the NTUH-iMD database. Ablation experiments confirm that each pretraining stage contributes measurable gains under linear probing. Benchmarked against five SSL baselines spanning general domain and medical domain pretraining under an identical evaluation protocol, our stroke model achieves 89.60 ± 1.39% accuracy and 89.55 ± 1.40% <italic>F</italic>1-score, outperforming all baselines including RAD-DINO, which uses nearly four times the parameters. Our heart failure model achieves 92.52 ± 1.57% accuracy, surpassing GLoRIA, CheXzero, and RAD-DINO. We further show that negative class composition critically affects performance and that normal CXR proportion should be interpreted in light of its effect on task difficulty. Attention visualizations confirm that both models focus on anatomically meaningful cardiac and pulmonary regions, supporting the clinical interpretability of the proposed framework.
      pubtype: Academic Journal
      doctype: Journal Article
      ougenre: Unknown
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N