Cross-site Validation of AI Segmentation and Harmonization in Breast MRI.

This work aims to perform a cross-site validation of automated segmentation for breast cancers in MRI and to compare the performance to radiologists. A three-dimensional (3D) U-Net was trained to segment cancers in dynamic contrast-enhanced axial MRIs using a large dataset from Site 1 (n = 15,266; 4...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Imaging Informatics in Medicine Vol. 38; no. 3; pp. 1642 - 1653
Autores principales: Huang, Yu, Leotta, Nicholas J., Hirsch, Lukas, Gullo, Roberto Lo, Hughes, Mary, Reiner, Jeffrey, Saphier, Nicole B., Myers, Kelly S., Panigrahi, Babita, Ambinder, Emily, Di Carlo, Philip, Grimm, Lars J., Lowell, Dorothy, Yoon, Sora, Ghate, Sujata V., Parra, Lucas C., Sutton, Elizabeth J.
Formato: diagnostic images research tables/charts Journal Article
Publicado: Springer Nature Jun2025
Acceso en línea:Ver este registro en EBSCOhost
Descripción
Sumario:This work aims to perform a cross-site validation of automated segmentation for breast cancers in MRI and to compare the performance to radiologists. A three-dimensional (3D) U-Net was trained to segment cancers in dynamic contrast-enhanced axial MRIs using a large dataset from Site 1 (n = 15,266; 449 malignant and 14,817 benign). Performance was validated on site-specific test data from this and two additional sites, and common publicly available testing data. Four radiologists from each of the three clinical sites provided two-dimensional (2D) segmentations as ground truth. Segmentation performance did not differ between the network and radiologists on the test data from Sites 1 and 2 or the common public data (median Dice score Site 1, network 0.86 vs. radiologist 0.85, n = 114; Site 2, 0.91 vs. 0.91, n = 50; common: 0.93 vs. 0.90). For Site 3, an affine input layer was fine-tuned using segmentation labels, resulting in comparable performance between the network and radiologist (0.88 vs. 0.89, n = 42). Radiologist performance differed on the common test data, and the network numerically outperformed 11 of the 12 radiologists (median Dice: 0.85–0.94, n = 20). In conclusion, a deep network with a novel supervised harmonization technique matches radiologists' performance in MRI tumor segmentation across clinical sites. We make code and weights publicly available to promote reproducible AI in radiology.