Robustness of Deep Networks for Mammography: Replication Across Public Datasets.
Deep neural networks have demonstrated promising performance in screening mammography with recent studies reporting performance at or above the level of trained radiologists on internal datasets. However, it remains unclear whether the performance of these trained models is robust and replicates acr...
| Publicado en: | Journal of Digital Imaging Vol. 37; no. 2; pp. 536 - 547 |
|---|---|
| Autores principales: | , , , |
| Formato: | pictorial research tables/charts Journal Article |
| Publicado: |
Springer Nature
Apr2024
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=177625993&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 177625993 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 08971889 DOQ jtl: Journal of Digital Imaging issn: 08971889 maglogo: N pubinfo: dt: Apr2024 vid: 37 iid: 2 pid: 237 pub: Springer Nature place: New York, New York artinfo: ui: 177625993 177625993 177625993 10.1007/s10278-023-00943-5 177625993 ppf: 536 ppct: 11 formats: fmt: – @attributes: type: T – @attributes: type: P tig: atl: Robustness of Deep Networks for Mammography: Replication Across Public Datasets. aug: au: Velarde, Osvaldo M. Lin, Clarissa Eskreis-Winkler, Sarah Parra, Lucas C. affil: https://ror.org/00wmhkr98 The Department of Biomedical Engineering, The City College of New York, 10030, New York, NY, USA sug: subj: Mammography Data Management Neural Networks (Computer) Deep Learning Human Radiologists ROC Curve Sensitivity and Specificity After Care Carcinoma in Situ Neoplasms Diagnosis Funding Source ab: Deep neural networks have demonstrated promising performance in screening mammography with recent studies reporting performance at or above the level of trained radiologists on internal datasets. However, it remains unclear whether the performance of these trained models is robust and replicates across external datasets. In this study, we evaluate four state-of-the-art publicly available models using four publicly available mammography datasets (CBIS-DDSM, INbreast, CMMD, OMI-DB). Where test data was available, published results were replicated. The best-performing model, which achieved an area under the ROC curve (AUC) of 0.88 on internal data from NYU, achieved here an AUC of 0.9 on the external CMMD dataset (N = 826 exams). On the larger OMI-DB dataset (N = 11,440 exams), it achieved an AUC of 0.84 but did not match the performance of individual radiologists (at a specificity of 0.92, the sensitivity was 0.97 for the radiologist and 0.53 for the network for a 1-year follow-up). The network showed higher performance for in situ cancers, as opposed to invasive cancers. Among invasive cancers, it was relatively weaker at identifying asymmetries and was relatively stronger at identifying masses. The three other trained models that we evaluated all performed poorly on external datasets. Independent validation of trained models is an essential step to ensure safe and reliable use. Future progress in AI for mammography may depend on a concerted effort to make larger datasets publicly available that span multiple clinical sites. pubtype: Academic Journal doctype: pictorial research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|