Robustness of Deep Networks for Mammography: Replication Across Public Datasets.

Deep neural networks have demonstrated promising performance in screening mammography with recent studies reporting performance at or above the level of trained radiologists on internal datasets. However, it remains unclear whether the performance of these trained models is robust and replicates acr...

Descripción completa

Detalles Bibliográficos
Publicado en:Journal of Digital Imaging Vol. 37; no. 2; pp. 536 - 547
Autores principales: Velarde, Osvaldo M., Lin, Clarissa, Eskreis-Winkler, Sarah, Parra, Lucas C.
Formato: pictorial research tables/charts Journal Article
Publicado: Springer Nature Apr2024
Acceso en línea:Ver este registro en EBSCOhost
fields @attributes:
  recordID: 1
pdfLink:
plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=177625993&site=ehost-live
header:
  @attributes:
    shortDbName: ccm
    uiTerm: 177625993
    longDbName: CINAHL Complete
    uiTag: AN
  controlInfo:
    bkinfo:
    dissinfo:
    jinfo:
      jid:
        08971889
        DOQ
      jtl: Journal of Digital Imaging
      issn: 08971889
      maglogo: N
    pubinfo:
      dt: Apr2024
      vid: 37
      iid: 2
      pid: 237
      pub: Springer Nature
      place: New York, New York
    artinfo:
      ui:
        177625993
        177625993
        177625993
        10.1007/s10278-023-00943-5
        177625993
      ppf: 536
      ppct: 11
      formats:
        fmt:
          – @attributes:
              type: T
          – @attributes:
              type: P
      tig:
        atl: Robustness of Deep Networks for Mammography: Replication Across Public Datasets.
      aug:
        au:
          Velarde, Osvaldo M.
          Lin, Clarissa
          Eskreis-Winkler, Sarah
          Parra, Lucas C.
        affil: https://ror.org/00wmhkr98 The Department of Biomedical Engineering, The City College of New York, 10030, New York, NY, USA
      sug:
        subj:
          Mammography
          Data Management
          Neural Networks (Computer)
          Deep Learning
          Human
          Radiologists
          ROC Curve
          Sensitivity and Specificity
          After Care
          Carcinoma in Situ
          Neoplasms Diagnosis
          Funding Source
      ab: Deep neural networks have demonstrated promising performance in screening mammography with recent studies reporting performance at or above the level of trained radiologists on internal datasets. However, it remains unclear whether the performance of these trained models is robust and replicates across external datasets. In this study, we evaluate four state-of-the-art publicly available models using four publicly available mammography datasets (CBIS-DDSM, INbreast, CMMD, OMI-DB). Where test data was available, published results were replicated. The best-performing model, which achieved an area under the ROC curve (AUC) of 0.88 on internal data from NYU, achieved here an AUC of 0.9 on the external CMMD dataset (N = 826 exams). On the larger OMI-DB dataset (N = 11,440 exams), it achieved an AUC of 0.84 but did not match the performance of individual radiologists (at a specificity of 0.92, the sensitivity was 0.97 for the radiologist and 0.53 for the network for a 1-year follow-up). The network showed higher performance for in situ cancers, as opposed to invasive cancers. Among invasive cancers, it was relatively weaker at identifying asymmetries and was relatively stronger at identifying masses. The three other trained models that we evaluated all performed poorly on external datasets. Independent validation of trained models is an essential step to ensure safe and reliable use. Future progress in AI for mammography may depend on a concerted effort to make larger datasets publicly available that span multiple clinical sites.
      pubtype: Academic Journal
      doctype:
        pictorial
        research
        tables/charts
        Journal Article
      ougenre: Article
    language: English
    refInfo:
    holdings:
      @attributes:
        islocal: N