Use of Expected Utility to Evaluate Artificial Intelligence–Enabled Rule-out Devices for Mammography Screening.
Background: An artificial intelligence (AI)–enabled rule-out device may autonomously remove patient images unlikely to have cancer from radiologist review. Many published studies evaluate this type of device by retrospectively applying the AI to large datasets and use sensitivity and specificity as...
| Publicado en: | Medical Decision Making Vol. 46; no. 3; pp. 321 - 334 |
|---|---|
| Autores principales: | , , , , |
| Formato: | equations & formulas research tables/charts Journal Article |
| Publicado: |
Sage Publications Inc.
Apr2026
|
| Acceso en línea: | Ver este registro en EBSCOhost |
| fields | @attributes: recordID: 1 pdfLink: plink: https://search.ebscohost.com/login.aspx?direct=true&db=ccm&AN=192206320&site=ehost-live header: @attributes: shortDbName: ccm uiTerm: 192206320 longDbName: CINAHL Complete uiTag: AN controlInfo: bkinfo: dissinfo: jinfo: jid: 0272989X DKI jtl: Medical Decision Making issn: 0272989X maglogo: Y pubinfo: dt: Apr2026 vid: 46 iid: 3 pid: 344 pub: Sage Publications Inc. place: Thousand Oaks, California artinfo: ui: 192206320 189511162 192206320 192206320 10.1177/0272989X251388665 192206320 ppf: 321 ppct: 13 formats: tig: atl: Use of Expected Utility to Evaluate Artificial Intelligence–Enabled Rule-out Devices for Mammography Screening. aug: au: Fan, Kwok Lung Thompson, Yee Lam Elim Chen, Weijie Abbey, Craig K. Samuelson, Frank W. affil: US Food and Drug Administration, Silver Spring, MD, USA sug: subj: Mammography Equipment and Supplies Breast Neoplasms Diagnosis Cancer Screening Artificial Intelligence Equipment and Supplies Artificial Intelligence Evaluation Radiographic Image Interpretation, Computer-Assisted Validity Human Radiologists Comparative Studies Radiographic Image Enhancement Workflow Decision Making, Computer Assisted Retrospective Design Record Review United States Early Detection of Cancer Detection Algorithms Sensitivity and Specificity Funding Source Descriptive Statistics Confidence Intervals Probability ab: Background: An artificial intelligence (AI)–enabled rule-out device may autonomously remove patient images unlikely to have cancer from radiologist review. Many published studies evaluate this type of device by retrospectively applying the AI to large datasets and use sensitivity and specificity as the performance metrics. However, these metrics have fundamental shortcomings because sensitivity will always be negatively affected in retrospective studies of rule-out applications of AI. Method: We reviewed 2 performance metrics to compare the screening performance between the radiologist-with-rule-out-device and radiologist-without-device workflows: positive/negative predictive values (PPV/NPV) and expected utility (EU). We applied both methods to a recent study that reported improved performance in the radiologist-with-device workflow using a retrospective US dataset. We then applied the EU method to a European study based on the reported recall and cancer detection rates at different AI thresholds to compare the potential utility among different thresholds. Results: For the US study, neither PPV/NPV nor EU can demonstrate significant improvement for any of the algorithm thresholds reported. For the study using European data, we found that EU is lower as AI rules out more patients including false-negative cases and reduces the overall screening performance. Conclusions: Due to the nature of the retrospective simulated study design, sensitivity and specificity can be ambiguous in evaluating a rule-out device. We showed that using PPV/NPV or EU can resolve the ambiguity. The EU method can be applied with only recall rates and cancer detection rates, which is convenient as ground truth is often unavailable for nonrecalled patients in screening mammography. Highlights: Sensitivity and specificity can be ambiguous metrics for evaluating a rule-out device in a retrospective setting. PPV and NPV can resolve the ambiguity but require the ground truth for all patients. Based on utility theory, expected utility (EU) is a potential metric that helps demonstrate improvement in screening performance due to a rule-out device using large retrospective datasets. We applied EU to a recent study that used a large retrospective mammography screening dataset from the United States. That study reported an improvement in specificity and decrease in sensitivity when using their AI as a rule-out device retrospectively. In terms of EU, we cannot conclude a significant improvement when the AI is used as a rule-out device. We applied the method to a European study that reported only recall rates and cancer detection rates. Since there is no established EU baseline value in European mammography screening workflow, we estimated the EU baseline using data from previous literature. We cannot conclude a significant improvement when the AI is used as a rule-out device for the European study. In this work, we investigated the use of EU to evaluate rule-out devices using large retrospective datasets. This metric, used with retrospective clinical data, could be used as supporting evidence for rule-out devices. pubtype: Academic Journal doctype: equations & formulas research tables/charts Journal Article ougenre: Article language: English refInfo: holdings: @attributes: islocal: N |
|---|