- 02.09.2026 - 08:00
Key message
A commercially available AI algorithm was retrospectively evaluated on 22,673 mammographies and showed significant performance differences among the three mammographies devices. A single general threshold notably reduced AI-based screening performance compared to standard double human reading. Device-specific thresholds partly compensated for this, for instance reducing the number of AI-flagged mammographies by almost a third, but remained poor on one device.
Background
AI algorithms are increasingly implemented to improve cancer detection rates and reduce human workload in organized mammography screening programs. Setting the AI threshold for flagging suspicious mammographies is a crucial task. Most published evaluations, however, rest on a single screening site with a single device from a single vendor, limiting their generalizability.
Methods
We retrospectively analyzed 22,673 mammographies of women aged 50–69 of years who participated in the “donna” program in the canton of St. Gallen in 2022–2023. These mammographies were acquired at six screening sites on three devices. All examinations were processed by a commercially available AI algorithm. The dataset included 115 screen-detected and 15 one-year interval breast cancers. Optimal thresholds were derived once across all devices and once per device, and screening performance was assessed using sensitivity, specificity, balanced accuracy and cancer detection rate.
Results and outlook
AI score distributions differed noticeably by device. For one device in particular, the median case score was significantly lower than for the other two. This gap persisted after adjusting for breast density, age and other factors. A general threshold flagged 18.6% of all mammographies; device-specific thresholds reduced this significantly to 12.7 while improving screening performance. On the device with significantly altered AI score distribution, however, screening performance remained poor.
Implications for Practice & Policy
AI algorithms should be validated on each device before implementation and monitored closely to maintain diagnostic accuracy. Device-specific thresholds are needed but may not suffice if the algorithm has not been adequately trained on a given device. Radiologists must be aware of varying AI score distributions across devices to mitigate interpretation biases.
Publication
Titel: AI performance varies considerably across mammography devices: a multi-site and multi-vendor retrospective study
Authors: Marcel Blum, Rudolf Morant, Alena Eichenberger, Alexander Geissler, Jonas Subelack, Justus Vogel, David Ehlig
Read the full article here: https://rdcu.be/fCeh4
