Ten Years of Artificial Intelligence in Screening Mammography: A Systematic Review and Meta-Analysis of Diagnostic Accuracy and Clinical Implementation (Literature Published 2015–2025)
Abstract
1. Introduction
2. Materials and Methods
2.1. Protocol and Registration
2.2. Eligibility Criteria
2.3. Information Sources and Search
2.4. Study Selection
2.5. Data Extraction
- Multiple AI systems in one study. Where a study evaluated several AI systems without nominating one as primary, the median-performing system entered the main analysis, with the best- and worst-performing systems carried into a sensitivity analysis. Taking the best performer, as our original submission did for Salim et al. [25], selects on the outcome and biases the pooled estimate upwards. Under the median rule, that study contributes AI-2, AUC 0.922 (0.910–0.934), rather than AI-1, AUC 0.956 (0.948–0.965).
- Multiple cohorts in one study. Where a study reported several independent cohorts, the largest entered the main analysis and the others were carried into a sensitivity analysis. Schaffter et al. [26] therefore contributes the US cohort (144,231 examinations, AUC 0.858) rather than the Swedish cohort (AUC 0.903).
2.6. Risk of Bias and Certainty of Evidence
2.7. Statistical Analysis
2.8. Cohort Overlap
2.9. Deviations from the Protocol
3. Results
3.1. Study Selection
3.2. Characteristics of Included Studies
3.3. Standalone AI Detection Accuracy (Analysis A)
| Study (Year) [Ref] | Design | AI AUC (95% CI) | Comparator in the Same Study | Comparator AUC |
|---|---|---|---|---|
| Rodriguez-Ruiz (2019) [34] | Enriched case–control | 0.840 (0.820–0.860) | Mean of 101 radiologists | 0.814 |
| Salim (2020) [25] | Enriched case–control | 0.922 (0.910–0.934) | Not reported on the AUC scale | – |
| Schaffter (2020) [26] | Population cohort | 0.858 (0.843–0.873) * | Not reported on the AUC scale | – |
| Kizildag Yirgin (2022) [35] | Enriched case–control | 0.853 (0.801–0.905) | Not reported on the AUC scale | – |
| Romero-Martin (2022) [36] | Population cohort | 0.930 (0.890–0.960) | Not reported on the AUC scale | – |
| Hsu (2022) [37] | Population cohort | 0.850 (0.840–0.870) | Not reported on the AUC scale | – |
| Marinovich (2023) [39] | Population cohort | 0.830 (0.812–0.848) * | Radiologists interpreting the same screens | 0.930 |
| Riveira-Martin (2023) [40] | Population cohort | 0.920 (0.890–0.950) | Not reported on the AUC scale | – |
| Kwon (2024) [41] | Population cohort | 0.800 (0.760–0.840) | Radiologist BI-RADS assessment | 0.740 (0.700–0.780) |
| Seker (2024) [42] | Population cohort | 0.896 (0.861–0.932) | Not reported on the AUC scale | – |
| Larsen (2024) [43] | Population cohort | 0.930 (0.920–0.930) | Not reported on the AUC scale | – |
| Graham-Knight (2025) [46] | Population cohort | 0.930 (0.920–0.940) | Not reported on the AUC scale | – |
| Yamaguchi (2025) [47] | Enriched case–control | 0.841 (0.822–0.859) | Not reported on the AUC scale | – |
| Retson (2022) [38] | Enriched case–control | 0.950 (0.940–0.960) | Not reported on the AUC scale | – |
| Pooled, random effects | 14 studies | 0.890 (0.858–0.915) | – | – |
| 95% prediction interval | 0.731–0.960 |
3.4. Sensitivity and Specificity (Analysis A2)
3.5. AI-Integrated Reading Versus Standard Reading (Analysis B)
3.6. Recall, Workload and Other Programme Outcomes (Analysis C)
3.7. Risk of Bias and Certainty of Evidence
4. Discussion
4.1. Limitations
4.2. Implications and Future Research
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial intelligence |
| AI-STREAM | Artificial Intelligence for Breast Cancer Screening in Mammography |
| AUC | Area under the receiver operating characteristic curve |
| BI-RADS | Breast Imaging-Reporting and Data System |
| CAD | Computer-aided detection |
| CDR | Cancer detection rate |
| CI | Confidence interval |
| DBT | Digital breast tomosynthesis |
| DLADS | Deep-learning-based automated diagnostic system |
| DM | Digital mammography |
| DREAM | Dialogue for Reverse Engineering Assessments and Methods |
| GRADE | Grading of Recommendations Assessment, Development and Evaluation |
| HSROC | Hierarchical summary receiver operating characteristic |
| MASAI | Mammography Screening with Artificial Intelligence |
| MDR | Medical Device Regulation |
| MRMC | Multi-reader multi-case |
| MRI | Magnetic resonance imaging |
| PIRD | Population, Index test, Reference standard, Diagnosis |
| PRISMA | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| PRISMA-DTA | PRISMA for Diagnostic Test Accuracy |
| PRISMA-S | PRISMA Search Reporting Extension |
| PROBAST+AI | Prediction model Risk-Of-Bias Assessment Tool for artificial intelligence |
| QUADAS-2 | Quality Assessment of Diagnostic Accuracy Studies, version 2 |
| QUADAS-C | QUADAS for comparative accuracy |
| REML | Restricted maximum likelihood |
| ROC | Receiver operating characteristic |
| TRIPOD + AI | Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis, AI extension |
| XAI | Explainable artificial intelligence |
References
- Sung, H.; Ferlay, J.; Siegel, R.L.; Laversanne, M.; Soerjomataram, I.; Jemal, A.; Bray, F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J. Clin. 2021, 71, 209–249. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Canelo-Aybar, C.; Ferreira, D.S.; Ballesteros, M.; Posso, M.; Montero, N.; Solà, I.; Saz-Parkinson, Z.; Lerda, D.; Rossi, P.G.; Duffy, S.W.; et al. Benefits and Harms of Breast Cancer Mammography Screening for Women at Average Risk of Breast Cancer: A Systematic Review for the European Commission Initiative on Breast Cancer. J. Med. Screen. 2021, 28, 389–404. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bennett, A.; Shaver, N.; Vyas, N.; Almoli, F.; Pap, R.; Douglas, A.; Kibret, T.; Skidmore, B.; Yaffe, M.; Wilkinson, A.; et al. Screening for Breast Cancer: A Systematic Review Update to Inform the Canadian Task Force on Preventive Health Care Guideline. Syst. Rev. 2024, 13, 304. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lauby-Secretan, B.; Scoccianti, C.; Loomis, D.; Benbrahim-Tallaa, L.; Bouvard, V.; Bianchini, F.; Straif, K. Breast-Cancer Screening: Viewpoint of the IARC Working Group. N. Engl. J. Med. 2015, 372, 2353–2358. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Seely, J.M.; Eby, P.R.; Yaffe, M.J. The Fundamental Flaws of the CNBSS Trials: A Scientific Review. J. Breast Imaging 2022, 4, 108–119. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gordon, P.B. The Impact of Dense Breasts on the Stage of Breast Cancer at Diagnosis: A Review and Options for Supplemental Screening. Curr. Oncol. 2022, 29, 3595–3636. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hinton, B.; Ma, L.; Mahmoudzadeh, A.P.; Malkov, S.; Fan, B.; Greenwood, H.; Joe, B.; Lee, V.; Kerlikowske, K.; Shepherd, J. Deep Learning Networks Find Unique Mammographic Differences in Previous Negative Mammograms between Interval and Screen-Detected Cancers: A Case-Case Study. Cancer Imaging 2019, 19, 41. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Giorgi Rossi, P.; Djuric, O.; Hélin, V.; Astley, S.; Mantellini, P.; Nitrosi, A.; Harkness, E.F.; Gauthier, E.; Puliti, D.; Balleyguier, C.; et al. Validation of a New Fully Automated Software for 2D Digital Mammographic Breast Density Evaluation in Predicting Breast Cancer Risk. Sci. Rep. 2021, 11, 19884. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oiwa, M.; Suda, N.; Morita, T.; Takahashi, Y.; Sato, Y.; Hayashi, T.; Kato, A.; Nishimura, R.; Ichihara, S.; Endo, T. Validity of Computed Mean Compressed Fibroglandular Tissue Thickness and Breast Composition for Stratification of Masking Risk in Japanese Women. Breast Cancer 2023, 30, 541–551. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Naik, S.; Varghese, A.P.; Asrar Ul Haq Andrabi, S.; Tivaskar, S.; Luharia, A.; Mishra, G.V. Addressing Global Gaps in Mammography Screening for Improved Breast Cancer Detection: A Review of the Literature. Cureus 2024, 16, e66198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lehman, C.D.; Wellman, R.D.; Buist, D.S.; Kerlikowske, K.; Tosteson, A.N.; Miglioretti, D.L.; for the Breast Cancer Surveillance Consortium. Diagnostic Accuracy of Digital Screening Mammography with and without Computer-Aided Detection. JAMA Intern. Med. 2015, 175, 1828–1837. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sechopoulos, I.; Teuwen, J.; Mann, R. Artificial Intelligence for Breast Cancer Detection in Mammography and Digital Breast Tomosynthesis: State of the Art. Semin. Cancer Biol. 2021, 72, 214–225. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ciurescu, S.; Cerbu, S.; Dima, C.N.; Borozan, F.; Pârvănescu, R.; Ilaș, D.G.; Cîtu, C.; Vernic, C.; Sas, I. AI in 2D Mammography: Improving Breast Cancer Screening Accuracy. Medicina 2025, 61, 809. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ciurescu, S.; Ciupici-Cladovan, M.; Buciu, V.B.; Ilaș, D.G.; Cîtu, C.; Sas, I. Systematic Review and Meta-Analysis of AI-Assisted Mammography and the Systemic Immune-Inflammation Index in Breast Cancer: Diagnostic and Prognostic Perspectives. Medicina 2025, 61, 1170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McKinney, S.M.; Sieniek, M.; Godbole, V.; Godwin, J.; Antropova, N.; Ashrafian, H.; Back, T.; Chesus, M.; Corrado, G.S.; Darzi, A.; et al. International Evaluation of an AI System for Breast Cancer Screening. Nature 2020, 577, 89–94. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lotter, W.; Diab, A.R.; Haslam, B.; Kim, J.G.; Grisot, G.; Wu, E.; Wu, K.; Onieva, J.O.; Boyer, Y.; Boxerman, J.L.; et al. Robust Breast Cancer Detection in Mammography and Digital Breast Tomosynthesis Using an Annotation-Efficient Deep Learning Approach. Nat. Med. 2021, 27, 244–249. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Freeman, K.; Geppert, J.; Stinton, C.; Todkill, D.; Johnson, S.; Clarke, A.; Taylor-Phillips, S. Use of Artificial Intelligence for Image Analysis in Breast Cancer Screening Programmes: Systematic Review of Test Accuracy. BMJ 2021, 374, n1872. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hickman, S.E.; Woitek, R.; Le, E.P.V.; Im, Y.R.; Mouritsen Luxhøj, C.; Aviles-Rivero, A.I.; Baxter, G.C.; MacKay, J.W.; Gilbert, F.J. Machine Learning for Workflow Applications in Screening Mammography: Systematic Review and Meta-Analysis. Radiology 2021, 302, 88–104. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, J.; Lei, J.; Ou, Y.; Zhao, Y.; Tuo, X.; Zhang, B.; Shen, M. Mammography Diagnosis of Breast Cancer Screening through Machine Learning: A Systematic Review and Meta-Analysis. Clin. Exp. Med. 2023, 23, 2341–2356. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yoon, J.H.; Strand, F.; Baltzer, P.A.T.; Conant, E.F.; Gilbert, F.J.; Lehman, C.D.; Morris, E.A.; Mullen, L.A.; Nishikawa, R.M.; Sharma, N.; et al. Standalone AI for Breast Cancer Detection at Screening Digital Mammography and Digital Breast Tomosynthesis: A Systematic Review and Meta-Analysis. Radiology 2023, 307, e222639. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lu, J.; Xu, X.; Zhang, Y.; Zhuang, K.; Fang, T.; Zhang, C.; Chen, K.; Huang, X.; Li, Y. Diagnostic Performance of AI-Assisted Radiologists in Breast Cancer Detection Using Digital Mammography: A Systematic Review and Meta-Analysis. Clin. Breast Cancer 2025, 26, 121–135. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. Int. J. Surg. 2021, 88, 105906. [Google Scholar] [CrossRef] [PubMed]
- McInnes, M.D.F.; Moher, D.; Thombs, B.D.; McGrath, T.A.; Bossuyt, P.M.; Clifford, T.; Cohen, J.F.; Deeks, J.J.; Gatsonis, C.; Hooft, L.; et al. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy Studies: The PRISMA-DTA Statement. JAMA 2018, 319, 388–396. [Google Scholar] [CrossRef] [Scilit]
- Rethlefsen, M.L.; Kirtley, S.; Waffenschmidt, S.; Ayala, A.P.; Moher, D.; Page, M.J.; Koffel, J.B.; PRISMA-S Group. PRISMA-S: An extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst. Rev. 2021, 10, 39. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Salim, M.; Wåhlin, E.; Dembrower, K.; Azavedo, E.; Foukakis, T.; Liu, Y.; Smith, K.; Eklund, M.; Strand, F. External Evaluation of 3 Commercial Artificial Intelligence Algorithms for Independent Assessment of Screening Mammograms. JAMA Oncol. 2020, 6, 1581–1588. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schaffter, T.; Buist, D.S.M.; Lee, C.I.; Nikulin, Y.; Ribli, D.; Guan, Y.; Lotter, W.; Jie, Z.; Du, H.; Wang, S.; et al. Evaluation of Combined Artificial Intelligence and Radiologist Assessment to Interpret Screening Mammograms. JAMA Netw. Open 2020, 3, e200265. [Google Scholar] [CrossRef] [Scilit]
- Whiting, P.F.; Rutjes, A.W.; Westwood, M.E.; Mallett, S.; Deeks, J.J.; Reitsma, J.B.; Leeflang, M.M.G.; Sterne, J.A.C.; Bossuyt, P.M.M.; QUADAS-2 Group. QUADAS-2: A Revised Tool for the Quality Assessment of Diagnostic Accuracy Studies. Ann. Intern. Med. 2011, 155, 529–536. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, B.; Mallett, S.; Takwoingi, Y.; Davenport, C.F.; Hyde, C.J.; Whiting, P.F.; Deeks, J.J.; Leeflang, M.M.; QUADAS-C Group. QUADAS-C: A Tool for Assessing Risk of Bias in Comparative Diagnostic Accuracy Studies. Ann. Intern. Med. 2021, 174, 1592–1599. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Collins, G.S.; Moons, K.G.M.; Dhiman, P.; Riley, R.D.; Beam, A.L.; Van Calster, B.; Ghassemi, M.; Liu, X.; Reitsma, J.B.; van Smeden, M.; et al. TRIPOD+AI Statement: Updated Guidance for Reporting Clinical Prediction Models That Use Regression or Machine Learning Methods. BMJ 2024, 385, e078378. [Google Scholar] [CrossRef] [Scilit]
- Moons, K.G.M.; Damen, J.A.A.; Kaul, T.; Hooft, L.; Andaur Navarro, C.; Dhiman, P.; Beam, A.L.; Van Calster, B.; Celi, L.A.; Denaxas, S.; et al. PROBAST+AI: An updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ 2025, 388, e082505. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schünemann, H.J.; Mustafa, R.A.; Brozek, J.; Steingart, K.R.; Leeflang, M.; Murad, M.H.; Bossuyt, P.; Glasziou, P.; Jaeschke, R.; Lange, S.; et al. GRADE guidelines: 21 part 1. Study design, risk of bias, and indirectness in rating the certainty across a body of evidence for test accuracy. J. Clin. Epidemiol. 2020, 122, 129–141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schünemann, H.J.; Mustafa, R.A.; Brozek, J.; Steingart, K.R.; Leeflang, M.; Murad, M.H.; Bossuyt, P.; Glasziou, P.; Jaeschke, R.; Lange, S.; et al. GRADE guidelines: 21 part 2. Test accuracy: Inconsistency, imprecision, publication bias, and other domains for rating the certainty of evidence and presenting it in evidence profiles and summary of findings tables. J. Clin. Epidemiol. 2020, 122, 142–152. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hanley, J.A.; McNeil, B.J. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 1982, 143, 29–36. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rodriguez-Ruiz, A.; Lång, K.; Gubern-Merida, A.; Broeders, M.; Gennaro, G.; Clauser, P.; Helbich, T.H.; Chevalier, M.; Tan, T.; Mertelmeier, T.; et al. Stand-Alone Artificial Intelligence for Breast Cancer Detection in Mammography: Comparison with 101 Radiologists. J. Natl. Cancer Inst. 2019, 111, 916–922. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kizildag Yirgin, I.; Koyluoglu, Y.O.; Seker, M.E.; Ozkan Gurdal, S.; Ozaydin, A.N.; Ozcinar, B.; Cabioğlu, N.; Ozmen, V.; Aribal, E. Diagnostic Performance of AI for Cancers Registered in A Mammography Screening Program: A Retrospective Analysis. Technol. Cancer Res. Treat. 2022, 21, 15330338221075172. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Romero-Martín, S.; Elías-Cabot, E.; Raya-Povedano, J.L.; Gubern-Mérida, A.; Rodríguez-Ruiz, A.; Álvarez-Benito, M. Stand-Alone Use of Artificial Intelligence for Digital Mammography and Digital Breast Tomosynthesis Screening: A Retrospective Evaluation. Radiology 2022, 302, 535–542. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hsu, W.; Hippe, D.S.; Nakhaei, N.; Wang, P.C.; Zhu, B.; Siu, N.; Ahsen, M.E.; Lotter, W.; Sorensen, A.G.; Naeim, A.; et al. External Validation of an Ensemble Model for Automated Mammography Interpretation by Artificial Intelligence. JAMA Netw. Open 2022, 5, e2242343. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Retson, T.A.; Watanabe, A.T.; Vu, H.; Chim, C.Y. Multicenter, Multivendor Validation of an FDA-Approved Algorithm for Mammography Triage. J. Breast Imaging 2022, 4, 488–495. [Google Scholar] [CrossRef] [Scilit]
- Marinovich, M.L.; Wylie, E.; Lotter, W.; Lund, H.; Waddell, A.; Madeley, C.; Pereira, G.; Houssami, N. Artificial intelligence (AI) for breast cancer screening: BreastScreen population-based cohort study of cancer detection. EBioMedicine 2023, 90, 104498. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Riveira-Martin, M.; Rodríguez-Ruiz, A.; Martí, R.; Chevalier, M. Multi-vendor robustness analysis of a commercial artificial intelligence system for breast cancer detection. J. Med. Imaging 2023, 10, 051807. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kwon, M.R.; Chang, Y.; Ham, S.Y.; Cho, Y.; Kim, E.Y.; Kang, J.; Park, E.K.; Kim, K.H.; Kim, M.; Kim, T.S.; et al. Screening mammography performance according to breast density: A comparison between radiologists versus standalone intelligence detection. Breast Cancer Res. 2024, 26, 68. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Seker, M.E.; Koyluoglu, Y.O.; Ozaydin, A.N.; Gurdal, S.O.; Ozcinar, B.; Cabioglu, N.; Ozmen, V.; Aribal, E. Diagnostic capabilities of artificial intelligence as an additional reader in a breast cancer screening program. Eur. Radiol. 2024, 34, 6145–6157. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Larsen, M.; Olstad, C.F.; Lee, C.I.; Hovda, T.; Hoff, S.R.; Martiniussen, M.A.; Mikalsen, K.Ø.; Lund-Hanssen, H.; Solli, H.S.; Silberhorn, M.; et al. Performance of an Artificial Intelligence System for Breast Cancer Detection on Screening Mammograms from BreastScreen Norway. Radiol. Artif. Intell. 2024, 6, e230375. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hickman, S.E.; Payne, N.R.; Black, R.T.; Huang, Y.; Priest, A.N.; Hudson, S.; Kasmai, B.; Juette, A.; Nanaa, M.; Gilbert, F.J. Deep Learning Algorithms for Breast Cancer Detection in a UK Screening Cohort: As Stand-alone Readers and Combined with Human Readers. Radiology 2024, 313, e233147. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kühl, J.; Elhakim, M.T.; Stougaard, S.W.; Rasmussen, B.S.B.; Nielsen, M.; Gerke, O.; Larsen, L.B.; Graumann, O. Population-wide evaluation of artificial intelligence and radiologist assessment of screening mammograms. Eur. Radiol. 2024, 34, 3935–3946. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Graham-Knight, J.B.; Liang, P.; Lin, W.; Wright, Q.; Shen, H.; Mar, C.; Sam, J.; Rajapakshe, R. External Testing of a Commercial AI Algorithm for Breast Cancer Detection at Screening Mammography. Radiol. Artif. Intell. 2025, 7, e240287. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yamaguchi, T.; Koyama, Y.; Inoue, K.; Ban, K.; Hirokaga, K.; Kujiraoka, Y.; Okanami, Y.; Shinohara, N.; Tsunoda, H.; Uematsu, T.; et al. Development of a Deep Learning-Based Automated Diagnostic System (DLADS) for Classifying Mammographic Lesions: A First Large-Scale Multi-Institutional Clinical Trial in Japan. Breast Cancer 2025, 32, 1115–1124. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dembrower, K.; Crippa, A.; Colón, E.; Eklund, M.; Strand, F. Artificial Intelligence for Breast Cancer Detection in Screening Mammography in Sweden: A Prospective, Population-Based, Paired-Reader, Non-Inferiority Study. Lancet Digit. Health 2023, 5, e703–e711. [Google Scholar] [CrossRef] [Scilit]
- Letter, H.; Peratikos, M.; Toledano, A.; Hoffmeister, J.; Nishikawa, R.; Conant, E.; Shisler, J.; Maimone, S.; de Villegas, H.D. Use of Artificial Intelligence for Digital Breast Tomosynthesis Screening: A Preliminary Real-World Experience. J. Breast Imaging 2023, 5, 258–266. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hernström, V.; Josefsson, V.; Sartor, H.; Schmidt, D.; Larsson, A.-M.; Hofvind, S.; Andersson, I.; Rosso, A.; Hagberg, O.; Lång, K. Screening Performance and Characteristics of Breast Cancer Detected in the Mammography Screening with Artificial Intelligence Trial (MASAI): A Randomised, Controlled, Parallel-Group, Non-Inferiority, Single-Blinded, Screening Accuracy Study. Lancet Digit. Health 2025, 7, e175–e183. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chang, Y.W.; Ryu, J.K.; An, J.K.; Choi, N.; Park, Y.M.; Ko, K.H.; Han, K. Artificial Intelligence for Breast Cancer Screening in Mammography (AI-STREAM): Preliminary Analysis of a Prospective Multicenter Cohort Study. Nat. Commun. 2025, 16, 2248. [Google Scholar] [CrossRef] [Scilit]
- Lauritzen, A.D.; Lillholm, M.; Lynge, E.; Nielsen, M.; Karssemeijer, N.; Vejborg, I. Early Indicators of the Impact of Using AI in Mammography Screening for Breast Cancer. Radiology 2024, 311, e232479. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elías-Cabot, E.; Romero-Martín, S.; Raya-Povedano, J.L.; Brehl, A.K.; Álvarez-Benito, M. Impact of real-life use of artificial intelligence as support for human reading in a population-based breast cancer screening program with mammography and tomosynthesis. Eur. Radiol. 2024, 34, 3958–3966. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nepute, J.A.; Peratikos, M.; Toledano, A.Y.; Salvas, J.P.; Delks, H.; Shisler, J.L.; Hoffmeister, J.W.; Madden, C.M. Improved Breast Cancer Detection with Artificial Intelligence in a Real-World Digital Breast Tomosynthesis Screening Program. Clin. Breast Cancer 2025, 25, 808–816.e5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Eisemann, N.; Bunk, S.; Mukama, T.; Baltus, H.; Elsner, S.A.; Gomille, T.; Hecht, G.; Heywang-Köbrunner, S.; Rathmann, R.; Siegmann-Luz, K.; et al. Nationwide Real-World Implementation of AI for Cancer Detection in Population-Based Mammography Screening. Nat. Med. 2025, 31, 917–924. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sandler Rahat, H.; Friehmann, T.; Shemesh, M.D.; Tamir, S.; Atar, E.; Shochat, T.; Makori, A.; Grubstein, A. Early Results of Using AI in Mammography Screening for Breast Cancer. J. Clin. Med. 2025, 14, 7886. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Veroniki, A.A.; Jackson, D.; Viechtbauer, W.; Bender, R.; Bowden, J.; Knapp, G.; Kuss, O.; Higgins, J.P.; Langan, D.; Salanti, G. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Res. Synth. Methods 2016, 7, 55–79. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- IntHout, J.; Ioannidis, J.P.; Borm, G.F. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med. Res. Methodol. 2014, 14, 25. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Riley, R.D.; Higgins, J.P.; Deeks, J.J. Interpretation of random effects meta-analyses. BMJ 2011, 342, d549. [Google Scholar] [CrossRef] [Scilit]
- Reitsma, J.B.; Glas, A.S.; Rutjes, A.W.; Scholten, R.J.; Bossuyt, P.M.; Zwinderman, A.H. Bivariate analysis of sensitivity and specificity produces informative summary measures in diagnostic reviews. J. Clin. Epidemiol. 2005, 58, 982–990. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Deeks, J.J.; Macaskill, P.; Irwig, L. The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed. J. Clin. Epidemiol. 2005, 58, 882–893. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lång, K.; Josefsson, V.; Larsson, A.-M.; Larsson, S.; Högberg, C.; Sartor, H.; Andersson, I.; Rosso, A. Artificial Intelligence-Supported Screen Reading versus Standard Double Reading in the Mammography Screening with Artificial Intelligence Trial (MASAI): A Clinical Safety Analysis. Lancet Oncol. 2023, 24, 936–944. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elhakim, M.T.; Stougaard, S.W.; Graumann, O.; Nielsen, M.; Lång, K.; Gerke, O.; Larsen, L.B.; Rasmussen, B.S.B. Breast cancer detection accuracy of AI in an entire screening population: A retrospective, multicentre study. Cancer Imaging 2023, 23, 127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Martiniussen, M.A.; Larsen, M.; Hovda, T.; Kristiansen, M.U.; Dahl, F.A.; Eikvil, L.; Brautaset, O.; Bjørnerud, A.; Kristensen, V.; Bergan, M.B.; et al. Performance of Two Deep Learning-Based AI Models for Breast Cancer Detection and Localization on Screening Mammograms from BreastScreen Norway. Radiol. Artif. Intell. 2025, 7, e240039. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lewis, S.; Clarke, M. Forest Plots: Trying to See the Wood and the Trees. BMJ 2001, 322, 1479–1480. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Salameh, J.P.; Bossuyt, P.M.; McGrath, T.A.; Thombs, B.D.; Hyde, C.J.; Macaskill, P.; Deeks, J.J.; Leeflang, M.; A Korevaar, D.; Whiting, P.; et al. Preferred Reporting Items for Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies (PRISMA-DTA): Explanation, Elaboration, and Checklist. BMJ 2020, 370, m2632. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, I.E.; Joines, M.M.; Capiro, N.; Dawar, R.; Sears, C.; Sayre, J.; Chalfant, J.; Fischer, C.; Hoyt, A.C.; Hsu, W.; et al. Commercial Artificial Intelligence Versus Radiologists: NPV and Recall Rate in Large Population-Based Digital Mammography and Tomosynthesis Screening Mammography Cohorts. AJR Am. J. Roentgenol. 2025, 225, e2532889. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lauritzen, A.D.; Rodríguez-Ruiz, A.; von Euler-Chelpin, M.C.; Lynge, E.; Vejborg, I.; Nielsen, M.; Karssemeijer, N.; Lillholm, M. An Artificial Intelligence-based Mammography Screening Protocol for Breast Cancer: Outcome and Radiologist Workload. Radiology 2022, 304, 41–49. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oberije, C.J.G.; Currie, R.; Leaver, A.; Redman, A.; Teh, W.; Sharma, N.; Fox, G.; Glocker, B.; Khara, G.; Nash, J.; et al. Assessing artificial intelligence in breast screening with stratified results on 306,839 mammograms across geographic regions, age, breast density and ethnicity: A Retrospective Investigation Evaluating Screening (ARIES) study. BMJ Health Care Inform. 2025, 32, e101318. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xavier, D.; Miyawaki, I.; Campello Jorge, C.A.; Freitas Silva, G.B.; Lloyd, M.; Moraes, F.; Patel, B.; Batalini, F. Artificial Intelligence for Triaging of Breast Cancer Screening Mammograms and Workload Reduction: A Meta-Analysis of a Deep Learning Software. J. Med. Screen. 2023, 31, 157–165. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- van Winkel, S.L.; Peters, J.; Janssen, N.; Kroes, J.; Loehrer, E.A.; Gommers, J.; Sechopoulos, I.; de Munck, L.; Teuwen, J.; Broeders, M.; et al. AI as an independent second reader in detection of clinically relevant breast cancers within a population-based screening programme in the Netherlands: A retrospective cohort study. Lancet Digit. Health 2025, 7, 100882. [Google Scholar] [CrossRef] [Scilit]
- Lång, K.; Hofvind, S.; Rodríguez-Ruiz, A.; Andersson, I. Can Artificial Intelligence Reduce the Interval Cancer Rate in Mammography Screening? Eur. Radiol. 2021, 31, 5940–5947. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McGuinness, L.A.; Higgins, J.P.T. Risk-of-Bias VISualization (robvis): An R Package and Shiny Web App for Visualizing Risk-of-Bias Assessments. Res. Synth. Methods 2021, 12, 55–61. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dratsch, T.; Chen, X.; Rezazade Mehrizi, M.; Kloeckner, R.; Mähringer-Kunz, A.; Püsken, M.; Baeßler, B.; Sauer, S.; Maintz, D.; dos Santos, D.P. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology 2023, 307, e222176. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Marinovich, M.L.; Lotter, W.; Waddell, A.; Houssami, N. Simulated Arbitration of Discordance between Radiologists and Artificial Intelligence Interpretation of Breast Cancer Screening Mammograms. J. Med. Screen. 2024, 32, 48–52. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pinto, M.C.; Rodriguez-Ruiz, A.; Pedersen, K.; Hofvind, S.; Wicklein, J.; Kappler, S.; Mann, R.M.; Sechopoulos, I. Impact of Artificial Intelligence Decision Support Using Deep Learning on Breast Cancer Screening Interpretation with Single-View Wide-Angle Digital Breast Tomosynthesis. Radiology 2021, 300, 529–536. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Elhakim, M.T.; Stougaard, S.W.; Graumann, O.; Nielsen, M.; Gerke, O.; Larsen, L.B.; Rasmussen, B.S.B. AI-Integrated Screening to Replace Double Reading of Mammograms: A Population-Wide Accuracy and Feasibility Study. Radiol. Artif. Intell. 2024, 6, e230529. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hickman, A.J.; Gomes, S.; Warren, L.M.; Smith, N.A.S.; Shenton-Taylor, C. Assessing the generalisation of artificial intelligence across mammography manufacturers. PLoS Digit. Health 2025, 4, e0000973. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- de Vries, C.F.; Colosimo, S.J.; Staff, R.T.; Dymiter, J.A.; Yearsley, J.; Dinneen, D.; Boyle, M.; Harrison, D.J.; Anderson, L.A.; Lip, G.; et al. Impact of Different Mammography Systems on Artificial Intelligence Performance in Breast Cancer Screening. Radiol. Artif. Intell. 2023, 5, e220146. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Soltan, A.; Washington, P. Challenges in Reducing Bias Using Post-Processing Fairness for Breast Cancer Stage Classification with Deep Learning. Algorithms 2024, 17, 141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, C.J.; Wang, S.; Kuo, P.C. Improving Fairness in Chest X-Ray Interpretation Models Using Attention-Driven Masked Image Modeling. In Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society, Orlando, FL, USA, 15–19 July 2024; pp. 1–4. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ghasemi, A.; Hashtarkhani, S.; Schwartz, D.L.; Shaban-Nejad, A. Explainable artificial intelligence in breast cancer detection and risk prediction: A systematic scoping review. Cancer Innov. 2024, 3, e136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lamb, L.R.; Lehman, C.D.; Do, S.; Kim, K.; Langarica, S.; Bahl, M. Artificial Intelligence (AI)-Based Computer-Assisted Detection and Diagnosis for Mammography: An Evidence-Based Review of Food and Drug Administration (FDA)-Cleared Tools for Screening Digital Breast Tomosynthesis (DBT). AI Precis. Oncol. 2024, 1, 195–206. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pesapane, F.; Volonté, C.; Codari, M.; Sardanelli, F. Artificial intelligence as a medical device in radiology: Ethical and regulatory issues in Europe and the United States. Insights Imaging 2018, 9, 745–753. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Morales Santos, Á.; Lojo Lendoiro, S.; Rovira Cañellas, M.; Valdés Solís, P. The legal regulation of artificial intelligence in the European Union: A practical guide for radiologists. Radiologia 2024, 66, 431–446. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Uwimana, A.; Gnecco, G.; Riccaboni, M. Artificial Intelligence for Breast Cancer Detection and Its Health Technology Assessment: A Scoping Review. Comput. Biol. Med. 2024, 184, 109391. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hill, H.; Roadevin, C.; Duffy, S.; Mandrik, O.; Brentnall, A. Cost-Effectiveness of AI for Risk-Stratified Breast Cancer Screening. JAMA Netw. Open 2024, 7, e2431715. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yala, A.; Mikhael, P.G.; Strand, F.; Lin, G.; Smith, K.; Wan, Y.L.; Lamb, L.; Hughes, K.; Lehman, C.; Barzilay, R. Toward Robust Mammography-Based Models for Breast Cancer Risk. Sci. Transl. Med. 2021, 13, eaba4373. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yala, A.; Mikhael, P.G.; Strand, F.; Lin, G.; Satuluru, S.; Kim, T.; Banerjee, I.; Gichoya, J.; Trivedi, H.; Lehman, C.D.; et al. Multi-Institutional Validation of a Mammography-Based Breast Cancer Risk Model. J. Clin. Oncol. 2022, 40, 1732–1740. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Qian, X.; Pei, J.; Han, C.; Liang, Z.; Zhang, G.; Chen, N.; Zheng, W.; Meng, F.; Yu, D.; Chen, Y.; et al. A Multimodal Machine Learning Model for the Stratification of Breast Cancer Risk. Nat. Biomed. Eng. 2025, 9, 356–370. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dembrower, K.; Liu, Y.; Azizpour, H.; Eklund, M.; Smith, K.; Lindholm, P.; Strand, F. AI-Based Selection of Individuals for Supplemental MRI in Population-Based Breast Cancer Screening: The Randomized ScreenTrustMRI Trial. Nat. Med. 2024, 30, 2547–2553. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Monticciolo, D.L. Digital Breast Tomosynthesis: A Decade of Practice in Review. J. Am. Coll. Radiol. 2023, 20, 127–133. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vilmun, B.M.; Napolitano, G.; Lillholm, M.; Winkel, R.R.; Lynge, E.; Nielsen, M.; Carlsen, J.F.; von Euler-Chelpin, M.; Vejborg, I. Introduction of One-View Tomosynthesis in Population-Based Mammography Screening: Impact on Detection Rate, Interval Cancer Rate and False-Positive Rate. J. Med. Screen. 2024, 32, 28–34. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kinkar, K.K.; Fields, B.K.K.; Yamashita, M.W.; Varghese, B.A. Empowering Breast Cancer Diagnosis and Radiology Practice: Advances in Artificial Intelligence for Contrast-Enhanced Mammography. Front. Radiol. 2024, 3, 1326831. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hu, Q.; Giger, M.L. Clinical Artificial Intelligence Applications: Breast Imaging. Radiol. Clin. N. Am. 2021, 59, 1027–1043. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Finlayson, S.G.; Subbaswamy, A.; Singh, K.; Bowers, J.; Kupke, A.; Zittrain, J.; Kohane, I.S.; Saria, S. The Clinician and Dataset Shift in Artificial Intelligence. N. Engl. J. Med. 2021, 385, 283–286. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sahiner, B.; Chen, W.; Samala, R.K.; Petrick, N. Data Drift in Medical Machine Learning: Implications and Potential Remedies. Br. J. Radiol. 2023, 96, 20220878. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hwang, T.J.; Kesselheim, A.S.; Vokinger, K.N. Lifecycle Regulation of Artificial Intelligence- and Machine Learning-Based Software Devices in Medicine. JAMA 2019, 322, 2285–2286. [Google Scholar] [CrossRef] [Scilit] [PubMed]





| Study (Year) [Ref] | Country | Design | Examinations | Cancers | AI System | Comparator | Analyses |
|---|---|---|---|---|---|---|---|
| Rodriguez-Ruiz (2019) [34] | Seven countries | MRMC reader study, enriched | 2652 | 653 | Transpara | 101 radiologists | A |
| Salim (2020) [25] | Sweden | Retrospective case–control | 8805 | 739 | AI-2 (median of three vendors) | First and second reader | A, A2 |
| Schaffter (2020) [26] | USA | Retrospective cohort (DREAM) | 144,231 | 952 | Top-ranked DREAM algorithm | Radiologists | A |
| McKinney (2020) [15] | UK, USA | Retrospective validation and reader study | 28,953 | Not reported | Google Health | Radiologists | — |
| Kizildag Yirgin (2022) [35] | Türkiye | Retrospective case–control | 211 | 110 | Commercial AI system | Two radiologists | A, A2 |
| Romero-Martin (2022) [36] | Spain | Retrospective consecutive cohort | 15,999 | 113 | Transpara | Single and double reading | A |
| Hsu (2022) [37] | USA | Retrospective external validation | 37,317 | Not reported | DREAM challenge ensemble model | Radiologists | A |
| Retson (2022) [38] | USA | Multicentre multivendor case–control | 1255 | 400 | Commercial triage algorithm (cmTriage) | BCSC reader benchmark | A, A2 |
| Marinovich (2023) [39] | Australia | Retrospective population cohort | 108,970 | 760 | Commercial AI algorithm | Radiologists | A, A2 |
| Riveira-Martin (2023) [40] | Spain | Retrospective population cohort | Not reported | Not reported | Commercial AI system | Programme readers | A |
| Kwon (2024) [41] | South Korea | Retrospective population cohort | 89,855 | 143 | Lunit INSIGHT MMG | Radiologist BI-RADS | A, A2 |
| Seker (2024) [42] | Türkiye | Retrospective programme cohort | 5136 | 105 | Commercial AI system | Two-reader programme | A, A2 |
| Larsen (2024) [43] | Norway | Retrospective national cohort | 661,695 | 4917 | Transpara | Programme readers | A |
| Hickman (2024) [44] | UK | Retrospective screening cohort | 26,722 | 760 | Three commercial algorithms | Single and double reading | A2 |
| Kuhl (2024) [45] | Denmark | Retrospective regional cohort | 249,402 | 2033 | Commercial AI system | First readers | A2 |
| Graham-Knight (2025) [46] | Canada | Retrospective provincial cohort | 136,700 | Not reported | Commercial AI algorithm | Radiologists | A |
| Yamaguchi (2025) [47] | Japan | Retrospective multi-institutional | 2059 | 500 | DLADS (SE-ResNet) | Pre-specified 80% target | A, A2 |
| ScreenTrustCAD (2023) [48] | Sweden | Prospective paired-reader trial | 55,581 | 261 | Lunit INSIGHT MMG | Two radiologists | B |
| Letter (2023) [49] | USA | Concurrent site-controlled comparison | Not reported | Not reported | iCAD ProFound AI v2.0 | Sites without AI | B, C |
| MASAI (2025) [50] | Sweden | Randomised controlled trial | 105,915 | 338 | Transpara 1.7.0 | Standard double reading | B, C |
| AI-STREAM (2025) [51] | South Korea | Prospective multicentre cohort | 24,543 | 140 | AI-CAD | Single reading, no AI | B |
| Lauritzen (2024) [52] | Denmark | Before-and-after cohorts | 118,997 | 480 | Commercial AI system | Double reading before AI | B, C |
| Elias-Cabot (2024) [53] | Spain | Before-and-after cohorts | 23,996 | 108 | Commercial AI system | Double reading before AI | B, C |
| Nepute (2025) [54] | USA | Before-and-after DBT interpretations | 16,729 | 39 | Deep-learning AI support | Conventional CAD | B, C |
| PRAIM (2025) [55] | Germany | Prospective real-world implementation | 463,094 | 1747 | Commercial AI system | Standard double reading | B, C |
| Sandler Rahat (2025) [56] | Israel | Retrospective audit before and after AI | 31,176 | Not reported | iCAD version 2.0 | Screening before AI | — |
| Study (Year) [Ref] | Rule Used to Set the Operating Point | Sensitivity % | Specificity % |
|---|---|---|---|
| Salim (2020) [25] | Cut point set at mean first-reader specificity (96.6%) | 67.0 | 96.6 |
| Kizildag Yirgin (2022) [35] | Youden-optimised risk-score cut-off (34.5%) | 72.8 | 88.3 |
| Marinovich (2023) [39] | Prospective vendor-recommended threshold | 67.0 | 81.0 |
| Kwon (2024) [41] | Probability-of-malignancy cut-off 10% | 67.1 | 93.0 |
| Seker (2024) [42] | Youden-optimised cut-off (30.44) | 72.4 | 92.9 |
| Hickman (2024) [44] | Preset specificity matched to a single human reader | 58.9 | 97.9 |
| Yamaguchi (2025) [47] | Heatmap concentration-gradient cut-off 15% | 83.5 | 84.7 |
| Kuhl (2024) [45] | Cut point matched to mean first-reader specificity | 62.6 | 97.7 |
| Retson (2022) [38] | Vendor default operating point | 93.0 | 76.3 |
| Bivariate summary, 9 studies | 73.3 (64.5–80.6) | 92.4 (86.8–95.8) |
| Study (Year) [Ref] | Tier | Role of AI | CDR with AI | CDR Standard | Detection Rate Ratio (95% CI) | Recall Ratio (95% CI) |
|---|---|---|---|---|---|---|
| MASAI (2025) [50] | Randomised/paired | AI triage plus detection support | 6.40 | 5.00 | 1.29 (1.09–1.51) | 1.08 (0.99–1.17) |
| ScreenTrustCAD (2023) [48] | Randomised/paired | AI replacing one of two readers | 4.70 | 4.50 | 1.04 (1.00–1.09) | Not reported |
| AI-STREAM (2025) [51] | Randomised/paired | AI-CAD concurrent support in single reading | 5.70 | 5.01 | 1.14 (0.89–1.45) | Not reported |
| Randomised/paired subgroup | 1.13 (0.83–1.55) | |||||
| Lauritzen (2024) [52] | Non-randomised | AI triage plus decision support | 8.24 | 6.96 | 1.18 (1.04–1.35) | 0.80 (0.74–0.85) |
| Elias-Cabot (2024) [53] | Non-randomised | AI concurrent support for double reading | 9.00 | 5.80 | 1.54 (1.14–2.08) | 1.13 (1.02–1.26) |
| Nepute (2025) [54] | Non-randomised | AI concurrent support for DBT reading | 6.10 | 3.70 | 1.65 (1.06–2.58) | 0.79 (0.70–0.89) |
| PRAIM (2025) [55] | Non-randomised | AI-supported double reading, radiologist opts in | 6.70 | 5.70 | 1.18 (1.06–1.31) | 0.97 (0.94–1.02) |
| Letter (2023) [49] | Non-randomised | AI concurrent support for DBT reading | 7.30 | 5.90 | 1.30 (0.95–1.79) | 1.00 (0.80–1.30) |
| Non-randomised subgroup | 1.22 (1.08–1.37) | |||||
| Recall, all 6 reporting studies | 0.95 (0.81–1.12) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ciurescu, S.; Buciu, V.; Ilaș, D.-G.; Pârvănescu, R.; Șerban, D. Ten Years of Artificial Intelligence in Screening Mammography: A Systematic Review and Meta-Analysis of Diagnostic Accuracy and Clinical Implementation (Literature Published 2015–2025). Diagnostics 2026, 16, 3045. https://doi.org/10.3390/diagnostics16183045
Ciurescu S, Buciu V, Ilaș D-G, Pârvănescu R, Șerban D. Ten Years of Artificial Intelligence in Screening Mammography: A Systematic Review and Meta-Analysis of Diagnostic Accuracy and Clinical Implementation (Literature Published 2015–2025). Diagnostics. 2026; 16(18):3045. https://doi.org/10.3390/diagnostics16183045
Chicago/Turabian StyleCiurescu, Sebastian, Victor Buciu, Diana-Gabriela Ilaș, Raluca Pârvănescu, and Denis Șerban. 2026. "Ten Years of Artificial Intelligence in Screening Mammography: A Systematic Review and Meta-Analysis of Diagnostic Accuracy and Clinical Implementation (Literature Published 2015–2025)" Diagnostics 16, no. 18: 3045. https://doi.org/10.3390/diagnostics16183045
APA StyleCiurescu, S., Buciu, V., Ilaș, D.-G., Pârvănescu, R., & Șerban, D. (2026). Ten Years of Artificial Intelligence in Screening Mammography: A Systematic Review and Meta-Analysis of Diagnostic Accuracy and Clinical Implementation (Literature Published 2015–2025). Diagnostics, 16(18), 3045. https://doi.org/10.3390/diagnostics16183045

