Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models
Abstract
1. Introduction
- Development of a unified reproducible pipeline for mammogram classification;
- Systematic evaluation of the effects of preprocessing, augmentation, and model architecture on performance;
- Demonstration of the advantages of combining general-purpose models with specialized, domain-specific pretrained models;
- Establishing a practical foundation for building scalable and transparent clinical decision-support systems.
2. Related Work
2.1. Data Quality and Preprocessing
2.2. Augmentation and Synthetic Data
2.3. Model Architectures and Pretraining
2.4. Gaps in Existing Research
- Most studies rely on single datasets, reducing the external validity of their results.
- Preprocessing, metadata standardization, and augmentation are tested in isolation, without controlled comparison across multiple datasets.
- There is no systematic benchmarking of different neural network architectures, including vision-language and foundation models as well as CNNs, under harmonized conditions.
3. Materials and Methods
3.1. Description of Data Used for Pipeline Training and Testing
- VinDr-Mammo is one of the largest and most recent datasets and was one of the key sources in this work. Class mapping in this dataset was based on the BI-RADS category.
- INBreast is a high-quality dataset. Class mapping in this dataset was based on the BI-RADS category.
- CMMD (Chinese Mammography Database), published in 2021, provides DICOM data and includes biopsy-confirmed binary cancer labels.
- CBIS-DDSM is a revised version of the classic DDSM dataset (Digital Database for Screening Mammography, 1999). The following class mapping was used: benign cases were mapped from NO_OBJECT, BENIGN_WITHOUT_CALLBACK, and BENIGN, and malignant cases were mapped from MALIGNANT.
- MosMed contains mammograms collected from 2018 to 2020. This dataset originally included a patient-level binary division into benign and malignant cases. During testing on MosMed, image-level predictions were aggregated to the patient level using max pooling across view and side.
3.2. Preprocessing and Cleaning of Anomalous Data
3.3. Selection of the Optimal Augmentation Set
3.4. Investigated Vision-Language Models
- OpenAI CLIP [42] is an initial vision-language model trained on 400 million image–text pairs. It is not specialized for any particular domain and is widely used across fields. Its visual encoder used the Vision Transformer architecture ViT-L/14. The input image resolution was limited to 336 × 336 px.
- Microsoft BiomedCLIP [43] is an adapted CLIP version with an enlarged Vision Transformer compared with the baseline version. It was trained on more than 4 million biomedical images and accompanying texts extracted from the PubMed Central Open Access Subset. ViT-B/16 [44] was used as the image encoder. The input image resolution was limited to 224 × 224 px.
- Mammo-CLIP [31] is a specialized CLIP version. Mammo-CLIP uses EfficientNet(EN)-B5 with ImageNet-pretrained weights as the image encoder. The Mammo-CLIP model was trained on more than 25,000 screening mammography images with textual descriptions of domain parameters, including BI-RADS annotations and imaging patterns. The input image resolution was 1520 × 912 px.
3.5. Software and Statistical Processing Methods
4. Results
4.1. Pooled Dataset
4.2. DICOM Intensity Analysis
4.3. Selection of an Effective Augmentation Set
4.4. Performance of Baseline Models
5. Discussion
5.1. Key Factors in Model Accuracy: Resolution and Domain Adaptation
5.2. Limitations of Transformers and Advantages of Convolutional Networks
5.3. Confirmation of the Hypothesis on the Effectiveness of VLMs
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AUROC | Area Under the Receiver Operating Characteristic curve |
| BI-RADS | Breast Imaging Reporting and Data System |
| ACR | American College of Radiology |
| VICC | Four open datasets: VinDr-Mammo, INBreast, CMMD, and CBIS-DDSM |
| CNN | Convolutional neural networks |
Appendix A
| Horizontal Flip | Shift Scale Rotate | Sharpen | Gauss Noise | Advanced Blur | Grid Dropout | Grid Distortion | Pixel Dropout | Coarse Dropout | Hue Saturation Value | Random Rotate 90 | Optical Distortion | Random Shadow | Summary Metrics | |||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| p = 0.5 | p = 1 | 0.0001 | 0.001 | 0.01 | ||||||||||||
| ✓ | ✓ | 0.840 ± 0.000 | ||||||||||||||
| ✓ | ✓ | 0.805 ± 0.012 | ||||||||||||||
| ✓ | ✓ | 0.817 ± 0.006 | ||||||||||||||
| ✓ | ✓ | 0.802 ± 0.009 | ||||||||||||||
| ✓ | ✓ | 0.820 ± 0.007 | ||||||||||||||
| ✓ | ✓ | ✓ | 0.792 ± 0.012 | |||||||||||||
| ✓ | ✓ | ✓ | 0.786 ± 0.000 | |||||||||||||
| ✓ | ✓ | ✓ | 0.820 ± 0.012 | |||||||||||||
| ✓ | ✓ | ✓ | 0.810 ± 0.004 | |||||||||||||
| ✓ | ✓ | ✓ | ✓ | 0.829 ± 0.004 | ||||||||||||
| ✓ | ✓ | ✓ | 0.797 ± 0.001 | |||||||||||||
| ✓ | ✓ | ✓ | 0.826 ± 0.000 | |||||||||||||
| ✓ | ✓ | ✓ | ✓ | 0.819 ± 0.001 | ||||||||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | 0.850 ± 0.009 | |||||||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.844 ± 0.003 | ||||||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.847 ± 0.000 | ||||||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.847 ± 0.001 | ||||||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.788 ± 0.010 | |||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.840 ± 0.006 | |||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.841 ± 0.004 | |||||||||
| ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | 0.846 ± 0.002 | |||||||||
References
- McKinney, S.M.; Sieniek, M.; Godbole, V.; Godwin, J.; Antropova, N.; Ashrafian, H.; Back, T.; Chesus, M.; Corrado, G.S.; Darzi, A.; et al. International evaluation of an AI system for breast cancer screening. Nature 2020, 577, 89–94. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lehman, C.D.; Yala, A.; Schuster, T.; Dontchos, B.; Bahl, M.; Swanson, K.; Barzilay, R. Mammographic breast density assessment using deep learning: Clinical implementation. Radiology 2019, 290, 52–58. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ibragimov, A.; Senotrusova, S.; Markova, K.; Karpulevich, E.; Ivanov, A.; Tyshchuk, E.; Grebenkina, P.; Stepanova, O.; Sirotskaya, A.; Kovaleva, A.; et al. Deep semantic segmentation of angiogenesis images. Int. J. Mol. Sci. 2023, 24, 1102. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, L. Mammography with deep learning for breast cancer detection. Front. Oncol. 2024, 14, 1281922. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rodriguez-Ruiz, A.; Lång, K.; Gubern-Merida, A.; Broeders, M.; Gennaro, G.; Clauser, P.; Helbich, T.H.; Chevalier, M.; Tan, T.; Mertelmeier, T.; et al. Stand-alone artificial intelligence for breast cancer detection in mammography: Comparison with 101 radiologists. JNCI J. Natl. Cancer Inst. 2019, 111, 916–922. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Amin, A.; U, D.A.; Koteshwara, P.; P C, S.; Mathew, S. A systematic literature review on mammography: Deep learning techniques for breast cancer detection with global and Asian perspectives. BMC Cancer 2025, 25, 1627. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hunde, B.R.; Woldeyohannes, A.D. Future prospects of computer-aided design (CAD)—A review from the perspective of artificial intelligence (AI), extended reality, and 3D printing. Results Eng. 2022, 14, 100478. [Google Scholar] [CrossRef] [Scilit]
- Larsen, M.; Aglen, C.F.; Lee, C.I.; Hoff, S.R.; Lund-Johansen, M.; Hofvind, S. AI performance by mammographic density in a retrospective cohort study of 99,489 participants in BreastScreen Norway. Eur. Radiol. 2024, 10, 6298–6308. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kretz, T.; Muller, K.R.; Schaeffter, T.; Elster, C. Mammography image quality assurance using deep learning. IEEE Trans. Biomed. Eng. 2020, 67, 3317–3326. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Oza, P.; Sharma, P.; Patel, S.; Adedoyin, F.; Bruno, A. Image augmentation techniques for mammogram analysis. J. Imaging 2022, 8, 141. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ibragimov, A.; Senotrusova, S.; Litvinov, A.; Ushakov, E.; Karpulevich, E.; Markin, Y. MamT4: Multi-View Attention Networks for Mammography Cancer Classification. In Proceedings of the IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), Osaka, Japan, 2–4 July 2024; pp. 1965–1970. [Google Scholar] [CrossRef] [Scilit]
- Kebaili, A.; Lapuyade-Lahorgue, J.; Ruan, S. Deep learning approaches for data augmentation in medical imaging: A review. J. Imaging 2023, 9, 81. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ushakov, E.; Naumov, A.; Fomberg, V.; Vishnyakova, P.; Asaturova, A.; Badlaeva, A.; Tregubova, A.; Karpulevich, E.; Sukhikh, G.; Fatkhudinov, T. EndoNet: A Model for the Automatic Calculation of H-Score on Histological Slides. Informatics 2023, 10, 90. [Google Scholar] [CrossRef] [Scilit]
- Makarchuk, A.; Asaturova, A.; Ushakov, E.; Tregubova, A.; Badlaeva, A.; Tabeeva, G.; Karpulevich, E.; Markin, Y. Artificial Intelligence (AI) Solution for Plasma Cells Detection. Program. Comput. Softw. 2023, 49, 873–880. [Google Scholar] [CrossRef] [Scilit]
- Al-Mnayyis, A.M.; Gharaibeh, H.; Amin, M.; Anakreh, D.; Akhdar, H.F.; Alshdaifat, E.H.; Nahar, K.M.O.; Nasayreh, A.; Gharaibeh, M.; Alsalman, N.; et al. (KAUH-BCMD) dataset: Advancing mammographic breast cancer classification with multi-fusion preprocessing and residual depth-wise network. Front. Big Data 2025, 8, 1529848. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hsu, Y. Beyond double reading: Multiple deep learning models enhancing radiologist-led breast screening. Radiol. Artif. Intell. 2025, 7, e250125. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Branco, P.E.S.C.; Franco, A.H.S.; Oliveira, A.P.; Carneiro, I.M.C.; Carvalho, L.M.C.; Souza, J.I.N.; Leandro, D.R. Artificial intelligence in mammography: A systematic review of external validation. Rev. Bras. Ginecol. Obstet. 2024, 46, E-rbgo71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gao, S.; Liu, J.; Li, L.; Yang, D.; Miao, Y.; Zhang, X.; Han, Q.; Shi, Y.; Wu, J.; Zhang, K. Application of deep learning technology in breast cancer: A systematic review of segmentation, detection, and classification approaches. Biomed. Eng. Online 2026, 25, 19. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bhalla, A.S.; Das, A.; Naranje, P.; Irodi, A.; Raj, V.; Goyal, A. Imaging protocols for CT chest: A recommendation. Indian J. Radiol. Imaging 2019, 29, 236–246. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gao, Y.E.; Lin, J.; Zhou, Y.; Lin, R. The application of traditional machine learning and deep learning techniques in mammography: A review. Front. Oncol. 2023, 13, 1213045. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mehrabi, M.; Salek, N. Enhancing diagnostic accuracy in breast cancer: Integrating novel machine learning approaches with enhanced image preprocessing for improved mammography analysis. Pol. J. Radiol. 2024, 89, E573. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Avci, H.; Karakaya, J. A novel medical image enhancement algorithm for breast cancer detection on mammography images using machine learning. Diagnostics 2023, 13, 348. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Garrucho, L.; Kushibar, K.; Jouide, S.; Diaz, O.; Igual, L.; Lekadir, K. Domain generalization in deep learning based mass detection in mammography: A large-scale multi-center study. Artif. Intell. Med. 2022, 132, 102386. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Panambur, A.B.; Yu, H.; Bhat, S.; Madhu, P.; Bayer, S.; Maier, A. Attention-guided erasing: Novel augmentation method for enhancing downstream breast density classification. In BVM Workshop; Springer Fachmedien Wiesbaden: Wiesbaden, Germany, 2024; pp. 13–18. [Google Scholar] [CrossRef] [Scilit]
- Montoya-del-Angel, R.; Sam-Millan, K.; Vilanova, J.C.; Marti, R. MAM-E: Mammographic Synthetic Image Generation with Diffusion Models. Sensors 2024, 24, 2076. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dhaliwal, B.K. Automated Breast Cancer Detection Systems: A Systematic Review of Machine Learning and Deep Learning Techniques (2015–2025). In Proceedings of the 2025 2nd International Conference on Artificial Intelligence for Innovations in Healthcare Industries (ICAIIHI), Raipur, India, 4–5 December 2025; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Rajpurkar, P.; Chen, E.; Banerjee, O.; Topol, E.J. AI in health and medicine. Nat. Med. 2022, 28, 31–38. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ameen, M. Explainable Mammogram Analysis with EfficientNetV2 and Grad-CAM++ for Robust Cancer Diagnosis. Diagnostics 2025, 16, 105. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hasan, M. Deep Learning for Breast Cancer Detection: Comparative Analysis of ConvNeXT and EfficientNet. In 2024 27th International Conference on Computer and Information Technology (ICCIT); IEEE: Piscataway, NJ, USA, 2024; pp. 1387–1391. [Google Scholar]
- Clancy, K.; Aboutalib, S.; Mohamed, A.; Sumkin, J.; Wu, S. Deep learning pre-training strategy for mammogram image classification: An evaluation study. J. Digit. Imaging 2020, 33, 1257–1265. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ghosh, S.; Poynton, C.B.; Visweswaran, S.; Batmanghelich, K. Mammo-CLIP: A vision language foundation model to enhance data efficiency and robustness in mammography. In International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer Nature Switzerland: Cham, Switzerland, 2024; pp. 632–642. [Google Scholar]
- Nguyen, H.T.; Nguyen, H.Q.; Pham, H.H.; Lam, K.; Le, L.T.; Dao, M.; Vu, V. VinDr-Mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography. Sci. Data 2023, 10, 277. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ines, C.M.; Igor, A.; Ines, D.; Antonio, C.; Maria, J.C.; Jaime, S.C. INbreast: Toward a full-field digital mammographic database. Acad. Radiol. 2012, 19, 236–248. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cai, H.; Wang, J.; Dan, T.; Li, J.; Fan, Z.; Yi, W.; Cui, C.; Jiang, X.; Li, L. An Online Mammography Database with Biopsy Confirmed Types. Sci. Data 2023, 10, 123. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, R.S.; Gimenez, F.; Hoogi, A.; Miyake, K.K.; Gorovoy, M.; Rubin, D.L. A curated mammography data set for use in computer-aided detection and diagnosis research. Sci. Data 2017, 4, 170177. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- MosMedData. Available online: https://mosmed.ai/datasets/datasets/mosmeddata-mmg-s-nalichiem-i-otsutstviem-priznakov-zlokachestvennih-novoobrazovanii-molochnoi-zhelezi-obogaschennii-klinicheskoi-informatsiei-v1/ (accessed on 25 August 2025).
- MosMedData. Available online: https://mosmed.ai/datasets/datasets/mosmeddata-mmg-s-nalichiem-i-otsutstviem-priznakov-zlokachestvennih-novoobrazovanii-molochnoi-zhelezi-obogaschennii-klinicheskoi-informatsiei-v2/ (accessed on 25 August 2025).
- Magny, S.J.; Shikhman, R.; Keppke, A.L. Breast Imaging Reporting and Data System. In StatPearls [Internet]; StatPearls Publishing: Treasure Island, FL, USA, 2023. [Google Scholar] [PubMed]
- Buslaev, A.; Iglovikov, V.I.; Khvedchenya, E.; Parinov, A.; Druzhinin, M.; Kalinin, A.A. Albumentations: Fast and flexible image augmentations. Information 2020, 11, 125. [Google Scholar] [CrossRef] [Scilit]
- Albumentations Development Team. Albumentations (Version 2.0.8) [Python Package]. PyPI. 2025. Available online: https://pypi.org/project/albumentations/2.0.8/ (accessed on 7 July 2026).
- Tan, M.; Le, Q. EfficientNet: Rethinking model scaling for convolutional neural networks. In International Conference on Machine Learning; PMLR: Los Angeles, CA, USA, 2019; pp. 6105–6114. [Google Scholar] [CrossRef] [Scilit]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In International Conference on Machine Learning; PMLR: Honolulu, HI, USA, 2021; pp. 8748–8763. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Xu, Y.; Usuyama, N.; Xu, H.; Bagga, J.; Tinn, R.; Preston, S.; Rao, R.; Wei, M.; Valluri, N.; et al. Biomedclip: A multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs. arXiv 2023, arXiv:2303.00915. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Wu, Z.; Agarwal, D.; Sun, J. Medclip: Contrastive learning from unpaired medical images and text. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 3876–3887. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Online, 11–17 October 2021; pp. 10012–10022. [Google Scholar] [CrossRef] [Scilit]
- Zech, J.R.; Badgeley, M.A.; Liu, M.; Costa, A.B.; Titano, J.J.; Oermann, E.K. Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study. PLoS Med. 2018, 15, e1002683. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Litvinov, A.; Ushakov, E.; Senotrusova, S.; Lukianov, K.; Markin, Y.; Mikhailova, L.; Karpulevich, E. Clinically Aware Learning: Ordinal Loss Improves Medical Image Classifiers. J. Clin. Med. 2026, 15, 365. [Google Scholar] [CrossRef] [Scilit] [PubMed]




| Dataset | Total Images | Pathology-Based Labels | Cancer Annotation | Cancer Fraction (%) 1 | Density | Finding Category |
|---|---|---|---|---|---|---|
| VinDr | 20,000 | 0 | BI-RADS | 4.9 | + | + |
| INBreast | 388 | 0 | BI-RADS | 25.9 | + | + |
| CMMD | 3744 | 2632 | Binary labels | 70 | - | + |
| CBIS-DDSM | 3103 | 1364 | Lesion type | 44 | + | + |
| MosMed | 779 | 391 | Binary labels | 50 | − | − |
| Filtered Dataset | Threshold | AUROC | F1 |
|---|---|---|---|
| Baseline | - | 0.808 ± 0.017 | 0.747 ± 0.009 |
| Without metadata anomalies | - | 0.818 ± 0.018 | 0.748 ±0.021 |
| Without white and black anomalies | Remove 1% | 0.810 ± 0.027 | 0.751 ± 0.028 |
| Remove 5% | 0.807 ± 0.025 | 0.746 ± 0.024 | |
| Without black anomalies | Remove 1% | 0.795 ± 0.043 | 0.74 ± 0.031 |
| Remove 5% | 0.823 ± 0.009 | 0.764 ± 0.012 | |
| Without white anomalies | Remove 1% | 0.799 ± 0.024 | 0.755 ± 0.022 |
| Remove 5% | 0.809 ± 0.039 | 0.758 ± 0.036 |
| Dataset | Metadata Anomalies | Black Anomalies | White Anomalies | Total Filtered | Train | Validation | Test |
|---|---|---|---|---|---|---|---|
| VinDr | 0 | 919 | 921 | 1840 | 12,711 | 1817 | 3632 |
| INBreast | 0 | 18 | 18 | 36 | 193 | 50 | 109 |
| CMMD | 0 | 188 | 189 | 377 | 1874 | 470 | 1023 |
| CBIS-DDSM | 272 | 156 | 159 | 587 | 1709 | 285 | 522 |
| Augmentation | VICC | Mosmed |
|---|---|---|
| Baseline | 0.815 ± 0.011 | 0.844 ± 0.008 |
| Advanced Blur | 0.810 ± 0.003 | 0.843 ± 0.005 |
| Affine Independent Axes | Not tested | 0.843 ± 0.01 |
| Clahe | 0.819 ± 0.017 | 0.833 ± 0.011 |
| Coarse Dropout | 0.816 ± 0.001 | 0.84 ± 0.002 |
| Gauss Noise | 0.791 ± 0.009 | 0.844 ± 0.003 |
| Grid Distortion | 0.826 ± 0.005 | 0.834 ± 0.009 |
| Grid Dropout | 0.816 ± 0.005 | 0.843 ± 0.006 |
| Horizontal Flip | 0.827 ± 0.002 | 0.847 ± 0.007 |
| Hue Saturation Value | 0.810 ± 0.007 | 0.855 ± 0.003 |
| Mask | Not tested | 0.852 ± 0.016 |
| Optical Distortion | 0.818 ± 0.002 | 0.835 ± 0.009 |
| Pixel Dropout | 0.822 ± 0.001 | 0.84 ± 0.005 |
| Random Brightness Contrast | Not tested | 0.852 ± 0.017 |
| Random Gamma | Not tested | 0.837 ± 0.001 |
| Random Rotate 90 | 0.829 ± 0.002 | 0.844 ± 0.01 |
| Random Tone Curve | Not tested | 0.840 ± 0.012 |
| Sharpen | 0.814 ± 0.003 | 0.846 ± 0.001 |
| Vertical Flip | Not tested | 0.847 ± 0.009 |
| “Best” Set | “Extended” Set | “All” Set | ||||
|---|---|---|---|---|---|---|
| Augmentation | VICC | Mosmed | VICC | Mosmed | VICC | Mosmed |
| Baseline | 0.850 ± 0.009 | 0.888 ± 0.006 | 0.841 ± 0.004 | 0.9 ± 0.006 | 0.788 ± 0.01 | 0.898 ± 0.004 |
| Advanced Blur | + | |||||
| Affine Independent Axes | + | + | ||||
| Clahe | + | + | ||||
| Coarse Dropout | + | |||||
| Gauss Noise | + | |||||
| Grid Distortion | + | + | ||||
| Grid Dropout | + | + | ||||
| Horizontal Flip | + | + | + | |||
| Hue Saturation Value | + | + | ||||
| Mask | + | + | + | |||
| Optical Distortion | + | + | ||||
| Pixel Dropout | + | + | ||||
| Random Brightness Contrast | + | + | + | |||
| Random Gamma | + | + | ||||
| Random Rotate 90 | + | + | ||||
| Random Tone Curve | + | |||||
| Sharpen | + | + | + | |||
| Vertical Flip | + | + | + | |||
| Model | AUROC | SE | 95% CI | Resolution |
|---|---|---|---|---|
| EfficientNet | 0.934 | 0.012 | [0.926, 0.972] | 1520 × 912 |
| Mammo-CLIP | 0.949 | 0.013 | [0.908, 0.960] | 1520 × 912 |
| MedCLIP | 0.856 | 0.019 | [0.818, 0.894] | 512 × 512 |
| BiomedCLIP | 0.594 | 0.028 | [0.592, 0.700] | 224 × 224 |
| OpenAI CLIP | 0.646 | 0.029 | [0.538, 0.650] | 336 × 336 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ushakov, E.; Zimina, S.; Litvinov, A.; Senotrusova, S.; Lukianov, K.; Gevorkyan, T.G.; Karpulevich, E. Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models. Informatics 2026, 13, 125. https://doi.org/10.3390/informatics13080125
Ushakov E, Zimina S, Litvinov A, Senotrusova S, Lukianov K, Gevorkyan TG, Karpulevich E. Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models. Informatics. 2026; 13(8):125. https://doi.org/10.3390/informatics13080125
Chicago/Turabian StyleUshakov, Egor, Sofya Zimina, Arsenii Litvinov, Sofia Senotrusova, Kirill Lukianov, Tigran G. Gevorkyan, and Evgeny Karpulevich. 2026. "Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models" Informatics 13, no. 8: 125. https://doi.org/10.3390/informatics13080125
APA StyleUshakov, E., Zimina, S., Litvinov, A., Senotrusova, S., Lukianov, K., Gevorkyan, T. G., & Karpulevich, E. (2026). Optimizing the Mammography AI Pipeline: From Data Filtering to Vision-Language Models. Informatics, 13(8), 125. https://doi.org/10.3390/informatics13080125

