CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity
Abstract
1. Introduction
2. Fundamentals of Convolutional Neural Networks
3. Applications of Convolutional Neural Networks to Four Biomedical Datasets
3.1. Mammography Dataset (CBIS-DDSM)
3.2. Chest X-Ray Dataset
3.3. Brain Tumor Dataset
3.4. ISIC Challenge 2017
4. Experimental Design and Sample-Size Reduction Strategy
4.1. Fixed Hyperparameters
4.2. CNN Architectures Considered
4.3. Systematic Variation of Training-Set Size
5. Sample-Size Sensitivity and Overfitting Analysis
5.1. Generalization Gap and Overfitting Onset
5.2. Architecture Ranking Reliability
5.3. Test Performance Monotonicity
6. Results
- (i)
- Overfitting-onset threshold
- (ii)
- Ranking-stabilization threshold
- (iii)
- Monotonicity of test performance
7. Discussion
8. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Thakur, G.K.; Thakur, A.; Kulkarni, S.; Khan, N.; Khan, S. Deep learning approaches for medical image analysis and diagnosis. Cureus 2024, 16, e59507. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khan, A.; Sohail, A.; Zahoora, U.; Qureshi, A.S. A survey of the recent architectures of deep convolutional neural networks. Artif. Intell. Rev. 2020, 53, 5455–5516. [Google Scholar] [CrossRef] [Scilit]
- Razzak, M.I.; Naz, S.; Zaib, A. Deep learning for medical image processing: Overview, challenges and the future. In Classification in BioApps: Automation of Decision Making; Springer: Cham, Switzerland, 2017; pp. 323–350. [Google Scholar]
- Tsuneki, M. Deep learning models in medical image analysis. J. Oral Biosci. 2022, 64, 312–320. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Reza, M.H.; Sibly, M.N.B.; Rabbani, S.G.; Sadi, S.H.; Ahamed, M.F.; Shafi, F.B.; Sarmun, R.; Chowdhury, M.E. A comprehensive review of convolutional neural networks: Foundations, enhancements and applications. Neural Comput. Appl. 2026, 38, 56. [Google Scholar] [CrossRef] [Scilit]
- Akhand, A.S. A Comparative Study of Custom CNNs, Pre-trained Models, and Transfer Learning Across Multiple Visual Datasets. arXiv 2026, arXiv:2601.02246. [Google Scholar]
- Alzubaidi, L.; Zhang, J.; Humaidi, A.J.; Al-Dujaili, A.; Duan, Y.; Al-Shamma, O.; Santamaría, J.; Fadhel, M.A.; Al-Amidie, M.; Farhan, L. Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions. J. Big Data 2021, 8, 53. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abhisheka, B.; Biswas, S.K.; Purkayastha, B.; Das, D.; Escargueil, A. Recent trend in medical imaging modalities and their applications in disease diagnosis: A review. Multimed. Tools Appl. 2024, 83, 43035–43070. [Google Scholar] [CrossRef] [Scilit]
- Iqbal, S.; Qureshi, A.N.; Ullah, A.; Li, J.; Mahmood, T. Improving the robustness and quality of biomedical cnn models through adaptive hyperparameter tuning. Appl. Sci. 2022, 12, 11870. [Google Scholar] [CrossRef] [Scilit]
- Bertrand, H. Hyper-Parameter Optimization in Deep Learning and Transfer Learning: Applications to Medical Imaging. Ph.D. Thesis, Université Paris Saclay (COmUE), Paris, France, 2019. [Google Scholar]
- Li, Y.; Gu, S.; Gool, L.V.; Timofte, R. Learning filter basis for convolutional neural network compression. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5623–5632. [Google Scholar]
- Ben Braiek, H.; Khomh, F. Testing feedforward neural networks training programs. ACM Trans. Softw. Eng. Methodol. 2023, 32, 105. [Google Scholar] [CrossRef] [Scilit]
- Ding, Y.; Huang, Z.; Shou, X.; Guo, Y.; Sun, Y.; Gao, J. Architecture-aware learning curve extrapolation via graph ordinary differential equation. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Volume 39, pp. 16289–16297. [Google Scholar]
- Le, M.; Nguyen, N.; Luong, N.H. Efficacy of Neural Prediction-Based Zero-Shot NAS. arXiv 2023, arXiv:2308.16775. [Google Scholar]
- Park, M. Data proxy generation for fast and efficient neural architecture search. arXiv 2019, arXiv:1911.09322. [Google Scholar]
- Nakkiran, P.; Kaplun, G.; Bansal, Y.; Yang, T.; Barak, B.; Sutskever, I. Deep double descent: Where bigger models and more data hurt. J. Stat. Mech. Theory Exp. 2021, 2021, 124003. [Google Scholar] [CrossRef] [Scilit]
- Nakkiran, P. More data can hurt for linear regression: Sample-wise double descent. arXiv 2019, arXiv:1912.07242. [Google Scholar]
- Stengel-Eskin, E.; Platanios, E.A.; Pauls, A.; Thomson, S.; Fang, H.; Van Durme, B.; Eisner, J.; Su, Y. When more data hurts: A troubling quirk in developing broad-coverage natural language understanding systems. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, 7–11 December 2022; pp. 11473–11487. [Google Scholar]
- Min, Y.; Chen, L.; Karbasi, A. The curious case of adversarially robust models: More data can help, double descend, or hurt generalization. In Proceedings of the Uncertainty in Artificial Intelligence; PMLR: Cambridge, MA, USA, 2021; pp. 129–139. [Google Scholar]
- Cho, J.; Lee, K.; Shin, E.; Choy, G.; Do, S. How much data is needed to train a medical image deep learning system to achieve necessary high accuracy? arXiv 2015, arXiv:1511.06348. [Google Scholar]
- Rokem, A.; Wu, Y.; Lee, A. Assessment of the need for separate test set and number of medical images necessary for deep learning: A sub-sampling study. BioRxiv 2017. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Zhang, H.; Jin, Y. A survey on computationally efficient neural architecture search. J. Autom. Intell. 2022, 1, 100002. [Google Scholar] [CrossRef] [Scilit]
- Kot, A. Minimum Fidelity for Reliable Architecture Ranking in Bayesian Nas for Object Detection. Electron. Control Syst. 2026, 88, 54. [Google Scholar] [CrossRef] [Scilit]
- Yu, Q. Deep Learning-based Multi-class Classification of Breast Cancer Pathology: A Comparative Study on Full-field Digital Mammograms Using the CBIS-DDSM Dataset. Bachelor’s Thesis, Satakunta University of Applied Sciences (SAMK), Pori, Finland, 2025. [Google Scholar]
- Gómez-Guzmán, M.A.; Jiménez-Beristaín, L.; García-Guerrero, E.E.; López-Bonilla, O.R.; Tamayo-Perez, U.J.; Esqueda-Elizondo, J.J.; Palomino-Vizcaino, K.; Inzunza-González, E. Classifying brain tumors on magnetic resonance imaging by using convolutional neural networks. Electronics 2023, 12, 955. [Google Scholar] [CrossRef] [Scilit]
- Das, H.S.; Das, A.; Neog, A.; Mallik, S.; Bora, K.; Zhao, Z. Breast cancer detection: Shallow convolutional neural network against deep convolutional neural networks based approach. Front. Genet. 2023, 13, 1097207. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zahoor, S.; Shoaib, U.; Lali, I.U. Breast cancer mammograms classification using deep neural network and entropy-controlled whale optimization algorithm. Diagnostics 2022, 12, 557. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kassem, M.A.; Hosny, K.M.; Fouad, M.M. Skin lesions classification into eight classes for ISIC 2019 using deep convolutional neural network and transfer learning. IEEE Access 2020, 8, 114822–114832. [Google Scholar] [CrossRef] [Scilit]
- DeVries, T.; Ramachandram, D. Skin lesion classification using deep multi-scale convolutional neural networks. arXiv 2017, arXiv:1703.01402. [Google Scholar]
- Katsamenis, I.; Protopapadakis, E.; Voulodimos, A.; Doulamis, A.; Doulamis, N. Transfer learning for COVID-19 pneumonia detection and classification in chest X-ray images. In Proceedings of the 24th Pan-Hellenic Conference on Informatics, Athens, Greece, 20–22 November 2020; pp. 170–174. [Google Scholar]
- Ahmed, W.S.; Karim, A.A.A. The impact of filter size and number of filters on classification accuracy in CNN. In Proceedings of the 2020 International Conference on Computer Science and Software Engineering (CSASE); IEEE: New York, NY, USA, 2020; pp. 88–93. [Google Scholar]
- Sawyer-Lee, R.; Gimenez, F.; Hoogi, A.; Rubin, D. Curated breast imaging subset of digital database for screening mammography (CBIS-DDSM). Sci. Data 2017, 4, 1–9. [Google Scholar] [CrossRef] [Scilit]
- Cheng, J. Brain Tumor Dataset, Version v8; Figshare: Cambridge, MA, USA, 2017. [CrossRef]
- Kermany, D.S.; Goldbaum, M.; Cai, W.; Valentim, C.C.; Liang, H.; Baxter, S.L.; McKeown, A.; Yang, G.; Wu, X.; Yan, F.; et al. Identifying medical diagnoses and treatable diseases by image-based deep learning. Cell 2018, 172, 1122–1131. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Codella, N.C.F.; Gutman, D.; Celebi, M.E.; Helba, B.; Marchetti, M.A.; Dusza, S.; Kalloo, A.; Liopyris, K.; Mishra, N.; Kittler, H.; et al. Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC). arXiv 2017, arXiv:1710.05006. [Google Scholar]
- Serena Low, W.C.; Chuah, J.H.; Tee, C.A.T.; Anis, S.; Shoaib, M.A.; Faisal, A.; Khalil, A.; Lai, K.W. An Overview of Deep Learning Techniques on Chest X-Ray and CT Scan Identification of COVID-19. Comput. Math. Methods Med. 2021, 2021, 5528144. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vineth Ligi, S.; Kundu, S.S.; Kumar, R.; Narayanamoorthi, R.; Lai, K.W.; Dhanalakshmi, S. Radiological Analysis of COVID-19 Using Computational Intelligence: A Broad Gauge Study. J. Healthc. Eng. 2022, 2022, 5998042. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khan, M.S.I.; Rahman, A.; Debnath, T.; Karim, M.R.; Nasir, M.K.; Band, S.S.; Mosavi, A.; Dehzangi, I. Accurate brain tumor detection using deep convolutional neural network. Comput. Struct. Biotechnol. J. 2022, 20, 4733–4745. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kaur, R.; GholamHosseini, H.; Sinha, R.; Lindén, M. Automatic lesion segmentation using atrous convolutional deep neural networks in dermoscopic skin cancer images. BMC Med. Imaging 2022, 22, 103. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abiwinanda, N.; Hanif, M.; Hesaputra, S.T.; Handayani, A.; Mengko, T.R. Brain tumor classification using convolutional neural network. In Proceedings of the World Congress on Medical Physics and Biomedical Engineering, Prague, Czech Republic, 3–8 June 2018; Springer: Berlin/Heidelberg, Germany, 2018; Volume 1, pp. 183–189. [Google Scholar]
- Kermany, D.; Zhang, K.; Goldbaum, M. Labeled Optical Coherence Tomography (OCT) and Chest X-Ray Images for Classification; Version V2; Mendeley Data: London, UK, 2018. [Google Scholar] [CrossRef]







| Class A | Class B | Class C | |
|---|---|---|---|
| Class A | |||
| Class B | |||
| Class C |
| Dataset | Class | ||
|---|---|---|---|
| Benign | Malignant | ||
| Mammography Dataset (CBIS-DDSM) 3032 images total Benign: 1685 Malignant: 1347 | ![]() | ![]() | |
| Chest X-Ray Dataset | Normal | Bacteria | Virus |
| 5856 images total Normal: 1583 Bacteria: 2780 Virus: 1493 | ![]() | ![]() | ![]() |
| Brain Tumor Dataset | Meningioma | Pituitary Tumor | Glioma |
| 3064 images total Meningioma: 708 Pituitary Tumor: 930 Glioma: 1426 | ![]() | ![]() | ![]() |
| ISIC Challenge 2017 | Melanoma | Seborrheic keratosis | Nevus |
| 2750 images total Melanoma: 521 Seborrheic keratosis: 386 Nevus: 1843 | ![]() | ![]() | ![]() |
| Parameter | Value |
|---|---|
| Dataset Split | 70% Train, 10% Validation, 20% Test |
| Data Augmentation | None |
| Input Normalization | [0, 1] range |
| Input Size | pixels |
| Convolutional Filter Size | |
| Padding | same, stride 1 |
| Batch Normalization | Applied |
| Activation Function | ReLU |
| Pooling | MaxPooling stride 2; GlobalAveragePooling2D() |
| Fully Connected Layer | 256 neurons |
| Dropout | 0.3 |
| Optimizer | Adam |
| Learning Rate | |
| Batch Size | 16 |
| Loss Function | Sparse Categorical Crossentropy |
| Epochs | 100 with early stopping after 20 epochs |
| Dataset | Monotonicity Violations | ||
|---|---|---|---|
| Brain-Tumor-H | 100% * | n.d. | 28/39 (72%) |
| CBIS-DDSM-H | 100% * | n.d. | 7/39 (18%) |
| Chest-X-Ray-2018 | 100% * | n.d. | 25/39 (64%) |
| ISIC-2017-H | 80% | n.d. | 2/39 (5%) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Sgarro, G.A.; Mendikowski, M.; Santoro, D.; Grilli, L. CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity. Bioengineering 2026, 13, 971. https://doi.org/10.3390/bioengineering13090971
Sgarro GA, Mendikowski M, Santoro D, Grilli L. CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity. Bioengineering. 2026; 13(9):971. https://doi.org/10.3390/bioengineering13090971
Chicago/Turabian StyleSgarro, Giacinto Angelo, Melle Mendikowski, Domenico Santoro, and Luca Grilli. 2026. "CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity" Bioengineering 13, no. 9: 971. https://doi.org/10.3390/bioengineering13090971
APA StyleSgarro, G. A., Mendikowski, M., Santoro, D., & Grilli, L. (2026). CNN Sample-Size Effects Across Biomedical Datasets: A Reliability Pattern in Overfitting, Ranking, and Monotonicity. Bioengineering, 13(9), 971. https://doi.org/10.3390/bioengineering13090971












