Explainability and Trust in Deep Learning for Cancer Imaging: Systematic Barriers, Clinical Misalignment, and a Translational Roadmap
Simple Summary
Abstract
1. Introduction
- RQ1. What are the key challenges limiting explainability and trust in deep learning–based cancer imaging analysis?
- RQ2. Do XAI methods measurably increase clinician trust in deep learning models for malignant tumour detection?
- RQ3. Which deep learning techniques and architectural modifications improve the interpretability of CNN-based cancer image classifiers?
- RQ4. How do interpretable AI methods align with clinical reasoning and decision-making in oncology imaging?
- RQ5. To what extent do XAI approaches improve diagnostic performance and clinician confidence compared with black-box models?
- Temporal and Sociotechnical Scope: We specifically synthesise developments from 2023 to 2025, moving beyond standard performance metrics to evaluate robustness, fairness, and human-centred trust.
- The “Trust-Critical Lifecycle” Framework: Unlike traditional technical overviews, this review conceptualises trust as a dynamic, system-level property shaped by every stage from data acquisition and annotation to post-deployment monitoring.
- Clinical–Epistemic Alignment: A core novelty is our investigation into whether algorithmic reasoning structurally matches the hierarchical and contextual reasoning used by clinicians, rather than just providing visual heatmaps.
- Translational Roadmap: We provide a strategic pathway that integrates technical architectures (intrinsic vs. post hoc) with regulatory evolution (e.g., FDA PCCP and EU AI Act), ethical accountability, and medico-legal liability.
- Research focusing on deep learning architectures (CNNs, Vision Transformers, GNNs) applied to oncology imaging.
- Studies addressing core computational tasks: tumour detection, segmentation, classification, and prognostic modelling.
- Papers evaluating interpretability methods and their impact on clinical reasoning or clinician confidence.
- Articles discussing regulatory frameworks (FDA, GDPR) and ethical considerations in AI deployment.
- General AI studies without specific application to cancer-specific imaging.
- Papers focusing exclusively on predictive accuracy without addressing explainability, robustness, or trust.
- Research published prior to 2023, unless providing a foundational technical lineage.
2. Foundations and Computational Tasks in Cancer Imaging
2.1. Imaging Modalities and Data Ecosystems
2.1.1. Radiological Interpretability
2.1.2. Digital Pathology Efficiency
2.1.3. Multimodal and Longitudinal Synthesis
2.2. Deep Learning Paradigms in the Clinical Context
2.2.1. Convolutional Architectures and Local Feature Hierarchies
2.2.2. Transformer-Based Models and Global Contextual Reasoning
2.2.3. Representation Learning Under Data Scarcity
2.2.4. Graph-Based Modelling of Tissue Architecture
2.2.5. Integrative Perspective
2.3. Core Clinical Tasks
2.3.1. Tumour Detection and Localization
2.3.2. Advanced Image Segmentation for Precision Oncology
2.3.3. Histopathological Classification and Grading
2.3.4. Prognosis and Dynamic Risk Stratification
2.3.5. Auxiliary Tasks
3. System-Level Barriers to Trustworthy AI
3.1. Interpretability, Explainability, and Clinical Trust
3.2. Data-Centric Barriers to Reliability
3.2.1. Label Integrity Crisis
3.2.2. Annotation Bottleneck
3.2.3. Imbalanced Datasets
3.2.4. Dataset Shift
3.3. Robustness, Fairness, and Generalization
3.3.1. Multicentre Variability
3.3.2. Adversarial Vulnerability
3.3.3. Algorithm Fairness and Demographic Equity
3.4. Failure Modes of Explainability Methods
3.4.1. Instability and Technical Inconsistency of Explanations
3.4.2. Misleading Visual Explanations and Shortcut Learning
3.4.3. Limitations to Fidelity and the Risk of False Confidence
3.4.4. Clinical Validation Gap
4. Explainable AI in Cancer Imaging: Methods and Clinical Alignment
4.1. Post Hoc Explainability Approaches
4.1.1. Gradient-Based and Saliency Mapping Techniques
4.1.2. Perturbation-Based and Model-Agnostic Approaches
4.1.3. Attention Visualization in Vision Transformers
4.2. Intrinsic and Hybrid Interpretable Models
4.2.1. Concept Bottleneck and Clinically Structured Architectures
4.2.2. Prototype-Based and Case-Based Reasoning
4.2.3. Biologically Constrained and Disentangled Representations
4.2.4. Neuro-Symbolic and Causality-Aware Hybrid Models
5. Human–AI Trust Calibration and Clinical Translation
5.1. Multidimensional Trust: Robustness, Uncertainty, Fairness
5.1.1. Robustness and Stability in Adverse Clinical Environments
5.1.2. Reliability Through Uncertainty Quantification
5.1.3. Transparency and the Interpretability
5.1.4. Fairness, Bias Mitigation, and Accountability
5.2. Clinician-Centred Explainability
5.2.1. Breast Ultrasound: Morphological and Signal-Based Logic
5.2.2. Prostate MRI: Multiparametric and Contextual Synthesis
5.2.3. Digital Pathology: Topological and Multiscale Hierarchy
5.2.4. Multimodal Oncology: Longitudinal and Integrative Reasoning
5.3. Human-in-the-Loop Systems in Oncology Imaging
5.4. Evaluating Explainable AI Beyond Accuracy Metrics
5.4.1. Limitations of Accuracy-Centric Benchmarking
5.4.2. Technical Evaluation of Explainability
5.5. Clinical Validation and Human–AI Interactions
5.6. Uncertainty Estimation and Trust Calibration
5.7. Case Studies of Prospective Validation and Clinical Challenges
5.7.1. Technical Fragility: Instability, Fidelity, and Computational Constraints
5.7.2. Clinical–Epistemic Misalignment: Lack of Alignment with Clinical Reasoning
5.7.3. Deployment Challenges: Human–AI Interaction and Cognitive Effects
5.7.4. Validation Crisis: Generalization, Domain Shift, and Dataset Bias
5.7.5. Hybrid Failure Modes: Interacting Limitations and Compounded Effects
5.8. Critical Limitations of XAI in Clinical Practice
5.8.1. Technical Fragility
- Sensitivity to Noise: The exploratory transparency layer is characterised by a significant ‘stability and fidelity gap’, where post hoc maps are highly sensitive to noise.
- The Disagreement Problem: Because the explanatory layer often produces conflicting rationales for the same image (the ‘disagreement problem’), it cannot yet be viewed as a ‘ground truth’ for medical conclusion-making.
- Fidelity vs. Plausibility: False confidence where an explanation appears clinically reasonable but does not actually reflect the model’s true global decision logic.
5.8.2. Deployment Challenges
- Increased Workload: Instead of streamlining workflows, AI-generated heatmaps can increase reporting time by 33% and delay dictation by over 40%. This cognitive load exploitation occurs because interpreting the explanation becomes more demanding than the original diagnostic task.
- Authority Modulation: The liability-driven behavior where clinicians feel compelled to scrutinize every AI-highlighted region to mitigate perceived risk, leads to fatigue.
- Automation Bias: This is the ethical risk where clinicians may over-rely on AI recommendations, leading to a loss of independent diagnostic vigilance.
5.8.3. Clinical–Epistemic Misalignment
- Lack of Contextual Reasoning: Models often fail to replicate the contextual reasoning used by experts, such as incorporating longitudinal patient history or prior biopsy results. For example, in prostate cancer, AI often lacks the ability to resolve equivocal lesions using the multi-parametric data points (like PSA levels) that radiologists routinely use.
- Shortcut Learning: Saliency maps often highlight spurious correlations (scanner artifacts, skin markings, or pen annotations) rather than true pathology. This “illusion of interpretability” can lead a clinician to trust a model that is “right for the wrong reasons”.
5.8.4. Validation Crisis
- Technical Proxies vs. Clinical Utility: Most XAI research relies on technical metrics (faithfulness, sparsity) rather than prospective clinical trials that measure improved patient outcomes or safer therapeutic choices.
- The Interpretability–Performance Paradox: The most accurate models (e.g., Vision Transformers) are often the least transparent, creating a barrier for high-stakes oncology.
5.9. Research Priorities and Future Directions
6. Bridging the Trust Gap
6.1. Regulatory Evolution and Lifecycle Management
6.2. Ethical and Legal Considerations
6.3. Integrative Synthesis
6.3.1. Addressing Key Challenges (RQ1)
6.3.2. Impact of XAI on Clinician Trust (RQ2)
6.3.3. Techniques for Improving Interpretability (RQ3)
6.3.4. Alignment with Clinical Reasoning (RQ4)
6.3.5. Impact on Diagnostic Performance and Confidence (RQ5)
6.3.6. A Framework for Evidence Maturity and Readiness
6.3.7. A Progressive Model of Trust Integration
6.4. Strategic Directions
6.4.1. Embedding Semantic Alignment into Model Architecture
6.4.2. Standardizing Trust and Calibration Metrics
6.4.3. Advancing Prospective and Multicentre Validation
6.4.4. Designing for Lifecycle Governance and Updating Transparency
6.4.5. Strengthening Ethical Robustness and Equity Safeguards
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Bray, F.; Laversanne, M.; Sung, H.; Ferlay, J.; Siegel, R.L.; Soerjomataram, I.; Jemal, A. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J. Clin. 2024, 74, 229–263. [Google Scholar] [CrossRef]
- Yao, I.Z.; Dong, M.; Hwang, W.Y.K. Deep Learning Applications in Clinical Cancer Detection: A Review of Implementation Challenges and Solutions. Mayo Clin. Proc. Digit. Health 2025, 3, 100253. [Google Scholar] [CrossRef]
- Zou, Y.; Miao, P. Explainable AI-Enabled Hybrid Deep Learning Architecture for Breast Cancer Detection. Front. Immunol. 2025, 16, 1658741. [Google Scholar] [CrossRef]
- Peng, Y.C.; Lee, W.J.; Chang, Y.C.; Chan, W.P.; Chen, S.J. Radiologist Burnout: Trends in Medical Imaging Utilization Under the National Health Insurance System with the Universal Code Bundling Strategy in an Academic Tertiary Medical Centre. Eur. J. Radiol. 2022, 157, 110596. [Google Scholar] [CrossRef]
- Olawuyi, O.; Viriri, S. Deep Learning Techniques for Prostate Cancer Analysis and Detection: Survey of the State of the Art. J. Imaging 2025, 11, 254. [Google Scholar] [CrossRef] [PubMed]
- Khalid, S.A.; Khaliq, T.; Rehman, Y.N.; Anwar, Z.; Mirani, W. Comparative Performance of Artificial Intelligence and Radiologists in Detecting Lung Nodules and Breast Lesions on CT and MRI: A Systematic Review. Cureus 2024, 16, e95943. [Google Scholar] [CrossRef]
- Erukude, S.T.; Marella, V.C.; Veluru, S.R. Explainable Deep Learning in Medical Imaging: Brain Tumor and Pneumonia Detection. arXiv 2025, arXiv:2510.21823. [Google Scholar] [CrossRef]
- Wang, L. Self-supervised learning and transformer-based technologies in breast cancer imaging. Front. Radiol. 2025, 5, 1684436. [Google Scholar] [CrossRef]
- Singh, Y.; Hathaway, Q.A.; Keishing, V.; Salehi, S.; Wei, Y.; Horvat, N.; Vera-Garcia, D.V.; Choudhary, A.; Mula Kh, A.; Quaia, E.; et al. Beyond Post hoc Explanations: A Comprehensive Framework for Accountable AI in Medical Imaging Through Transparency, Interpretability, and Explainability. Bioengineering 2025, 12, 879. [Google Scholar] [CrossRef]
- Wang, H.; Hou, J.; Chen, H. Concept Complement Bottleneck Model for Interpretable Medical Image Diagnosis. arXiv 2024, arXiv:2410.15446. [Google Scholar] [CrossRef]
- Cheng, C.H.; Shi, S. Artificial intelligence in cancer: Applications, challenges, and future perspectives. Mol. Cancer 2025, 24, 274. [Google Scholar] [CrossRef]
- Chaddad, A.; Hassan, L.; Desrosiers, C.; Toews, M.; Tanougast, C. Generalizable and Explainable Deep Learning for Medical Image Computing: An Overview. Curr. Opin. Biomed. Eng. 2025, 33, 100567. [Google Scholar] [CrossRef]
- Ahuchogu, M.C.; Saha, H.; Saha, G.C. Explainable AI in Medical Imaging: Improving Clinical Trust in Deep Learning Model. Eksplorium 2025, 46, 144. [Google Scholar] [CrossRef]
- Düsing, C.; Cimiano, P.; Rehberg, S.; Scherer, C.; Kaup, O.; Köster, C.; Hellmich, S.; Herrmann, D.; Meier, K.L.; Claßen, S.; et al. Integrating federated learning for improved counterfactual explanations in clinical decision support systems for sepsis therapy. Artif. Intell. Med. 2024, 157, 102982. [Google Scholar] [CrossRef] [PubMed]
- Abbas, Q.; Jeong, W.; Lee, S.W. Explainable AI in Clinical Decision Support Systems: A Meta-Analysis of Methods, Applications, and Usability Challenges. Healthcare 2025, 13, 2154. [Google Scholar] [CrossRef] [PubMed]
- Hildt, E. What Is the Role of Explainability in Medical Artificial Intelligence? A Case-Based Approach. Bioengineering 2025, 12, 375. [Google Scholar] [CrossRef] [PubMed]
- Noë, A.; Bouhouita-Guermech, S.; Zawati, M.H. The Right to Explanation in AI: In a Lonely Place. J. Med. Internet Res. 2025, 27, e64482. [Google Scholar] [CrossRef]
- Waqas, A.; Tripathi, A.; Rasool, G. Multimodal data integration for oncology in the era of deep neural networks: A review. Front. Artif. Intell. 2024, 7, 1408843. [Google Scholar] [CrossRef]
- Mastoi, Q.U.; Latif, S.; Brohi, S.; Ahmad, J.; Alqhatani, A.; Alshehri, M.S.; Al Mazroa, A.; Ullah, R. Explainable AI in medical imaging: An interpretable and collaborative federated learning model for brain tumour classification. Front. Oncol. 2025, 15, 1535478. [Google Scholar] [CrossRef]
- de Almeida, J.G.; Rodrigues, N.M.; Verde, A.S.C.; Gaivão, A.M.; Bilreiro, C.; Santiago, I.; Ip, J.; Belião, S.; Matos, C.; Silva, S.; et al. Impact of Scanner Manufacturer, Endorectal Coil Use, and Clinical Variables on Deep Learning–assisted Prostate Cancer Classification Using Multiparametric MRI. Radiol. Artif. Intell. 2025, 7, e230555. [Google Scholar] [CrossRef]
- Kersting, D.; Borys, K.; Küper, A.; Kim, M.; Haubold, J.; Goerttler, T.; Umutlu, L.; Costa, P.F.; Kleesiek, J.; Rischpler, C.; et al. Staging of prostate Cancer with ultrafast PSMA-PET scans enhanced by AI. Eur. J. Nucl. Med. Mol. Imaging 2025, 52, 1658–1670. [Google Scholar] [CrossRef]
- Dimitriou, N.; Arandjelović, O.; Caie, P.D. Deep Learning for Whole Slide Image Analysis: An Overview. Front. Med. 2019, 6, 264. [Google Scholar] [CrossRef] [PubMed]
- Tang, W.; Qin, R.; Fang, H.; Zhou, F.; Chen, H.; Li, X.; Cheng, M.-M. Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology. arXiv 2025, arXiv:2506.02408. [Google Scholar] [CrossRef]
- Hashemian, S.; Bidgoli, A.A. EvoPS: Evolutionary Patch Selection for Whole Slide Image Analysis in Computational Pathology. arXiv 2025, arXiv:2511.07560. [Google Scholar] [CrossRef]
- Komura, D.; Ochi, M.; Ishikawa, S. Machine learning methods for histopathological image analysis: Updates in 2024. Comput. Struct. Biotechnol. J. 2025, 27, 383–400. [Google Scholar] [CrossRef] [PubMed]
- Song, B.; Leroy, A.; Yang, K.; Dam, T.; Wang, X.; Maurya, H.; Pathak, T.; Lee, J.; Stock, S.; Li, X.T.; et al. Deep learning informed multimodal fusion of radiology and pathology to predict outcomes in HPV-associated oropharyngeal squamous cell carcinoma. EBioMedicine 2025, 114, 105663. [Google Scholar] [CrossRef] [PubMed]
- Kim, H.; Karaman, B.K.; Zhao, Q.; Wang, A.Q.; Sabuncu, M.R. Alzheimer’s Disease Neuroimaging Initiative. Learning-based inference of longitudinal image changes: Applications in embryo development, wound healing, and aging brain. Proc. Natl. Acad. Sci. USA 2025, 122, e2411492122. [Google Scholar] [CrossRef] [PubMed]
- Zhuang, L.; Park, S.H.; Skates, S.J.; Prosper, A.E.; Aberle, D.R.; Hsu, W. Advancing precision oncology through modelling of longitudinal and multimodal data. IEEE Rev. Biomed. Eng. 2025, 19, 182–200. [Google Scholar] [CrossRef]
- Aburass, S.; Dorgham, O.; Al Shaqsi, J.; Abu Rumman, M.; Al-Kadi, O. Vision Transformers in Medical Imaging: A Comprehensive Review of Advancements and Applications Across Multiple Diseases. J. Imaging Inform. Med. 2025, 38, 3928–3971. [Google Scholar] [CrossRef]
- Lyu, Y.; Tian, X. MWG-UNet++: Hybrid Transformer U-Net Model for Brain Tumor Segmentation in MRI Scans. Bioengineering 2025, 12, 140. [Google Scholar] [CrossRef]
- Chen, M.; Wang, K.; Wang, J. Vision Transformer-Based Multilabel Survival Prediction for Oropharynx Cancer After Radiation Therapy. Int. J. Radiat. Oncol. Biol. Phys. 2024, 118, 1123–1134. [Google Scholar] [CrossRef]
- Kim, J.W.; Khan, A.U.; Banerjee, I. Systematic Review of Hybrid Vision Transformer Architectures for Radiological Image Analysis. J. Imaging Inform. Med. 2025, 38, 3248–3262. [Google Scholar] [CrossRef]
- Nawaz, A.; Khan, S.S.; Ahmad, A. Ensemble of Autoencoders for Anomaly Detection in Biomedical Data: A Narrative Review. IEEE Access 2024, 12, 17273–17289. [Google Scholar] [CrossRef]
- Cai, Y.; Chen, H.; Cheng, K.T. Rethinking Autoencoders for Medical Anomaly Detection from a Theoretical Perspective. arXiv 2024, arXiv:2403.09303. [Google Scholar] [CrossRef]
- Brussee, S.; Buzzanca, G.; Schrader, A.M.R.; Kers, J. Graph neural networks in histopathology: Emerging trends and future directions. Med. Image Anal. 2025, 101, 103444. [Google Scholar] [CrossRef]
- Pati, P.; Jaume, G.; Foncubierta-Rodríguez, A.; Feroce, F.; Anniciello, A.M.; Scognamiglio, G.; Brancati, N.; Fiche, M.; Dubruc, E.; Riccio, D.; et al. Hierarchical graph representations in digital pathology. Med. Image Anal. 2022, 75, 102264. [Google Scholar] [CrossRef]
- Bui, D.C.; Song, B.; Kim, K.; Kwak, J.T. Spatially Constrained and -Unconstrained Bi-Graph Interaction Network for Multi-Organ Pathology Image Classification. IEEE Trans. Med. Imaging 2025, 44, 194–206. [Google Scholar] [CrossRef]
- Wang, X.; Su, C. Deep learning for cancer detection based on genomic and imaging data: A comprehensive review. Cancer Manag. Res. 2025, 17, 2089–2125. [Google Scholar] [CrossRef] [PubMed]
- Han, W.; Dong, X.; Wang, G.; Ding, Y.; Yang, A. Application and improvement of YOLO11 for brain tumour detection in medical images. Front. Oncol. 2025, 15, 1643208. [Google Scholar] [CrossRef] [PubMed]
- Chourib, I. From Detection to Diagnosis: An Advanced Transfer Learning Pipeline Using YOLO11 with Morphological Post-Processing for Brain Tumor Analysis for MRI Images. J. Imaging 2025, 11, 282. [Google Scholar] [CrossRef]
- Jiang, X.; Hu, Z.; Wang, S.; Zhang, Y. Deep learning for medical image-based cancer diagnosis. Cancers 2023, 15, 3608. [Google Scholar] [CrossRef]
- Mostafa, A.M.; Alaerjan, A.S.; Aldughayfiq, B.; Allahem, H.; Mahmoud, A.A.; Said, W.; Shabana, H.; Ezz, M. Optimized YOLOv8 for enhanced breast tumour segmentation in ultrasound imaging. Discov. Oncol. 2025, 16, 1152. [Google Scholar] [CrossRef]
- Xu, H.-L.; Gong, T.-T.; Song, X.-J.; Chen, Q.; Bao, Q.; Yao, W.; Xie, M.-M.; Li, C.; Grzegorzek, M.; Shi, Y.; et al. Artificial intelligence performance in image-based cancer identification: Umbrella review of systematic reviews. J. Med. Internet Res. 2025, 27, e53567. [Google Scholar] [CrossRef]
- Hu, R.; Xie, Y.; Zhang, L.; Liu, L.; Luo, H.; Wu, R.; Luo, D.; Liu, Z.; Hu, Z. A two-stage deep-learning framework for CT denoising based on a clinically structure-unaligned paired dataset. Quant. Imaging Med. Surg. 2024, 14, 335–351. [Google Scholar] [CrossRef]
- Abrar, M.; Salam, A.; Ullah, F.; Ullah, F.; Al Ghamdi, A.S. Enhancing brain tumour segmentation using attention-based convolutional U-Net on MRI images. Sci. Rep. 2025, 15, 36603. [Google Scholar] [CrossRef] [PubMed]
- Gad, E.; Soliman, S.; Darweesh, M.S. Advancing brain tumour segmentation via attention-based 3D U-Net architecture and digital image processing. In Proceedings of the 12th International Conference on Model and Data Engineering; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2023; Volume 14396, pp. 245–258. [Google Scholar] [CrossRef]
- Wang, Q.; Bi, Q.; Qu, L.; Deng, Y.; Wang, X.; Zheng, Y.; Li, C.; Meng, Q.; Miao, K. MAMILNet: Advancing precision oncology with multiscale attentional multi-instance learning for whole slide image analysis. Front. Oncol. 2024, 14, 1275769. [Google Scholar] [CrossRef]
- Melanthota, S.K.; Spandana, K.U.; Raghavendra, U.; Rai, S.; Nayak, R.; Kistenev, Y.V.; Shil, S.; Mahato, K.K.; Mazumder, N. Machine learning based multiclass classification and grading of squamous cell carcinoma in optical microscopy. Microsc. Res. Technol. 2025, 88, e70016. [Google Scholar] [CrossRef]
- Piedimonte, S.; Mohamed, M.; Rosa, G.; Gerstl, B.; Vicus, D. Predicting Response to Treatment and Survival in Advanced Ovarian Cancer Using Machine Learning and Radiomics: A Systematic Review. Cancers 2025, 17, 336. [Google Scholar] [CrossRef] [PubMed]
- Mesinovic, M.; Watkinson, P.; Zhu, T. DySurv: Dynamic deep learning model for survival analysis with conditional variational inference. arXiv 2023, arXiv:2310.18681. [Google Scholar] [CrossRef] [PubMed]
- Kanakarajan, H.; Zhou, J.; Lobo Gomes, A.; Kalendralis, P.; Liang, W.; Tohidinezhad, F.; Dekker, A.; De Baene, W.; Sitskoorn, M. Predicting overall survival of NSCLC patients with clinical, radiomics and deep learning features. medRxiv 2025, preprint. [Google Scholar] [CrossRef]
- Liu, X.; Gao, K.; Liu, B.; Pan, C.; Liang, K.; Yan, L.; Ma, J.; He, F.; Zhang, S.; Pan, S.; et al. Advances in deep learning-based medical image analysis. Health Data Sci. 2021, 2021, 8786793. [Google Scholar] [CrossRef]
- Hu, M.; Pan, S.; Chang, C.-W.; Qiu, R.L.J.; Peng, J.; Wang, T.; Roper, J.; Mao, H.; Yu, D.; Yang, X. Cross-modality 3D MRI synthesis via cycle-guided denoising diffusion probability model. J. Med. Imaging 2025, 12, 064003. [Google Scholar] [CrossRef] [PubMed]
- Yu, M.; Xu, Z.; Lukasiewicz, T. A general survey on medical image super-resolution via deep learning. Comput. Biol. Med. 2025, 193, 110345. [Google Scholar] [CrossRef]
- Rezaeian, O.; Bayrak, A.E.; Asan, O. Explainability and AI Confidence in Clinical Decision Support Systems: Effects on Trust, Diagnostic Performance, and Cognitive Load in Breast Cancer Care. arXiv 2025, arXiv:2501.16693. [Google Scholar] [CrossRef]
- Ahmed, M.; Bibi, T.; Khan, R.A.; Nasir, S. Enhancing breast cancer diagnosis in mammography: Evaluation and integration of convolutional neural networks and explainable AI. In Proceedings of the 2024 26th International Multi-Topic Conference (INMIC); IEEE: Karachi, Pakistan, 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Gao, S.; Wang, S.; Gao, Y.; Wang, B.; Zhuang, X.; Warren, A.; Stewart, G.; Jones, J.O.; Crispin-Ortuzar, M. Evaluating foundation models with pathological concept learning for kidney cancer. arXiv 2025, arXiv:2509.25552. [Google Scholar] [CrossRef]
- Hassan, M.R.; Islam, M.F.; Uddin, M.Z.; Ghoshal, G.; Hassan, M.M.; Huda, S.; Fortino, G. Prostate cancer classification from ultrasound and MRI images using deep learning based explainable artificial intelligence. Future Gener. Comput. Syst. 2022, 127, 462–472. [Google Scholar] [CrossRef]
- Alshammri, G.H.; Obayya, M.; Negm, N.; AlAqil, M.A.; Alshahrani, M.M.; Alsini, R.; Alsuhaibani, R.; Alhashmi, A.A. Towards enhanced bladder cancer detection using explainable artificial intelligence with a hybrid feature engineering framework on biomedical images. Eng. Appl. Artif. Intell. 2026, 167, 113729. [Google Scholar] [CrossRef]
- Talukder, M.A. An improved XAI-based DenseNet model for breast cancer detection using reconstruction and fine-tuning. Results Eng. 2025, 26, 104802. [Google Scholar] [CrossRef]
- Khanna, A.; Singh, P.; Kaur, I.; Ul Hassan, I.; Kumar, V. Explainable Deep Learning Framework for Histopathological Cancer Detection. In Proceedings of the 2025 8th International Conference on Circuit, Power & Computing Technologies (ICCPCT), Kollam, India, 7–8 August 2025; pp. 1400–1406. [Google Scholar] [CrossRef]
- de Sousa, I.P.; Vellasco, M.M.B.R.; da Silva, E.C. Local Interpretable Model-Agnostic Explanations for Classification of Lymph Node Metastases. Sensors 2019, 19, 2969. [Google Scholar] [CrossRef] [PubMed]
- Sun, Z.; Wang, K.; Kong, Z.; Xing, Z.; Chen, Y.; Luo, N.; Yu, Y.; Song, B.; Wu, P.; Wang, X.; et al. A multicenter study of artificial intelligence-aided software for detecting visible clinically significant prostate cancer on mpMRI. Insights Into Imaging 2023, 14, 72. [Google Scholar] [CrossRef]
- Ahn, S.H.; Baek, S.; Park, J.; Kim, J.; Rhee, H.; Chung, Y.E.; Kim, H.; Lee, Y.H. Uncertainty Quantification in Automated Detection of Vertebral Metastasis Using Ensemble Monte Carlo Dropout. J. Imaging Inform. Med. 2025, 38, 2700–2715. [Google Scholar] [CrossRef]
- Eisemann, N.; Bunk, S.; Mukama, T.; Baltus, H.; Elsner, S.A.; Gomille, T.; Hecht, G.; Heywang-Köbrunner, S.; Rathmann, R.; Siegmann-Luz, K.; et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nat. Med. 2025, 31, 917–924. [Google Scholar] [CrossRef] [PubMed]
- Cheng, N.; Shen, C.; Yang, J.; Lou, B.; Toffanin, S.; Ge, Y. AI-driven SERS Biosensing Chip for Digital Biomarker Detection in Diagnosis and Monitoring of Early-Stage HCC. Chem. Eng. J. 2026, 527, 171519. [Google Scholar] [CrossRef]
- Jozi, N.S.; Al-Suhail, G.A. LCxNet: An Explainable CNN Framework for Lung Cancer Detection in CT Images Using Multi-Optimizer and Visual Interpretability. Appl. Syst. Innov. 2025, 8, 153. [Google Scholar] [CrossRef]
- Prakash, P.S.; Rao, P.K.; Pasha, M.J.; Algarni, A.; Ayadi, M.; Cho, Y.; Nam, Y. CausalX-Net: A causality-guided explainable segmentation network for brain tumours. Front. Med. 2025, 12, 1693603. [Google Scholar] [CrossRef]
- Jones, C.K.; Wang, G.; Yedavalli, V.; Sair, H. Direct quantification of epistemic and aleatoric uncertainty in 3D U-net segmentation. J. Med. Imaging 2022, 9, 034002. [Google Scholar] [CrossRef]
- Rahman, M.A. HyFormer-Net: A Synergistic CNN-Transformer with Interpretable Multi-Scale Fusion for Breast Lesion Segmentation and Classification in Ultrasound Images. arXiv 2025, arXiv:2511.01013. [Google Scholar] [CrossRef]
- Gallée, L.; Lisson, C.S.; Ropinski, T.; Beer, M.; Götz, M. Proto-Caps: Interpretable Medical Image Classification Using Prototype Learning and Privileged Information. PeerJ Comput. Sci. 2025, 11, e2908. [Google Scholar] [CrossRef]
- Hasan, M.J.; Hasan, M.; Akter, S.; Siddique Mahi, A.B.; Uddin, M.P. Enhancing Brain Tumor Classification with a Novel Attention-Based Explainable Deep Learning Framework. Biomed. Signal Process. Control 2026, 112, 108636. [Google Scholar] [CrossRef]
- Zeineldin, R.A.; Karar, M.E.; Elshaer, Z.; Coburger, J.; Wirtz, C.R.; Burgert, O.; Mathis-Ullrich, F. Explainability of deep neural networks for MRI analysis of brain tumours. Int. J. Comput. Assist. Radiol. Surg. 2022, 17, 1673–1683. [Google Scholar] [CrossRef]
- Wei, Y.; Tam, R.; Tang, X. MProtoNet: A Case-Based Interpretable Model for Brain Tumor Classification with 3D Multi-Parametric Magnetic Resonance Imaging. arXiv 2023, arXiv:2304.06258. [Google Scholar] [CrossRef]
- Muhammad, D.; Bendechache, M. More than just a heatmap: Elevating XAI with rigorous evaluation metrics. Front. Med. Technol. 2025, 7, 1674343. [Google Scholar] [CrossRef]
- Zhang, B.; Vakanski, A.; Xian, M. BI-RADS-Net: An Explainable Multitask Learning Approach for Cancer Diagnosis in Breast Ultrasound Images. In Proceedings of the 2021 IEEE 31st International Workshop on Machine Learning for Signal Processing (MLSP), Gold Coast, Australia, 25–28 October 2021; pp. 1–6. [Google Scholar] [CrossRef]
- Hussain, S.M.; Buongiorno, D.; Altini, N.; Berloco, F.; Prencipe, B.; Moschetta, M.; Bevilacqua, V.; Brunetti, A. Shape-Based Breast Lesion Classification Using Digital Tomosynthesis Images: The Role of Explainable Artificial Intelligence. Appl. Sci. 2022, 12, 6230. [Google Scholar] [CrossRef]
- Burgos, D.; Morshed, A.; Rashid, M.M.; Mandala, S. A comparison of machine learning models to deep learning models for cancer image classification and explainability of classification. In Proceedings of the 2024 International Conference on Data Science and Its Applications (ICoDSA 2024), Kuta, Bali, Indonesia, 10–11 July 2024; pp. 386–390. [Google Scholar] [CrossRef]
- Stacke, K.; Eilertsen, G.; Unger, J.; Lundstrom, C. Measuring Domain Shift for Deep Learning in Histopathology. IEEE J. Biomed. Health Inform. 2021, 25, 325–336. [Google Scholar] [CrossRef]
- López-Miguel, N.; Díaz-Hernández, R.; Altamirano-Robles, L. Confidence Calibration of CNNs in Medical Image Databases. Comput. Sist. 2025, 29, 7–14. [Google Scholar] [CrossRef]
- Emam, M.M.; Ibrahim, D.S.; Abdel Samee, N.; Houssein, E.H. An Efficient Explainable Deep Learning Model for Multiclass Classification of Gynecological Cancers. Knowl.-Based Syst. 2026, 334, 115109. [Google Scholar] [CrossRef]
- Merabet, A.; Saighi, A.; Laboudi, Z.; Ferradji, M.A.; Harous, S.; Mohamed, A.W.; Mousavirad, S.J. Few-shot Learning and Explainable AI for Colon Cancer Histopathology: A Prototypical Network with Multi-Technique Interpretability. Int. J. Med. Inform. 2026, 206, 106167. [Google Scholar] [CrossRef] [PubMed]
- Alkhalaf, S.; Alturise, F.; Bahaddad, A.A.; Elamin Elnaim, B.M.; Shabana, S.; Abdel-Khalek, S.; Mansour, R.F. Adaptive Aquila Optimizer with Explainable Artificial Intelligence-Enabled Cancer Diagnosis on Medical Imaging. Cancers 2023, 15, 1492. [Google Scholar] [CrossRef] [PubMed]
- Dolezal, J.M.; Srisuwananukorn, A.; Karpeyev, D.; Ramesh, S.; Kochanny, S.; Cody, B.; Mansfield, A.S.; Rakshit, S.; Bansal, R.; Bois, M.C.; et al. Uncertainty-informed deep learning models enable high-confidence predictions for digital histopathology. Nat. Commun. 2022, 13, 6572. [Google Scholar] [CrossRef]
- Siam, A.S.M.; Hasan, M.d.M.; Arafat, Y.; Chowdhury, M.d.M.; Jobayer, S.H.; Hafiz, F.; Azim, R. FVCM-Net: Interpretable Privacy-Preserved Attention-Driven Lung Cancer Detection from CT Scan Images with Explainable HiRes-CAM Attribution Map and Ensemble Learning. Biomed. Signal Process. Control. 2026, 114, 108719. [Google Scholar] [CrossRef]
- Sangwan, H. Quantifying Explainable AI Methods in Medical Diagnosis: A Study in Skin Cancer. medRxiv 2024. [Google Scholar] [CrossRef]
- Pandala, M.L.; Periyanayagi, S. Optimal explainable vision transformer framework for skin cancer diagnosis with neural architecture search feature learning. Biomed. Signal Process. Control 2026, 112, 108723. [Google Scholar] [CrossRef]
- Graceline, H.H.; Sulochana, C.H. XceSCNN: An interpretable deep learning framework for accurate skin cancer classification using enhanced segmentation and SHAP explanations. Biomed. Signal Process. Control 2026, 117, 109560. [Google Scholar] [CrossRef]
- Metta, C.; Beretta, A.; Guidotti, R.; Yin, Y.; Gallinari, P.; Rinzivillo, S.; Giannotti, F. Improving trust and confidence in medical skin lesion diagnosis through explainable deep learning. Int. J. Data Sci. Anal. 2023, 17, 183–195. [Google Scholar] [CrossRef]
- Abraham, S.S.; Peter, J.D. Retrospective Analysis of an Interpretable Skin Cancer Classification using Deep Learning Models. In Proceedings of the 2024 International Conference on Cognitive Robotics and Intelligent Systems (ICC-ROBINS 2024), Coimbatore, India, 17–19 April 2024; pp. 327–332. [Google Scholar] [CrossRef]
- Park, S.; Ayana, G.; Wako, B.D.; Jeong, K.C.; Yoon, S.-D.; Choe, S. Vision Transformers for Low-Quality Histopathological Images: A Case Study on Squamous Cell Carcinoma Margin Classification. Diagnostics 2025, 15, 260. [Google Scholar] [CrossRef] [PubMed]
- Gupta, S.P.A.; Reddy, G.H.V.; Natarajan, K.; Srivastava, V.; Goda, J. Artificial Intelligence in Radiology: Transforming Cancer Detection and Diagnosis. Cureus 2025, 17, e96518. [Google Scholar] [CrossRef]
- Hrinivich, W.T.; Wang, T.; Wang, C. Editorial: Interpretable and Explainable Machine Learning Models in Oncology. Front. Oncol. 2023, 13, 1184428. [Google Scholar] [CrossRef]
- Khade, P.S. Explainable AI (XAI) in Healthcare: Building Trust in Medical Diagnosis Systems. Int. J. Sci. Res. Sci. Technol. 2025, 12, 07–22. [Google Scholar] [CrossRef]
- Bettinger, H.; Lenczner, G.; Guigui, J.; Rotenberg, L.; Zerbib, E.; Attia, A.; Vidal, J.; Beaumel, P. Evaluation of the Performance of an Artificial Intelligence (AI) Algorithm in Detecting Thoracic Pathologies on Chest Radiographs. Diagnostics 2024, 14, 1183. [Google Scholar] [CrossRef]
- Zhou, S.K.; Greenspan, H.; Davatzikos, C.; Duncan, J.S.; Van Ginneken, B.; Madabhushi, A.; Prince, J.L.; Rueckert, D.; Summers, R.M. A review of deep learning in medical imaging: Imaging traits, technology trends, case studies with progress highlights, and future promises. Proc. IEEE 2021, 109, 820–838. [Google Scholar] [CrossRef]
- Rajendran, P.; Safari, M.; He, W.; Hu, M.; Wang, S.; Zhou, J.; Yang, X. Foundation Models in Medical Image Analysis: A Systematic Review and Meta-Analysis. arXiv 2025, arXiv:2510.16973. [Google Scholar] [CrossRef]
- Shobayo, O.; Saatchi, R. Developments in Deep Learning Artificial Neural Network Techniques for Medical Image Analysis and Interpretation. Diagnostics 2025, 15, 1072. [Google Scholar] [CrossRef]
- Wei, Y.; Deng, Y.; Sun, C.; Lin, M.; Jiang, H.; Peng, Y. Deep learning with noisy labels in medical prediction problems: A scoping review. J. Am. Med. Inform. Assoc. 2024, 31, 1596–1607. [Google Scholar] [CrossRef]
- Ma, Y.; Hou, J.; Zhang, C.; Zhou, Y.; Ge, Z.; Xie, H.; Ju, L. Benchmarking Real-World Medical Image Classification with Noisy Labels: Challenges, Practice, and Outlook. arXiv 2025, arXiv:2512.09315. [Google Scholar] [CrossRef]
- Karimi, D.; Dou, H.; Warfield, S.K.; Gholipour, A. Deep learning with noisy labels: Exploring techniques and remedies in medical image analysis. Med. Image Anal. 2020, 65, 101759. [Google Scholar] [CrossRef] [PubMed]
- Rajaraman, S.; Zamzmi, G.; Antani, S.K. Novel loss functions for ensemble-based medical image classification. PLoS ONE 2021, 16, e0261307. [Google Scholar] [CrossRef] [PubMed]
- van Veldhuizen, V.; Botha, V.; Lu, C.; Erdal Cesur, M.; Groot Lipman, K.; de Jong, E.D.; Horlings, H.; Sanchez, C.I.; Snoek, C.G.M.; Wessels, L.; et al. Foundation Models in Medical Imaging: A Review and Outlook. arXiv 2025, arXiv:2506.09095. [Google Scholar] [CrossRef]
- Guo, J.; Li, M.; Su, H.; López, S.; Fan, L.; Kim, D.; Katsaggelos, A.K. Vision–Language Enhanced Foundation Model for Semi-Supervised Medical Image Segmentation. arXiv 2025, arXiv:2511.19759. [Google Scholar] [CrossRef]
- Guo, X.; Chai, W.; Li, S.-Y.; Wang, G. LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound. arXiv 2024, arXiv:2410.15074. [Google Scholar] [CrossRef]
- Akinci D’Antonoli, T.; Bluethgen, C.; Cuocolo, R.; Klontzas, M.E.; Ponsiglione, A.; Kocak, B. Foundation models for radiology: Fundamentals, applications, opportunities, challenges, risks, and prospects. Diagn. Interv. Radiol. 2025. online ahead of print. [Google Scholar] [CrossRef]
- Kim, Y.; Jeong, H.; Chen, S.; Li, S.S.; Lu, M.; Alhamoud, K.; Mun, J.; Grau, C.; Jung, M.; Gameiro, R.; et al. Medical Hallucination in Foundation Models and Their Impact on Healthcare. medRxiv 2025, preprint. [Google Scholar] [CrossRef]
- Jiao, R.; Zhang, Y.; Zhang, Y.; Ding, L.; Ding, L.; Xue, B.; Xue, B.; Zhang, J.; Cai, R.; Cheng, J.; et al. Learning with limited annotations: A survey on deep semi-supervised learning for medical image segmentation. Comput. Biol. Med. 2024, 169, 107840. [Google Scholar] [CrossRef] [PubMed]
- Chaitanya, K.; Erdil, E.; Karani, N.; Konukoglu, E. Local contrastive loss with pseudo-label based self-training for semi-supervised medical image segmentation. Med. Image Anal. 2023, 87, 102792. [Google Scholar] [CrossRef] [PubMed]
- Koçak, B.; Ponsiglione, A.; Stanzione, A.; Bluethgen, C.; Santinha, J.; Ugga, L.; Huisman, M.; Klontzas, M.E.; Cannella, R.; Cuocolo, R. Bias in artificial intelligence for medical imaging: Fundamentals, detection, avoidance, mitigation, challenges, ethics, and prospects. Diagn. Interv. Radiol. 2025, 31, 75–88. [Google Scholar] [CrossRef] [PubMed]
- Oakden-Rayner, L.; Dunnmon, J.A.; Carneiro, G.; Ré, C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. Proc. ACM Conf. Health Inference Learn. 2020, 2020, 151–159. [Google Scholar] [CrossRef]
- Krishna, A.; Wang, G.; Mueller, K. Multi-Conditioned Denoising Diffusion Probabilistic Model (mDDPM) for Medical Image Synthesis. arXiv 2024, arXiv:2409.04670. [Google Scholar] [CrossRef]
- Kidder, B.L. Advanced image generation for cancer using diffusion models. Biol. Methods Protoc. 2024, 9, bpae062. [Google Scholar] [CrossRef]
- Montoya-del-Angel, R.; Sam-Millan, K.; Vilanova, J.C.; Martí, R. MAM-E: Mammographic Synthetic Image Generation with Diffusion Models. Sensors 2024, 24, 2076. [Google Scholar] [CrossRef]
- Hosseini, A.; Serag, A. Is Synthetic Data Generation Effective in Maintaining Clinical Biomarkers? Investigating Diffusion Models Across Diverse Imaging Modalities. Front. Artif. Intell. 2025, 8, 1454441. [Google Scholar] [CrossRef]
- Suleman, M.U.; Mursaleen, M.; Khalil, U.; Saboor, A.; Bilal, M.; Khan, S.A.; Subhani, M.A.; Hussnain, M.A.; Tabassum, S.N.; Tahir, M. Assessing the generalizability of artificial intelligence in radiology: A systematic review of performance across different clinical settings. Ann. Med. Surg. 2025, 87, 8803–8811. [Google Scholar] [CrossRef]
- Hasanzadeh, F.; Josephson, C.B.; Waters, G.; Adedinsewo, D.; Azizi, Z.; White, J.A. Bias recognition and mitigation strategies in artificial intelligence healthcare applications. NPJ Digit. Med. 2025, 8, 154. [Google Scholar] [CrossRef]
- Drenkow, N.; Pavlak, M.; Harrigian, K.; Zirikly, A.; Subbaswamy, A.; Farhangi, M.M.; Petrick, N.; Unberath, M. Detecting Dataset Bias in Medical AI: A Generalized and Modality-Agnostic Auditing Framework. arXiv 2025, arXiv:2503.09969. [Google Scholar] [CrossRef]
- Tzortzis, I.N.; Gutierrez-Torre, A.; Sykiotis, S.; Agullo Lopez, F.; Bakalos, N.; Doulamis, A.; Doulamis, N.; Berral, J.L. Towards Generalizable Federated Learning in Medical Imaging: A Real-World Case Study on Mammography Data. Comput. Struct. Biotechnol. J. 2025, 28, 106–117. [Google Scholar] [CrossRef] [PubMed]
- Rehman, M.H.U.; Hugo Lopez Pinaya, W.; Nachev, P.; Teo, J.T.; Ourselin, S.; Cardoso, M.J. Federated Learning for Medical Imaging Radiology: A Review. Br. J. Radiol. 2023, 96, 20220890. [Google Scholar] [CrossRef]
- Akyüz, U.; Katircioglu-Öztürk, D.; Süslü, E.K.; Keleş, B.; Kaya, M.C.; Durhan, G.; Akpınar, M.; Demirkazık, F.B.; Akar, G.B. DoSReMC: Domain Shift Resilient Mammography Classification using Batch Normalization Adaptation. arXiv 2025, arXiv:2508.15452. [Google Scholar] [CrossRef]
- Castillo, J.M.T.; Starmans, M.P.A.; Arif, M.; Niessen, W.J.; Klein, S.; Bangma, C.H.; Schoots, I.G.; Veenland, J.F. A Multi-Center, Multi-Vendor Study to Evaluate the Generalizability of a Radiomics Model for Classifying Prostate Cancer: High Grade vs. Low Grade. Diagnostics 2021, 11, 369. [Google Scholar] [CrossRef]
- Joseph, J. The Protocol Genome: A Self-Supervised Learning Framework from DICOM Headers. arXiv 2025, arXiv:2509.06995. [Google Scholar] [CrossRef]
- Joel, M.Z.; Umrao, S.; Chang, E.; Choi, R.; Yang, D.X.; Duncan, J.S.; Omuro, A.; Herbst, R.; Krumholz, H.M.; Aneja, S. Using Adversarial Images to Assess the Robustness of Deep Learning Models Trained on Diagnostic Images in Oncology. JCO Clin. Cancer Inform. 2022, 6, e2100170. [Google Scholar] [CrossRef] [PubMed]
- Gichoya, J.W.; Banerjee, I.; Bhimireddy, A.R.; Burns, J.L.; Celi, L.A.; Chen, L.C.; Correa, R.; Dullerud, N.; Ghassemi, M.; Huang, S.C.; et al. AI Recognition of Patient Race in Medical Imaging: A Modelling Study. Lancet Digit. Health 2022, 4, e406–e414. [Google Scholar] [CrossRef] [PubMed]
- Lin, S.-Y.; Tsai, P.-C.; Su, F.-Y.; Chen, C.-Y.; Li, F.; Zhao, J.; Ho, Y.Y.; Lee, T.-M.; Healey, E.; Lin, P.J.; et al. Contrastive Learning Enhances Fairness in Pathology Artificial Intelligence Systems. Cell Rep. Med. 2025, 6, 102527. [Google Scholar] [CrossRef]
- Parikh, A.; Das, S.; Feragen, A. Investigating Label Bias and Representational Sources of Age-Related Disparities in Medical Segmentation. arXiv 2025, arXiv:2511.00477. [Google Scholar] [CrossRef]
- D’Amiano, A.J.; Cheunkarndee, T.; Azoba, C.; Chen, K.Y.; Mak, R.H.; Perni, S. Transparency and Representation in Clinical Research Utilizing Artificial Intelligence in Oncology: A Scoping Review. Cancer Med. 2025, 14, e70728. [Google Scholar] [CrossRef] [PubMed]
- Hou, J.; Liu, S.; Bie, Y.; Wang, H.; Tan, A.; Luo, L.; Chen, H. Self-eXplainable AI for Medical Image Analysis: A Survey and New Outlooks. arXiv 2024, arXiv:2410.02331. [Google Scholar] [CrossRef]
- Narkhede, J. Comparative Evaluation of Post-Hoc Explainability Methods in AI: LIME, SHAP, and Grad-CAM. In Proceedings of the 2024 4th International Conference on Sustainable Expert Systems (ICSES 2024), Chennai, India, 12–13 December 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 826–830. [Google Scholar] [CrossRef]
- Chien, J.-C.; Lee, J.-D.; Hu, C.-S.; Wu, C.-T. The Usefulness of Gradient-Weighted CAM in Assisting Medical Diagnoses. Appl. Sci. 2022, 12, 7748. [Google Scholar] [CrossRef]
- Cheng, Z.; Wu, Y.; Li, Y.; Cai, L.; Ihnaini, B. A Comprehensive Review of Explainable Artificial Intelligence (XAI) in Computer Vision. Sensors 2025, 25, 4166. [Google Scholar] [CrossRef]
- Talaat, F.M.; Gamel, S.A.; El-Balka, R.M.; Shehata, M.; Eldin, H.Z. Grad-CAM Enabled Breast Cancer Classification with a 3D Inception-ResNet V2: Empowering Radiologists with Explainable Insights. Cancers 2024, 16, 3668. [Google Scholar] [CrossRef]
- Ye, T.; Li, S.; Zhang, Y. Genomic Pan-Cancer Classification Using Image-Based Deep Learning. Comput. Struct. Biotechnol. J. 2021, 19, 835–846. [Google Scholar] [CrossRef] [PubMed]
- Arravalli, T.; Chadaga, K.; Muralikrishna, H.; Sampathila, N.; Cenitta, D.; Rajagopala Chadaga, R.; Swathi, K.S. Detection of Breast Cancer Using Machine Learning and Explainable Artificial Intelligence. Sci. Rep. 2025, 15, 26931. [Google Scholar] [CrossRef]
- Moldovanu, S.; Munteanu, D.; Biswas, K.C.; Moraru, L. Breast Lesion Detection Using Weakly Dependent Customized Features and Machine Learning Models with Explainable Artificial Intelligence. J. Imaging 2025, 11, 135. [Google Scholar] [CrossRef]
- Sharma, S.; Singh, M.; McDaid, L.; Bhattacharyya, S. XAI-based Data Visualization in Multimodal Medical Data. bioRxiv 2025. [Google Scholar] [CrossRef]
- Chhetri, B.; Rathish Kumar, B.V. Bridging Accuracy and Interpretability: Deep Learning with XAI for Breast Cancer Detection. arXiv 2025, arXiv:2510.21780. [Google Scholar] [CrossRef]
- Jiang, X.; Wang, S.; Zhang, Y. Vision Transformer Promotes Cancer Diagnosis: A Comprehensive Review. Expert Syst. Appl. 2024, 252, 124113. [Google Scholar] [CrossRef]
- Takagi, Y.; Hashimoto, N.; Masuda, H.; Miyoshi, H.; Ohshima, K.; Hontani, H.; Takeuchi, I. Transformer-based Personalized Attention Mechanism for Medical Images with Clinical Records. J. Pathol. Inform. 2023, 14, 100185. [Google Scholar] [CrossRef] [PubMed]
- Merjane, V.; Perin, D.M.P.; El Bacha, P.M.G.; Maranhão Miranda, B.M.; Bitencourt, A.G.V.; Iared, W. Breast Imaging Reporting and Data System (BI-RADS®®): A Success History and Particularities of Its Use in Brazil. Rev. Bras. Ginecol. Obstet. 2024, 46, e-rbgo6. [Google Scholar] [CrossRef]
- Räz, T.; Pahud De Mortanges, A.; Reyes, M. Explainable AI in Medicine: Challenges of Integrating XAI into the Future Clinical Routine. Front. Radiol. 2025, 5, 1627169. [Google Scholar] [CrossRef]
- Patrício, C.; Teixeira, L.F.; Neves, J.C. A two-step concept-based approach for enhanced interpretability and trust in skin lesion diagnosis. Comput. Struct. Biotechnol. J. 2025, 28, 71–79. [Google Scholar] [CrossRef]
- Koh, P.W.; Nguyen, T.; Tang, Y.S.; Mussmann, S.; Pierson, E.; Kim, B.; Liang, P. Concept Bottleneck Models. arXiv 2020, arXiv:2007.04612. [Google Scholar] [CrossRef]
- Yang, Y.; Gandhi, M.; Wang, Y.; Wu, Y.; Yao, M.S.; Callison-Burch, C.; Gee, J.C.; Yatskar, M. A textbook remedy for domain shifts: Knowledge priors for medical image analysis. arXiv 2024, arXiv:2405.14839. [Google Scholar] [CrossRef]
- Sun, S.; Tessier, L.; Meeuwsen, F.; Grisi, C.; van Midden, D.; Litjens, G.; Baumgartner, C.F. Label-free Concept Based Multiple Instance Learning for Gigapixel Histopathology. arXiv 2025, arXiv:2501.02922. [Google Scholar] [CrossRef]
- Kim, C.; Gadgil, S.U.; DeGrave, A.J.; Omiye, J.A.; Cai, Z.R.; Daneshjou, R.; Lee, S.-I. Transparent medical image AI via an image–text foundation model grounded in medical literature. Nat. Med. 2024, 30, 1154–1165. [Google Scholar] [CrossRef]
- Cheng, X.; Niu, Z.; Jiang, Z.; Li, L. Enhancing bottleneck concept learning in image classification. Sensors 2025, 25, 2398. [Google Scholar] [CrossRef]
- Parisini, E.; Chakraborti, T.; Harbron, C.; MacArthur, B.D.; Banerji, C.R.S. Leakage and interpretability in concept-based models. arXiv 2025, arXiv:2504.14094. [Google Scholar] [CrossRef]
- Moffett, L.; Barnett, A.J.; Donnelly, J.; Schwartz, F.R.; Trivedi, H.; Lo, J.; Rudin, C. Multi-site validation of an interpretable model to analyse breast masses. PLoS ONE 2025, 20, e0320091. [Google Scholar] [CrossRef]
- De Santi, L.A.; Piparo, F.I.; Bargagna, F.; Santarelli, M.F.; Celi, S.; Positano, V. Part-Prototype Models in Medical Imaging: Applications and Current Challenges. BioMedInformatics 2024, 4, 2149–2172. [Google Scholar] [CrossRef]
- Rives, G.; Lopez, O.; Bousquet, N. WTNN: Weibull-Tailored Neural Networks for survival analysis. arXiv 2025, arXiv:2512.09163. [Google Scholar] [CrossRef]
- Gupta, M.; Cotter, A.; Pfeifer, J.; Voevodski, K.; Canini, K.; Mangylov, A.; Moczydlowski, W.; van Esbroeck, A. Monotonic calibrated interpolated look-up tables. arXiv 2015, arXiv:1505.06378. [Google Scholar] [CrossRef]
- Gatsak, T.; Abhishek, K.; Ben Yedder, H.; Asgari Taghanaki, S.; Hamarneh, G. Disentangled PET lesion segmentation. arXiv 2024, arXiv:2411.01758. [Google Scholar] [CrossRef]
- Bercea, C.I.; Wiestler, B.; Rueckert, D.; Albarqouni, S. Federated disentangled representation learning for unsupervised brain anomaly detection. Nat. Mach. Intell. 2022, 4, 685–695. [Google Scholar] [CrossRef]
- Murthy, R.S.; Stassen, S.V.; Siu, D.M.D.; Lo, M.C.K.; Yip, G.G.K.; Tsia, K. Generalizable morphological profiling of cells by interpretable unsupervised learning. Nat. Commun. 2025, 16, 11465. [Google Scholar] [CrossRef] [PubMed]
- García-Barragán, A.; Sakor, A.; Vidal, M.-E.; Menasalvas, E.; Sanchez Gonzalez, J.C.; Provencio, M.; Robles, V. NSSC: A neuro-symbolic AI system for enhancing accuracy of named entity recognition and linking from oncologic clinical notes. Med. Biol. Eng. Comput. 2025, 63, 749–772. [Google Scholar] [CrossRef]
- Goisauf, M.; Cano Abadía, M.; Akyüz, K.; Bobowicz, M.; Buyx, A.; Colussi, I.; Fritzsche, M.C.; Lekadir, K.; Marttinen, P.; Mayrhofer, M.T.; et al. Trust, Trustworthiness, and the Future of Medical AI: Outcomes of an Interdisciplinary Expert Workshop. J. Med. Internet Res. 2025, 27, e71236. [Google Scholar] [CrossRef] [PubMed]
- Study Finds Most People Trust Doctors More Than AI but See Its Potential for Cancer Diagnosis, Source: EurekAlert! News Release 1106860, Published 8 December 2025. Available online: https://www.eurekalert.org/news-releases/1106860 (accessed on 19 February 2026).
- Dhar, T.; Dey, N.; Borra, S.; Sherratt, R.S. Challenges of deep learning in medical image analysis—Improving explainability and trust. IEEE Trans. Technol. Soc. 2023, 4, 68–75. [Google Scholar] [CrossRef]
- Lekadir, K.; Frangi, A.F.; Porras, A.R.; Glocker, B.; Cintas, C.; Langlotz, C.P.; Weicken, E.; Asselbergs, F.W.; Prior, F.; Collins, G.S.; et al. FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ 2025, 388, e081554. [Google Scholar] [CrossRef]
- Singh, Y.; Andersen, J.B.; Hathaway, Q.; Venkatesh, S.K.; Gores, G.J.; Erickson, B. Deep learning-based uncertainty quantification for quality assurance in hepatobiliary imaging-based techniques. Oncotarget 2025, 16, 249–255. [Google Scholar] [CrossRef]
- MacDonald, S.; Foley, H.; Yap, M.; Johnston, R.L.; Steven, K.; Koufariotis, L.T.; Sharma, S.; Wood, S.; Addala, V.; Pearson, J.V.; et al. Generalizing uncertainty improves accuracy and safety of deep learning analytics applied to oncology. Sci. Rep. 2023, 13, 7395. [Google Scholar] [CrossRef]
- Mumuni, F.; Mumuni, A. Explainable artificial intelligence (XAI): From inherent explainability to large language models. arXiv 2025, arXiv:2501.09967. [Google Scholar] [CrossRef]
- Cestonaro, C.; Delicati, A.; Marcante, B.; Caenazzo, L.; Tozzo, P. Defining medical liability when artificial intelligence is applied on diagnostic algorithms: A systematic review. Front. Med. 2023, 10, 1305756. [Google Scholar] [CrossRef]
- Xue, P.; Si, M.; Qin, D.; Wei, B.; Seery, S.; Ye, Z.; Chen, M.; Wang, S.; Song, C.; Zhang, B.; et al. Unassisted clinicians versus deep learning–assisted clinicians in image-based cancer diagnostics: Systematic review with meta-analysis. J. Med. Internet Res. 2023, 25, e43832. [Google Scholar] [CrossRef]
- Sanford, T.; Harmon, S.A.; Turkbey, E.B.; Kesani, D.; Tuncer, S.; Madariaga, M.; Yang, C.; Sackett, J.; Mehralivand, S.; Yan, P.; et al. Deep-learning-based artificial intelligence for PI-RADS classification to assist multiparametric prostate MRI interpretation: A development study. J. Magn. Reson. Imaging 2020, 52, 1499–1507. [Google Scholar] [CrossRef] [PubMed]
- Gavade, A.B.; Gavade, P.A.; Nerli, R.B.; Cooper, D.C.; Sztandera, L.; Mehta, U. Artificial intelligence in prostate cancer diagnosis: A systematic review of advances in Gleason grade and PI-RADS classification. Imaging 2025, 17, 91–103. [Google Scholar] [CrossRef]
- Saha, A.; van Ginneken, B.; Bjartell, A.; Bonekamp, D.; Villeirs, G.; Salomon, G.; Giannarini, G.; Kalpathy-Cramer, J.; Barentsz, J.; Rusu, M.; et al. Artificial intelligence and radiologists in prostate cancer detection on MRI (PI-CAI): An international, paired, non-inferiority, confirmatory study. Lancet Oncol. 2024, 25, 879–887. [Google Scholar] [CrossRef]
- Wang, X.; Wang, Q.; Ding, G.; Wang, J.; Tang, Y.; Feng, Y. Artificial intelligence in multidisciplinary tumour boards enhancing decision making and clinical outcomes in oncology. iScience 2025, 28, 114082. [Google Scholar] [CrossRef] [PubMed]
- Cheung, J.L.S.; Ali, A.; Abdalla, M.; Fine, B. U”AI” Testing: User interface and usability testing of a chest X-ray AI tool in a simulated real-world workflow. Can. Assoc. Radiol. J. 2023, 74, 314–325. [Google Scholar] [CrossRef] [PubMed]
- Sauter, D.; Lodde, G.; Nensa, F.; Schadendorf, D.; Livingstone, E.; Kukuk, M. Validating automatic concept-based explanations for AI-based digital histopathology. Sensors 2022, 22, 5346. [Google Scholar] [CrossRef]
- Levins, H. In the Loop or on the Loop: The Conundrum of AI Clinical Decision Support 2025 Penn Nudges in Health Care Symposium Focuses on the Human–Machine Interface. Available online: https://ldi.upenn.edu/our-work/research-updates/in-the-loop-or-on-the-loop-the-conundrum-of-ai-clinical-decision-support/ (accessed on 19 February 2026).
- Luo, M.; Yousefirizi, F.; Rouzrokh, P.; Jin, W.; Alberts, I.; Gowdy, C.; Bouchareb, Y.; Hamarneh, G.; Klyuzhin, I.; Rahmim, A. Physician-in-the-Loop Active Learning in Radiology Artificial Intelligence Workflows: Opportunities, Challenges, and Future Directions. AJR Am. J. Roentgenol. 2025, 225, e2533364. [Google Scholar] [CrossRef]
- Aresta, G.; Ferreira, C.; Pedrosa, J.; Araújo, T.; Rebelo, J.; Negrão, E.; Morgado, M.; Alves, F.; Cunha, A.; Ramos, I. Automatic lung nodule detection combined with gaze information improves radiologists’ screening performance. IEEE J. Biomed. Health Inform. 2020, 24, 2894–2901. [Google Scholar] [CrossRef]
- Available online: https://eurohealthobservatory.who.int/monitors/pace/case-studies/pace/pace-bulgaria-2025/propa-360-ai-tele-oncology-platform-(shemha-health-s-apci-clinical-dec (accessed on 31 January 2026).
- Prince, E.W.; Mirsky, D.M.; Hankinson, T.C.; Görg, C. Impact of AI decision support on clinical experts’ radiographic interpretation of Adamantinomatous Craniopharyngioma. AMIA Annu. Symp. Proc. 2025, 2024, 930–939. [Google Scholar] [PubMed] [PubMed Central]
- Chen, D.; Parsa, R.; Swanson, K.; Nunez, J.-J.; Critch, A.; Bitterman, D.S.; Liu, F.-F.; Raman, S. Large language models in oncology: A review. BMJ Oncol. 2025, 4, e000759. [Google Scholar] [CrossRef]
- Bader, T. Artificial intelligence in healthcare diagnostics: A literature review. Open J. Appl. Sci. 2025, 15, 4110–4133. [Google Scholar] [CrossRef]
- Behzad, S.; Tabatabaei, S.M.H.; Lu, M.Y.; Eibschutz, L.S.; Gholamrezanezhad, A. Pitfalls in interpretive applications of artificial intelligence in radiology. AJR Am. J. Roentgenol. 2024, 223, e2431493. [Google Scholar] [CrossRef]
- Hicks, S.A.; Strümke, I.; Thambawita, V.; Hammou, M.; Riegler, M.A.; Halvorsen, P.; Parasa, S. On evaluation metrics for medical applications of artificial intelligence. Sci. Rep. 2022, 12, 5979. [Google Scholar] [CrossRef] [PubMed]
- Jin, W.; Li, X.; Hamarneh, G. Evaluating explainable AI on a multi-modal medical imaging task: Can existing algorithms fulfil clinical requirements? arXiv 2022, arXiv:2203.06487. [Google Scholar] [CrossRef]
- Nastoska, A.; Jancheska, B.; Rizinski, M.; Trajanov, D. Evaluating trustworthiness in AI: Risks, metrics, and applications across industries. Electronics 2025, 14, 2717. [Google Scholar] [CrossRef]
- Budzyń, K.; Romańczyk, M.; Kitala, D.; Kołodziej, P.; Bugajski, M.; Adami, H.O.; Blom, J.; Buszkiewicz, M.; Halvorsen, N.; Hassan, C.; et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: A multicentre, observational study. Lancet Gastroenterol. Hepatol. 2025, 10, 896–903. [Google Scholar] [CrossRef] [PubMed]
- Mello-Thoms, C.; Mello, C.A.B. Clinical applications of artificial intelligence in radiology. Br. J. Radiol. 2023, 96, 20221031. [Google Scholar] [CrossRef] [PubMed]
- Datta, S.; Buchireddygari, D.; Kaza, L.V.; Bhalke, M.; Singh, K.; Pandey, A.; Vasipalli, S.S.; Karnwal, U.; Bhatti, H.B.S.; Maroo, B.R.; et al. Radiology’s Last Exam (RadLE): Benchmarking frontier multimodal AI against human experts and a taxonomy of visual reasoning errors in radiology. arXiv 2025, arXiv:2509.25559. [Google Scholar] [CrossRef]
- Lång, K.; Josefsson, V.; Larsson, A.-M.; Larsson, S.; Högberg, C.; Sartor, H.; Hofvind, S.; Andersson, I.; Rosso, A. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): A clinical safety analysis of a randomized, controlled, non-inferiority, single-blinded, screening accuracy study. Lancet Oncol. 2023, 24, 936–944. [Google Scholar] [CrossRef]
- De Koning, H.J.; Van Der Aalst, C.M.; De Jong, P.A.; Scholten, E.T.; Nackaerts, K.; Heuvelmans, M.A.; Lammers, J.-W.J.; Weenink, C.; Yousaf-Khan, U.; Horeweg, N.; et al. Reduced lung-cancer mortality with volume CT screening in a randomized trial. N. Engl. J. Med. 2020, 382, 503–513. [Google Scholar] [CrossRef]
- Song, B.; Sunny, S.; Li, S.; Gurushanth, K.; Mendonca, P.; Mukhia, N.; Patrick, S.; Gurudath, S.; Raghavan, S.; Tsusennaro, I.; et al. Bayesian deep learning for reliable oral cancer image classification. Biomed. Opt. Express. 2021, 12, 6422–6430. [Google Scholar] [CrossRef]
- Lambert, B.; Forbes, F.; Doyle, S.; Dehaene, H.; Dojat, M. Trustworthy clinical AI solutions: A unified review of uncertainty quantification in deep learning models for medical image analysis. Artif. Intell. Med. 2024, 150, 102830. [Google Scholar] [CrossRef]
- Gal, Y.; Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML 2016), New York, NY, USA, 20–22 June 2016; pp. 1050–1059. [Google Scholar]
- Delporte, G. Robustness Analysis of a Deep Learning–Based sCT Generation Algorithm. Master Thesis, University of Liège, Liège, Belgium, 2025. Available online: https://matheo.uliege.be/handle/2268.2/23235 (accessed on 30 January 2026).
- Milanés-Hermosilla, D.; Trujillo Codorniú, R.; López-Baracaldo, R.; Sagaró-Zamora, R.; Delisle-Rodriguez, D.; Villarejo-Mayor, J.J.; Núñez-Álvarez, J.R. Monte Carlo dropout for uncertainty estimation and motor imagery classification. Sensors 2021, 21, 7241. [Google Scholar] [CrossRef]
- Kurz, A.; Hauser, K.; Mehrtens, H.A.; Krieghoff-Henning, E.; Hekler, A.; Kather, J.N.; Fröhling, S.; von Kalle, C.; Brinker, T.J. Uncertainty estimation in medical image classification: Systematic review. JMIR Med. Inform. 2022, 10, e36427. [Google Scholar] [CrossRef] [PubMed]
- Choudhury, K.; Roy, S.; Chanda, A.; Biswas, S.; Kuiry, S. Improving predictive confidence in medical imaging via online label smoothing. BIO Web. Conf. 2025, 204, 01019. [Google Scholar] [CrossRef]
- Sinaga, K.P.; Nair, A.S. Calibration meets reality: Making machine learning predictions trustworthy. arXiv 2025, arXiv:2509.23665. [Google Scholar] [CrossRef]
- Li, S.; Yuan, M.; Dai, X.; Zhang, C. Evaluation of uncertainty estimation methods in medical image segmentation: Exploring the usage of uncertainty in clinical deployment. Comput. Med. Imaging Graph. 2025, 124, 102574. [Google Scholar] [CrossRef]
- Odesola, P.A.; Adegoke, A.A.; Babalola, I. Model uncertainty quantification: A post hoc calibration approach for heart disease prediction. medRxiv 2025. [Google Scholar] [CrossRef]
- Tomani, C.; Cremers, D.; Buettner, F. Parameterized temperature scaling for boosting the expressive power in post-hoc uncertainty calibration. arXiv 2021, arXiv:2102.12182. [Google Scholar] [CrossRef]
- Hong, Z.; Yue, Y.; Chen, Y.; Lin, H.; Luo, Y.; Wang, M.H.; Wang, W.; Xu, J.; Yang, X.; Li, Z.; et al. Out-of-distribution detection in medical image analysis: A survey. arXiv 2024, arXiv:2404.18279. [Google Scholar] [CrossRef]
- Mehta, R.; Shui, C.; Arbel, T. Evaluating the fairness of deep learning uncertainty estimates in medical image analysis. arXiv 2023, arXiv:2303.03242. [Google Scholar] [CrossRef]
- Yu, Y.; Gomez-Cabello, C.A.; Haider, S.A.; Genovese, A.; Prabha, S.; Trabilsy, M.; Collaco, B.G.; Wood, N.G.; Bagaria, S.; Tao, C.; et al. Enhancing clinician trust in AI diagnostics: A dynamic framework for confidence calibration and transparency. Diagnostics 2025, 15, 2204. [Google Scholar] [CrossRef]
- Gulum, M.A.; Trombley, C.M.; Kantardzic, M. Multiple Interpretations Improve Deep Learning Transparency for Prostate Lesion Detection. In Heterogeneous Data Management, Polystores, and Analytics for Healthcare; Gadepally, V., Mattson, T., Stonebraker, M., Kraska, T., Wang, F., Luo, G., Kong, J., Dubovitskaya, A., Eds.; Springer: Cham, Switzerland, 2021; Volume 12633, pp. 120–137. [Google Scholar] [CrossRef]
- Malafaia, M.; Silva, F.; Neves, I.; Pereira, T.; Oliveira, H.P. Robustness Analysis of Deep Learning-Based Lung Cancer Classification Using Explainable Methods. IEEE Access 2022, 10, 112731–112741. [Google Scholar] [CrossRef]
- Hägele, M.; Seegerer, P.; Lapuschkin, S.; Bockmayr, M.; Samek, W.; Klauschen, F.; Müller, K.-R.; Binder, A. Resolving Challenges in Deep Learning-Based Analyses of Histopathological Images Using Explanation Methods. Sci. Rep. 2020, 10, 6423. [Google Scholar] [CrossRef] [PubMed]
- ISO/IEC TS 6254:2025; Artificial Intelligence—Objectives and Approaches for Explainability and Interpretability of Machine Learning (ML) Models and Artificial Intelligence (AI) Systems. ISO/IEC JTC 1/SC 42: Geneva, Switzerland, 2025.
- ISO/IEC TR 24028:2020; Information Technology—Artificial Intelligence—Overview of Trustworthiness in Artificial Intelligence. ISO/IEC JTC 1/SC 42: Geneva, Switzerland, 2020.


| Ref. | Imaging Modality | Cancer Site | Clinical Task | Model Architecture | Performance Metrics | Challenges and Limitations | Dataset Used | Dataset Category |
|---|---|---|---|---|---|---|---|---|
| Clinical Decision Support (CDSS) and Interpretation | ||||||||
| [55] | Ultrasound | Breast | Clinical decision support | CNN + U-Net | Accuracy: 81% | Automation bias: Risk of overreliance; Cognitive load: Visual complexity of localization | Breast ultrasound (780 images) | Public |
| [56] | Mammography | Breast | Enhance diagnosis and interpretation | VGG/Inception/ResNet | Test Accuracy: 76% | Method instability: LIME varies across runs; Inconsistency: Lower accuracy of explanations | CBIS-DDSM | Public |
| [57] | Whole Slide Imaging | Kidney | Concept learning/Survival analysis | GNNs | AUC: 0.789; C-Index: 0.725 | Concept overlap: Difficulty defining distinct concepts; Rarity: Rare concepts show lower performance | TCGA-RCC | Public |
| [58] | Ultrasound, MRI | Prostate | Early-stage detection; explanation of decision rationale | Pretrained DL (VGG-16, ResNet, Xception, etc.) + shallow ML (SVM, RF, etc.) | US Acc: 99%; MRI Acc: 87.5% | Local surrogacy: Requirement for local data; Input quality: Dependence on image quality | Prostate-MRI-US-Biopsy | Public |
| Image Recognition and Object Detection | ||||||||
| [59] | Biomedical images | Bladder | Automatic image recognition, grade/contour recognition | Deep Belief Network (DBN), ConvNeXt, RegNet, MaxViT | Accuracy: 98.75%; F1: 98.57% | Interpretability thresholds: Gaps in existing models; Constraints: Real-world noise applicability | Bladder cancer classification | Public |
| [39] | MRI | Brain | Automated object detection | YOLOv11 with Attention | mAP50: 96.8% | Intrinsic Limitation: Study lacks dedicated XAI (LIME/SHAP); Hardware: Constraints on inference | Brain Tumor (Kaggle) | Public |
| [60] | Histopathology | Breast | Detection through feature extraction | Modified DenseNet169 | Accuracy: 99.50%; F1: 99.42% | Interpretability gap: Absence of advanced features; Case sensitivity: Difficulty with subtle benign cases | BreakHis, BACH, BHI | Public |
| [61] | Histopathology | Lung | Cancer detection | EfficientNetB3, Custom CNN | Accuracy: 95.1%; F1: 89.5% | Coarseness: Grad-CAM explanations can be non-specific; Sensitivity: To non-malignant features. | TCGA and LC25000 | Public |
| [62] | Whole Slide Imaging | Lymph node | Tumour tissue detection | Custom CNN, VGG19 | Accuracy: 0.9683 | Unreliability: Methods sensitive to superpixel parameters; Complexity: High computational intensity | PatchCamelyon dataset (P-CAM) | Public |
| [63] | mpMRI | Prostate | Lesion detection and characterization | Cascaded FCN | sensitivity: 93.9% | Redundancy: Inefficiency for obvious cases; Bias: Reference standard bias in multi-institutional studies | 11 MR devices; 3 institutions | Multi-institutional clinical data |
| [64] | Abdominal CT | Vertebra | Metastasis classification | EMCD + DenseNet201 | Accuracy: 85.79%; AUC: 0.93 | Alignment: Uncertainty maps may lack clinical alignment; Calibration: MCDO/DE limitations | Severance Hospital cohort | Private/Internal |
| Screening, Early Detection and Mortality Prediction | ||||||||
| [65] | Mammography | Breast | Organized screening | Deep CNNs (Vara MG) | BCDR: 6.7/1000; PPV: 17.9% | Lack of software interoperability, Binary confidence constraint, Initial feature gaps | Description of the 463,094 women | Multi-institutional clinical data |
| [66] | SERS Biosensing (SEARCH Chip) | Liver | Early detection and staging | Self-Learning CNN | AUC: 0.97; Accuracy: 0.87 | Thresholds: Dimensionality reduction below 20 features reduces accuracy; Actionability gap | Serum samples (300 subjects) | Multi-population cohort consisting of serum samples from 300 subjects |
| [67] | CT | Lung | Early detection and CAD | LCxNet (Custom CNN) | Accuracy: 99.39%; F1: 99.40% | Ambiguity: Grad-CAM heatmaps can be diffuse; Validation: Needs multi-metric expert review | IQ-OTH/NCCD | IQ-OTH/NCCD lung cancer dataset |
| Segmentation and Analysis | ||||||||
| [68] | Multimodal MRI | Brain | Segmentation and analysis | CausalX-Net (DL + SCM) | DSC: 93.2%; HD95: 0.91 mm | Assumptions: High dependency on causal assumptions; Complexity: Implementation difficulty | BraTS 2021; ISLES 2017 | Multi-institutional clinical data |
| [69] | MRI | Brain | Segmentation | 3D U-Net | Dice Coefficient: 0.73 | Constraint Gap: Lack of spatial constraints; Bias: Reliance on signal intensity | Institutional brain tumour data | Private/Internal |
| [70] | Ultrasound | Breast | Segmentation and classification | HyFormer-Net | Accuracy: 93.2%; Dice: 0.902 | Baseline Diffuseness: Grad-CAM limitations; Validation: Lack of quantitative validation | BUSI; BUS-UCLM | Public |
| [42] | Ultrasound | Breast | Image segmentation | LIME/SHAP | Accuracy: 72.0% | Inconsistent feature attribution across models | Breast ultrasound images | Public |
| [71] | CT, Chest X-ray | Lung | Classification and segmentation | Proto-Caps | Accuracy: 95.3% | Persuasiveness Risk: Users may be misled if model is wrong; Complexity: High memory demands | LIDC-IDRI, CheXpert | Public |
| Tumor Classification and Grading | ||||||||
| [72] | MRI | Brain | Multiclass classification | SSPANet | Accuracy: 97%; Kappa: 95% | Coarseness: Grad-CAM coarse focus; Disconnect: XAI often disconnected from core architecture | Figshare Brain Tumor | Public |
| [73] | MRI | Brain | Glioma grading and localization | ResNet-50, 3D DeepSeg | Accuracy: 98.62%; Dice: 92 | Voxel Gaps: Grad-CAM low-resolution; Noisy: Vanilla Gradient produces noisy visualisations | BraTS 2019/2021 | Public/multi-institutional clinical data |
| [74] | 3D mpMRI | Brain | 3D brain tumour classification | MProtoNet | Bal. Accuracy: 0.870 | Localization Coherence: Poor Grad-CAM performance; Complexity: 3D localization issues | BraTS 2020 | Public |
| [40] | MRI | Brain | Detection and Classification | YOLOv11 (Two-stage) | Accuracy: 92.6%; F1: 0.899 | Sensitivity: To MRI variations; Manual Dependency: Gap in automated bridging | BTDM, BTDS datasets | Public |
| [75] | FLAIR MRI, US | Brain, Breast | Tumour classification | SpikeNet (Hybrid) | Accuracy: 98.23%; AUC: 0.996 | Visual Inaccuracy: Grad-CAM boundary spillover; Noise: Fragmented SHAP/LIME maps | TCGA–LGG; BUSI | Public/multi-institutional clinical data |
| [3] | Ultrasound | Breast | Classification | Hybrid Model Fusion | Accuracy: 97.14%; F1: 97.18% | Visual Constraint: Lacks automated reasoning support; Reasoning gap: Traditional XAI limits | Ultrasound breast images | Public |
| [76] | Ultrasound | Breast | Classifying tumours; associations with clinical descriptors | Multitask DL (VGG/ResNet encoders) | Accuracy 88.9%, Sensitivity 83.8%, and Specificity 92.3% | Clinical Trustworthiness Inter-observer Variability Alignment with Medical Practice Information Insufficiency | BUSIS dataset | Public |
| [36] | H&E Stained Images | Breast | Subtyping and classification | HACT-Net | Weighted F1: 84.15%; Weighted Accuracy: 63.21%; Concordance: 90% | Sensitive to entity detection accuracy | BRACS; BACH | Public/multi-institutional |
| [77] | Tomosynthesis (DBT) | Breast | Shape-based classification of lesions | 8 Pretrained CNNs | AUC: 98.2% | Coarseness: Grad-CAM coarseness; Unclear: LIME visually less clear | 39 breast DBT exams | Private/Internal |
| [78] | Mammography | Breast | Binary classification | CNN | Accuracy: 0.9675; AUC: 0.9937 | Partial Explanation: Lack of proof vs. confidence; Opacity: “Black Box” problem | RSNA-Breast-Cancer | Public |
| [79] | Histopathology | Breast, Colon | Tumour classification | Simple/Mini-GoogLeNet | Accuracy: (High) | Calibration: Uncertainty miscalibration; Latent space: “Black Box” latent representation | CAMELYON17; AIDA-LNCO | Public/multi-institutional |
| [80] | Histopathology; Tomosynthesis; X-ray | Breast, Lung | Classification | ResNet (18/34/50) | Accuracy 97.61 Expected Calibration Error (ECE) 0.0095 | Overconfidence: Confidence gap bias; Scale: Dataset scale sensitivity | BreakHis, BCS-DBT, Lung | Public/multi-institutional |
| [81] | Histopathological (Herlev) and Cytological (CIVa) images | Cervix, Ovary | Multiclass classification | RIRXEnsemble | Accuracy: 99.88%; AUC: 1.00 | Perturbation: LIME reliance on random samples; Logic: Grad-CAM lacks clinical reasoning. | Herlev, CIVa, Mendeley | Public |
| [82] | Histopathology | Colon | Colon cancer diagnosis | Few-shot (ProtoNet) | Accuracy: 98.5%; AUC: 1.00 | LIME Stability: Random perturbation reliance; Resolution: Grad-CAM coarse highlighting. | LC25000; EBHI | Public/multi-institutional |
| [37] | Tissue Images/WSIs | Multi-organ | Classification and grading | SCUBa-Net (GNN) | Accuracy: 93.0%; F1: 0.841 | Latency: High computational complexity; Resolution: Explanation resolution differences | Colorectal, Prostate, Gastric | Public/multi-institutional |
| [83] | Whole Slide Imaging | Colorectal, Bone | Classification | AAOXAI-CD (Ensemble) | Accuracy: 9.42%; F-Score: 98.87% | Trade-off: Accuracy-interpretability trade-off; Dependency: Perturbation dependency | Warwick-QU; Osteosarcoma | Public |
| [84] | Histopathology | Lung | Classification | Bayesian Xception | Accuracy: 94.1%; AUC: 0.94–0.96 | Threshold Sensitivity: OOD false confidence; Coverage: Reduced data coverage during abstention | TCGA, CPTAC, Mayo Clinic | Public/Private/Multi-inst. |
| [85] | CT | Lung | Identification and classification | FVCM-Net (Federated) | Accuracy: 97.66%; AUC: 0.995 | Tuning Sensitivity: Manual tuning requirements; Bias: Static client weighting. | LIDC-IDRI, IQ-OTH/NCCD | Public |
| [20] | Biparametric MRI | Prostate | Grading aggressiveness | 3D VGG/ResNet/ViT | AUC: 0.73 | Context Gap: Lack of spatial context; Variability: Protocol/scanner variation. | ProstateNet dataset | Public |
| [86] | Dermoscopy | Skin | Lesion classification | DenseNet, ResNet | Accuracy: 93.28%; AUC: 99.64% | Semantic Gap: Pixel-based maps lack reasoning; Sensitivity: High hyperparameter dependency | HAM10000 | Public |
| [87] | Dermoscopy | Skin | Detection and classification | VT-CNN | Accuracy: 99.89%; F1: 99.389% | Interference: Grad-CAM background noise; Compensation: Lack of HOA-XAI compensation | ISIC-2019, HAM10000 | Public |
| [88] | Dermoscopy | Skin | Multiclass classification | XceSCNN | Accuracy: 92.643% | Approximation: Requirement for SHAP approximation; Subjectivity: Interpretation difficulty | ISIC | Public |
| [89] | Dermoscopy | Skin | Recognize lesion types | ResNet + ABELE | Bal. Accuracy: 0.838 | Latency: Time-consuming extraction; Quality: Dependent on autoencoder reconstruction | ISIC 2019 | Public |
| [90] | Dermoscopy | Skin | Categorization | ResNet-50/VGG16 | Accuracy: 87–96% | Approximation: Local emphasis approximation issues; Robustness: Concerns in manual correction | ISIC and HAM10000 | Public |
| [91] | Histopathology | Skin | Margin classification | ViT Transfer Learning | Accuracy: 0.928 | Opacity: Intrinsic non-interpretable nature; Trust Gap: Subjective interpretation | Histopathological slides (50 pts) | Public |
| Broad Obstacle Domain | Subcategory | Technical Constraints | Trade-Offs | Clinical Implication |
|---|---|---|---|---|
| Data availability, Quality and bias | Dataset scarcity and imbalance | Small sample sizes for rare cancers; minority underrepresentation; fragmented health records | Statistical power vs. generalisability | Inflated performance in development cohorts; failure in real-world deployment |
| Annotation burden and subjectivity | Manual segmentation; pixelwise labelling; inter/intraobserver variability | Annotation fidelity vs. scalability | Limits dataset growth: inconsistent ground truth undermines trust | |
| Acquisition heterogeneity | Scanner/vendor bias; staining variability; resolution differences | Dataset diversity vs. distributional stability | Domain shift across institutions; degraded external validation | |
| Model architecture and Computational complexity | Model scale and capacity | Large parameter counts; quadratic self-attention; long training times | Expressivity vs. feasibility | High-performing models impractical for clinical infrastructure |
| Memory and efficiency constraints | Gigapixel WSIs; 3D context handling; client-side memory limits | Spatial resolution vs. deployability | Reduced resolution or patching compromises spatial reasoning | |
| Optimization stability | Vanishing gradients; unstable losses; manual hyperparameter tuning | Training stability vs. architectural flexibility | Reproducibility challenges across centres | |
| Generalization and robustness | Distribution shift and OOD data | Non-IID clinical data; temporal drift; confounding variables | Robustness vs. dataset specificity | Performance decay after deployment |
| Sensitivity to perturbations | Noise, artifacts, color statistics shifts | Sensitivity vs. feature richness | Erratic predictions in routine clinical settings | |
| Cross-site variability | Single-centre studies; single-vendor training | Controlled performance vs. external validity | Limited regulatory acceptance | |
| Interpretability, Explainability and human factors | Black-box opacity | Deep feature abstraction; unclear decision logic | Predictive accuracy vs. interpretability | Clinician reluctance to adopt AI outputs |
| Explanation instability | Saliency inconsistency; scattered attention maps | Transparency vs. reliability | Undermines confidence in safety-critical decisions | |
| Human subjectivity | Subjective visual validation; low interrater agreement | Human intuition vs. algorithmic rigor | Inconsistent evaluation standards | |
| Clinical integration and workflow constraints | Workflow disruption | Manual pre/postprocessing; lack of automation | Model precision vs. usability | Poor adoption despite strong benchmarks |
| Device and infrastructure limits | Hardware constraints; lack of clinical-grade devices | Model sophistication vs. bedside deployment | Restricts real-time or point-of-care use | |
| Contextual incompleteness | Image-only analysis; missing clinical metadata | Task simplicity vs. clinical realism | Reduced decision relevance | |
| Regulatory, Ethical and Legal barriers | Privacy and governance | compliance; restricted data sharing | Privacy vs. reproducibility | Limits multi-institutional validation |
| Accountability and liability | Medico-legal responsibility; risk of automation bias | Autonomy vs. oversight | Slows approval of autonomous systems | |
| Regulatory lag | High certification costs; evolving standards | Innovation speed vs. safety assurance | Delays translation into standard of care | |
| Evaluation, validation and trustworthiness | Calibration and uncertainty | Overconfident predictions; epistemic uncertainty | Sensitivity vs. reliability | Unsafe decision-making in high-risk oncology |
| Metric inadequacy | Lack of standard reliability metrics | Benchmark simplicity vs. clinical relevance | Misleading performance claims | |
| External validation gaps | Retrospective designs; selection bias | Study control vs. real-world validity | Weak evidence for deployment |
| Category | Approach | Strengths | Limitations | Suitability in Oncology |
|---|---|---|---|---|
| Post hoc explainability | Gradient-based (Grad-CAM, Grad-CAM++) | Computationally efficient; architecture-agnostic; visually intuitive localization | Coarse spatial resolution; sensitive to input noise; may highlight non-pathological artifacts (e.g., surgical markers, scanner bias) | Screening and initial tumour localization (e.g., mammography, ultrasound) where rapid visual verification is required |
| Perturbation-based (LIME, SHAP) | Strong local fidelity (LIME); theoretically grounded feature attribution (SHAP); supports multimodal inputs | Computationally expensive; sensitive to segmentation granularity; instability across repeated runs | Multimodal analysis combining imaging with genomic/clinical data to identify prognostic drivers | |
| Attention-based (Vision Transformers) | Captures global context; does not require pixel-level annotations; effective for large images | Limited interpretability of attention weights; unclear causal relationship between attention and prediction | Whole-slide histopathology analysis where tissue architecture is critical | |
| Intrinsic/hybrid interpretable models | Concept Bottleneck Models (CBMs) | Clinically interpretable; supports human-in-the-loop refinement; aligns with medical ontologies | High annotation burden; risk of concept leakage; dependent on concept quality | Staging and prognosis tasks (e.g., TNM-based cancer assessment) |
| Prototype-based (ProtoPNet, MProtoNet) | Provides intuitive “this looks like that” reasoning; links decision to exemplars | Limited flexibility with fixed prototypes; challenges in multimodal fusion | Diagnostic classification (MRI, mammography) aligned with radiologist reasoning patterns | |
| Disentangled representations | Improves robustness to domain shift; enables interpretable feature isolation | Latent factors may lack direct clinical meaning; interpretability not always guaranteed | Biomarker discovery and cross-site generalization (e.g., PET, MRI) |
| Ref. | Clinical Objective | Deep Learning Architecture | Explainability Method | Reported Validation of Explainability | Performance Metrics | Reported Explainability Limitations | Dataset Used | Dataset Category |
|---|---|---|---|---|---|---|---|---|
| [3] | Early cancer detection | Hybrid Deep Fusion: VGG16 + DenseNet121 + Xception | Grad-CAM++ | Visually consistent with expert assessment | Accuracy: 97.14%; F1: 97.18% | Reasoning Gap: Lacks automated clinical reasoning support; limited degree of interpretability beyond visual overlays | Ultrasound breast images | Public |
| [7] | Explainable diagnosis and trust | ResNet50 and DenseNet121 (Pretrained) | Grad-CAM | Localized core pathological regions | Accuracy: 94.3%; AUC: 0.99 | Coarseness: Diffuse activations; Semantic inconsistency: High probability of highlighting non-salient features | Brain MRI; Chest X-ray | Public |
| [37] | Histopathology classification | SCUBa-Net (GCN + Transformer) | Grad-CAM | Activation maps aligned with histopathology | Accuracy: 93.0%; F1: 0.833 | Computational overhead: Slow inference generation; Propagation bias: Sensitivity to graph node aggregation logic | Colorectal, Prostate, Gastric, Bladder | Public/multi-institutional |
| [36] | Breast tumour subtyping | HACT-Net (Hierarchical GNN) | GraphGradCAM | Feature attribution aligned with pathology | Weighted F1: 84.15%; Accuracy: 63.21% | Structural dependency: Explanation fidelity sensitive to the accuracy of the underlying entity detection nodes | BRACS; BACH | Public/multi-institutional |
| [59] | Efficient bladder screening | ConvNeXt + RegNet X + MaxViT + DBN | SHAP | Identified individual feature contributions | Accuracy: 98.75%; F1: 98.57% | Computational complexity: SHAP approximation requirements reduce global explanation fidelity | Bladder cancer classification | Public |
| [86] | XAI effectiveness in skin cancer | DenseNet, ResNet, MobileNet | Integrated Gradients, SHAP, LIME | Mathematical saliency regions vs. clinical meaningfulness | Accuracy: 93.28%; AUC: 99.64% | Semantic Gap: Pixel-based maps do not translate to semantic ABCDE criteria; Robustness: High hyperparameter sensitivity | HAM10000 | Public |
| [56] | Trust enhancement via diagnostics | ResNet50 (fine-tuned) | Grad-CAM, LIME, SHAP | Hausdorff distance to expert ROIs | Test Accuracy: 76% | Method Instability: LIME is unstable across runs; Inconsistency: Lower accuracy of explanations compared to model | CBIS-DDSM (Mammography) | Public |
| [88] | Multiclass skin classification | XceSCNN (XCovNet + SCNN + ELANet) | SHAP | Global SHAP plots of feature influence | Accuracy: 92.643%; F1: 96.47% | Complexity: SHAP becomes inefficient for large models; Approximation: Requirement for approximation reduces fidelity | ISIC Dataset | Public |
| [81] | Early multiclass cell classification | RIRXEnsemble (ResNet + InceptionResNet + Xception) | Grad-CAM, LIME | Confidence Drop; 3D visualization (t-SNE) | Accuracy: 99.88%; AUC: 1.00 | Perturbation Bias: LIME reliance on random samples; Logic Gap: Grad-CAM lacks complete clinical reasoning logic | Herlev, CIVa, Mendeley LBC | Public |
| [62] | Tumor detection explanation | Custom CNN + VGG19 | LIME (SLIC, FHA, etc.) | Heatmaps aligned with expert knowledge | Accuracy: 0.9683 | Unreliability: Super pixel methods are sensitive to parameters; Granularity: May lack global contextual reasoning | Patch Camelyon (P-CAM) | Public |
| [73] | Transparent brain diagnosis | ResNet-50 + 3D DeepSeg | NeuroXAI (VG, GBP, IG, GIG, SmoothGrad, Grad-CAM) | Network inspection of hierarchical detection | Accuracy: 98.62%; Dice: 92 | Noisy visuals: Vanilla Gradient is noisy; Resolution: Grad-CAM suffers from low-resolution heatmaps | BraTS 2019 and 2021 | Public/multi-institutional |
| [61] | Interpretable pathology support | Bespoke CNN + EfficientNetB3 | Grad-CAM | Alignment between focus and annotations (0.78) | Accuracy: 95.1%; F1: 89.5% | Spatial resolution gap: Diffuse heatmaps lacking pathological specificity | TCGA and LC25000 | Public |
| [60] | Feature-driven cancer detection | Modified DenseNet | CAM, Saliency Map | Comprehensive reasoning via CAM/Saliency | Accuracy: 99.50%; F1: 99.42% | Resolution constraint: Inability to delineate fine morphological features required for subtle case differentiation | BreakHis, BACH, BHI | Public |
| [83] | Explainable cancer classification | Faster SqueezeNet + RNN Ensemble | LIME | Clear explanations for black-box predictions | Accuracy: 99.42%; F1: 98.87% | Perturbation dependency: Results depend on sampling; Trade-off: Accuracy-interpretability trade-off | Warwick-QU, Osteosarcoma | Public |
| [72] | Context-aware detection | SSPANet | Grad-CAM, Grad-CAM++, EigenGradCAM | Noise-free heatmaps aligned with anatomy | Accuracy: 97%; Kappa: 95% | Architectural disconnect: Disparity between XAI output and core feature design; Low-rank approximation limits | Figshare Brain Tumor | Public |
| [85] | Privacy-preserving detection | FVCM-Net (VGG16 + CBAM) | SHAP, HiRes-CAM | Boundaries confirmed by radiologist | Accuracy: 97.66%; AUC: 0.995 | Manual parametrization: High sensitivity to kernel width and sampling hyperparameters | LIDC-IDRI, IQ-OTH/NCCD | Public |
| [50] | Personalized survival risk | CVAE-based DySurv with LSTM | Permutation importance | Time-dependent concordance metrics | C-Index: 70.4%; IBS: 0.122 | Interaction blindness: Failure to account for non-linear feature interactions in risk attribution | MIMIC-IV and eICU | Public |
| [66] | Monitoring of early-stage hepatocellular carcinoma | Self-Learning CNN | SHAP (Feature extraction) | Reduced dimensionality by 95% | AUC: 0.97; Accuracy: 0.87 | Thresholds: Accuracy drops if reduction is too aggressive; Actionability: Gap between SHAP values and clinical action | Serum samples (300 subjects) | Multi-institutional |
| Ref. | Clinical Objective | Deep Learning Architecture | Explainability Method | Reported Validation of Explainability | Performance Metrics | Reported Explainability Limitations | Dataset Used | Dataset Category |
|---|---|---|---|---|---|---|---|---|
| [76] | Clinically aligned CAD systems | Multitask learning backbone | BI-RADS descriptors + tumour class | Descriptors consistent with clinical practice | Accuracy: 88.9%; | Lexicon restriction: Explanations constrained by a predefined medical vocabulary | BUSIS; BUSI | Public |
| [89] | Exemplar-based trust enhancement | ResNet + ABELE (PGAAE) | ABELE (Exemplars and Saliency) | +22% confidence increase after correcting errors | Balanced Accuracy: 0.838; RMSE: 0.08–0.24 | Latency: Explanation extraction is time-consuming; Quality: Highly dependent on autoencoder reconstruction fidelity | ISIC 2019 | Public |
| [10] | High-performance explainable diagnosis | Concept Complement Bottleneck (CCBM) | Ante-hoc concept-based model | Faithfulness via concept intervention | AUC: 93.96%; ACC: 88.15% | Concept divergence: Discovered concepts may lack clinical semantic alignment; Potential concept leakage | Derm7pt, Skincon, BrEaST, LIDC-IDRI | Public/External |
| [82] | Explainable few-shot cancer diagnosis | ProtoNet (ConvNeXt-Tiny) | Grad-CAM, LIME, prototypes | Validated by panel of 4 medical professionals | Accuracy: 98.5%; ROC-AUC: 1.000 | LIME Instability: Reliance on random perturbations; Grad-CAM Resolution: Lack of complete textual reasoning | LC25000; EBHI | Public/multi-institutional |
| [68] | Causality-aware tumour segmentation | CausalX-Net (3D U-Net + SCM) | Structural Causal Models | Counterfactual maps identified 81% of edema errors | Dice (WT): 93.2%; HD95: 0.91 mm | Complexity: High implementation difficulty; Assumptions: High dependency on causal assumptions | BraTS 2021; ISLES 2017 | Public/multi-institutional |
| [71] | lung nodule classification | Proto-Caps (Capsule + Prototype) | Visual prototypes + attribute scores | Alignment confirmed via expert rater study | Accuracy: 95.3%; Faithfulness: 0.62 | Persuasiveness Risk: Users may be misled if model is wrong; Complexity: High memory/computational demands | LIDC-IDRI; CheXpert | Public |
| [74] | 3D case-based tumour classification | MProtoNet (3D ResNet + online-CAM) | Case-based reasoning (prototypes) | Improved localization over Grad-CAM | Balanced Accuracy: 0.870 ± 0.021 | Fixed Assignments: Prototype assignments are fixed; Fusion Gaps: Difficulty analyzing individual modalities | BraTS 2020 | Public |
| [57] | RCC subtyping and survival prediction | GNN-based Concept Learning | Concept Bottleneck Models (CBMs) | High-risk concepts linked to mortality | Balanced Acc: 0.682; C-Index: 0.725 | Concept overlap: Difficulty defining distinct local concepts; Rarity: Rare concepts show lower performance | TCGA-RCC | Public |
| [156] | Bias-aware interpretable profiling | MorphoGenie (VAE + GAN) | Disentangled learning | Outperformed baseline VAEs in reconstruction | AUC (Actin): 0.87; Nucleoli: 0.83 | Interpretation Gap: Latent features lack clear morphological interpretation; Reductionist Framework | LC, CPA, CCy, EMT | Public/Private |
| [87] | Automated skin cancer detection | OXAI-SCC-Net | Grad-CAM guided NAS; HOA-XAI | Highlighted medically relevant borders | Accuracy: 99.89%; F1: 99.389% | Interference: Grad-CAM includes background noise; Lack of compensation: Missing HOA-XAI compensation | ISIC-2019; HAM10000 | Public |
| [70] | Joint lesion segmentation/classification | HyFormer-Net | Intrinsic Attention + Grad-CAM | Mean IoU 0.86 (attention-boundary alignment) | Accuracy: 93.2 ± 0.6%; Dice: 0.902 | Baseline diffuseness: Coarse localization of Grad-CAM; Attention over-smoothing | BUSI; BUS-UCLM | Public |
| [73] | Voxelwise brain tumour segmentation | 3D U-Net + Multinomial Dirichlet | NeuroXAI (VG, GBP, IG, Grad-CAM) | Uncertainty maps identified model ignorance | Accuracy: 98.62%; Dice: 92 | Noisy visuals: Vanilla Gradient noise; Resolution: Grad-CAM suffers from low-resolution heatmaps | BraTS 2019/2021 | Public/multi-institutional |
| [64] | Vertebral metastasis detection | EMCD (YOLOv5m + DenseNet201) | Uncertainty-CAM | Accuracy improved to 95.68% (50% retention) | Accuracy: 85.79%; AUC: 0.93 | Clinical alignment: Uncertainty maps may not perfectly align clinically; MCDO/DE limitations | Severance Hospital cohort | Private/Internal |
| [84] | Confidence-calibrated prediction | DCNN (Xception/ResNet) + MCDO | Uncertainty thresholding | Experts confirmed biological relevance of maps | AUROC: 0.945–0.966; Acc: 94.1% | Reduced coverage: ~25% cases abstain; Sensitivity: High threshold sensitivity and cost | TCGA, CPTAC, Mayo | Public/Private/multi-institutional |
| [67] | Deployable early cancer detection | LCxNet (Custom CNN) | Grad-CAM, t-SNE | t-SNE proved linear separability of features | Accuracy: 99.39%; F1: 99.40% | Spatial ambiguity: Low-resolution Grad-CAM heatmaps; Lack of structured evidence | IQ-OTH/NCCD; LC25000 | Public |
| [77] | Morphological shape classification | Pretrained CNN Ensemble | Grad-CAM, LIME, t-SNE | Correct identification improved performance | AUC: 98.2% | Coarseness: Grad-CAM produces coarse maps; Visual clarity: LIME super pixels are visually less clear | 39 breast DBT exams | Private/Internal |
| Domain | Clinically Meaningful Explanation (Requirement) | Decision Pathway (Alignment) | Primary Failure Mode (Misalignment) |
|---|---|---|---|
| Breast ultrasound | Localization of microcalcifications and spiculated margins | Alignment with BI-RADS descriptors for benign/malignant classification | High sensitivity to visual noise and artifacts, leading to false positives |
| Prostate MRI | Highlighting diffusion restriction and T2-weighted signal changes | Integration of PI-RADS v2.1 logic and contextual data (PSA, patient history) | Scanner/vendor variability and inability to resolve equivocal PI-RADS 3 lesions without context |
| Digital pathology | Translation of pixels into cellular/glandular constructs | Multiscale reasoning (5× to 20×) from cellular morphology to tissue architecture | Shortcut learning from staining protocols and annotation sparsity in gigapixel slides |
| Multimodal oncology | Correlation between metabolic uptake (PET) and anatomical detail (CT) | Alignment of macroscale patterns with genomic/longitudinal trajectories | Overfitting on small, paired datasets and missing temporal data points |
| Attribute | Visually Persuasive Explanations | Trust Calibrating Explanations |
|---|---|---|
| Common methods | Grad-CAM, Saliency maps | Concept bottleneck models, Uncertainty maps |
| Primary goal | Highlighting where the model looks | Clarifying why the model is certain or uncertain |
| Cognitive impact | Increases load through visual clutter and “authority modulation” | Streamlines triage by flagging high-entropy cases for review |
| Trust outcome | Risk of “illusion of interpretability” and automation bias | Promotes epistemic alignment with clinical reasoning |
| Success metric | Plausibility (looks right to the eye) | Faithfulness/Fidelity (reflects true model logic) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Borra, S.; Dey, N.; Fong, S.; Sherratt, R.S.; Shi, F. Explainability and Trust in Deep Learning for Cancer Imaging: Systematic Barriers, Clinical Misalignment, and a Translational Roadmap. Cancers 2026, 18, 1361. https://doi.org/10.3390/cancers18091361
Borra S, Dey N, Fong S, Sherratt RS, Shi F. Explainability and Trust in Deep Learning for Cancer Imaging: Systematic Barriers, Clinical Misalignment, and a Translational Roadmap. Cancers. 2026; 18(9):1361. https://doi.org/10.3390/cancers18091361
Chicago/Turabian StyleBorra, Surekha, Nilanjan Dey, Simon Fong, R. Simon Sherratt, and Fuqian Shi. 2026. "Explainability and Trust in Deep Learning for Cancer Imaging: Systematic Barriers, Clinical Misalignment, and a Translational Roadmap" Cancers 18, no. 9: 1361. https://doi.org/10.3390/cancers18091361
APA StyleBorra, S., Dey, N., Fong, S., Sherratt, R. S., & Shi, F. (2026). Explainability and Trust in Deep Learning for Cancer Imaging: Systematic Barriers, Clinical Misalignment, and a Translational Roadmap. Cancers, 18(9), 1361. https://doi.org/10.3390/cancers18091361

