Bidirectional Perceptual Multimodal Interaction Network Based on Contrastive Learning for Breast Cancer pCR Prediction
Simple Summary
Abstract
1. Introduction
- We propose the BiCMA fusion mechanism to address modality heterogeneity and semantic misalignment through dynamic bidirectional calibration, comprising IGC-Attention and CGI-Attention. This mechanism establishes stable cross-modal semantic associations by mining complementary inter-modal information, thereby bridging the gap between high-dimensional imaging features and discrete clinical indicators for pCR prediction.
- We design the MCFE module as a key component tightly integrated into the pCR-oriented contrastive learning framework. Integrating multimodal perceptual dynamic calibration, semantic selection, dual-path refinement, and a feedback mechanism leveraging intra-class similarity for adaptive sample activation, it enhances feature discriminability and generalization by strengthening the activation of hard-to-classify samples and establishes a unique closed-loop interaction between feature refinement and contrastive optimization to ensure robustness to challenging pCR cases.
- Experimental results on two public multicenter breast cancer datasets reveal that BPMINet consistently outperforms existing approaches, confirming the effectiveness of its multimodal interaction mechanism for superior pCR prediction.
2. Related Work
3. Materials and Methods
3.1. Datasets
3.2. DCE-MRI Image Preprocessing
3.3. Clinical Information Preprocessing
3.4. Method Overview
3.4.1. 3D Vision Transformer for Global Encoding of DCE-MRI
3.4.2. Adaptive Clinical Semantic Representation Learning
3.4.3. Bidirectional Cross-Modal Attention Fusion Mechanism
3.4.4. Multimodal Contrast-Aware Feature Enhancement Module
3.4.5. pCR-Oriented Contrastive Learning
- Positive sample set: comprises all samples in the batch sharing the same pCR label as sample i, excluding the sample itself.
- Scaled similarity: , where is the temperature parameter used to adjust the similarity distribution.
- Non-self samples: , representing all samples in the batch except for sample i itself.
- Intra-class similarity: , which denotes the average scaled similarity between sample i and its positive samples.
3.4.6. Dual-Loss Collaborative Optimization
4. Experiments
4.1. Implementation Details
- Optimizer: The AdamW optimizer [40] was employed with a weight decay of 0.01. A layer-wise learning rate strategy was implemented: the initial learning rate for the 3D ViT encoder was set to , while the learning rate for other non-pre-trained modules was set to . This distinction was made to preserve the pre-trained parameters of the ViT model.
- Learning Rate Scheduler: The CosineAnnealingLR scheduler [41] was utilized to perform periodic adaptive adjustment of the learning rate.
- Hyperparameters: The batch size was set to 4, and the total number of training epochs was set to 50.
4.2. Evaluation Metrics
- Accuracy (ACC): Represents the proportion of correctly classified pCR and non-pCR samples among all cases:
- Positive Predictive Value (PPV): Also known as Precision, it represents the proportion of samples that the model predicts as pCR and which are actually pCR:
- Negative Predictive Value (NPV): Evaluates the proportion of samples that the model predicts as non-pCR and which are actually non-pCR:
- Sensitivity (SEN): Also known as Recall, it quantifies the model’s ability to correctly identify true pCR cases:
- Specificity (SPE): Quantifies the model’s ability to correctly identify true non-pCR cases:
- F1 Score (F1): The harmonic mean of PPV and SEN, providing a balanced assessment especially under class imbalance:
- Area Under the Curve (AUC): Refers to the area under the Receiver Operating Characteristic (ROC) curve, which plots the True Positive Rate () against the False Positive Rate (). Mathematically, AUC quantifies the overall discriminative performance by integrating the TPR over the full range of FPR:A higher AUC value signifies a superior ability to distinguish between pCR and non-pCR samples, with a value of 1.0 representing a perfect classifier.
5. Results
5.1. Performance Comparison with Other Methods
- CI-UM: Unimodal methods leveraging only clinical information.
- DCE-MRI-UM: Unimodal methods leveraging only DCE-MRI images.
- Multimodality: Multimodal methods that leverage multimodal data for predictive tasks.
5.1.1. Unimodal Methods for Clinical Information
5.1.2. Unimodal Methods for DCE-MRI Images
5.1.3. Multimodal Methods
5.1.4. Confidence Interval and Statistical Analysis
5.2. Ablation Studies
5.3. Validation of Clinical Information Selection
5.4. Hyperparameter Sensitivity Analysis for Contrastive Learning
5.5. Visual Explainability of BPMINet via Grad-CAM
6. Discussion
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| pCR | Pathological Complete Response |
| NAC | Neoadjuvant Chemotherapy |
| DCE-MRI | Dynamic Contrast-Enhanced Magnetic Resonance Imaging |
| MRI | Magnetic Resonance Imaging |
| CNN | Convolutional Neural Network |
| 3D | Three-Dimensional |
| 2D | Two-Dimensional |
| ADC | Apparent Diffusion Coefficient |
| TCIA | The Cancer Imaging Archive |
| TCGA | The Cancer Genome Atlas |
| Z-score | Standard Score |
| ER | Estrogen Receptor |
| PR | Progesterone Receptor |
| HR | Hormone Receptor |
| HER2 | Human Epidermal Growth Factor Receptor 2 |
| MLP | Multi-Layer Perceptron |
| LayerNorm | Layer Normalization |
| ReLU | Rectified Linear Unit |
| GlobalAvgPool | Global Average Pooling |
| GPU | Graphics Processing Unit |
| AdamW | Adam with Weight Decay |
| CosineAnnealingLR | Cosine Annealing Learning Rate Scheduler |
| CE | Cross-Entropy |
| AUC | Area Under the Receiver Operating Curve |
| ACC | Accuracy |
| F1 | F1 Score |
| PPV | Positive Predictive Value |
| NPV | Negative Predictive Value |
| SPE | Specificity |
| SEN | Sensitivity |
| SD | Standard Deviation |
| 5-FCV | 5-Fold Cross-Validation |
| CI | Confidence Interval |
| BPMINet | Bidirectional Perceptual Multimodal Interactive Network |
| BiCMA | Bidirectional Cross-Modal Attention |
| MCFE | Multimodal Contrast-Aware Feature Enhancement |
| CL | Contrastive Learning |
| MCFE-CL | MCFE Module Integrated into Contrastive Learning |
| W/o | Without |
| Grad-CAM | Gradient-Weighted Class Activation Mapping |
| IGC-Attention | Imaging-Guided Clinical Attention |
| CGI-Attention | Clinical-Guided Imaging Attention |
Appendix A
| Methods | AUC (95% CI) | ACC (95% CI) | SEN (95% CI) | SPE (95% CI) | F1 (95% CI) | PPV (95% CI) | NPV (95% CI) |
|---|---|---|---|---|---|---|---|
| BERT [43] | 0.6909 (0.5842, 0.7936) | 0.6968 (0.6124, 0.7831) | 0.5555 (0.4218, 0.6842) | 0.7546 (0.6632, 0.8415) | 0.5147 (0.4023, 0.6214) | 0.4882 (0.3741, 0.6053) | 0.8065 (0.7328, 0.8742) |
| MLP [44] | 0.7030 (0.6014, 0.8021) | 0.7226 (0.6482, 0.7915) | 0.4889 (0.3724, 0.6018) | 0.8182 (0.7135, 0.9126) | 0.5069 (0.4125, 0.5982) | 0.5504 (0.4231, 0.6742) | 0.7966 (0.7314, 0.8583) |
| DenseNet [45] | 0.6389 (0.5214, 0.7486) | 0.6387 (0.5421, 0.7284) | 0.4000 (0.2814, 0.5132) | 0.7364 (0.6415, 0.8247) | 0.3912 (0.2831, 0.4952) | 0.3873 (0.2741, 0.4936) | 0.7498 (0.6712, 0.8235) |
| ConVit [46] | 0.4889 (0.4021, 0.5732) | 0.6709 (0.5846, 0.7512) | 0.2222 (0.1142, 0.3325) | 0.8545 (0.7712, 0.9324) | 0.2826 (0.1654, 0.3981) | 0.4124 (0.2641, 0.5582) | 0.7283 (0.6652, 0.7891) |
| ResNet-50 [47] | 0.6243 (0.5142, 0.7315) | 0.6129 (0.5214, 0.7012) | 0.6445 (0.5126, 0.7732) | 0.6000 (0.4823, 0.7145) | 0.4914 (0.3952, 0.5831) | 0.4043 (0.3124, 0.4952) | 0.8065 (0.7341, 0.8732) |
| ViT [34] | 0.6495 (0.5512, 0.7431) | 0.6839 (0.6012, 0.7642) | 0.4889 (0.3621, 0.6084) | 0.7636 (0.6642, 0.8561) | 0.4762 (0.3721, 0.5742) | 0.4782 (0.3614, 0.5892) | 0.7834 (0.7126, 0.8514) |
| SIDLN [30] | 0.6803 (0.5912, 0.7645) | 0.6774 (0.5842, 0.7621) | 0.6222 (0.4614, 0.7752) | 0.7000 (0.5842, 0.8124) | 0.5161 (0.4124, 0.6152) | 0.4719 (0.3712, 0.5684) | 0.8317 (0.7512, 0.9042) |
| BERT-ViT | 0.7162 (0.6214, 0.8052) | 0.7742 (0.6912, 0.8512) | 0.6000 (0.4532, 0.7412) | 0.8454 (0.7612, 0.9242) | 0.5978 (0.4831, 0.7052) | 0.6278 (0.5124, 0.7381) | 0.8430 (0.7712, 0.9124) |
| Interactive-Model [13] | 0.7515 (0.6631, 0.8324) | 0.6710 (0.5812, 0.7564) | 0.6222 (0.5012, 0.7382) | 0.6909 (0.5842, 0.7931) | 0.5231 (0.4214, 0.6215) | 0.4588 (0.3531, 0.5594) | 0.8185 (0.7423, 0.8872) |
| TMSS [48] | 0.7010 (0.6123, 0.7842) | 0.6710 (0.5814, 0.7562) | 0.3556 (0.2124, 0.4952) | 0.8000 (0.7012, 0.8941) | 0.3541 (0.2312, 0.4721) | 0.5382 (0.3842, 0.6852) | 0.7546 (0.6712, 0.8342) |
| Integrated-Model [49] | 0.7626 (0.6742, 0.8461) | 0.6581 (0.5742, 0.7362) | 0.5556 (0.4321, 0.6742) | 0.7000 (0.5912, 0.8042) | 0.4816 (0.3812, 0.5794) | 0.4586 (0.3524, 0.5614) | 0.7992 (0.7214, 0.8712) |
| MRI-RNA [24] | 0.7586 (0.6632, 0.8514) | 0.7871 (0.7123, 0.8582) | 0.6000 (0.4741, 0.7214) | 0.8636 (0.7688, 0.9573) | 0.6144 (0.5012, 0.7231) | 0.6491 (0.5312, 0.7624) | 0.8431 (0.7714, 0.9124) |
| CITR-Net [25] | 0.7788 (0.6912, 0.8624) | 0.6968 (0.6124, 0.7782) | 0.4444 (0.3214, 0.5632) | 0.8000 (0.7042, 0.8912) | 0.4311 (0.3214, 0.5384) | 0.4742 (0.3612, 0.5842) | 0.7882 (0.7124, 0.8612) |
| AER-SwinT [32] | 0.7434 (0.6532, 0.8312) | 0.6710 (0.5824, 0.7542) | 0.5111 (0.3921, 0.6254) | 0.7364 (0.6412, 0.8272) | 0.4591 (0.3512, 0.5624) | 0.4484 (0.3342, 0.5582) | 0.7947 (0.7231, 0.8624) |
| BPMINet (ours) | 0.8475 (0.7356, 0.9594) | 0.8452 (0.7366, 0.9538) | 0.7778 (0.5470, 0.9989) | 0.8727 (0.7756, 0.9698) | 0.7406 (0.5503, 0.9309) | 0.7190 (0.5469, 0.8911) | 0.9088 (0.8220, 0.9956) |
| Methods | AUC (95% CI) | ACC (95% CI) | SEN (95% CI) | SPE (95% CI) | F1 (95% CI) | PPV (95% CI) | NPV (95% CI) |
|---|---|---|---|---|---|---|---|
| BERT [43] | 0.6503 (0.5986, 0.7022) | 0.6321 (0.5816, 0.6826) | 0.5054 (0.4274, 0.5832) | 0.6893 (0.6324, 0.7388) | 0.4608 (0.3815, 0.5406) | 0.4234 (0.3461, 0.5012) | 0.7553 (0.6854, 0.8258) |
| MLP [44] | 0.6513 (0.5823, 0.7218) | 0.6756 (0.6187, 0.7291) | 0.3871 (0.2911, 0.4853) | 0.8058 (0.7375, 0.8655) | 0.4260 (0.3484, 0.5073) | 0.4737 (0.3500, 0.5938) | 0.7444 (0.6794, 0.8049) |
| DenseNet [45] | 0.5432 (0.4722, 0.6140) | 0.6187 (0.5487, 0.6871) | 0.2796 (0.1932, 0.3647) | 0.7718 (0.7150, 0.8284) | 0.3133 (0.2293, 0.3981) | 0.3562 (0.2641, 0.4607) | 0.7035 (0.6473, 0.7534) |
| ConVit [46] | 0.5451 (0.4748, 0.6163) | 0.6321 (0.5821, 0.6853) | 0.4194 (0.3401, 0.5035) | 0.7282 (0.6649, 0.7872) | 0.4149 (0.3478, 0.4865) | 0.4105 (0.3377, 0.4901) | 0.7353 (0.6683, 0.8021) |
| ResNet-50 [47] | 0.5359 (0.4521, 0.6197) | 0.6087 (0.5452, 0.6722) | 0.2043 (0.1241, 0.2847) | 0.7913 (0.7261, 0.8568) | 0.2452 (0.1821, 0.3089) | 0.3065 (0.1379, 0.4828) | 0.6878 (0.6290, 0.7410) |
| ViT [34] | 0.6053 (0.5392, 0.6724) | 0.6455 (0.5591, 0.7322) | 0.4086 (0.3027, 0.5143) | 0.7524 (0.6632, 0.8452) | 0.4176 (0.3358, 0.4998) | 0.4270 (0.2500, 0.6043) | 0.7381 (0.6625, 0.8132) |
| SIDLN [30] | 0.6688 (0.6050, 0.7295) | 0.6421 (0.5741, 0.7114) | 0.4946 (0.4251, 0.5622) | 0.7087 (0.6444, 0.7650) | 0.4623 (0.3864, 0.5385) | 0.4340 (0.3711, 0.4993) | 0.7565 (0.6604, 0.8431) |
| BERT-ViT | 0.5854 (0.5036, 0.6682) | 0.6589 (0.5842, 0.7128) | 0.3763 (0.2764, 0.4731) | 0.7864 (0.7383, 0.8351) | 0.4070 (0.2538, 0.5617) | 0.4430 (0.1975, 0.6875) | 0.7364 (0.6712, 0.8014) |
| Interactive-Model [13] | 0.5423 (0.4849, 0.6020) | 0.6555 (0.5875, 0.7215) | 0.2903 (0.2086, 0.3741) | 0.8204 (0.7667, 0.8744) | 0.3439 (0.2877, 0.4089) | 0.4219 (0.3498, 0.4885) | 0.7191 (0.6364, 0.7908) |
| TMSS [48] | 0.6170 (0.5406, 0.6845) | 0.6187 (0.5617, 0.6724) | 0.4624 (0.3636, 0.5647) | 0.6893 (0.6321, 0.7391) | 0.4300 (0.3632, 0.5002) | 0.4019 (0.3234, 0.4787) | 0.7396 (0.6776, 0.8000) |
| Integrated-Model [49] | 0.6938 (0.6423, 0.7456) | 0.6823 (0.6165, 0.7469) | 0.3763 (0.3101, 0.4470) | 0.8204 (0.7750, 0.8653) | 0.4242 (0.3529, 0.5061) | 0.4861 (0.3279, 0.6445) | 0.7445 (0.6880, 0.8015) |
| MRI-RNA [24] | 0.6905 (0.6222, 0.7484) | 0.6923 (0.6418, 0.7461) | 0.4516 (0.3727, 0.5326) | 0.8010 (0.7482, 0.8530) | 0.4773 (0.3989, 0.5568) | 0.5060 (0.3953, 0.6203) | 0.7639 (0.7040, 0.8165) |
| CITR-Net [25] | 0.6946 (0.6274, 0.7533) | 0.7124 (0.6573, 0.7638) | 0.4624 (0.3516, 0.5652) | 0.8252 (0.7691, 0.8798) | 0.5000 (0.4196, 0.5794) | 0.5443 (0.4742, 0.6151) | 0.7727 (0.7072, 0.8306) |
| AER-SwinT [32] | 0.7008 (0.6329, 0.7637) | 0.7023 (0.6455, 0.7525) | 0.4946 (0.3928, 0.5976) | 0.7961 (0.7418, 0.8498) | 0.5083 (0.4099, 0.5923) | 0.5227 (0.4167, 0.6292) | 0.7773 (0.7158, 0.8291) |
| BPMINet (ours) | 0.7370 (0.6730, 0.7953) | 0.7391 (0.6854, 0.7862) | 0.5376 (0.4338, 0.6324) | 0.8301 (0.7812, 0.8814) | 0.5618 (0.4687, 0.6428) | 0.5882 (0.4820, 0.6956) | 0.7991 (0.7401, 0.8487) |
References
- Polyak, K. Heterogeneity in Breast Cancer. J. Clin. Investig. 2011, 121, 3786–3788. [Google Scholar] [CrossRef] [Scilit]
- Mieog, J.S.D.; Van der Hage, J.A.; Van De Velde, C.J.H. Neoadjuvant Chemotherapy for Operable Breast Cancer. J. Br. Surg. 2007, 94, 1189–1200. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cortazar, P.; Zhang, L.; Untch, M.; Mehta, K.; Costantino, J.P.; Wolmark, N.; Bonnefoi, H.; Cameron, D.; Gianni, L.; Valagussa, P.; et al. Pathological Complete Response and Long-Term Clinical Benefit in Breast Cancer: The CTNeoBC Pooled Analysis. Lancet 2014, 384, 164–172. [Google Scholar] [CrossRef] [Scilit]
- Qi, Y.-J.; Su, G.-H.; You, C.; Zhang, X.; Xiao, Y.; Jiang, Y.-Z.; Shao, Z.-M. Radiomics in Breast Cancer: Current Advances and Future Directions. Cell Rep. Med. 2024, 5, 101719. [Google Scholar] [CrossRef] [Scilit]
- Moslemi, A.; Osapoetra, L.O.; Dasgupta, A.; Halstead, S.; Alberico, D.; Trudeau, M.; Gandhi, S.; Eisen, A.; Wright, F.; Look-Hong, N.; et al. Prediction of Chemotherapy Response in Locally Advanced Breast Cancer Patients at Pre-Treatment Using CT Textural Features and Machine Learning: Comparison of Feature Selection Methods. Tomography 2025, 11, 33. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ogut, Z.; Karaduman, M.; Yildirim, M. Clinically Focused Computer-Aided Diagnosis for Breast Cancer Using SE and CBAM with Multi-Head Attention. Tomography 2025, 11, 138. [Google Scholar] [CrossRef] [Scilit]
- Xiong, Z.; Zhao, K.; Ji, L.; Shu, X.; Long, D.; Chen, S.; Yang, F. Multi-modality 3D CNN Transformer for Assisting Clinical Decision in Intracerebral Hemorrhage. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Marrakesh, Morocco, 6–10 October 2024; pp. 522–531. [Google Scholar]
- Bibault, J.-E.; Giraud, P.; Housset, M.; Durdux, C.; Taieb, J.; Berger, A.; Coriat, R.; Chaussade, S.; Dousset, B.; Nordlinger, B.; et al. Deep Learning and Radiomics Predict Complete Response after Neo-adjuvant Chemoradiation for Locally Advanced Rectal Cancer. Sci. Rep. 2018, 8, 12611. [Google Scholar] [CrossRef] [Scilit]
- Qu, Y.-H.; Zhu, H.-T.; Cao, K.; Li, X.-T.; Ye, M.; Sun, Y.-S. Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy in Breast Cancer Using a Deep Learning (DL) Method. Thorac. Cancer 2020, 11, 651–658. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Luo, L.; Wang, X.; Lin, Y.; Ma, X.; Tan, A.; Chan, R.; Vardhanabhuti, V.; Chu, W.C.W.; Cheng, K.-T.; Chen, H. Deep Learning in Breast Cancer Imaging: A Decade of Progress and Future Directions. IEEE Rev. Biomed. Eng. 2024, 18, 130–151. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Liu, Y.; Liu, X.; Liu, Y.; Zhang, J. Prognoses of Patients with Hormone Receptor-Positive and Human Epidermal Growth Factor Receptor 2-Negative Breast Cancer Receiving Neoadjuvant Chemotherapy Before Surgery: A Retrospective Analysis. Cancers 2023, 15, 1157. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y.; Ventura-Diaz, S.; Wang, X.; He, M.; Xu, Z.; Weir, A.; Zhou, H.-Y.; Zhang, T.; van Duijnhoven, F.H.; Han, L.; et al. An Explainable Longitudinal Multi-Modal Fusion Model for Predicting Neoadjuvant Therapy Response in Women with Breast Cancer. Nat. Commun. 2024, 15, 9613. [Google Scholar] [CrossRef] [Scilit]
- Duanmu, H.; Huang, P.B.; Brahmavar, S.; Lin, S.; Ren, T.; Kong, J.; Wang, F.; Duong, T.Q. Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy in Breast Cancer Using Deep Learning with Integrative Imaging, Molecular and Demographic Data. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Lima, Peru, 4–8 October 2020; pp. 242–252. [Google Scholar]
- Li, Y.; Fan, Y.; Xu, D.; Li, Y.; Zhong, Z.; Pan, H.; Huang, B.; Xie, X.; Yang, Y.; Liu, B. Deep Learning Radiomic Analysis of DCE-MRI Combined with Clinical Characteristics Predicts Pathological Complete Response to Neoadjuvant Chemotherapy in Breast Cancer. Front. Oncol. 2023, 12, 1041142. [Google Scholar] [CrossRef] [Scilit]
- Syed, A.; Adam, R.; Ren, T.; Lu, J.; Maldjian, T.; Duong, T.Q. Machine Learning with Textural Analysis of Longitudinal Multiparametric MRI and Molecular Subtypes Accurately Predicts Pathologic Complete Response in Patients with Invasive Breast Cancer. PLoS ONE 2023, 18, e0280320. [Google Scholar] [CrossRef] [Scilit]
- Herrero Vicent, C.; Tudela, X.; Moreno Ruiz, P.; Pedralva, V.; Jimenez Pastor, A.; Ahicart, D.; Novella, S.R.; Meneu, I.; Albuixech, Á.M.; Santamaria, M.Á.; et al. Machine Learning Models and Multiparametric Magnetic Resonance Imaging for the Prediction of Pathologic Response to Neoadjuvant Chemotherapy in Breast Cancer. Cancers 2022, 14, 3508. [Google Scholar] [CrossRef] [Scilit]
- Huang, S.-C.; Pareek, A.; Seyyedi, S.; Banerjee, I.; Lungren, M.P. Fusion of Medical Imaging and Electronic Health Records Using Deep Learning: A Systematic Review and Implementation Guidelines. NPJ Digit. Med. 2020, 3, 136. [Google Scholar] [CrossRef] [Scilit]
- Liang, X.; Yu, X.; Gao, T. Machine Learning with Magnetic Resonance Imaging for Prediction of Response to Neoadjuvant Chemotherapy in Breast Cancer: A Systematic Review and Meta-Analysis. Eur. J. Radiol. 2022, 150, 110247. [Google Scholar] [CrossRef] [Scilit]
- Cui, C.; Yang, H.; Wang, Y.; Zhao, S.; Asad, Z.; Coburn, L.A.; Wilson, K.T.; Landman, B.A.; Huo, Y. Deep Multimodal Fusion of Image and Non-image Data in Disease Diagnosis and Prognosis: A Review. Prog. Biomed. Eng. 2023, 5, 022001. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.P.; Polosukhin, I. Attention Is All You Need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; p. 30. [Google Scholar]
- Wang, J.; Liu, X.; Gong, Z.; Yang, L.; Zhang, H.; Long, Y.; Fan, Y.; Jiang, Y.; Duan, X.; Zhao, W. HARM3-Fusion: Hierarchical Attentional Representation Learning of Multi-modal, Multi-temporal, and Multi-sequence Fusion for Pathological Complete Response Prediction of Head and Neck Squamous Cell Carcinoma. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Daejeon, Republic of Korea, 23–27 September 2025; pp. 246–255. [Google Scholar]
- Maruf, N.A.; Basuhail, A.; Ramzan, M.U. Enhanced Breast Cancer Diagnosis Using Multimodal Feature Fusion with Radiomics and Transfer Learning. Diagnostics 2025, 15, 2170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised Contrastive Learning. Adv. Neural Inf. Process. Syst. 2020, 33, 18661–18673. [Google Scholar]
- Li, H.; Zhao, Y.; Duan, J.; Gu, J.; Liu, Z.; Zhang, H.; Zhang, Y.; Li, Z.-C. MRI and RNA-seq Fusion for Prediction of Pathological Response to Neoadjuvant Chemotherapy in Breast Cancer. Displays 2024, 83, 102698. [Google Scholar] [CrossRef] [Scilit]
- Liu, T.; Wang, H.; Feng, F.; Li, W.; Zheng, F.; Wu, K.; Yu, S.; Sun, Y. Integrating Clinicopathologic Information and Dynamic Contrast-Enhanced MRI for Augmented Prediction of Neoadjuvant Chemotherapy Response in Breast Cancer. Biomed. Signal Process. Control 2025, 103, 107385. [Google Scholar] [CrossRef] [Scilit]
- Zhang, S.; Du, S.; Sun, C.; Li, B.; Shao, L.; Zhang, L.; Wang, K.; Liu, Z.; Tian, J. M2Fusion: Multi-time Multimodal Fusion for Prediction of Pathological Complete Response in Breast Cancer. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Marrakesh, Morocco, 6–10 October 2024; pp. 458–468. [Google Scholar]
- Guo, M.; Luo, Z.; Liu, J.; Zhou, R. Mathematically-Grounded Multimodal Attention Network for Breast Cancer Prognosis. In Proceedings of the 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), Lisboa, Portugal, 3–6 December 2024; pp. 4529–4536. [Google Scholar]
- Newitt, D.; Hylton, N.; on behalf of the I-SPY 1 Network and ACRIN 6657 Trial Team. Multi-Center Breast DCE-MRI Data and Segmentations from Patients in the I-SPY 1/ACRIN 6657 Trials. The Cancer Imaging Archive (TCIA). 2016. Available online: https://www.cancerimagingarchive.net/collection/ispy1/ (accessed on 12 May 2026).
- Chitalia, R.; Pati, S.; Bhalerao, M.; Thakur, S.P.; Jahani, N.; Belenky, V.; McDonald, E.S.; Gibbs, J.; Newitt, D.C.; Hylton, N.M.; et al. Expert Tumor Annotations and Radiomics for Locally Advanced Breast Cancer in DCE-MRI for ACRIN 6657/ISPY1. Sci. Data 2022, 9, 440. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gao, Y.; Ding, D.-W.; Zeng, H. A Self-Interpretable Deep Learning Network for Early Prediction of Pathologic Complete Response to Neoadjuvant Chemotherapy Based on Breast Pre-Treatment Dynamic Contrast-Enhanced Magnetic Resonance Imaging. Eng. Appl. Artif. Intell. 2024, 138, 109431. [Google Scholar] [CrossRef] [Scilit]
- Garrucho, L.; Kushibar, K.; Reidel, C.-A.; Joshi, S.; Osuala, R.; Tsirikoglou, A.; Bobowicz, M.; Del Riego, J.; Catanese, A.; Gwo’zdziewicz, K.; et al. A Large-Scale Multicenter Breast Cancer DCE-MRI Benchmark Dataset with Expert Segmentations. Sci. Data 2025, 12, 453. [Google Scholar] [CrossRef] [Scilit]
- Sang, S.; Sun, Z.; Zheng, W.; Wang, W.; Islam, M.T.; Chen, Y.; Yuan, Q.; Cheng, C.; Xi, S.; Han, Z.; et al. TME-Guided Deep Learning Predicts Chemotherapy and Immunotherapy Response in Gastric Cancer with Attention-Enhanced Residual Swin Transformer. Cell Rep. Med. 2025, 6, 102242. [Google Scholar] [CrossRef] [Scilit]
- Comes, M.C.; Fanizzi, A.; Bove, S.; Didonna, V.; Diotiaiuti, S.; Fadda, F.; La Forgia, D.; Giotta, F.; Latorre, A.; Nardone, A.; et al. Explainable 3D CNN Based on Baseline Breast DCE-MRI to Give an Early Prediction of Pathological Complete Response to Neoadjuvant Chemotherapy. Comput. Biol. Med. 2024, 172, 108132. [Google Scholar] [CrossRef] [Scilit]
- Dosovitskiy, A. An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Glorot, X.; Bordes, A.; Bengio, Y. Deep Sparse Rectifier Neural Networks. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Fort Lauderdale, FL, USA, 11–13 April 2011; pp. 315–323. [Google Scholar]
- Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer Normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar] [CrossRef] [Scilit]
- Zeng, X.; Li, L.; Liang, Y.; Chen, W.; Lei, B. Multiview Feature Fusion and Contrastive Learning for Drug-Target Interaction Prediction. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Daejeon, Republic of Korea, 23–27 September 2025; pp. 386–395. [Google Scholar]
- Li, W.; Liu, T.; Feng, F.; Yu, S.; Wang, H.; Sun, Y. BTSSPro: Prompt-Guided Multimodal Co-Learning for Breast Cancer Tumor Segmentation and Survival Prediction. IEEE J. Biomed. Health Inform. 2024, 28, 7322–7331. [Google Scholar] [CrossRef] [Scilit]
- Paszke, A.; Gross, S.; Chintala, S.; Chanan, G.; Yang, E.; DeVito, Z.; Lin, Z.; Desmaison, A.; Antiga, L.; Lerer, A. Automatic Differentiation in PyTorch. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 2017), Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Loshchilov, I.; Hutter, F. Decoupled Weight Decay Regularization. arXiv 2017, arXiv:1711.05101. [Google Scholar]
- Loshchilov, I.; Hutter, F. Sgdr: Stochastic Gradient Descent with Warm Restarts. arXiv 2016, arXiv:1608.03983. [Google Scholar]
- Efron, B. Bootstrap Methods: Another Look at the Jackknife. In Breakthroughs in Statistics: Methodology and Distribution; Springer: New York, NY, USA, 1992; pp. 569–593. [Google Scholar]
- Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA, 2–7 June 2019; pp. 4171–4186. [Google Scholar]
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning Representations by Back-Propagating Errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
- Huang, G.; Liu, Z.; Van Der Maaten, L.; Weinberger, K.Q. Densely Connected Convolutional Networks. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 4700–4708. [Google Scholar]
- d’Ascoli, S.; Touvron, H.; Leavitt, M.L.; Morcos, A.S.; Biroli, G.; Sagun, L. Convit: Improving Vision Transformers with Soft Convolutional Inductive Biases. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 2286–2296. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Saeed, N.; Sobirov, I.; Al Majzoub, R.; Yaqub, M. TMSS: An End-to-End Transformer-Based Multimodal Network for Segmentation and Survival Prediction. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Singapore, 18–22 September 2022; Springer: Cham, Switzerland, 2022; pp. 319–329. [Google Scholar]
- Dammu, H.; Ren, T.; Duong, T.Q. Deep Learning Prediction of Pathological Complete Response, Residual Cancer Burden, and Progression-Free Survival in Breast Cancer Patients. PLoS ONE 2023, 18, e0280148. [Google Scholar] [CrossRef] [Scilit]
- DeLong, E.R.; DeLong, D.M.; Clarke-Pearson, D.L. Comparing the Areas under Two or More Correlated Receiver Operating Characteristic Curves: A Nonparametric Approach. Biometrics 1988, 44, 837–845. [Google Scholar] [CrossRef] [Scilit]
- Guo, J.; Chen, B.; Cao, H.; Dai, Q.; Qin, L.; Zhang, J.; Zhang, Y.; Zhang, H.; Sui, Y.; Chen, T.; et al. Cross-Modal Deep Learning Model for Predicting Pathologic Complete Response to Neoadjuvant Chemotherapy in Breast Cancer. NPJ Precis. Oncol. 2024, 8, 189. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A Simple Framework for Contrastive Learning of Visual Representations. In Proceedings of the International Conference on Machine Learning, Virtual, 13–18 July 2020; pp. 1597–1607. [Google Scholar]
- Bergstra, J.; Bengio, Y. Random Search for Hyper-Parameter Optimization. J. Mach. Learn. Res. 2012, 13, 281–305. [Google Scholar]
- Zhang, H.; Liu, X.; Huang, S.; Yuan, Y.; Zhang, D.; Zhang, L. Multi-view Graph Contrastive Learning with Dynamic Self-aware and Cross-Sample Topology Augmentation for Brain Disorder Diagnosis. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Daejeon, Republic of Korea, 23–27 September 2025; pp. 532–542. [Google Scholar]
- Li, H.; Li, Z.; Mao, Y.; Ding, Z.; Huang, Z. DC-Seg: Disentangled Contrastive Learning for Brain Tumor Segmentation with Missing Modalities. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Daejeon, Republic of Korea, 23–27 September 2025; pp. 138–148. [Google Scholar]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-cam: Visual Explanations from Deep Networks via Gradient-based Localization. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 618–626. [Google Scholar]







| Datasets | Demography | Clinicopathology |
|---|---|---|
| ISPY1 | ER | |
| Age | PR | |
| Race | HR | |
| HER2 | ||
| Molecular subtype | ||
| MAMA-MIA | ER | |
| Age | PR | |
| Race | HR | |
| HER2 | ||
| Molecular subtype |
| Feature | Missing (n) | Missing (%) |
|---|---|---|
| Age | 3 | 0.20% |
| Race | 16 | 1.07% |
| ER | 996 | 66.80% |
| PR | 996 | 66.80% |
| HR | 16 | 1.07% |
| HER2 | 22 | 1.48% |
| Molecular subtype | 26 | 1.74% |
| Datasets | Categories | pCR | Non-pCR | Total |
|---|---|---|---|---|
| ISPY1 | Train | 34 | 86 | 120 |
| Test | 9 | 22 | 31 | |
| 151 * | ||||
| MAMA-MIA | Train | 347 | 838 | 1185 |
| Test | 93 | 213 | 306 | |
| 1491 * |
| Folds | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|
| 1 | 0.9545 | 0.9355 | 1 | 0.9091 | 0.9 | 0.8182 | 1 |
| 2 | 0.8838 | 0.871 | 0.7778 | 0.9091 | 0.7778 | 0.7778 | 0.9091 |
| 3 | 0.7172 | 0.7419 | 0.4444 | 0.8636 | 0.5 | 0.5714 | 0.7917 |
| 4 | 0.7677 | 0.7419 | 0.7778 | 0.7273 | 0.6364 | 0.5385 | 0.8889 |
| 5 | 0.9141 | 0.9355 | 0.8889 | 0.9545 | 0.8889 | 0.8889 | 0.9545 |
| Mean ± SD | 0.8475 ± 0.0901 | 0.8452 ± 0.0875 | 0.7778 ± 0.1859 | 0.8727 ± 0.0782 | 0.7406 ± 0.1533 | 0.719 ± 0.1386 | 0.9088 ± 0.0699 |
| Methods | AUC (95% CI) | ACC (95% CI) | SEN (95% CI) | SPE (95% CI) | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|
| BERT [43] | 0.6909 (0.5842, 0.7936) | 0.6968 (0.6124, 0.7831) | 0.5555 (0.4218, 0.6842) | 0.7546 (0.6632, 0.8415) | 0.5147 | 0.4882 | 0.8065 |
| MLP [44] | 0.7030 (0.6014, 0.8021) | 0.7226 (0.6482, 0.7915) | 0.4889 (0.3724, 0.6018) | 0.8182 (0.7135, 0.9126) | 0.5069 | 0.5504 | 0.7966 |
| DenseNet [45] | 0.6389 (0.5214, 0.7486) | 0.6387 (0.5421, 0.7284) | 0.4000 (0.2814, 0.5132) | 0.7364 (0.6415, 0.8247) | 0.3912 | 0.3873 | 0.7498 |
| ConVit [46] | 0.4889 (0.4021, 0.5732) | 0.6709 (0.5846, 0.7512) | 0.2222 (0.1142, 0.3325) | 0.8545 (0.7712, 0.9324) | 0.2826 | 0.4124 | 0.7283 |
| ResNet-50 [47] | 0.6243 (0.5142, 0.7315) | 0.6129 (0.5214, 0.7012) | 0.6445 (0.5126, 0.7732) | 0.6000 (0.4823, 0.7145) | 0.4914 | 0.4043 | 0.8065 |
| ViT [34] | 0.6495 (0.5512, 0.7431) | 0.6839 (0.6012, 0.7642) | 0.4889 (0.3621, 0.6084) | 0.7636 (0.6642, 0.8561) | 0.4762 | 0.4782 | 0.7834 |
| SIDLN [30] | 0.6803 (0.5912, 0.7645) | 0.6774 (0.5842, 0.7621) | 0.6222 (0.4614, 0.7752) | 0.7000 (0.5842, 0.8124) | 0.5161 | 0.4719 | 0.8317 |
| BERT-ViT | 0.7162 (0.6214, 0.8052) | 0.7742 (0.6912, 0.8512) | 0.6000 (0.4532, 0.7412) | 0.8454 (0.7612, 0.9242) | 0.5978 | 0.6278 | 0.8430 |
| Interactive-Model [13] | 0.7515 (0.6631, 0.8324) | 0.6710 (0.5812, 0.7564) | 0.6222 (0.5012, 0.7382) | 0.6909 (0.5842, 0.7931) | 0.5231 | 0.4588 | 0.8185 |
| TMSS [48] | 0.7010 (0.6123, 0.7842) | 0.6710 (0.5814, 0.7562) | 0.3556 (0.2124, 0.4952) | 0.8000 (0.7012, 0.8941) | 0.3541 | 0.5382 | 0.7546 |
| Integrated-Model [49] | 0.7626 (0.6742, 0.8461) | 0.6581 (0.5742, 0.7362) | 0.5556 (0.4321, 0.6742) | 0.7000 (0.5912, 0.8042) | 0.4816 | 0.4586 | 0.7992 |
| MRI-RNA [24] | 0.7586 (0.6632, 0.8514) | 0.7871 (0.7123, 0.8582) | 0.6000 (0.4741, 0.7214) | 0.8636 (0.7688, 0.9573) | 0.6144 | 0.6491 | 0.8431 |
| CITR-Net [25] | 0.7788 (0.6912, 0.8624) | 0.6968 (0.6124, 0.7782) | 0.4444 (0.3214, 0.5632) | 0.8000 (0.7042, 0.8912) | 0.4311 | 0.4742 | 0.7882 |
| AER-SwinT [32] | 0.7434 (0.6532, 0.8312) | 0.6710 (0.5824, 0.7542) | 0.5111 (0.3921, 0.6254) | 0.7364 (0.6412, 0.8272) | 0.4591 | 0.4484 | 0.7947 |
| BPMINet (ours) | 0.8475 (0.7356, 0.9594) | 0.8452 (0.7366, 0.9538) | 0.7778 (0.5470, 0.9989) | 0.8727 (0.7756, 0.9698) | 0.7406 | 0.7190 | 0.9088 |
| Methods | AUC (95% CI) | ACC (95% CI) | SEN (95% CI) | SPE (95% CI) | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|
| BERT [43] | 0.6503 (0.5986, 0.7022) | 0.6321 (0.5816, 0.6826) | 0.5054 (0.4274, 0.5832) | 0.6893 (0.6324, 0.7388) | 0.4608 | 0.4234 | 0.7553 |
| MLP [44] | 0.6513 (0.5823, 0.7218) | 0.6756 (0.6187, 0.7291) | 0.3871 (0.2911, 0.4853) | 0.8058 (0.7375, 0.8655) | 0.4260 | 0.4737 | 0.7444 |
| DenseNet [45] | 0.5432 (0.4722, 0.6140) | 0.6187 (0.5487, 0.6871) | 0.2796 (0.1932, 0.3647) | 0.7718 (0.7150, 0.8284) | 0.3133 | 0.3562 | 0.7035 |
| ConVit [46] | 0.5451 (0.4748, 0.6163) | 0.6321 (0.5821, 0.6853) | 0.4194 (0.3401, 0.5035) | 0.7282 (0.6649, 0.7872) | 0.4149 | 0.4105 | 0.7353 |
| ResNet-50 [47] | 0.5359 (0.4521, 0.6197) | 0.6087 (0.5452, 0.6722) | 0.2043 (0.1241, 0.2847) | 0.7913 (0.7261, 0.8568) | 0.2452 | 0.3065 | 0.6878 |
| ViT [34] | 0.6053 (0.5392, 0.6724) | 0.6455 (0.5591, 0.7322) | 0.4086 (0.3027, 0.5143) | 0.7524 (0.6632, 0.8452) | 0.4176 | 0.4270 | 0.7381 |
| SIDLN [30] | 0.6688 (0.6050, 0.7295) | 0.6421 (0.5741, 0.7114) | 0.4946 (0.4251, 0.5622) | 0.7087 (0.6444, 0.7650) | 0.4623 | 0.4340 | 0.7565 |
| BERT-ViT | 0.5854 (0.5036, 0.6682) | 0.6589 (0.5842, 0.7128) | 0.3763 (0.2764, 0.4731) | 0.7864 (0.7383, 0.8351) | 0.4070 | 0.4430 | 0.7364 |
| Interactive-Model [13] | 0.5423 (0.4849, 0.6020) | 0.6555 (0.5875, 0.7215) | 0.2903 (0.2086, 0.3741) | 0.8204 (0.7667, 0.8744) | 0.3439 | 0.4219 | 0.7191 |
| TMSS [48] | 0.6170 (0.5406, 0.6845) | 0.6187 (0.5617, 0.6724) | 0.4624 (0.3636, 0.5647) | 0.6893 (0.6321, 0.7391) | 0.4300 | 0.4019 | 0.7396 |
| Integrated-Model [49] | 0.6938 (0.6423, 0.7456) | 0.6823 (0.6165, 0.7469) | 0.3763 (0.3101, 0.4470) | 0.8204 (0.7750, 0.8653) | 0.4242 | 0.4861 | 0.7445 |
| MRI-RNA [24] | 0.6905 (0.6222, 0.7484) | 0.6923 (0.6418, 0.7461) | 0.4516 (0.3727, 0.5326) | 0.8010 (0.7482, 0.8530) | 0.4773 | 0.5060 | 0.7639 |
| CITR-Net [25] | 0.6946 (0.6274, 0.7533) | 0.7124 (0.6573, 0.7638) | 0.4624 (0.3516, 0.5652) | 0.8252 (0.7691, 0.8798) | 0.5000 | 0.5443 | 0.7727 |
| AER-SwinT [32] | 0.7008 (0.6329, 0.7637) | 0.7023 (0.6455, 0.7525) | 0.4946 (0.3928, 0.5976) | 0.7961 (0.7418, 0.8498) | 0.5083 | 0.5227 | 0.7773 |
| BPMINet (ours) | 0.7370 (0.6730, 0.7953) | 0.7391 (0.6854, 0.7862) | 0.5376 (0.4338, 0.6324) | 0.8301 (0.7812, 0.8814) | 0.5618 | 0.5882 | 0.7991 |
| Model | Backbone | BiCMA | MCFE-CL | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|---|---|---|
| Base model | ✓ | × | × | 0.6444 | 0.6516 | 0.4889 | 0.7182 | 0.4524 | 0.4294 | 0.7731 |
| Model-1 | ✓ | ✓ | × | 0.7556 | 0.8 | 0.6667 | 0.8545 | 0.6504 | 0.6517 | 0.8671 |
| Model-2 | ✓ | × | ✓ | 0.7495 | 0.7613 | 0.6 | 0.8273 | 0.5916 | 0.6077 | 0.8379 |
| BPMINet | ✓ | ✓ | ✓ | 0.8475 | 0.8452 | 0.7778 | 0.8727 | 0.7406 | 0.719 | 0.9088 |
| Model | Backbone | BiCMA | MCFE-CL | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|---|---|---|
| Base model | ✓ | × | × | 0.4864 | 0.6187 | 0.2796 | 0.7718 | 0.3133 | 0.3562 | 0.7035 |
| Model-1 | ✓ | ✓ | × | 0.6979 | 0.689 | 0.4301 | 0.8058 | 0.4624 | 0.5 | 0.758 |
| Model-2 | ✓ | × | ✓ | 0.6911 | 0.6856 | 0.4409 | 0.7961 | 0.4659 | 0.494 | 0.7593 |
| BPMINet | ✓ | ✓ | ✓ | 0.737 | 0.7391 | 0.5376 | 0.8301 | 0.5618 | 0.5882 | 0.7991 |
| Datasets | Setting | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|---|
| ISPY1 | Concat | 0.6444 | 0.6516 | 0.4889 | 0.7182 | 0.4524 | 0.4294 | 0.7731 |
| IGC-Attention | 0.6788 | 0.6968 | 0.4667 | 0.7909 | 0.47 | 0.475 | 0.7844 | |
| CGI-Attention | 0.5243 | 0.6 | 0.2889 | 0.7273 | 0.2841 | 0.2864 | 0.7166 | |
| BiCMA | 0.7556 | 0.8 | 0.6667 | 0.8545 | 0.6504 | 0.6517 | 0.8671 | |
| MAMA-MIA | Concat | 0.4864 | 0.6187 | 0.2796 | 0.7718 | 0.3133 | 0.3562 | 0.7035 |
| IGC-Attention | 0.6723 | 0.6388 | 0.4194 | 0.7379 | 0.4194 | 0.4194 | 0.7379 | |
| CGI-Attention | 0.4515 | 0.5151 | 0.2903 | 0.6165 | 0.2714 | 0.2547 | 0.658 | |
| BiCMA | 0.6979 | 0.689 | 0.4301 | 0.8058 | 0.4624 | 0.5 | 0.758 |
| Datasets | MCFE Configuration | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|---|
| ISPY1 | W/o CL | 0.7556 | 0.8 | 0.6667 | 0.8545 | 0.6504 | 0.6517 | 0.8671 |
| Conventional contrastive learning | 0.7283 | 0.7677 | 0.6889 | 0.8 | 0.6298 | 0.5874 | 0.866 | |
| W/o multimodal dynamic calibration | 0.7667 | 0.8129 | 0.7334 | 0.8454 | 0.6912 | 0.6614 | 0.8887 | |
| W/o dual-path refinement | 0.7677 | 0.8065 | 0.6667 | 0.8636 | 0.6594 | 0.6683 | 0.8674 | |
| W/o contrast-aware dynamic activation | 0.8192 | 0.7161 | 0.7334 | 0.7091 | 0.602 | 0.5191 | 0.8677 | |
| MCFE | 0.8475 | 0.8452 | 0.7778 | 0.8727 | 0.7406 | 0.719 | 0.9088 | |
| MAMA-MIA | W/o CL | 0.6979 | 0.689 | 0.4301 | 0.8058 | 0.4624 | 0.5 | 0.758 |
| Conventional contrastive learning | 0.6822 | 0.6756 | 0.4409 | 0.7816 | 0.4581 | 0.4767 | 0.7559 | |
| W/o multimodal dynamic calibration | 0.691 | 0.6957 | 0.4516 | 0.8058 | 0.48 | 0.5122 | 0.765 | |
| W/o dual-path refinement | 0.6867 | 0.6823 | 0.4301 | 0.7961 | 0.4571 | 0.4878 | 0.7558 | |
| W/o contrast-aware dynamic activation | 0.6933 | 0.689 | 0.4731 | 0.7864 | 0.4862 | 0.5 | 0.7678 | |
| MCFE | 0.737 | 0.7391 | 0.5376 | 0.8301 | 0.5618 | 0.5882 | 0.7991 |
| Dataset | Setting | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|---|
| MAMA-MIA | w/o ER+PR | 0.6867 | 0.699 | 0.4194 | 0.8252 | 0.4643 | 0.52 | 0.7589 |
| w/ ER+PR (ours) | 0.737 | 0.7391 | 0.5376 | 0.8301 | 0.5618 | 0.5882 | 0.7991 |
| Datasets | Setting | AUC | ACC | SEN | SPE | F1 | PPV | NPV |
|---|---|---|---|---|---|---|---|---|
| ISPY1 | Clinical 1 | 0.6313 | 0.6387 | 0.4889 | 0.7 | 0.4358 | 0.413 | 0.772 |
| Clinical 2 | 0.7646 | 0.8 | 0.6445 | 0.8636 | 0.6453 | 0.6777 | 0.8602 | |
| Clinical 1 & 2 (ours) | 0.8475 | 0.8452 | 0.7778 | 0.8727 | 0.7406 | 0.719 | 0.9088 | |
| MAMA-MIA | Clinical 1 | 0.5168 | 0.5987 | 0.3763 | 0.699 | 0.3684 | 0.3608 | 0.7129 |
| Clinical 2 | 0.7065 | 0.6823 | 0.5269 | 0.7524 | 0.5078 | 0.49 | 0.7789 | |
| Clinical 1 & 2 (ours) | 0.737 | 0.7391 | 0.5376 | 0.8301 | 0.5618 | 0.5882 | 0.7991 |
| AUC | ACC | SEN | SPE | F1 | PPV | NPV | ||
|---|---|---|---|---|---|---|---|---|
| 0.1 | 0.8182 | 0.7226 | 0.7556 | 0.7091 | 0.6102 | 0.5134 | 0.8793 | |
| 0.3 | 0.7566 | 0.7871 | 0.6222 | 0.8545 | 0.6203 | 0.6433 | 0.8513 | |
| 0.5 | 0.7859 | 0.7807 | 0.6889 | 0.8182 | 0.6428 | 0.6142 | 0.8682 | |
| 0.7 | 0.7636 | 0.7613 | 0.5333 | 0.8545 | 0.5426 | 0.6221 | 0.8277 | |
| 0.9 | 0.8475 | 0.8452 | 0.7778 | 0.8727 | 0.7406 | 0.719 | 0.9088 | |
| 1 | 0.7778 | 0.8129 | 0.7334 | 0.8454 | 0.6912 | 0.6614 | 0.8887 | |
| 0.1 | 0.7616 | 0.742 | 0.7333 | 0.7455 | 0.6181 | 0.5364 | 0.8756 | |
| 0.3 | 0.7737 | 0.729 | 0.5556 | 0.8 | 0.4852 | 0.7082 | 0.8412 | |
| 0.5 | 0.803 | 0.729 | 0.6889 | 0.7455 | 0.5889 | 0.527 | 0.863 | |
| 0.7 | 0.8364 | 0.7161 | 0.7111 | 0.7182 | 0.5895 | 0.5248 | 0.8696 | |
| 0.9 | 0.7727 | 0.7871 | 0.6889 | 0.8273 | 0.639 | 0.6858 | 0.8792 | |
| 1 | 0.7939 | 0.6774 | 0.4667 | 0.7636 | 0.4024 | 0.3555 | 0.7876 | |
| 0.1 | 0.7818 | 0.8 | 0.7334 | 0.8273 | 0.6817 | 0.641 | 0.8838 | |
| 0.3 | 0.8182 | 0.7226 | 0.6222 | 0.7636 | 0.5641 | 0.5552 | 0.8365 | |
| 0.5 | 0.8111 | 0.6839 | 0.7334 | 0.6636 | 0.5724 | 0.4758 | 0.8625 | |
| 0.7 | 0.8081 | 0.7161 | 0.6667 | 0.7364 | 0.5806 | 0.5416 | 0.8462 | |
| 0.9 | 0.7949 | 0.7226 | 0.7556 | 0.7091 | 0.6164 | 0.5365 | 0.8816 | |
| 1 | 0.7374 | 0.7742 | 0.7111 | 0.8 | 0.6306 | 0.5828 | 0.8801 |
| AUC | ACC | SEN | SPE | F1 | PPV | NPV | ||
|---|---|---|---|---|---|---|---|---|
| 0.1 | 0.7084 | 0.7191 | 0.5054 | 0.8155 | 0.5281 | 0.5529 | 0.785 | |
| 0.3 | 0.7004 | 0.6923 | 0.5161 | 0.7718 | 0.5106 | 0.5053 | 0.7794 | |
| 0.5 | 0.6809 | 0.6823 | 0.4624 | 0.7816 | 0.4751 | 0.4886 | 0.763 | |
| 0.7 | 0.7014 | 0.6957 | 0.5161 | 0.7767 | 0.5134 | 0.5106 | 0.7805 | |
| 0.9 | 0.737 | 0.7391 | 0.5376 | 0.8301 | 0.5618 | 0.5882 | 0.7991 | |
| 1 | 0.706 | 0.6923 | 0.4194 | 0.8155 | 0.4588 | 0.5065 | 0.7568 | |
| 0.1 | 0.7241 | 0.7157 | 0.4731 | 0.8252 | 0.5087 | 0.55 | 0.7763 | |
| 0.3 | 0.6879 | 0.6823 | 0.4516 | 0.7864 | 0.4693 | 0.4884 | 0.7606 | |
| 0.5 | 0.6889 | 0.689 | 0.4731 | 0.7864 | 0.4862 | 0.5 | 0.7678 | |
| 0.7 | 0.6805 | 0.6555 | 0.5269 | 0.7136 | 0.4876 | 0.4537 | 0.7696 | |
| 0.9 | 0.6917 | 0.6957 | 0.4624 | 0.801 | 0.4859 | 0.5119 | 0.7674 | |
| 1 | 0.6883 | 0.6856 | 0.3871 | 0.8204 | 0.4337 | 0.4932 | 0.7478 | |
| 0.1 | 0.7158 | 0.7023 | 0.4946 | 0.7961 | 0.5083 | 0.5227 | 0.7773 | |
| 0.3 | 0.6714 | 0.6722 | 0.4086 | 0.7913 | 0.4368 | 0.4691 | 0.7477 | |
| 0.5 | 0.7187 | 0.7258 | 0.5269 | 0.8155 | 0.5444 | 0.5632 | 0.7925 | |
| 0.7 | 0.6879 | 0.699 | 0.4516 | 0.8107 | 0.4828 | 0.5185 | 0.7661 | |
| 0.9 | 0.7194 | 0.7291 | 0.5161 | 0.8252 | 0.5424 | 0.5714 | 0.7907 | |
| 1 | 0.7225 | 0.709 | 0.5161 | 0.7961 | 0.5246 | 0.5333 | 0.7847 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Feng, J.; Jiang, Z.; Zhang, J. Bidirectional Perceptual Multimodal Interaction Network Based on Contrastive Learning for Breast Cancer pCR Prediction. Tomography 2026, 12, 74. https://doi.org/10.3390/tomography12050074
Feng J, Jiang Z, Zhang J. Bidirectional Perceptual Multimodal Interaction Network Based on Contrastive Learning for Breast Cancer pCR Prediction. Tomography. 2026; 12(5):74. https://doi.org/10.3390/tomography12050074
Chicago/Turabian StyleFeng, Jingjing, Zongli Jiang, and Jinli Zhang. 2026. "Bidirectional Perceptual Multimodal Interaction Network Based on Contrastive Learning for Breast Cancer pCR Prediction" Tomography 12, no. 5: 74. https://doi.org/10.3390/tomography12050074
APA StyleFeng, J., Jiang, Z., & Zhang, J. (2026). Bidirectional Perceptual Multimodal Interaction Network Based on Contrastive Learning for Breast Cancer pCR Prediction. Tomography, 12(5), 74. https://doi.org/10.3390/tomography12050074
