Anti-Forgetting Adaptive Teacher-Driven Knowledge Distillation for Medical Image Classification
Abstract
1. Introduction
- We propose A2T-KD, an adaptive knowledge distillation framework that updates the teacher during student training while constraining the loss of knowledge acquired during teacher pretraining.
- We introduce MITR and DSDO to constrain cross-epoch changes in teacher representations and predictions, thereby supporting the retention of pretraining knowledge during teacher adaptation.
- We evaluate A2T-KD on nine medical imaging datasets and multiple teacher–student configurations. A2T-KD achieved competitive predictive performance relative to the evaluated teacher-update baselines.
2. Related Work
3. Method
3.1. Notations
3.2. Key Modules in A2T-KD
3.2.1. Mutual Information Temporal Regularization (MITR)

3.2.2. Discriminative Subspace Decoupling Optimization (DSDO)
3.2.3. Structure-Guided Knowledge Distillation (SGKD)
3.3. Model Optimization
- Updating the teacher model while freezing the student model:
- Updating the student model while freezing the teacher model:
4. Experiments
4.1. Experiment Setup
4.1.1. Datasets
4.1.2. Evaluation Metrics
4.1.3. Experimental Protocol and Data Partitioning
4.1.4. Implementation Details
4.1.5. Comparison Methods
4.2. Experimental Results and Analysis
4.2.1. Performance Comparison
4.2.2. Teacher Knowledge Preservation, Teacher Updating, and Student Performance

4.3. Ablation Study
4.3.1. Effectiveness of MITR
4.3.2. Effectiveness of DSDO
4.3.3. Effectiveness of SGKD
4.4. Backbone Configurations
4.5. Hyperparameter Analysis ( and )
5. Discussion
6. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Feng, W.; Zhou, S.; Jiang, Y.; Tang, F.; Ge, Z. Neighbor-Guided Unbiased Framework for Generalized Category Discovery in Medical Image Classification. IEEE J. Biomed. Health Inform. 2025, 29, 5736–5747. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, D.; Zhao, J.; Guo, R.; Feng, Z.; Huo, C.; Jin, D.; Pedrycz, W.; Zhang, W. Distill & Contrast: A New Graph Self-Supervised Method With Approximating Nature Data Relationships. IEEE Trans. Knowl. Data Eng. 2025, 37, 3284–3297. [Google Scholar] [CrossRef] [Scilit]
- Ling, Y.; Nie, F.; Yu, W.; Li, X. Self-Labeling and Self-Knowledge Distillation Unsupervised Feature Selection. IEEE Trans. Knowl. Data Eng. 2025, 37, 4270–4284. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Jin, Y.; Stoyanov, D.; Wang, L. FedDP: Dual Personalization in Federated Medical Image Segmentation. IEEE Trans. Med. Imaging 2024, 43, 297–308. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tan, T.; Li, Z.; Sun, Y.; Wu, S. Guest Editorial: Multi-Modal Joint Learning in Healthcare Imaging. IEEE J. Biomed. Health Inform. 2025, 29, 3083–3085. [Google Scholar] [CrossRef] [Scilit]
- Solatidehkordi, Z.; Zualkernan, I. Survey on recent trends in medical image classification using semi-supervised learning. Appl. Sci. 2022, 12, 12094. [Google Scholar] [CrossRef] [Scilit]
- Zhang, T.; Dai, W.; Chen, Z.; Yang, S.; Liu, F.; Zheng, H. Few-shot image classification via mutual distillation. Appl. Sci. 2023, 13, 13284. [Google Scholar] [CrossRef] [Scilit]
- Luo, X.; Wu, J.; Yang, J.; Chen, H.; Li, Z.; Peng, H.; Zhou, C. Knowledge distillation guided interpretable brain subgraph neural networks for brain disorder exploration. IEEE Trans. Neural Netw. Learn. Syst. 2024, 36, 3559–3572. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wu, L.; Lin, H.; Gao, Z.; Zhao, G.; Li, S.Z. A teacher-free graph knowledge distillation framework with dual self-distillation. IEEE Trans. Knowl. Data Eng. 2024, 36, 4375–4385. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Zhang, T.; Chen, C.P. Cyclic Data Distillation Semi-supervised Learning For Multi-modal Emotion Recognition. IEEE Trans. Knowl. Data Eng. 2025, 37, 5078–5092. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.J.; Dai, X.; Ma, C.Y.; Liu, Y.C.; Chen, K.; Wu, B.; He, Z.; Kitani, K.; Vajda, P. Cross-Domain Adaptive Teacher for Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 7581–7590. [Google Scholar]
- Li, L.; Jin, Z. Shadow Knowledge Distillation: Bridging Offline and Online Knowledge Transfer. Adv. Neural Inf. Process. Syst. 2022, 35, 635–649. [Google Scholar] [CrossRef] [Scilit]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the Knowledge in a Neural Network. arXiv 2015, arXiv:1503.02531. [Google Scholar]
- Li, Y.; Yang, C.; Zeng, H.; Dong, Z.; An, Z.; Xu, Y.; Tian, Y.; Wu, H. Frequency-Aligned Knowledge Distillation for Lightweight Spatiotemporal Forecasting. In Proceedings of the 2025 IEEE/CVF International Conference on Computer Vision (ICCV), Honolulu, HI, USA, 19–23 October 2025; pp. 7262–7272. [Google Scholar]
- Sun, S.; Ren, W.; Li, J.; Wang, R.; Cao, X. Logit Standardization in Knowledge Distillation. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 19–21 June 2024; pp. 15731–15740. [Google Scholar]
- Sahoo, N.N.; Sachidanand, V.; Gayathri, M.N.; Murugesan, B.; Ram, K.; Joseph, J.; Sivaprakasam, M. KDPhys: An attention guided 3D to 2D knowledge distillation for real-time video-based physiological measurement. Biomed. Signal Process. Control 2025, 107, 107797. [Google Scholar] [CrossRef] [Scilit]
- Xiang, Z.; Cui, S.; Shang, C.; Jiang, J.; Zhang, L. GMoD: Graph-driven momentum distillation framework with active perception of disease severity for radiology report generation. In International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Berlin/Heidelberg, Germany, 2024; pp. 295–305. [Google Scholar]
- Shu, T.; Shi, J.; Sun, D.; Jiang, Z.; Zheng, Y. SlideGCD: Slide-based graph collaborative training with knowledge distillation for whole slide image classification. In International Conference on Medical Image Computing and Computer-Assisted Intervention; Springer: Berlin/Heidelberg, Germany, 2024; pp. 470–480. [Google Scholar]
- Qi, Y.; Zhang, W.; Wang, X.; You, X.; Hu, S.; Chen, J. Efficient knowledge distillation for brain tumor segmentation. Appl. Sci. 2022, 12, 11980. [Google Scholar] [CrossRef] [Scilit]
- Lei, Y.; Chen, X.; Wang, Y.; Tang, R.; Zhang, B. A lightweight knowledge-distillation-based model for the detection and classification of impacted mandibular third molars. Appl. Sci. 2023, 13, 9970. [Google Scholar] [CrossRef] [Scilit]
- Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; Veness, J.; Desjardins, G.; Rusu, A.A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al. Overcoming catastrophic forgetting in neural networks. Proc. Natl. Acad. Sci. USA 2017, 114, 3521–3526. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, Z.; Hoiem, D. Learning without forgetting. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 2935–2947. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, S.; Chen, Z.; Liu, Y.; Wang, Y.; Yang, D.; Zhao, Z.; Zhou, Z.; Yi, X.; Li, W.; Zhang, W.; et al. Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement Augmentation. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2023; pp. 23436–23446. [Google Scholar]
- Chen, Z.; Ngiam, J.; Huang, Y.; Luong, T.; Kretzschmar, H.; Chai, Y.; Anguelov, D. Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout. Adv. Neural Inf. Process. Syst. 2020, 33, 2039–2050. [Google Scholar]
- Liu, B.; Liu, X.; Jin, X.; Stone, P.; Liu, Q. Conflict-Averse Gradient Descent for Multi-task learning. Adv. Neural Inf. Process. Syst. 2021, 34, 18878–18890. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016. [Google Scholar]
- Wang, J.; Lu, L.; Chi, M.; Chen, J. MDR: Multi-stage Decoupled Relational Knowledge Distillation with Adaptive Stage Selection. In Proceedings of the 32nd ACM International Conference on Multimedia, Melbourne, Australia, 28 October–1 November 2024; pp. 2175–2183. [Google Scholar]
- Yang, Z.; Zeng, A.; Li, Z.; Zhang, T.; Yuan, C.; Li, Y. From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft Labels. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 17185–17194. [Google Scholar]
- Hao, Z.; Guo, J.; Han, K.; Tang, Y.; Hu, H.; Wang, Y.; Xu, C. One-for-All: Bridge the Gap Between Heterogeneous Architectures in Knowledge Distillation. Adv. Neural Inf. Process. Syst. 2023, 36, 79570–79582. [Google Scholar] [CrossRef] [Scilit]
- Zhou, S.; Liu, W.; Hu, C.; Zhou, S.; Ma, C. UniDistill: A Universal Cross-Modality Knowledge Distillation Framework for 3D Object Detection in Bird’s-Eye View. In Proceedings of the2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2023; pp. 5116–5125. [Google Scholar]
- Ji, Z.; Tian, X.; Liu, Y. AFFAKT: A Hierarchical Optimal Transport Based Method for Affective Facial Knowledge Transfer in Video Deception Detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025; Association for the Advancement of Artificial Intelligence: Washington, DC, USA, 2025; Volume 39, pp. 1336–1344. [Google Scholar]
- Yang, Y.; Wang, C.; Gong, L.; Wu, M.; Chen, Z.; Zhou, X. FG-KD: A Novel Forward Gradient-Based Framework for Teacher Knowledge Augmentation. IEEE Trans. Artif. Intell. 2025, 7, 439–454. [Google Scholar] [CrossRef] [Scilit]

| Dataset | #C | #N | #S | Modality | Split Unit | Disease |
|---|---|---|---|---|---|---|
| LC25000 https://www.kaggle.com/datasets/javaidahmadwani/lc25000 | 5 | 25,000 | – | Histopathology | Original Source Image | Lung Cancer |
| HAM10000 https://www.kaggle.com/datasets/vrindaat/ham10000-dataset | 7 | 10,015 | 7470 | Dermoscopy | Lesion | Skin Lesions |
| Chaoyang https://github.com/bupt-ai-cz/HSA-NRL/tree/main | 4 | 6160 | – | Histopathology | Image | Colorectal Cancer |
| Multiple Myeloma | 4 | 1816 | 63 | MRI | Patient | Multiple Myeloma |
| Brain Tumor https://www.kaggle.com/datasets/masoudnickparvar/brain-tumor-mri-dataset | 4 | 3264 | – | MRI | Image | Brain Tumor |
| Kidney https://github.com/Ritesh18117/Detection-and-Classification-of-Kidney-Diseases-Using-CT-Scanned-Image/tree/master | 4 | 3956 | – | CT | Image | Kidney Disease |
| Breast Tumor https://www.kaggle.com/datasets/sabahesaraki/breast-ultrasound-images-dataset | 3 | 780 | 600 | Ultrasound | Patient | Breast Cancer |
| Cataract https://www.kaggle.com/datasets/jr2ngb/cataractdataset | 4 | 601 | – | Retinal photography | Image | Cataract |
| PAPILA https://figshare.com/articles/dataset/PAPILA/14798004 | 3 | 488 | 244 | Fundus photography | Patient | Glaucoma |
| Method | LC25000 | HAM10000 | Chaoyang | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | Macro-F1 | AUC-OVO | AUC-OVR | ACC | Macro-F1 | AUC-OVO | AUC-OVR | ACC | Macro-F1 | AUC-OVO | AUC-OVR | |
| Baseline 1 | 78.6 ± 3.0 | 77.7 ± 1.6 | 91.3 ± 3.2 | 90.5 ± 0.6 | 71.4 ± 0.9 | 24.7 ± 7.3 | 60.9 ± 3.0 | 69.9 ± 2.2 | 60.4 ± 0.5 | 37.1 ± 0.3 | 79.0 ± 0.8 | 82.0 ± 0.5 |
| Baseline 2 | 95.1 ± 0.2 | 94.9 ± 0.6 | 98.8 ± 0.0 | 98.7 ± 0.1 | 72.7 ± 0.8 | 30.1 ± 1.2 | 66.1 ± 1.0 | 77.0 ± 2.7 | 73.9 ± 2.2 | 68.3 ± 1.5 | 87.1 ± 0.9 | 87.3 ± 1.0 |
| Vanilla KD | 99.1 ± 0.5 | 99.6 ± 0.5 | 93.8 ± 3.1 | 99.7 ± 0.3 | 80.2 ± 1.1 | 68.1 ± 1.5 | 90.3 ± 0.7 | 91.0 ± 0.7 | 80.6 ± 1.1 | 76.4 ± 1.6 | 89.3 ± 0.7 | 92.8 ± 0.6 |
| MDR | 97.0 ± 0.4 | 98.5 ± 0.5 | 99.2 ± 0.2 | 99.3 ± 0.3 | 81.9 ± 0.9 | 43.0 ± 1.7 | 69.3 ± 1.4 | 68.1 ± 1.6 | 77.6 ± 0.9 | 74.9 ± 1.1 | 92.0 ± 0.6 | 92.8 ± 0.6 |
| USKD | 82.2 ± 1.2 | 84.5 ± 1.1 | 94.8 ± 0.5 | 95.7 ± 1.0 | 75.2 ± 1.4 | 41.2 ± 2.1 | 73.6 ± 1.3 | 82.4 ± 1.1 | 72.4 ± 1.0 | 67.0 ± 1.7 | 87.6 ± 0.7 | 90.9 ± 0.8 |
| LSKD | 65.8 ± 2.1 | 61.7 ± 2.3 | 92.0 ± 0.7 | 95.9 ± 0.5 | 78.8 ± 1.1 | 43.5 ± 3.1 | 82.2 ± 0.9 | 84.9 ± 0.6 | 73.9 ± 1.4 | 60.2 ± 2.2 | 82.9 ± 1.2 | 88.3 ± 0.6 |
| OFAKD | 85.0 ± 0.9 | 87.1 ± 0.9 | 97.2 ± 0.3 | 99.0 ± 0.4 | 78.0 ± 1.4 | 51.9 ± 2.7 | 83.2 ± 0.9 | 86.3 ± 0.6 | 72.4 ± 2.4 | 69.2 ± 2.8 | 84.6 ± 0.9 | 87.6 ± 0.6 |
| UniDistill | 92.9 ± 0.7 | 95.8 ± 0.6 | 98.7 ± 0.5 | 99.6 ± 0.2 | 79.2 ± 1.4 | 28.8 ± 3.3 | 66.0 ± 1.9 | 60.2 ± 2.3 | 76.6 ± 1.3 | 74.3 ± 1.4 | 87.6 ± 0.7 | 91.9 ± 0.5 |
| AFFAKT | 98.9 ± 0.4 | 99.1 ± 0.3 | 99.7 ± 0.1 | 99.2 ± 0.8 | 83.0 ± 0.7 | 65.2 ± 2.5 | 87.6 ± 0.5 | 92.7 ± 0.4 | 78.5 ± 1.2 | 73.3 ± 1.5 | 89.9 ± 0.9 | 91.9 ± 0.7 |
| FG-KD | 97.1 ± 0.5 | 97.8 ± 0.5 | 98.4 ± 0.4 | 98.7 ± 0.2 | 80.9 ± 0.9 | 63.8 ± 2.3 | 86.6 ± 1.0 | 92.0 ± 0.6 | 80.3 ± 0.9 | 74.2 ± 2.1 | 91.5 ± 0.5 | 92.9 ± 0.4 |
| A2T-KD | 99.5 ± 0.2 | 99.3 ± 0.6 | 99.8 ± 0.1 | 99.9 ± 0.1 | 84.8 ± 0.6 | 70.4 ± 1.2 | 93.0 ± 0.5 | 96.5 ± 0.3 | 83.6 ± 1.2 | 77.9 ± 1.1 | 92.2 ± 0.4 | 93.4 ± 0.4 |
| Method | Multiple Myeloma | Brain Tumor | Kidney | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | Macro-F1 | AUC-OVO | AUC-OVR | ACC | Macro-F1 | AUC-OVO | AUC-OVR | ACC | Macro-F1 | AUC-OVO | AUC-OVR | |
| Baseline 1 | 52.2 ± 3.0 | 52.3 ± 3.0 | 71.7 ± 1.7 | 74.4 ± 1.8 | 79.1 ± 1.4 | 81.0 ± 1.6 | 88.6 ± 0.9 | 90.2 ± 0.7 | 93.5 ± 0.6 | 90.8 ± 1.1 | 97.5 ± 0.3 | 98.5 ± 0.2 |
| Baseline 2 | 56.6 ± 2.2 | 57.8 ± 2.8 | 78.1 ± 1.4 | 79.6 ± 1.3 | 82.2 ± 1.0 | 80.8 ± 1.2 | 91.6 ± 0.7 | 92.3 ± 0.5 | 96.2 ± 0.3 | 97.6 ± 0.5 | 98.8 ± 0.2 | 99.2 ± 0.1 |
| Vanilla KD | 60.7 ± 2.3 | 56.8 ± 2.6 | 80.4 ± 1.8 | 80.1 ± 2.1 | 79.9 ± 1.6 | 78.0 ± 1.8 | 91.9 ± 1.3 | 92.0 ± 0.7 | 96.6 ± 0.6 | 95.5 ± 0.7 | 98.9 ± 0.3 | 99.2 ± 0.2 |
| MDR | 47.7 ± 2.6 | 30.4 ± 2.4 | 70.1 ± 1.9 | 71.9 ± 2.2 | 81.4 ± 1.4 | 81.2 ± 2.8 | 90.8 ± 1.1 | 91.5 ± 0.7 | 97.9 ± 0.5 | 96.8 ± 0.6 | 98.4 ± 0.4 | 99.0 ± 0.3 |
| USKD | 55.0 ± 2.3 | 53.3 ± 2.2 | 80.4 ± 1.1 | 77.8 ± 1.5 | 80.2 ± 1.7 | 75.9 ± 2.6 | 91.9 ± 0.6 | 88.0 ± 4.8 | 86.8 ± 2.2 | 68.6 ± 2.3 | 95.5 ± 0.5 | 97.9 ± 0.3 |
| LSKD | 55.2 ± 2.3 | 42.8 ± 2.3 | 74.0 ± 2.2 | 73.1 ± 2.3 | 79.1 ± 1.5 | 80.9 ± 1.5 | 93.8 ± 0.7 | 94.0 ± 0.6 | 91.6 ± 0.9 | 91.8 ± 1.1 | 98.2 ± 0.4 | 99.7 ± 0.2 |
| OFAKD | 56.2 ± 3.0 | 56.7 ± 2.8 | 83.3 ± 1.2 | 82.3 ± 1.4 | 73.7 ± 2.4 | 75.3 ± 2.1 | 92.8 ± 0.6 | 93.1 ± 0.6 | 89.8 ± 1.2 | 88.9 ± 2.0 | 99.8 ± 0.2 | 98.5 ± 0.4 |
| UniDistill | 50.6 ± 3.4 | 39.2 ± 3.3 | 60.3 ± 3.3 | 58.8 ± 3.0 | 82.1 ± 1.2 | 80.7 ± 1.5 | 88.1 ± 2.1 | 89.7 ± 1.4 | 98.4 ± 0.3 | 97.5 ± 0.6 | 97.7 ± 0.4 | 99.5 ± 0.3 |
| AFFAKT | 61.2 ± 2.4 | 56.4 ± 2.6 | 76.8 ± 2.0 | 76.9 ± 1.9 | 81.3 ± 1.4 | 79.6 ± 1.7 | 89.7 ± 1.5 | 90.5 ± 1.1 | 95.3 ± 0.6 | 93.7 ± 1.0 | 97.6 ± 0.5 | 99.1 ± 0.8 |
| FG-KD | 58.8 ± 2.2 | 57.4 ± 2.7 | 77.7 ± 1.9 | 78.7 ± 2.2 | 78.5 ± 1.6 | 77.7 ± 1.9 | 91.1 ± 1.1 | 91.2 ± 0.8 | 94.8 ± 0.9 | 93.6 ± 1.0 | 97.9 ± 0.4 | 99.7 ± 0.1 |
| A2T-KD | 64.9 ± 2.8 | 61.2 ± 2.6 | 80.3 ± 1.9 | 82.7 ± 1.3 | 85.6 ± 1.2 | 83.6 ± 1.5 | 94.7 ± 0.7 | 93.7 ± 0.7 | 97.2 ± 0.4 | 99.0 ± 0.3 | 99.1 ± 0.1 | 98.8 ± 0.2 |
| Method | Breast Tumor | Cataract | PAPILA | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ACC | Macro-F1 | AUC-OVO | AUC-OVR | ACC | Macro-F1 | AUC-OVO | AUC-OVR | ACC | Macro-F1 | AUC-OVO | AUC-OVR | |
| Baseline 1 | 76.7 ± 2.4 | 69.1 ± 2.2 | 86.2 ± 1.1 | 85.8 ± 0.8 | 69.1 ± 2.6 | 59.5 ± 3.3 | 80.7 ± 1.5 | 79.7 ± 1.3 | 79.6 ± 1.8 | 65.0 ± 3.5 | 80.9 ± 1.7 | 78.7 ± 1.6 |
| Baseline 2 | 68.0 ± 2.6 | 53.2 ± 3.1 | 84.7 ± 1.3 | 86.1 ± 1.2 | 56.8 ± 3.2 | 32.0 ± 2.5 | 61.7 ± 2.3 | 66.6 ± 2.7 | 62.9 ± 2.5 | 48.9 ± 2.4 | 80.7 ± 1.6 | 77.2 ± 1.7 |
| Vanilla KD | 83.8 ± 2.0 | 83.1 ± 2.3 | 93.5 ± 0.7 | 95.9 ± 0.6 | 71.6 ± 2.1 | 58.0 ± 2.2 | 81.8 ± 1.1 | 83.4 ± 1.0 | 83.1 ± 1.1 | 75.4 ± 1.9 | 85.1 ± 0.9 | 88.9 ± 0.8 |
| MDR | 81.3 ± 2.0 | 77.7 ± 2.3 | 89.9 ± 0.8 | 87.6 ± 1.1 | 70.7 ± 2.3 | 60.5 ± 2.4 | 81.0 ± 1.6 | 77.5 ± 2.1 | 69.9 ± 2.0 | 32.1 ± 3.1 | 63.6 ± 2.9 | 64.1 ± 2.1 |
| USKD | 79.9 ± 2.2 | 73.7 ± 2.9 | 89.2 ± 1.1 | 88.0 ± 1.1 | 61.9 ± 2.9 | 49.3 ± 4.0 | 68.2 ± 1.9 | 70.1 ± 1.9 | 75.8 ± 2.0 | 66.2 ± 2.1 | 83.8 ± 1.3 | 83.2 ± 1.4 |
| LSKD | 66.8 ± 2.1 | 45.3 ± 2.6 | 89.2 ± 0.9 | 87.8 ± 1.0 | 63.4 ± 2.9 | 42.2 ± 2.5 | 74.2 ± 2.5 | 74.3 ± 2.0 | 73.5 ± 1.8 | 43.4 ± 2.7 | 80.2 ± 2.1 | 82.9 ± 1.6 |
| OFAKD | 75.4 ± 2.0 | 65.9 ± 2.3 | 83.2 ± 1.3 | 83.4 ± 1.4 | 71.6 ± 2.7 | 65.3 ± 2.2 | 84.2 ± 1.1 | 82.1 ± 1.6 | 73.3 ± 2.1 | 42.8 ± 2.7 | 76.9 ± 2.5 | 78.4 ± 2.1 |
| UniDistill | 73.2 ± 2.0 | 68.1 ± 3.2 | 89.6 ± 1.1 | 86.8 ± 1.3 | 65.2 ± 2.4 | 39.3 ± 2.6 | 57.5 ± 2.9 | 58.6 ± 2.7 | 67.9 ± 2.7 | 25.4 ± 2.5 | 50.3 ± 2.5 | 52.6 ± 2.4 |
| AFFAKT | 81.3 ± 1.8 | 75.6 ± 1.8 | 91.5 ± 0.6 | 92.4 ± 0.7 | 64.0 ± 2.4 | 43.2 ± 2.6 | 79.9 ± 1.5 | 79.7 ± 1.4 | 71.4 ± 2.2 | 40.7 ± 2.8 | 79.4 ± 1.6 | 81.7 ± 1.4 |
| FG-KD | 77.6 ± 2.2 | 77.0 ± 2.3 | 93.6 ± 0.5 | 90.5 ± 0.7 | 68.7 ± 2.1 | 53.9 ± 2.4 | 79.2 ± 1.6 | 77.3 ± 2.3 | 79.0 ± 2.0 | 66.4 ± 2.3 | 86.1 ± 1.2 | 87.0 ± 1.0 |
| A2T-KD | 87.2 ± 1.7 | 84.1 ± 2.0 | 96.2 ± 1.0 | 95.8 ± 0.4 | 72.2 ± 1.9 | 58.0 ± 2.0 | 80.9 ± 1.1 | 84.4 ± 1.1 | 85.8 ± 0.9 | 76.9 ± 2.1 | 87.6 ± 1.0 | 89.6 ± 0.8 |
| Method | Training Time (×) | Peak Memory (×) | Parameters (M) | Training FLOPs (×) |
|---|---|---|---|---|
| Vanilla KD | 1.00 | 1.00 | 11.18 | 1.00 |
| LSKD | 1.01 | 1.00 | 11.18 | 1.00 |
| UniDistill | 1.18 | 1.15 | 11.85 | 1.11 |
| OFAKD | 1.30 | 1.23 | 12.30 | 1.20 |
| MDR | 5.20 | 2.35 | 15.50 | 4.70 |
| A2T-KD | 2.65 | 1.95 | 13.99 | 2.85 |
| Model Variant | Chaoyang | Kidney | Breast Tumor |
|---|---|---|---|
| Direct Smoothing | |||
| w/o MI | |||
| MLP-Enc | |||
| w/o Stab | |||
| MITR (Ours) | 83.6 ± 1.2 | 97.2 ± 0.4 | 87.2 ± 1.7 |
| Model Variant | LC25000 | Brain Tumor | Cataract |
|---|---|---|---|
| w/o CLC | |||
| w/o Ang | |||
| w/o Dec | |||
| DSDO (Ours) | 99.5 ± 0.2 | 85.6 ± 1.2 | 72.2 ± 1.9 |
| Method Variant | LC25000 | Kidney | Breast Tumor |
|---|---|---|---|
| Feature Only | |||
| Logit Only | |||
| Full SGKD (Ours) | 99.5 ± 0.2 | 97.2 ± 0.4 | 87.2 ± 1.7 |
| Teacher → Student | LC25000 | Multiple Myeloma | Breast Tumor |
|---|---|---|---|
| ResNet101 → ResNet18 | 99.8 ± 0.1 | 80.3 ± 1.9 | 96.2 ± 1.0 |
| ResNet50 → ResNet18 | |||
| ResNet101 → MobileNetV2 | |||
| ResNet101 → MobileNetV3 | |||
| Swin-Tiny → ResNet18 | |||
| ViT-S → ResNet18 | |||
| ResNet50 → ShuffleNetV2 | |||
| DenseNet121 → EfficientNet-B0 |
| ACC | Macro-F1 | AUC-OVO | AUC-OVR | |
|---|---|---|---|---|
| 0.1 | ||||
| 0.5 | ||||
| 1 | 84.8 ± 0.6 | 70.4 ± 1.2 | 93.0 ± 0.5 | 96.5 ± 0.3 |
| 2 | ||||
| 4 |
| ACC | Macro-F1 | AUC-OVO | AUC-OVR | |
|---|---|---|---|---|
| 1 | ||||
| 2 | 83.6 ± 1.2 | 77.9 ± 1.1 | 92.2 ± 0.4 | 93.4 ± 0.4 |
| 4 | ||||
| 8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, T.; Zhou, C.; Wang, Y.; Hadjiiski, L.M.; Dong, Q. Anti-Forgetting Adaptive Teacher-Driven Knowledge Distillation for Medical Image Classification. Appl. Sci. 2026, 16, 8756. https://doi.org/10.3390/app16178756
Chen T, Zhou C, Wang Y, Hadjiiski LM, Dong Q. Anti-Forgetting Adaptive Teacher-Driven Knowledge Distillation for Medical Image Classification. Applied Sciences. 2026; 16(17):8756. https://doi.org/10.3390/app16178756
Chicago/Turabian StyleChen, Tao, Chuan Zhou, Yifan Wang, Lubomir M. Hadjiiski, and Qian Dong. 2026. "Anti-Forgetting Adaptive Teacher-Driven Knowledge Distillation for Medical Image Classification" Applied Sciences 16, no. 17: 8756. https://doi.org/10.3390/app16178756
APA StyleChen, T., Zhou, C., Wang, Y., Hadjiiski, L. M., & Dong, Q. (2026). Anti-Forgetting Adaptive Teacher-Driven Knowledge Distillation for Medical Image Classification. Applied Sciences, 16(17), 8756. https://doi.org/10.3390/app16178756

