Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions
Abstract
1. Introduction
1.1. Literature Search Methodology
1.1.1. Database-Specific Search Strategy
1.1.2. Quality Assessment
2. Existing Continual Learning Surveys and Remaining Gaps
2.1. Continual Learning
2.1.1. Continual Learning vs. Transfer Learning
2.1.2. Continual Learning vs. Multi-Task Learning
2.1.3. Continual Learning vs. Online Learning
2.1.4. Scope and Contribution of This Review
2.2. Setup
2.3. Basic Formulation of Continual Learning
3. Types of Continual Learning
3.1. Task-Incremental Learning
3.1.1. Applications of Task-Incremental Learning
3.1.2. Examples of Task-Incremental Learning
3.1.3. Challenges
3.2. Domain-Incremental Learning
3.2.1. Examples of Domain-Incremental Learning
3.2.2. Challenges
3.3. Class-Incremental Learning
3.3.1. Examples of Class-Incremental Learning
3.3.2. Challenges in Class-Incremental Learning
3.4. Data-Incremental Learning
- Incremental data arrival: The model receives data sequentially, one instance or batch at a time, without knowledge of whether the data introduces new classes or extends existing ones.
- No explicit task boundaries: Unlike task-based scenarios, data-incremental learning does not provide information about task transitions, requiring the model to infer patterns and adjust its learning dynamically.
- Challenges of catastrophic forgetting: As new data arrives, the model’s parameters may be updated in ways that overwrite knowledge of previously learned classes, leading to catastrophic forgetting.
- Adaptation and generalization: The model must generalize well to new instances and classes while preserving accuracy on old ones, requiring a balance between plasticity (learning new data) and stability (retaining old knowledge).
3.4.1. Examples of Data-Incremental Learning
3.4.2. Challenges in Data-Incremental Learning
- Unstructured data streams: The lack of clear task boundaries increases the difficulty of organizing and processing data effectively.
- Memory constraints: Retaining past data or features for replay becomes resource-intensive as the volume of data grows.
- Class imbalance: Incrementally arriving data may introduce imbalanced class distributions, skewing the model’s performance.
- Replay mechanisms: Retaining a subset of past data or using generative models to recreate previous data for rehearsal as discussed in Section 6.2 in detail.
- Dynamic networks: Expanding model capacity incrementally to accommodate new data without overwriting existing knowledge.
- Regularization methods: Penalizing changes to parameters critical for previously learned data to mitigate forgetting as discussed in Section 6.1.
3.5. Other Emerging Paradigms in Continual Learning
3.5.1. Few-Shot Continual Learning
3.5.2. Unsupervised Continual Learning
3.5.3. Meta-Continual Learning
3.5.4. Federated Continual Learning
3.5.5. Multi-Agent Continual Learning
4. Theoretical Foundations of Continual Learning
4.1. Stability–Plasticity Dilemma
Formal Definition of Catastrophic Forgetting
4.2. Catastrophic Forgetting
4.3. Forward and Backward Transfer
4.4. Negative Transfer
4.5. Representation Learning
4.6. Gradient Interference and Optimization Geometry
4.7. Neuroscientific Motivation
4.8. Mathematical Frameworks
4.8.1. Regularization-Based Models
4.8.2. Replay and Memory Models
4.8.3. Dynamic Architectures
4.8.4. Bayesian Models
4.8.5. Information-Theoretic Models
4.9. Unique Challenges in Foundation-Model Continual Learning
4.10. Recent Theoretical Developments in Continual Learning
4.11. Worked Mathematical Example: Stability–Plasticity Trade-Off
5. The Catastrophic Forgetting Problem
5.1. Why Neural Networks Forget
5.2. Weight Updates and Parameter Drift
5.3. Factors Exacerbating Catastrophic Forgetting
5.4. Mitigation Strategies
6. Method Taxonomy in CL
6.1. Regularization-Based Methods
6.1.1. Elastic Weight Consolidation
6.1.2. Information-Geometric Interpretation of Elastic Weight Consolidation
6.1.3. Synaptic Intelligence and Related Methods
6.2. Replay-Based Methods
6.2.1. Experience Replay
6.2.2. Generative Replay
6.3. Architecture-Based Methods
6.4. Optimization-Based Methods
6.5. Representation-Learning Methods
7. Foundation-Model Adaptation Through Prompt Learning and Parameter-Efficient Fine-Tuning
7.1. Parameter-Efficient and Prompt-Based Continual Learning
7.1.1. Replay-Free Adaptation of Pretrained Vision Models
7.1.2. Adapter-Based and Parameter-Efficient Continual Learning
7.1.3. Continual Instruction Tuning for Large Language Models
7.1.4. Critical Discussion and Open Challenges
7.2. Practical Comparison of Continual Learning Paradigms
8. Evaluation Protocols, Benchmarks, and Metrics in CL
8.1. Benchmark Datasets
8.2. Continual Learning Evaluation Settings
8.3. Task Construction and Data Splits
8.4. Evaluation Metrics
8.4.1. Average Accuracy
8.4.2. Forgetting Measure
8.4.3. Forward Transfer
8.4.4. Backward Transfer
8.4.5. Memory and Computational Efficiency
8.5. Challenges in Continual Learning Evaluation
8.6. Toward a Standardized Evaluation Protocol for Continual Learning
8.7. Representative Continual Learning Benchmarks
8.8. Quantitative Comparison of Representative Continual Learning Methods
9. Comparative Analysis of CL Methods
9.1. Overview of CL Method Categories
9.2. Regularization-Based Methods
9.3. Replay-Based Methods
9.4. Architecture-Based Methods
9.5. Optimization-Based Methods
9.6. Representation Learning Approaches
9.7. Prompt-Based and Parameter-Efficient CL
9.8. Comparison Across CL Settings
9.9. Memory, Scalability, and Computational Trade-Offs
9.10. Training Memory, Inference Memory, and Long-Term Storage
10. Applications of CL
10.1. Applications in Healthcare and Medical Imaging
10.2. Applications in Robotics and Autonomous Systems
10.3. Application in Natural Language Processing
10.4. Recommender Systems
10.5. Cybersecurity
11. Open Challenges and Future Directions
11.1. Catastrophic Forgetting and Long-Term Knowledge Retention
11.2. Scalability and Realistic Streaming Benchmarks
11.3. Memory, Computation, and Deployment Constraints
11.4. Continual Learning for Foundation Models
11.5. Parameter-Efficient Continual Adaptation
11.6. Multimodal Continual Learning
11.7. Privacy-Preserving and Federated Continual Learning
11.8. CL in Medical Imaging Under Domain Shift
11.9. Ethical, Fairness, and Safety Considerations
11.10. Toward Robust Lifelong AI
11.11. Critical Analysis of Current Continual Learning Approaches
11.12. Summary
12. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Russell, S.; Norvig, P. Artificial Intelligence: A Modern Approach, 4th ed.; Pearson: Hoboken, NJ, USA, 2020. [Google Scholar]
- Annan, R.; Qingge, L. Artificial intelligence in COVID-19 research: A comprehensive survey of innovations, challenges, and future directions. Comput. Sci. Rev. 2025, 57, 100751. [Google Scholar] [CrossRef] [Scilit]
- Herrera, F. Reflections and attentiveness on eXplainable Artificial Intelligence (XAI). The journey ahead from criticisms to human–AI collaboration. Inf. Fusion 2025, 121, 103133. [Google Scholar] [CrossRef] [Scilit]
- Utomo, S.; Pratap, A.; Karthikeyan, P.; Ayeelyan, J.; Hsu, H.C.; Hsiung, P.A. When explainable artificial intelligence meets data governance: Enhancing trustworthiness in multimodal gas classification. Inf. Fusion 2025, 125, 103440. [Google Scholar]
- Longo, L.; Brcic, M.; Cabitza, F.; Choi, J.; Confalonieri, R.; Del Ser, J.; Guidotti, R.; Hayashi, Y.; Herrera, F.; Holzinger, A.; et al. Explainable Artificial Intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions. Inf. Fusion 2024, 106, 102301. [Google Scholar] [CrossRef] [Scilit]
- Naser, M. From failure to fusion: A survey on learning from bad machine learning models. Inf. Fusion 2025, 120, 103122. [Google Scholar] [CrossRef] [Scilit]
- Escovedo, T.; Koshiyama, A.; da Cruz, A.A.; Vellasco, M. Neuroevolutionary learning in nonstationary environments. Appl. Intell. 2020, 50, 1590–1608. [Google Scholar] [CrossRef] [Scilit]
- Criado, M.F.; Casado, F.E.; Iglesias, R.; Regueiro, C.V.; Barro, S. Non-iid data and continual learning processes in federated learning: A long road ahead. Inf. Fusion 2022, 88, 263–280. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, C.V.; Achille, A.; Lam, M.; Hassner, T.; Mahadevan, V.; Soatto, S. Toward understanding catastrophic forgetting in continual learning. arXiv 2019, arXiv:1908.01091. [Google Scholar]
- Parisi, G.; Kemker, R.; Part, J.; Kanan, C.; Wermter, S. Continual lifelong learning with neural networks: A review. Neural Netw. Off. J. Int. Neural Netw. Soc. 2019, 113, 54–71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, L.; Zhang, X.; Su, H.; Zhu, J. A comprehensive survey of continual learning: Theory, method and application. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5362–5383. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; Veness, J.; Desjardins, G.; Rusu, A.A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al. Overcoming catastrophic forgetting in neural networks. Proc. Natl. Acad. Sci. USA 2017, 114, 3521–3526. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xu, X.; Chen, J.; Thakur, D.; Hong, D. Multi-modal disease segmentation with continual learning and adaptive decision fusion. Inf. Fusion 2025, 118, 102962. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Li, Z.; Gao, Y.; Chiclana, F.; Chen, X.; Dong, Y. An endogenous and continual learning approach to personalize individual semantics to support linguistic consensus reaching. Inf. Fusion 2025, 114, 102640. [Google Scholar] [CrossRef] [Scilit]
- Shahrivari, S. Beyond batch processing: Towards real-time and streaming big data. Computers 2014, 3, 117–129. [Google Scholar] [CrossRef] [Scilit]
- Parisi, G.I.; Lomonaco, V. Online continual learning on sequences. In Proceedings of the Recent Trends in Learning From Data: Tutorials from the INNS Big Data and Deep Learning Conference (INNSBDDL2019); Springer: Berlin/Heidelberg, Germany, 2020; pp. 197–221. [Google Scholar]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- van de Ven, G.M.; Tuytelaars, T.; Tolias, A.S. Three types of incremental learning. Nat. Mach. Intell. 2022, 4, 1185–1197. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bidaki, S.A.; Mohammadkhah, A.; Rezaee, K.; Hassani, F.; Eskandari, S.; Salahi, M.; Ghassemi, M.M. Online continual learning: A systematic literature review of approaches, challenges, and benchmarks. arXiv 2025, arXiv:2501.04897. [Google Scholar]
- Zhou, D.W.; Wang, Q.W.; Qi, Z.H.; Ye, H.J.; Zhan, D.C.; Liu, Z. Class-incremental learning: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 9851–9873. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wickramasinghe, B.; Saha, G.; Roy, K. Continual learning: A review of techniques, challenges, and future directions. IEEE Trans. Artif. Intell. 2023, 5, 2526–2546. [Google Scholar]
- Thrun, S.; Mitchell, T.M. Lifelong robot learning. Robot. Auton. Syst. 1995, 15, 25–46. [Google Scholar] [CrossRef] [Scilit]
- Tan, A.; Wang, Y.; Wu, W.Z.; Ding, W.; Liang, J. Multi-View Fusion Graph Attention Network for multilabel class incremental learning. Inf. Fusion 2025, 123, 103309. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Wang, T.; Chen, J.; Kawaguchi, K.; Lian, C.; Zeng, Z. Multi-view class incremental learning. Inf. Fusion 2024, 102, 102021. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.; Zhang, X.; Tian, Z.; Du, S. Enhancing few-shot lifelong learning through fusion of cross-domain knowledge. Inf. Fusion 2025, 115, 102730. [Google Scholar] [CrossRef] [Scilit]
- Mehta, S.V.; Patil, D.; Chandar, S.; Strubell, E. An empirical investigation of the role of pre-training in lifelong learning. J. Mach. Learn. Res. 2023, 24, 1–50. [Google Scholar] [CrossRef] [Scilit]
- Kanakis, M.; Bruggemann, D.; Saha, S.; Georgoulis, S.; Obukhov, A.; Van Gool, L. Reparameterizing convolutions for incremental multi-task learning without task interference. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part XX 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 689–707. [Google Scholar]
- Vödisch, N.; Cattaneo, D.; Burgard, W.; Valada, A. Covio: Online continual learning for visual-inertial odometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 2464–2473. [Google Scholar]
- Ullah, Z.; Usman, M.; Gwak, J. MTSS-AAE: Multi-task semi-supervised adversarial autoencoding for COVID-19 detection based on chest X-ray images. Expert Syst. Appl. 2023, 216, 119475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bonicelli, L.; Boschini, M.; Frascaroli, E.; Porrello, A.; Pennisi, M.; Bellitto, G.; Palazzo, S.; Spampinato, C.; Calderara, S. On the effectiveness of equivariant regularization for robust online continual learning. arXiv 2023, arXiv:2305.03648. [Google Scholar]
- Ali, S.; Abuhmed, T.; El-Sappagh, S.; Muhammad, K.; Alonso-Moral, J.M.; Confalonieri, R.; Guidotti, R.; Del Ser, J.; Díaz-Rodríguez, N.; Herrera, F. Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence. Inf. Fusion 2023, 99, 101805. [Google Scholar] [CrossRef] [Scilit]
- Abbass, H. What is artificial intelligence? IEEE Trans. Artif. Intell. 2021, 2, 94–95. [Google Scholar] [CrossRef] [Scilit]
- Smith, P.D. Hands-On Artificial Intelligence for Beginners: An Introduction to AI Concepts, Algorithms, and Their Implementation; Packt Publishing Ltd.: Birmingham, UK, 2018. [Google Scholar]
- Chen, Z.; Liu, B. Lifelong Machine Learning; Morgan & Claypool Publishers: San Rafael, CA, USA, 2018. [Google Scholar]
- Yang, Y.; Zhou, J.; Ding, X.; Huai, T.; Liu, S.; Chen, Q.; Xie, Y.; He, L. Recent advances of foundation language models-based continual learning: A survey. ACM Comput. Surv. 2025, 57, 1–38. [Google Scholar] [CrossRef] [Scilit]
- Kudithipudi, D.; Aguilar-Simon, M.; Babb, J.; Bazhenov, M.; Blackiston, D.; Bongard, J.; Brna, A.P.; Chakravarthi Raja, S.; Cheney, N.; Clune, J.; et al. Biological underpinnings for lifelong learning machines. Nat. Mach. Intell. 2022, 4, 196–210. [Google Scholar] [CrossRef] [Scilit]
- Hadsell, R.; Rao, D.; Rusu, A.A.; Pascanu, R. Embracing change: Continual learning in deep neural networks. Trends Cogn. Sci. 2020, 24, 1028–1040. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Qu, H.; Rahmani, H.; Xu, L.; Williams, B.; Liu, J. Recent advances of continual learning in computer vision: An overview. IET Comput. Vis. 2025, 19, e70013. [Google Scholar] [CrossRef] [Scilit]
- Masana, M.; Twardowski, B.; Van de Weijer, J. On class orderings for incremental learning. arXiv 2020, arXiv:2007.02145. [Google Scholar]
- Biesialska, M.; Biesialska, K.; Costa-Jussa, M.R. Continual lifelong learning in natural language processing: A survey. In Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain, 8–13 December 2020; pp. 6523–6541. [Google Scholar]
- Ke, Z.; Liu, B. Continual learning of natural language processing tasks: A survey. arXiv 2022, arXiv:2211.12701. [Google Scholar]
- Khetarpal, K.; Riemer, M.; Rish, I.; Precup, D. Towards continual reinforcement learning: A review and perspectives. J. Artif. Intell. Res. 2022, 75, 1401–1476. [Google Scholar] [CrossRef] [Scilit]
- Ghosh, S. Dynamic vaes with generative replay for continual zero-shot learning. arXiv 2021, arXiv:2104.12468. [Google Scholar]
- Singh, P.; Mazumder, P.; Rai, P.; Namboodiri, V.P. Rectification-based knowledge retention for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 15282–15291. [Google Scholar]
- Tao, X.; Hong, X.; Chang, X.; Dong, S.; Wei, X.; Gong, Y. Few-shot class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 12183–12192. [Google Scholar]
- Wang, L.; Yang, K.; Li, C.; Hong, L.; Li, Z.; Zhu, J. Ordisco: Effective and efficient usage of incremental unlabeled data for semi-supervised continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 5383–5392. [Google Scholar]
- Joseph, K.; Khan, S.; Khan, F.S.; Balasubramanian, V.N. Towards open world object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 5830–5840. [Google Scholar]
- Wang, Q.F.; Geng, X.; Lin, S.X.; Xia, S.Y.; Qi, L.; Xu, N. Learngene: From open-world to your learning task. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2022; Volume 36, pp. 8557–8565. [Google Scholar]
- Hu, D.; Yan, S.; Lu, Q.; Hong, L.; Hu, H.; Zhang, Y.; Li, Z.; Wang, X.; Feng, J. How well does self-supervised pre-training perform with streaming data? arXiv 2021, arXiv:2104.12081. [Google Scholar]
- Rao, D.; Visin, F.; Rusu, A.; Pascanu, R.; Teh, Y.W.; Hadsell, R. Continual unsupervised representation learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Ruvolo, P.; Eaton, E. ELLA: An efficient lifelong learning algorithm. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2013; pp. 507–515. [Google Scholar]
- Masse, N.Y.; Grant, G.D.; Freedman, D.J. Alleviating catastrophic forgetting using context-dependent gating and synaptic stabilization. Proc. Natl. Acad. Sci. USA 2018, 115, E10467–E10475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ramesh, R.; Chaudhari, P. Model zoo: A growing “brain” that learns continually. arXiv 2021, arXiv:2106.03027. [Google Scholar]
- PourKeshavarzi, M.; Zhao, G.; Sabokrou, M. Looking back on learned experiences for class/task incremental learning. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Newton Highlands, MA, USA, 2021. [Google Scholar]
- Xie, X.; Xu, J.; Hu, P.; Zhang, W.; Huang, Y.; Zheng, W.; Wang, R. Task-incremental medical image classification with task-specific batch normalization. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV); Springer: Berlin/Heidelberg, Germany, 2023; pp. 309–320. [Google Scholar]
- Feng, F.; Chan, R.H.; Shi, X.; Zhang, Y.; She, Q. Challenges in task incremental learning for assistive robotics. IEEE Access 2019, 8, 3434–3441. [Google Scholar]
- Lopez-Paz, D.; Ranzato, M. Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Vogelstein, J.T.; Dey, J.; Helm, H.S.; LeVine, W.; Mehta, R.D.; Tomita, T.M.; Xu, H.; Geisa, A.; Wang, Q.; van de Ven, G.M.; et al. A Simple Lifelong Learning Approach. arXiv 2020, arXiv:2004.12908. [Google Scholar]
- Ke, Z.; Liu, B.; Xu, H.; Shu, L. CLASSIC: Continual and contrastive learning of aspect sentiment classification tasks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 6871–6883. [Google Scholar]
- Mirza, M.J.; Masana, M.; Possegger, H.; Bischof, H. An efficient domain-incremental learning approach to drive in all weather conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 3001–3011. [Google Scholar]
- Aljundi, R.; Chakravarty, P.; Tuytelaars, T. Expert gate: Lifelong learning with a network of experts. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 3366–3375. [Google Scholar]
- Von Oswald, J.; Henning, C.; Grewe, B.F.; Sacramento, J. Continual learning with hypernetworks. arXiv 2019, arXiv:1906.00695. [Google Scholar]
- Lomonaco, V.; Maltoni, D. Core50: A new dataset and benchmark for continuous object recognition. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2017; pp. 17–26. [Google Scholar]
- Garg, P.; Saluja, R.; Balasubramanian, V.N.; Arora, C.; Subramanian, A.; Jawahar, C. Multi-domain incremental learning for semantic segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2022; pp. 761–771. [Google Scholar]
- Capuano, N.; Greco, L.; Ritrovato, P.; Vento, M. Sentiment analysis for customer relationship management: An incremental learning approach. Appl. Intell. 2021, 51, 3339–3352. [Google Scholar]
- Rebuffi, S.A.; Kolesnikov, A.; Sperl, G.; Lampert, C.H. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 2001–2010. [Google Scholar]
- Shin, H.; Lee, J.K.; Kim, J.; Kim, J. Continual learning with deep generative replay. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- van de Ven, G.M.; Siegelmann, H.T.; Tolias, A.S. Brain-inspired replay for continual learning with artificial neural networks. Nat. Commun. 2020, 11, 4069. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhou, D.W.; Yang, Y.; Zhan, D.C. Learning to classify with incremental new class. IEEE Trans. Neural Netw. Learn. Syst. 2021, 33, 2429–2443. [Google Scholar]
- Belouadah, E.; Popescu, A.; Kanellos, I. A comprehensive study of class incremental learning algorithms for visual tasks. Neural Netw. 2021, 135, 38–54. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Masana, M.; Liu, X.; Twardowski, B.; Menta, M.; Bagdanov, A.D.; Van De Weijer, J. Class-incremental learning: Survey and performance evaluation on image classification. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 5513–5533. [Google Scholar]
- Channappayya, S.; Tamma, B.R. Augmented memory replay-based continual learning approaches for network intrusion detection. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2023; Volume 36, pp. 17156–17169. [Google Scholar]
- Li, X.; Wang, S.; Sun, J.; Xu, Z. Variational data-free knowledge distillation for continual learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12618–12634. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2012; Volume 25. [Google Scholar]
- Shmelkov, K.; Schmid, C.; Alahari, K. Incremental learning of object detectors without catastrophic forgetting. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 3400–3409. [Google Scholar]
- Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2015; pp. 1440–1448. [Google Scholar]
- Ramakrishnan, K.; Panda, R.; Fan, Q.; Henning, J.; Oliva, A.; Feris, R. Relationship matters: Relation guided knowledge transfer for incremental learning of object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2020; pp. 250–251. [Google Scholar]
- Paik, I.; Oh, S.; Kwak, T.; Kim, I. Overcoming catastrophic forgetting by neuron-level plasticity control. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2020; Volume 34, pp. 5339–5346. [Google Scholar]
- Zhou, X.; Wang, D.; Krähenbühl, P. Objects as points. arXiv 2019, arXiv:1904.07850. [Google Scholar]
- Li, D.; Tasci, S.; Ghosh, S.; Zhu, J.; Zhang, J.; Heck, L. RILOD: Near real-time incremental learning for object detection at the edge. In Proceedings of the 4th ACM/IEEE Symposium on Edge Computing; ACM: New York, NY, USA, 2019; pp. 113–126. [Google Scholar]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
- Feng, T.; Wang, M.; Yuan, H. Overcoming catastrophic forgetting in incremental object detection via elastic response distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 9427–9436. [Google Scholar]
- Li, X.; Wang, W.; Wu, L.; Chen, S.; Hu, X.; Li, J.; Tang, J.; Yang, J. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 21002–21012. [Google Scholar]
- Hao, Y.; Fu, Y.; Jiang, Y.G.; Tian, Q. An end-to-end architecture for class-incremental object detection with knowledge distillation. In Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME); IEEE: New York, NY, USA, 2019; pp. 1–6. [Google Scholar]
- Peng, C.; Zhao, K.; Lovell, B.C. Faster ilod: Incremental learning for object detectors based on faster rcnn. Pattern Recognit. Lett. 2020, 140, 109–115. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Zhang, J.; Ghosh, S.; Li, D.; Tasci, S.; Heck, L.; Zhang, H.; Kuo, C.C.J. Class-incremental learning via deep model consolidation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2020; pp. 1131–1140. [Google Scholar]
- Dong, N.; Zhang, Y.; Ding, M.; Lee, G.H. Bridging non co-occurrence with unlabeled in-the-wild data for incremental object detection. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 30492–30503. [Google Scholar]
- Joseph, K.; Rajasegaran, J.; Khan, S.; Khan, F.S.; Balasubramanian, V.N. Incremental object detection via meta-learning. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 9209–9216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2015; Volume 28. [Google Scholar]
- Zhao, N.; Lee, G.H. Static-dynamic co-teaching for class-incremental 3d object detection. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2022; Volume 36, pp. 3436–3445. [Google Scholar]
- Wang, J.; Wang, X.; Shang-Guan, Y.; Gupta, A. Wanderlust: Online continual object detection in the real world. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 10829–10838. [Google Scholar]
- Perez-Rua, J.M.; Zhu, X.; Hospedales, T.M.; Xiang, T. Incremental few-shot object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13846–13855. [Google Scholar]
- Feng, J.; Phillips, R.V.; Malenica, I.; Bishara, A.; Hubbard, A.E.; Celi, L.A.; Pirracchio, R. Clinical artificial intelligence quality improvement: Towards continual monitoring and updating of AI algorithms in healthcare. npj Digit. Med. 2022, 5, 66. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Franklin, S. Autonomous agents as embodied AI. Cybern. Syst. 1997, 28, 499–520. [Google Scholar] [CrossRef] [Scilit]
- Shi, G.; Wu, Y.; Liu, J.; Wan, S.; Wang, W.; Lu, T. Incremental few-shot semantic segmentation via embedding adaptive-update and hyper-class representation. In Proceedings of the 30th ACM International Conference on Multimedia; ACM: New York, NY, USA, 2022; pp. 5547–5556. [Google Scholar]
- Ganea, D.A.; Boom, B.; Poppe, R. Incremental few-shot instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 1185–1194. [Google Scholar]
- Jin, X.; Lin, B.Y.; Rostami, M.; Ren, X. Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 714–729. [Google Scholar]
- Cossu, A.; Carta, A.; Passaro, L.; Lomonaco, V.; Tuytelaars, T.; Bacciu, D. Continual pre-training mitigates forgetting in language and vision. Neural Netw. 2024, 179, 106492. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Madaan, D.; Yoon, J.; Li, Y.; Liu, Y.; Hwang, S.J. Representational continuity for unsupervised continual learning. arXiv 2021, arXiv:2110.06976. [Google Scholar]
- Riemer, M.; Cases, I.; Ajemian, R.; Liu, M.; Rish, I.; Tu, Y.; Tesauro, G. Learning to learn without forgetting by maximizing transfer and minimizing interference. arXiv 2018, arXiv:1810.11910. [Google Scholar]
- Guo, Q.; Zhao, W.; Lyu, Z.; Zhao, T. A GAN enhanced meta-deep reinforcement learning approach for DCN routing optimization. Inf. Fusion 2025, 121, 103160. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Zhong, Z.; Yang, F.; Luo, Z.; Lin, Y.; Li, S.; Sebe, N. Learning to generalize unseen domains via memory-based multi-source meta-learning for person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 6277–6286. [Google Scholar]
- Javed, K.; White, M. Meta-learning representations for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Beaulieu, S.; Frati, L.; Miconi, T.; Lehman, J.; Stanley, K.O.; Clune, J.; Cheney, N. Learning to continually learn. In ECAI 2020; IOS Press: Amsterdam, The Netherlands, 2020; pp. 992–1001. [Google Scholar]
- Lee, E.; Huang, C.H.; Lee, C.Y. Few-shot and continual learning with attentive independent mechanisms. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 9455–9464. [Google Scholar]
- Rajasegaran, J.; Khan, S.; Hayat, M.; Khan, F.S.; Shah, M. itaml: An incremental task-agnostic meta-learning approach. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13588–13597. [Google Scholar]
- Gupta, G.; Yadav, K.; Paull, L. Look-ahead meta learning for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 11588–11598. [Google Scholar]
- Caccia, L.; Belilovsky, E.; Caccia, M.; Pineau, J. Online learned continual compression with adaptive quantization modules. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; pp. 1240–1250. [Google Scholar]
- KJ, J.; N Balasubramanian, V. Meta-consolidation for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 14374–14386. [Google Scholar]
- Henning, C.; Cervera, M.; D’Angelo, F.; Von Oswald, J.; Traber, R.; Ehret, B.; Kobayashi, S.; Grewe, B.F.; Sacramento, J. Posterior meta-replay for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 14135–14149. [Google Scholar]
- Hurtado, J.; Raymond, A.; Soto, A. Optimizing reusable knowledge for continual learning via metalearning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 14150–14162. [Google Scholar]
- Wang, R.; Bao, Y.; Zhang, B.; Liu, J.; Zhu, W.; Guo, G. Anti-retroactive interference for lifelong learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 163–178. [Google Scholar]
- McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; y Arcas, B.A. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2017; pp. 1273–1282. [Google Scholar]
- Yoon, J.; Jeong, W.; Lee, G.; Yang, E.; Hwang, S.J. Federated continual learning with weighted inter-client transfer. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 12073–12086. [Google Scholar]
- Usmanova, A.; Portet, F.; Lalanda, P.; Vega, G. A distillation-based approach integrating continual learning and federated learning for pervasive services. arXiv 2021, arXiv:2109.04197. [Google Scholar]
- Park, T.J.; Kumatani, K.; Dimitriadis, D. Tackling dynamics in federated incremental learning with variational embedding rehearsal. arXiv 2021, arXiv:2110.09695. [Google Scholar]
- Mermillod, M.; Bugaiska, A.; Bonin, P. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects. Front. Psychol. 2013, 4, 504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Grossberg, S. Adaptive Resonance Theory: How a brain learns to consciously attend, learn, and recognize a changing world. Neural Netw. 2013, 37, 1–47. [Google Scholar] [PubMed]
- Abraham, W.C.; Robins, A. Memory retention–the synaptic stability versus plasticity dilemma. Trends Neurosci. 2005, 28, 73–78. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hebb, D.O. The Organization of Behavior: A Neuropsychological Theory; Psychology Press: Hove, UK, 2005. [Google Scholar]
- Power, J.D.; Schlaggar, B.L. Neural plasticity across the lifespan. Wiley Interdiscip. Rev. Dev. Biol. 2017, 6, e216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bennani, M.A.; Doan, T.; Sugiyama, M. Generalisation guarantees for continual learning with orthogonal gradient descent. arXiv 2020, arXiv:2006.11942. [Google Scholar]
- Doan, T.; Bennani, M.A.; Mazoure, B.; Rabusseau, G.; Alquier, P. A theoretical analysis of catastrophic forgetting through the ntk overlap matrix. In Proceedings of the International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2021; pp. 1072–1080. [Google Scholar]
- McCloskey, M.; Cohen, N.J. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of Learning and Motivation; Elsevier: Amsterdam, The Netherlands, 1989; Volume 24, pp. 109–165. [Google Scholar]
- Ratcliff, R. Connectionist models of recognition memory: Constraints imposed by learning and forgetting functions. Psychol. Rev. 1990, 97, 285. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, J.; Zhang, X.; Zhao, B.; Hu, W.; Diao, T.; Wang, L.; Zhong, Y.; Li, Q. Genetic dissection of mutual interference between two consecutive learning tasks in Drosophila. eLife 2023, 12, e83516. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hayashi-Takagi, A.; Yagishita, S.; Nakamura, M.; Shirai, F.; Wu, Y.I.; Loshbaugh, A.L.; Kuhlman, B.; Hahn, K.M.; Kasai, H. Labelling and optical erasure of synaptic memory traces in the motor cortex. Nature 2015, 525, 333–338. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, G.; Pan, F.; Gan, W.B. Stably maintained dendritic spines are associated with lifelong memories. Nature 2009, 462, 920–924. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, X.; Li, Q.; Wang, L.; Liu, Z.J.; Zhong, Y. Active protection: Learning-activated Raf/MAPK activity protects labile memory from Rac1-independent forgetting. Neuron 2018, 98, 142–155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huszár, F. Note on the quadratic penalties in elastic weight consolidation. Proc. Natl. Acad. Sci. USA 2018, 115, E2496–E2497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McNaughton, B.L.; O’Reilly, R.C. Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of. Psychol. Rev. 1995, 102, 419–457. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Graves, L.; Nagisetty, V.; Ganesh, V. Does AI remember? Neural Networks and the Right to be Forgotten. Master’s Thesis University of Waterloo, Waterloo, ON, Canada, 2020. [Google Scholar]
- Ding, M.; Ji, K.; Wang, D.; Xu, J. Understanding forgetting in continual learning with linear regression. arXiv 2024, arXiv:2405.17583. [Google Scholar]
- Aljundi, R.; Lin, M.; Goujaud, B.; Bengio, Y. Gradient based sample selection for online continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Ritter, H.; Botev, A.; Barber, D. Online structured laplace approximations for overcoming catastrophic forgetting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
- Schwarz, J.; Czarnecki, W.; Luketina, J.; Grabska-Barwinska, A.; Teh, Y.W.; Pascanu, R.; Hadsell, R. Progress & compress: A scalable framework for continual learning. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; pp. 4528–4537. [Google Scholar]
- Gou, J.; Yu, B.; Maybank, S.J.; Tao, D. Knowledge distillation: A survey. Int. J. Comput. Vis. 2021, 129, 1789–1819. [Google Scholar] [CrossRef] [Scilit]
- Dhar, P.; Singh, R.V.; Peng, K.C.; Wu, Z.; Chellappa, R. Learning without memorizing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 5138–5146. [Google Scholar]
- Iscen, A.; Zhang, J.; Lazebnik, S.; Schmid, C. Memory-efficient incremental learning through feature adaptation. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 699–715. [Google Scholar]
- Li, Z.; Hoiem, D. Learning without forgetting. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 2935–2947. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Castro, F.M.; Marín-Jiménez, M.J.; Guil, N.; Schmid, C.; Alahari, K. End-to-end incremental learning. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 233–248. [Google Scholar]
- Douillard, A.; Cord, M.; Ollion, C.; Robert, T.; Valle, E. Podnet: Pooled outputs distillation for small-tasks incremental learning. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part XX 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 86–102. [Google Scholar]
- Hou, S.; Pan, X.; Loy, C.C.; Wang, Z.; Lin, D. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 831–839. [Google Scholar]
- Wu, C.; Herranz, L.; Liu, X.; Van De Weijer, J.; Raducanu, B. Memory replay gans: Learning to generate new categories without forgetting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
- Liu, X.; Masana, M.; Herranz, L.; Van de Weijer, J.; Lopez, A.M.; Bagdanov, A.D. Rotate your networks: Better weight consolidation and less catastrophic forgetting. In Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR); IEEE: New York, NY, USA, 2018; pp. 2262–2268. [Google Scholar]
- Benzing, F. Unifying importance based regularisation methods for continual learning. In Proceedings of the International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2022; pp. 2372–2396. [Google Scholar]
- Lee, S.W.; Kim, J.H.; Jun, J.; Ha, J.W.; Zhang, B.T. Overcoming catastrophic forgetting by incremental moment matching. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Chaudhry, A.; Rohrbach, M.; Elhoseiny, M.; Ajanthan, T.; Dokania, P.K.; Torr, P.H.; Ranzato, M. On tiny episodic memories in continual learning. arXiv 2019, arXiv:1902.10486. [Google Scholar]
- Vitter, J.S. Random sampling with a reservoir. ACM Trans. Math. Softw. (TOMS) 1985, 11, 37–57. [Google Scholar] [CrossRef] [Scilit]
- Borsos, Z.; Mutny, M.; Krause, A. Coresets via bilevel optimization for continual learning and streaming. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 14879–14890. [Google Scholar]
- Yoon, J.; Madaan, D.; Yang, E.; Hwang, S.J. Online coreset selection for rehearsal-based continual learning. arXiv 2021, arXiv:2106.01085. [Google Scholar]
- Shim, D.; Mai, Z.; Jeong, J.; Sanner, S.; Kim, H.; Jang, J. Online class-incremental continual learning with adversarial shapley value. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 9630–9638. [Google Scholar]
- Bang, J.; Kim, H.; Yoo, Y.; Ha, J.W.; Choi, J. Rainbow memory: Continual learning with a memory of diverse samples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 8218–8227. [Google Scholar]
- Tiwari, R.; Killamsetty, K.; Iyer, R.; Shenoy, P. Gcr: Gradient coreset based replay buffer selection for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 99–108. [Google Scholar]
- Van Den Oord, A.; Vinyals, O. Neural discrete representation learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
- Wang, L.; Zhang, X.; Yang, K.; Yu, L.; Li, C.; Hong, L.; Zhang, S.; Li, Z.; Zhong, Y.; Zhu, J. Memory replay with data compression for continual learning. arXiv 2022, arXiv:2202.06592. [Google Scholar]
- Kulesza, A.; Taskar, B. Determinantal point processes for machine learning. Found. Trends® Mach. Learn. 2012, 5, 123–286. [Google Scholar] [CrossRef] [Scilit]
- Kumari, L.; Wang, S.; Zhou, T.; Bilmes, J.A. Retrospective adversarial replay for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 28530–28544. [Google Scholar]
- Zhang, H.; Cisse, M.; Dauphin, Y.N.; Lopez-Paz, D. mixup: Beyond empirical risk minimization. arXiv 2017, arXiv:1710.09412. [Google Scholar]
- Belouadah, E.; Popescu, A. Il2m: Class incremental learning with dual memory. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 583–592. [Google Scholar]
- Ebrahimi, S.; Petryk, S.; Gokul, A.; Gan, W.; Gonzalez, J.E.; Rohrbach, M.; Darrell, T. Remembering for the right reasons: Explanations reduce catastrophic forgetting. Appl. AI Lett. 2021, 2, e44. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Su, Y.; Liu, A.A.; Schiele, B.; Sun, Q. Mnemonics training: Multi-class incremental learning without forgetting. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 12245–12254. [Google Scholar]
- Jin, X.; Sadhu, A.; Du, J.; Ren, X. Gradient-based editing of memory examples for online task-free continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 29193–29205. [Google Scholar]
- Chaudhry, A.; Ranzato, M.; Rohrbach, M.; Elhoseiny, M. Efficient Lifelong Learning with A-GEM. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2019. [Google Scholar]
- Tang, S.; Chen, D.; Zhu, J.; Yu, S.; Ouyang, W. Layerwise optimization by gradient decomposition for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 9634–9643. [Google Scholar]
- Sun, Q.; Lyu, F.; Shang, F.; Feng, W.; Wan, L. Exploring example influence in continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 27075–27086. [Google Scholar]
- Aljundi, R.; Belilovsky, E.; Tuytelaars, T.; Charlin, L.; Caccia, M.; Lin, M.; Page-Caccia, L. Online continual learning with maximal interfered retrieval. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Chaudhry, A.; Gordo, A.; Dokania, P.; Torr, P.; Lopez-Paz, D. Using hindsight to anchor past knowledge in continual learning. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 6993–7001. [Google Scholar]
- Wu, Y.; Chen, Y.; Wang, L.; Ye, Y.; Liu, Z.; Guo, Y.; Fu, Y. Large scale incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 374–382. [Google Scholar]
- Zhao, B.; Xiao, X.; Gan, G.; Zhang, B.; Xia, S.T. Maintaining discrimination and fairness in class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13208–13217. [Google Scholar]
- Ahn, H.; Kwak, J.; Lim, S.; Bang, H.; Kim, H.; Moon, T. Ss-il: Separated softmax for incremental learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 844–853. [Google Scholar]
- Cha, H.; Lee, J.; Shin, J. Co2l: Contrastive continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 9516–9525. [Google Scholar]
- Simon, C.; Koniusz, P.; Harandi, M. On learning the geodesic path for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 1591–1600. [Google Scholar]
- Joseph, K.; Khan, S.; Khan, F.S.; Anwer, R.M.; Balasubramanian, V.N. Energy-based latent aligner for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 7452–7461. [Google Scholar]
- Kurmi, V.K.; Patro, B.N.; Subramanian, V.K.; Namboodiri, V.P. Do not forget to attend to uncertainty while mitigating catastrophic forgetting. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2021; pp. 736–745. [Google Scholar]
- Ashok, A.; Joseph, K.; Balasubramanian, V.N. Class-incremental learning with cross-space clustering and controlled transfer. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 105–122. [Google Scholar]
- Hu, X.; Tang, K.; Miao, C.; Hua, X.S.; Zhang, H. Distilling causal effect of data in class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 3957–3966. [Google Scholar]
- Bhat, P.; Zonooz, B.; Arani, E. Task-aware information routing from common representation space in lifelong learning. arXiv 2023, arXiv:2302.11346. [Google Scholar]
- Hou, S.; Pan, X.; Loy, C.C.; Wang, Z.; Lin, D. Lifelong learning via progressive distillation and retrospection. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 437–452. [Google Scholar]
- Wang, F.Y.; Zhou, D.W.; Ye, H.J.; Zhan, D.C. Foster: Feature boosting and compression for class-incremental learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 398–414. [Google Scholar]
- Chaudhry, A.; Dokania, P.K.; Ajanthan, T.; Torr, P.H. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 532–547. [Google Scholar]
- Wang, L.; Zhang, M.; Jia, Z.; Li, Q.; Bao, C.; Ma, K.; Zhu, J.; Zhong, Y. Afec: Active forgetting of negative transfer in continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 22379–22391. [Google Scholar]
- Verwimp, E.; De Lange, M.; Tuytelaars, T. Rehearsal revealed: The limits and merits of revisiting samples in continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 9385–9394. [Google Scholar]
- Bonicelli, L.; Boschini, M.; Porrello, A.; Spampinato, C.; Calderara, S. On the effectiveness of lipschitz-driven rehearsal in continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 31886–31901. [Google Scholar]
- Yu, L.; Hu, T.; Hong, L.; Liu, Z.; Weller, A.; Liu, W. Continual learning by modeling intra-class variation. arXiv 2022, arXiv:2210.05398. [Google Scholar]
- Buzzega, P.; Boschini, M.; Porrello, A.; Abati, D.; Calderara, S. Dark experience for general continual learning: A strong, simple baseline. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 15920–15930. [Google Scholar]
- Boschini, M.; Bonicelli, L.; Buzzega, P.; Porrello, A.; Calderara, S. Class-incremental continual learning into the extended der-verse. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 5497–5512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Prabhu, A.; Torr, P.H.; Dokania, P.K. Gdumb: A simple approach that questions our progress in continual learning. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part II 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 524–540. [Google Scholar]
- Ayub, A.; Wagner, A.R. EEC: Learning to encode and regenerate images for continual learning. arXiv 2021, arXiv:2101.04904. [Google Scholar]
- Ostapenko, O.; Puscas, M.; Klein, T.; Jahnichen, P.; Nabi, M. Learning to remember: A synaptic plasticity driven framework for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 11321–11329. [Google Scholar]
- Kemker, R.; Kanan, C. Fearnet: Brain-inspired model for incremental learning. arXiv 2017, arXiv:1711.10563. [Google Scholar]
- Riemer, M.; Klinger, T.; Bouneffouf, D.; Franceschini, M. Scalable recollections for continual lifelong learning. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2019; Volume 33, pp. 1352–1359. [Google Scholar]
- Rostami, M.; Kolouri, S.; Pilly, P.K. Complementary learning for overcoming catastrophic forgetting using experience replay. In Proceedings of the 28th International Joint Conference on Artificial Intelligence; International Joint Conferences on Artificial Intelligence: Marina Del Rey, CA, USA, 2019; pp. 3339–3345. [Google Scholar]
- Pfülb, B.; Gepperth, A.; Bagus, B. Continual learning with fully probabilistic models. arXiv 2021, arXiv:2104.09240x. [Google Scholar]
- Gopalakrishnan, S.; Singh, P.R.; Fayek, H.; Ramasamy, S.; Ambikapathi, A. Knowledge capture and replay for continual learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2022; pp. 10–18. [Google Scholar]
- Ye, F.; Bors, A.G. Learning latent representations across multiple data domains using lifelong VAEGAN. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 777–795. [Google Scholar]
- Nguyen, C.V.; Li, Y.; Bui, T.D.; Turner, R.E. Variational continual learning. arXiv 2017, arXiv:1710.10628. [Google Scholar]
- Seff, A.; Beatson, A.; Suo, D.; Liu, H. Continual learning in generative adversarial nets. arXiv 2017, arXiv:1705.08395. [Google Scholar]
- He, C.; Wang, R.; Shan, S.; Chen, X. Exemplar-supported generative reproduction for class incremental learning. In Proceedings of the BMVC; BMVA (British Machine Vision Association): Durham, UK, 2018; Volume 1, p. 2. [Google Scholar]
- Xiang, Y.; Fu, Y.; Ji, P.; Huang, H. Incremental learning using conditional adversarial networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 6619–6628. [Google Scholar]
- Cong, Y.; Zhao, M.; Li, J.; Wang, S.; Carin, L. Gan memory with no forgetting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 16481–16494. [Google Scholar]
- Liu, X.; Wu, C.; Menta, M.; Herranz, L.; Raducanu, B.; Bagdanov, A.D.; Jui, S.; de Weijer, J.v. Generative feature replay for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2020; pp. 226–227. [Google Scholar]
- Ostapenko, O.; Lesort, T.; Rodriguez, P.; Arefin, M.R.; Douillard, A.; Rish, I.; Charlin, L. Continual learning with foundation models: An empirical study of latent replay. In Proceedings of the Conference on Lifelong Learning Agents; PMLR: Cambridge, MA, USA, 2022; pp. 60–91. [Google Scholar]
- Wang, Z.; Liu, L.; Kong, Y.; Guo, J.; Tao, D. Online continual learning with contrastive vision transformer. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 631–650. [Google Scholar]
- Wang, Y.; Huang, Z.; Hong, X. S-prompts learning with pre-trained transformers: An occam’s razor for domain incremental learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 5682–5695. [Google Scholar]
- Wang, Z.; Zhang, Z.; Ebrahimi, S.; Sun, R.; Zhang, H.; Lee, C.Y.; Ren, X.; Su, G.; Perot, V.; Dy, J.; et al. Dualprompt: Complementary prompting for rehearsal-free continual learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 631–648. [Google Scholar]
- Wang, Z.; Zhang, Z.; Lee, C.Y.; Zhang, H.; Sun, R.; Ren, X.; Su, G.; Perot, V.; Dy, J.; Pfister, T. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 139–149. [Google Scholar]
- Smith, J.S.; Karlinsky, L.; Gutta, V.; Cascante-Bonilla, P.; Kim, D.; Arbelle, A.; Panda, R.; Feris, R.; Kira, Z. Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 11909–11919. [Google Scholar]
- McDonnell, M.D.; Gong, D.; Parvaneh, A.; Abbasnejad, E.; van den Hengel, A. RanPAC: Random Projections and Pre-trained Models for Continual Learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); NeurIPS: San Diego, CA, USA, 2023; Volume 36, pp. 12022–12053. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the The Tenth International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2022. [Google Scholar]
- Zenke, F.; Poole, B.; Ganguli, S. Continual Learning Through Synaptic Intelligence. In Proceedings of the 34th International Conference on Machine Learning (ICML); Precup, D., Teh, Y.W., Eds.; PMLR (Proceedings of Machine Learning Research): Cambridge, MA, USA, 2017; Volume 70, pp. 3987–3995. [Google Scholar]
- Aljundi, R.; Babiloni, F.; Elhoseiny, M.; Rohrbach, M.; Tuytelaars, T. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 139–154. [Google Scholar]
- Rolnick, D.; Ahuja, A.; Schwarz, J.; Lillicrap, T.P.; Wayne, G. Experience Replay for Continual Learning. In Proceedings of the Advances in Neural Information Processing Systems 32 (NeurIPS 2019); Wallach, H.M., Larochelle, H., Beygelzimer, A., d’Alché Buc, F., Fox, E., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2019; pp. 348–358. [Google Scholar]
- Rusu, A.A.; Rabinowitz, N.C.; Desjardins, G.; Soyer, H.; Kirkpatrick, J.; Kavukcuoglu, K.; Pascanu, R.; Hadsell, R. Progressive Neural Networks. arXiv 2016, arXiv:1606.04671. [Google Scholar]
- Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised contrastive learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 18661–18673. [Google Scholar]
- Fini, E.; Da Costa, V.G.T.; Alameda-Pineda, X.; Ricci, E.; Alahari, K.; Mairal, J. Self-Supervised Models are Continual Learners. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022. [Google Scholar]
- Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-Efficient Transfer Learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (ICML); PMLR (Proceedings of Machine Learning Research): Cambridge, MA, USA, 2019; Volume 97, pp. 2790–2799. [Google Scholar]
- Wang, X.; Zhang, Y.; Chen, T.; Gao, S.; Jin, S.; Yang, X.; Xi, Z.; Zheng, R.; Zou, Y.; Gui, T.; et al. TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models. arXiv 2023, arXiv:2310.06762. [Google Scholar]
- Chen, H.; Wu, Z.; Han, X.; Jia, M.; Jiang, Y.G. Promptfusion: Decoupling stability and plasticity for continual learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 196–212. [Google Scholar]
- Park, C.W.; Seo, S.W.; Kang, N.; Ko, B.; Choi, B.W.; Park, C.M.; Chang, D.K.; Kim, H.; Kim, H.; Lee, H.; et al. Artificial intelligence in health care: Current applications and issues. J. Korean Med. Sci. 2020, 35, e379. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, D.; Bu, Q.; Zhu, Z.; Zhang, Y.; Wang, Z. Advancing autonomy through lifelong learning: A survey of autonomous intelligent systems. Front. Neurorobotics 2024, 18, 1385778. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ciupek, D.; Malawski, M.; Pieciak, T. Federated Learning: A new frontier in the exploration of multi-institutional medical imaging data. arXiv 2025, arXiv:2503.20107. [Google Scholar]
- Thakur, G.K.; Thakur, A.; Kulkarni, S.; Khan, N.; Khan, S. Deep learning approaches for medical image analysis and diagnosis. Cureus 2024, 16. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jeon, J.; Kim, J.; Kim, J.; Kim, K.; Mohaisen, A.; Kim, J.K. Privacy-preserving deep learning computation for geo-distributed medical big-data platforms. In Proceedings of the 2019 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks–Supplemental Volume (DSN-S); IEEE: New York, NY, USA, 2019; pp. 3–4. [Google Scholar]
- Pianykh, O.S.; Langs, G.; Dewey, M.; Enzmann, D.R.; Herold, C.J.; Schoenberg, S.O.; Brink, J.A. Continuous learning AI in radiology: Implementation principles and early applications. Radiology 2020, 297, 6–14. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pinto-Coelho, L. How artificial intelligence is shaping medical imaging technology: A survey of innovations and applications. Bioengineering 2023, 10, 1435. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhu, Z.; Sun, Y.; Honarvar Shakibaei Asli, B. Early Breast Cancer Detection Using Artificial Intelligence Techniques Based on Advanced Image Processing Tools. Electronics 2024, 13, 3575. [Google Scholar] [CrossRef] [Scilit]
- Lee, C.S.; Lee, A.Y. Applications of Continual Learning Machine Learning in Clinical Practice. Lancet Digit. Health 2020, 2, e279–e281. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Derakhshani, M.M.; Najdenkoska, I.; van Sonsbeek, T.; Zhen, X.; Mahapatra, D.; Worring, M.; Snoek, C.G.M. LifeLonger: A Benchmark for Continual Disease Classification. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention (MICCAI); Springer: Berlin/Heidelberg, Germany, 2022. [Google Scholar]
- Perkonigg, M.; Hofmanninger, J.; Herold, C.J.; Brink, J.A.; Pianykh, O.; Prosch, H.; Langs, G. Dynamic Memory to Alleviate Catastrophic Forgetting in Continual Learning with Medical Imaging. Nat. Commun. 2021, 12, 5678. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- da Silva Motta, D.; Badaró, R.; Santos, A.; Kirchner, F. Use of Artificial Intelligence on the Control of Vector-Borne Diseases; IntechOpen: London, UK, 2018. [Google Scholar]
- Dasari, S.; Ebert, F.; Tian, S.; Nair, S.; Bucher, B.; Schmeckpeper, K.; Singh, S.; Levine, S.; Finn, C. Robonet: Large-scale multi-robot learning. arXiv 2019, arXiv:1910.11215. [Google Scholar]
- Haque, N. Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks. arXiv 2025, arXiv:2504.01241. [Google Scholar]
- Yao, Y.; González-Vélez, H. AI-Powered System to Facilitate Personalized Adaptive Learning in Digital Transformation. Appl. Sci. 2025, 15, 4989. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Chen, Z.; Cho, E.; Hao, J.; Liu, X.; Xing, F.; Guo, C.; Liu, Y. Overcoming catastrophic forgetting during domain adaptation of seq2seq language generation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 5441–5454. [Google Scholar]
- Liu, T.; Ungar, L.; Sedoc, J. Continual learning for sentence representations using conceptors. arXiv 2019, arXiv:1904.09187. [Google Scholar]
- Monaikul, N.; Castellucci, G.; Filice, S.; Rokhlenko, O. Continual learning for named entity recognition. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 13570–13577. [Google Scholar]
- Li, G.; Zhai, Y.; Chen, Q.; Gao, X.; Zhang, J.; Zhang, Y. Continual few-shot intent detection. In Proceedings of the 29th International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 333–343. [Google Scholar]
- Liu, Q.; Yu, X.; He, S.; Liu, K.; Zhao, J. Lifelong intent detection via multi-strategy rebalancing. arXiv 2021, arXiv:2108.04445. [Google Scholar]
- Varshney, V.; Patidar, M.; Kumar, R.; Shroff, G.; Vig, L. Prompt Augmented Generative Replay via Supervised Contrastive Training for Lifelong Intent Detection. U.S. Patent App. 18/215,972, 11 January 2024. [Google Scholar]
- Qin, C.; Joty, S. Lfpt5: A unified framework for lifelong few-shot language learning based on prompt tuning of t5. arXiv 2021, arXiv:2110.07298. [Google Scholar]
- Sun, J.; Wang, S.; Zhang, J.; Zong, C. Distill and replay for continual language learning. In Proceedings of the 28th International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 3569–3579. [Google Scholar]
- Cao, Y.; Wei, H.R.; Chen, B.; Wan, X. Continual learning for neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 3964–3974. [Google Scholar]
- Shao, C.; Feng, Y. Overcoming catastrophic forgetting beyond continual learning: Balanced training for neural machine translation. arXiv 2022, arXiv:2203.03910. [Google Scholar]
- Qin, Y.; Zhang, J.; Lin, Y.; Liu, Z.; Li, P.; Sun, M.; Zhou, J. Elle: Efficient lifelong pre-training for emerging data. arXiv 2022, arXiv:2203.06311. [Google Scholar]
- Huang, Y.; Zhang, Y.; Chen, J.; Wang, X.; Yang, D. Continual learning for text classification with information disentanglement based regularization. arXiv 2021, arXiv:2104.05489. [Google Scholar]
- de Masson D’Autume, C.; Ruder, S.; Kong, L.; Yogatama, D. Episodic memory in lifelong language learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
- Wang, Z.; Mehta, S.V.; Póczos, B.; Carbonell, J. Efficient meta lifelong-learning with limited memory. arXiv 2020, arXiv:2010.02500. [Google Scholar]
- Xu, K.; Verma, S.; Finn, C.; Levine, S. Continual learning of control primitives: Skill discovery via reset-games. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 4999–5010. [Google Scholar]
- Mi, F.; Chen, L.; Zhao, M.; Huang, M.; Faltings, B. Continual learning for natural language generation in task-oriented dialog systems. arXiv 2020, arXiv:2010.00910. [Google Scholar]
- Li, Z.; Qu, L.; Haffari, G. Total recall: A customized continual learning method for neural semantic parsers. arXiv 2021, arXiv:2109.05186. [Google Scholar]
- Sun, F.K.; Ho, C.H.; Lee, H.Y. Lamol: Language modeling for lifelong language learning. arXiv 2019, arXiv:1909.03329. [Google Scholar]
- Zhang, Y.; Wang, X.; Yang, D. Continual sequence generation with adaptive compositional modules. arXiv 2022, arXiv:2203.10652. [Google Scholar]
- Wang, R.; Yu, T.; Zhao, H.; Kim, S.; Mitra, S.; Zhang, R.; Henao, R. Few-shot class-incremental learning for named entity recognition. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 571–582. [Google Scholar]
- Geng, B.; Yuan, F.; Xu, Q.; Shen, Y.; Xu, R.; Yang, M. Continual learning for task-oriented dialogue system with iterative network pruning, expanding and masking. arXiv 2021, arXiv:2107.08173. [Google Scholar]
- Shen, Y.; Zeng, X.; Jin, H. A progressive model to enable continual learning for semantic slot filling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 1279–1284. [Google Scholar]
- Wang, C.; Pan, H.; Liu, Y.; Chen, K.; Qiu, M.; Zhou, W.; Huang, J.; Chen, H.; Lin, W.; Cai, D. Mell: Large-scale extensible user intent classification for dialogue systems with meta lifelong learning. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining; ACM: New York, NY, USA, 2021; pp. 3649–3659. [Google Scholar]
- Wu, T.; Li, X.; Li, Y.F.; Haffari, G.; Qi, G.; Zhu, Y.; Xu, G. Curriculum-meta learning for order-robust continual relation extraction. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 10363–10369. [Google Scholar]
- Madotto, A.; Lin, Z.; Zhou, Z.; Moon, S.; Crook, P.A.; Liu, B.; Yu, Z.; Cho, E.; Fung, P.; Wang, Z. Continual learning in task-oriented dialogue systems. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 7452–7467. [Google Scholar]
- Ermis, B.; Zappella, G.; Wistuba, M.; Rawal, A.; Archambeau, C. Memory efficient continual learning with transformers. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 10629–10642. [Google Scholar]
- Zhu, Q.; Li, B.; Mi, F.; Zhu, X.; Huang, M. Continual prompt tuning for dialog state tracking. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 1124–1137. [Google Scholar]
- Liu, M.; Chang, S.; Huang, L. Incremental prompting: Episodic memory prompt for lifelong event detection. In Proceedings of the 29th International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 2157–2165. [Google Scholar]
- Yin, W.; Li, J.; Xiong, C. Contintin: Continual learning from task instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 3062–3072. [Google Scholar]
- Xia, C.; Yin, W.; Feng, Y.; Yu, P.S. Incremental few-shot text classification with multi-round new classes: Formulation, dataset and system. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 1351–1360. [Google Scholar]
- Garcia, X.; Constant, N.; Parikh, A.; Firat, O. Towards continual learning for multilingual machine translation via vocabulary substitution. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 1184–1192. [Google Scholar]
- Yan, S.; Hong, L.; Xu, H.; Han, J.; Tuytelaars, T.; Li, Z.; He, X. Generative negative text replay for continual vision-language pretraining. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 22–38. [Google Scholar]
- Greco, C.; Plank, B.; Fernández, R.; Bernardi, R. Psycholinguistics meets continual learning: Measuring catastrophic forgetting in visual question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 3601–3605. [Google Scholar]
- Martínez-Plumed, F.; Ferri, C.; Hernández-Orallo, J.; Ramírez-Quintana, M.J. Forgetting and consolidation for incremental and cumulative knowledge acquisition systems. arXiv 2015, arXiv:1502.05615. [Google Scholar]
- Christakopoulou, K.; Lalama, A.; Adams, C.; Qu, I.; Amir, Y.; Chucri, S.; Vollucci, P.; Soldo, F.; Bseiso, D.; Scodel, S.; et al. Large language models for user interest journeys. arXiv 2023, arXiv:2305.15498. [Google Scholar]
- Wang, X.J.; Lee, C.P.; Mutlu, B. LearnMate: Enhancing Online Education with LLM-Powered Personalized Learning Plans and Support. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 2025; pp. 1–10. [Google Scholar]
- Sabeima, M.; Lamolle, M.; Nanne, M.F. Towards personalized adaptive learning in e-learning recommender systems. Int. J. Adv. Comput. Sci. Appl. 2022, 13, 14–20. [Google Scholar] [CrossRef] [Scilit]
- Joy, J.; Raj, N.S.; VG, R. Ontology-based E-learning content recommender system for addressing the pure cold-start problem. ACM J. Data Inf. Qual. 2021, 13, 1–27. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Soljacic, M.; Hou, T.; Tegmark, M. KAN: Kolmogorov–arnold networks. In Proceedings of the International Conference on Learning Representations; ICLR: Appleton, WI, USA, 2025; Volume 2025, pp. 70367–70413. [Google Scholar]
- Bountouni, N.; Koussouris, S.; Vasileiou, A.; Kazazis, S.A. A Holistic Framework for Safeguarding of SMEs: A Case Study. In Proceedings of the 2023 19th International Conference on the Design of Reliable Communication Networks (DRCN); IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar]
- Asmar, M.; Tuqan, A. Integrating machine learning for sustaining cybersecurity in digital banks. Heliyon 2024, 10, e37571. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ahmed, U.; Nazir, M.; Sarwar, A.; Ali, T.; Aggoune, E.H.M.; Shahzad, T.; Khan, M.A. Signature-based intrusion detection using machine learning and deep learning approaches empowered with fuzzy clustering. Sci. Rep. 2025, 15, 1726. [Google Scholar] [CrossRef] [Scilit]
- Dohare, S.; Hernandez-Garcia, J.F.; Lan, Q.; Rahman, P.; Mahmood, A.R.; Sutton, R.S. Loss of plasticity in deep continual learning. Nature 2024, 632, 768–774. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jones, R.; Omar, M.; Mohammed, D.; Nobles, C.; Dawson, M. Harnessing the speed and accuracy of machine learning to advance cybersecurity. In Proceedings of the 2023 Congress in Computer Science, Computer Engineering, & Applied Computing (CSCE); IEEE: New York, NY, USA, 2023; pp. 418–421. [Google Scholar]
- Rahul-Vigneswaran, K.; Poornachandran, P.; Soman, K. A compendium on network and host based intrusion detection systems. In Proceedings of the ICDSMLA 2019: Proceedings of the 1st International Conference on Data Science, Machine Learning and Applications; Springer: Berlin/Heidelberg, Germany, 2020; pp. 23–30. [Google Scholar]
- Stokes, J.W.; Wang, D.; Marinescu, M.; Marino, M.; Bussone, B. Attack and defense of dynamic analysis-based, adversarial neural malware detection models. In Proceedings of the MILCOM 2018 IEEE Military Communications Conference (MILCOM); IEEE: New York, NY, USA, 2018; pp. 1–8. [Google Scholar]
- Sameen, M.; Han, K.; Hwang, S.O. PhishHaven-An efficient real-time AI phishing URLs detection system. IEEE Access 2020, 8, 83425–83443. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Chen, M.; Bukharin, A.; He, P.; Cheng, Y.; Chen, W.; Zhao, T. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2023. [Google Scholar]
- Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); NeurIPS: San Diego, CA, USA, 2023. [Google Scholar]
- Liu, S.Y.; Wang, C.Y.; Yin, H.; Molchanov, P.; Wang, Y.C.F.; Cheng, K.T.; Chen, M.H. DoRA: Weight-Decomposed Low-Rank Adaptation. In Proceedings of the 41st International Conference on Machine Learning (ICML); OpenReview.net: Newton Highlands, MA, USA, 2024. [Google Scholar]
- Chen, Y.; Qian, S.; Tang, H.; Lai, X.; Liu, Z.; Han, S.; Jia, J. LongLoRA: Efficient Fine-Tuning of Long-Context Large Language Models. In International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2024. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]







| Database | Representative Search Query |
|---|---|
| IEEE Xplore | (“continual learning” OR “lifelong learning” OR “incremental learning”) AND (“catastrophic forgetting” OR “experience replay” OR “regularization” OR “knowledge distillation”) |
| ACM Digital Library | (“continual learning” OR “lifelong learning”) AND (“foundation models” OR “prompt learning” OR “Parameter-Efficient Fine-Tuning”) |
| SpringerLink | (“continual learning”) AND (“vision transformer” OR “foundation model” OR “multimodal learning”) |
| ScienceDirect | (“continual learning”) AND (“experience replay” OR “regularization” OR “benchmark”) |
| Web of Science | TS = (“continual learning” OR “incremental learning”) AND TS = (“catastrophic forgetting” OR “foundation model”) |
| Scopus | TITLE-ABS-KEY (“continual learning” OR “lifelong learning”) AND (“multimodal” OR “prompt learning” OR “PEFT”) |
| Google Scholar | “continual learning” “foundation model” “prompt learning” “Parameter-Efficient Fine-Tuning” |
| Representative Study Type | Reason for Exclusion |
|---|---|
| Transfer learning studies | Focused exclusively on transfer learning without continual or lifelong adaptation. |
| Multi-task learning studies | Addressed simultaneous multi-task optimization rather than sequential continual learning. |
| Hardware implementation studies | Primarily focused on hardware acceleration or system implementation without methodological contributions to continual learning. |
| Workshop papers | Excluded because they lacked sufficient methodological detail or comprehensive experimental validation. |
| Duplicate publications | Earlier preprint versions were excluded when a corresponding peer-reviewed journal or conference paper was available. |
| Studies outside the review scope | Focused on unrelated topics that did not directly address continual learning algorithms, evaluation, theory, or applications. |
| Survey | Year | Main Focus | CL Coverage | Modern Trends | Main Limitation |
|---|---|---|---|---|---|
| Wang et al. [11] | 2024 | General CL theory and methods | TIL, DIL, and CIL | Limited discussion of prompting and PEFT | Minimal focus on foundation-model adaptation and modern multimodal CL |
| Van de Ven et al. [18] | 2022 | Taxonomy of CL scenarios | TIL, DIL, and CIL | Does not cover recent CL trends | Primarily focused on conceptual categorization of CL settings |
| Bidaki et al. [19] | 2025 | Online CL | Streaming and online CL | Benchmark-oriented discussion | Narrow scope centered on online learning settings |
| Zhou et al. [20] | 2024 | Class-incremental learning | Mainly CIL | Limited multimodal and foundation-model discussion | Restricted primarily to CIL strategies and benchmarks |
| Wickramasinghe et al. [21] | 2023 | Overview of CL methods | General CL settings | Covers traditional CL methods | Limited synthesis of transformer-, prompt-, and foundation-model-based continual learning methods |
| This review | 2026 | Comprehensive review of modern continual learning | TIL, DIL, CIL, data-incremental, online, multimodal, and federated CL | Foundation models, prompt learning, PEFT, diffusion models, and multimodal continual learning | Introduces an expanded taxonomy of continual learning scenarios and methodological categories, provides a unified analysis of evaluation protocols, benchmark fragmentation, and reproducibility, critically examines foundation-model-based continual learning, and synthesizes emerging research directions for scalable and deployment-oriented lifelong learning |
| Feature | CL | Transfer Learning | Multi-Task Learning | Online Learning |
|---|---|---|---|---|
| Task Availability | Sequential | One-time transfer | Simultaneous | Single task |
| Focus | Learning without forgetting | Knowledge transfer | Shared representation | Incremental updates |
| Addresses Forgetting | Yes | No | No | No |
| Data Distribution | Non-stationary | Varies | Varies | Stationary |
| Scenario | Task Label Know | Output Space | Data Distribution | Example Application |
|---|---|---|---|---|
| Task-Incremental | Yes | Varies | Changes | Multi-task NLP, robotics |
| Domain-Incremental | NO | Same | Changes | Handwriting recognition, IoT sensors |
| Class-Incremental | No | Expands | Changes | Image classification, object detection |
| Instance- Incremental | N/A | Same | Same (new data) | Spam filtering, online analytics |
| Unsupervised/Other | N/A | N/A | Changes | Clustering, RL in dynamic settings |
| Aspect | Details |
|---|---|
| Definition | Models learn a sequence of distinct tasks, with task identity provided during both training and inference. |
| Core Challenge | Maintaining task-specific performance without interference between tasks (catastrophic forgetting). |
| Inference Requirement | Task identity is known, allowing the model to use task-specific components (e.g., separate output heads). |
| Key Techniques | - Task-specific output heads. - Parameter isolation (dedicated parameters for each task). - Regularization to preserve important parameters. |
| Advantages | - Robust retention of task-specific knowledge. - Simplified learning due to known task boundaries and identities. |
| Challenges | - Scalability issues with a growing number of tasks. - Limited knowledge transfer between tasks. |
| Example Applications | - Sequential learning of different object categories (e.g., animals, vehicles). - Robotics: Learning distinct tasks like grasping and navigation. - Diagnostic systems for different modalities (e.g., X-rays, MRIs). |
| Evaluation Metrics | - Task-specific accuracy. - Memory and computational efficiency for handling multiple tasks. |
| Future Directions | - Modular architectures with adaptive parameter sharing and efficient knowledge transfer. |
| Aspect | Details |
|---|---|
| Definition | Models learn to adapt to new data distributions (domains) over time while maintaining the same task objective. |
| Core Challenge | Adapting to new domains without forgetting knowledge of previously learned domains (catastrophic forgetting). |
| Inference Requirement | Task identity is unknown; the model must generalize across domains without explicit domain information. |
| Key Techniques | - Domain adaptation methods (e.g., feature alignment). - Regularization techniques to retain domain-invariant features. - Memory replay or dynamic models to balance old and new knowledge. |
| Advantages | - Allows systems to handle non-stationary data distributions. - Maintains consistent task performance across multiple domains. |
| Challenges | - Catastrophic forgetting when adapting to new domains. - Handling domain-specific biases while ensuring generalization. - Computational and memory constraints as new domains increase. |
| Example Applications | - Object recognition in different environmental conditions (e.g., sunny, foggy, rainy). - Medical imaging systems adapting to scans from different hospitals or devices. - NLP tasks such as sentiment analysis across different domains (e.g., movie reviews, product reviews). |
| Evaluation Metrics | - Performance consistency across domains. -Forgetting rate for previously learned domains. - Domain generalization ability on unseen domains. |
| Future Directions | - Efficient methods for domain adaptation without overfitting to new domains. - Scalable approaches to handle increasing numbers of domains. - Techniques to balance domain-specific and domain-invariant learning. |
| Aspect | Details |
|---|---|
| Definition | Models learn new classes sequentially, and the task identity is not provided during inference. |
| Core Challenge | Catastrophic forgetting-new learning overwrites knowledge of previously learned classes. |
| Inference Requirement | Model must classify inputs across all learned classes without knowledge of task identity. |
| Key Techniques | - Memory replay (storing/replaying previous class examples). - KD (preserving learned representations). - Dynamic architecture (expanding capacity for new classes). |
| Advantages | - Enables incremental learning without full retraining. - Efficient handling of scenarios where new class data is available over time. |
| Challenges | - Handling class imbalance, as new classes often have fewer examples. - Managing memory and computational costs as the number of classes increases. |
| Example Applications | - Extending image classifiers with new object categories. - Autonomous vehicles learning new traffic signs and objects. - Healthcare models adapting to diagnose new diseases. |
| Evaluation Metrics | - Accuracy across all classes (old and new). - Forgetting rate (performance drop on previously learned classes). |
| Future Directions | - Scalable memory-efficient replay methods. - Adaptive architectures that balance stability and plasticity. - Improved algorithms for mitigating class imbalance and preserving older class knowledge. |
| Aspect | Details |
|---|---|
| Definition | Models learn incrementally from a stream of data instances, which may belong to existing or new classes, without explicit task boundaries. |
| Core Challenge | Adapting to new data while retaining knowledge of previously learned data, especially without clear transitions or task identities. |
| Inference Requirement | The model must classify instances across all learned classes without explicit knowledge of when new data or classes were introduced. |
| Key Techniques | - Memory replay (storing or generating past data). - Regularization techniques to preserve critical parameters. - Dynamic architectures for flexible capacity adjustment. |
| Advantages | - Handles continuously evolving data streams. - Allows for learning without task-specific information or retraining. |
| Challenges | - Managing catastrophic forgetting as new data arrives. - Handling class imbalance and unstructured data streams. - Resource efficiency for memory and computational costs. |
| Example Applications | - Object recognition systems that adapt to new categories dynamically. - Recommendation systems updating preferences with new user data and items. - Continuous monitoring systems in healthcare, incorporating evolving signals from wearable devices. |
| Evaluation Metrics | - Accuracy across all classes (old and new). - Forgetting rate (performance drop on previously learned data). - Adaptation speed to new data. |
| Future Directions | - Hybrid methods combining memory replay with adaptive architectures. - Scalable solutions for handling large and imbalanced data streams. - Techniques for efficient data prioritization and representation learning. |
| Paradigm | Description | Key Challenges | Key Techniques | Example Applications |
|---|---|---|---|---|
| Few-Shot CL | Models learn new tasks or classes with minimal labeled data while retaining prior knowledge. | - Adapting with limited data. - Avoiding catastrophic forgetting. | - Meta-learning. - Episodic memory. - Generative replay. | - Rare disease diagnosis. - Few-shot object recognition. |
| Unsupervised CL | Models learn from data streams without explicit labels by discovering patterns or structures. | - Extracting meaningful features from unlabeled data. - Balancing old and new pattern representations. | - Self-supervised learning. - Contrastive learning. - Clustering methods. | - Video surveillance anomaly detection. - Social media trend analysis. |
| Meta-Continual Learning | Combines meta-learning with CL to enable rapid adaptation to new tasks. | - Balancing fast adaptation with knowledge retention. - Stability–plasticity trade-off. | - Gradient-based meta-learning. - Memory-augmented neural networks. | - Personalized AI assistants. - Adaptive recommendation systems. |
| Federated CL | Models learn incrementally across distributed nodes while preserving privacy. | - Handling heterogeneous data distributions across nodes. - Avoiding forgetting across distributed devices. - Privacy concerns. | - Decentralized learning algorithms. - Secure aggregation protocols. - Adaptive synchronization methods. | - Personalized healthcare monitoring. - Mobile device personalization. |
| Multi-Agent CL | Multiple agents learn and adapt in a shared environment while interacting and collaborating. | - Coordinating knowledge transfer between agents. - Managing inter-agent dependencies and scalability. | - Communication protocols. - Shared memory systems. - Ensemble learning. | - Collaborative robotics. - Distributed sensor networks. |
| Concept | Core Idea | Main Challenge | Representative Strategies |
|---|---|---|---|
| Stability–Plasticity Dilemma | Balancing retention of prior knowledge with adaptation to new information | Excessive stability limits adaptation, while excessive plasticity causes forgetting | Regularization, replay mechanisms, adaptive architectures |
| Catastrophic Forgetting | Learning new tasks degrades performance on earlier tasks | Parameter interference and overlapping representations | Replay methods, parameter isolation, knowledge distillation, regularization |
| Forward and Backward Transfer | Leveraging previous knowledge to improve future learning and vice versa | Avoiding negative transfer across tasks | Shared representations, multi-task learning, transferable feature learning |
| Representation Learning | Learning reusable and task-invariant feature representations | Separating task-specific and generalizable features | Self-supervised learning, contrastive learning, feature disentanglement |
| Neuroscientific Inspiration | Drawing inspiration from biological memory and adaptation mechanisms | Translating biological principles into scalable AI systems | Synaptic consolidation, rehearsal mechanisms, dynamic expansion |
| Practical and Ethical Considerations | Ensuring reliable and responsible continual adaptation | Resource constraints, fairness, privacy, and safety | Lightweight models, federated learning, fairness-aware training |
| Aspect | Description |
|---|---|
| Catastrophic Forgetting | Learning new tasks reduces performance on previously learned tasks. |
| Negative Transfer | Knowledge from previous tasks impairs learning or generalization on future tasks. |
| Primary Cause | Parameter overwriting and representation drift. |
| Primary Cause of Negative Transfer | Conflicting task representations and gradient interference. |
| Typical Solutions | Replay, regularization, parameter isolation, knowledge distillation. |
| Negative Transfer Mitigation | Gradient projection, modular networks, prompt tuning, task similarity estimation, parameter-efficient adaptation. |
| Challenge | Cause | Open Research Direction |
|---|---|---|
| Prompt interference | Competition among task-specific prompts during long continual learning sequences | Prompt routing, prompt composition, dynamic prompt allocation |
| Multimodal alignment collapse | Drift in shared embedding space during continual multimodal adaptation | Alignment-preserving continual optimization and modality-aware replay |
| Adapter accumulation | Increasing number of task-specific adapters and LoRA modules | Adapter compression, parameter sharing, dynamic module selection |
| Instruction drift | Instruction-following ability changes after continual fine-tuning | Instruction-aware continual learning and alignment regularization |
| Long-term reasoning degradation | Sequential updates alter pretrained reasoning capabilities | Knowledge consolidation and reasoning-aware continual adaptation |
| Evaluation inconsistency | Lack of benchmarks for foundation-model continual learning | Unified benchmarks including multimodal reasoning and long-term adaptation |
| Aspect | Description |
|---|---|
| Definition | The significant loss of performance on previously learned tasks when a neural network learns new tasks. |
| Cause | Overwriting of neural network parameters due to global updates during training on new tasks. |
| Key Mechanism | Parameter Drift: Critical parameters for previous tasks are modified to optimize new task learning. |
| Factors Exacerbating Forgetting | - Overlapping representations shared by different tasks. - Sequential data access without revisiting earlier tasks. - Lack of task awareness during inference in class-/domain-incremental settings. |
| Examples | - A model trained to classify animals forgetting how to classify vehicles after learning new classes. - An object detection model in autonomous driving failing to recognize stop signs after adapting to new road signs. |
| Mitigation Strategy | Description | Examples |
|---|---|---|
| Regularization Methods | Introduce constraints during training to prevent significant updates to parameters crucial for earlier tasks. | - EWC: Penalizes parameter changes. - Synaptic Intelligence: Tracks parameter importance. |
| Replay-Based Methods | Retain and replay data from previous tasks during training on new tasks. | - Experience Replay: Stores a subset of prior task data. - Generative Replay: Generates synthetic data from past tasks. |
| Dynamic Architectures | Expand or adapt the network architecture to allocate new resources for each task. | - Progressive Neural Networks: Adds new parameters per task. - Dynamically expandable networks. |
| Representation Learning | Learns generalizable features that can be reused across tasks, reducing task-specific interference. | - Self-supervised pretraining. - Disentangled representations. |
| Hybrid Approaches | Combine multiple strategies, such as regularization with replay or dynamic architectures. | - Replay with EWC to balance plasticity and stability. |
| Evaluation Metrics | Description | |
| Forgetting Rate | Measures the drop in performance on previously learned tasks after learning new ones. | |
| Accuracy | Assesses performance across all tasks (old and new). | |
| Knowledge Transfer | Evaluates how well the model uses previous knowledge to improve learning on new tasks. | |
| Aspect | Experience Replay | Generative Replay |
|---|---|---|
| Memory usage | Stores selected raw samples or compressed examples. | Stores a generative model that synthesizes previous data. |
| Privacy | May be problematic when previous data are sensitive. | Avoids direct storage of raw samples but may still leak information if not properly controlled. |
| Replay quality | High fidelity because original samples are replayed. | Depends on generator quality, diversity, and label consistency. |
| Computational cost | Relatively low compared with training a generator. | Higher due to training and maintaining a generative model. |
| Best suited for | Class-incremental learning and reinforcement learning when memory is available. | Privacy-sensitive or memory-constrained settings where raw data cannot be stored. |
| Method Category | Memory | Comp. | Scalability | Benchmark Performance | Suitable CL Scenarios | Foundation Models | Major Limitation |
|---|---|---|---|---|---|---|---|
| Regularization-Based | Low | Low | High | Medium | TIL, DIL | Limited | Performance degrades under long task sequences and severe domain shifts. |
| Replay-Based | High | Medium | Medium | High | CIL, DIL, Online | Moderate | Requires large replay memory and raises privacy concerns. |
| Architecture-Based | High | Medium | Low | High | TIL, CIL | Moderate | Model size continuously increases as new tasks are added. |
| Optimization-Based | Low | High | Medium | Medium | Online, CIL | Limited | High optimization cost due to gradient conflict management. |
| Representation Learning | Medium | Medium | High | High | DIL, SSL, Multimodal | High | Sensitive to representation drift under significant distribution shifts. |
| Prompt-Based | Low | Low | High | High | CIL, Multimodal | Excellent | Prompt interference may occur during long continual adaptation. |
| PEFT | Low | Low | High | High | Foundation-Model CL | Excellent | Accumulation of adapters increases long-term management complexity. |
| Federated Continual Learning | Medium | High | Medium | Medium | Distributed CL | Moderate | Communication overhead and heterogeneous client distributions. |
| Method Category | Final Accuracy | Forgetting | Memory Cost | Training Cost | Scalability | Official Code | Representative Methods |
|---|---|---|---|---|---|---|---|
| Regularization | Medium | High | Very Low | Low | High | Yes | EWC [12], SI [211], and MAS [212] |
| Replay | High | Low | High | Medium | Medium | Yes | ER [213], DER++ [186], and GDumb [188] |
| Architecture-Based | High | Very Low | Very High | Medium | Low | Partial | PNN [214] and Expert Gate [61] |
| Optimization-Based | Medium–High | Medium | Low | High | Medium | Yes | GEM [57] and A-GEM [164] |
| Representation Learning | High | Medium | Medium | Medium | High | Yes | SupCon [215] and self-supervised CL [216] |
| Prompt Learning | High | Low | Very Low | Low | High | Yes | L2P [207], DualPrompt [206], and CODA-Prompt [208] |
| PEFT | High | Low | Low | Low | High | Yes | LoRA [210] and adapters [217] |
| Foundation Models | Very High | Medium | Medium | Very High | Medium | Yes | RanPAC [209] and TRACE [218] |
| Method | Split CIFAR-100 | Split TinyImageNet | ImageNet-R | Average Forgetting |
|---|---|---|---|---|
| EWC [12] | 58–65 | 42–48 | 35–40 | High |
| SI [211] | 60–66 | 43–49 | 36–41 | High |
| ER [213] | 70–78 | 56–63 | 48–54 | Low |
| DER++ [186] | 73–81 | 60–67 | 52–58 | Very Low |
| A-GEM [164] | 63–70 | 48–55 | 40–46 | Medium |
| L2P [207] | 82–86 | 70–75 | 62–66 | Very Low |
| DualPrompt [206] | 84–88 | 72–77 | 64–69 | Very Low |
| LoRA-based CL [210] | 83–87 | 71–76 | 63–68 | Low |
| Evaluation Component | Recommended Reporting Practice |
|---|---|
| Continual learning scenario | Clearly specify whether the study follows task-, domain-, class-, online-, or data-incremental learning. |
| Benchmark datasets | Report all datasets, task order, number of tasks, number of classes, and train/test splits. |
| Model architecture | Describe backbone network, parameter initialization, pretrained weights, and continual learning components. |
| Memory budget | Report replay memory size, storage strategy, memory sampling policy, and whether memory size is fixed or adaptive. |
| Evaluation metrics | Report final average accuracy (FAA), average forgetting (AF), backward transfer (BWT), forward transfer (FWT), and task-wise performance whenever applicable. |
| Computational efficiency | Report training time, inference time, model parameters, FLOPs, GPU memory usage, and computational overhead. |
| Statistical analysis | Report mean and standard deviation over multiple random seeds together with statistical significance tests whenever appropriate. |
| Baseline comparison | Compare against representative methods from replay-based, regularization-based, architecture-based, optimization-based, and parameter-efficient continual learning. |
| Implementation details | Provide optimizer, learning rate, batch size, number of epochs, hardware configuration, and software framework. |
| Reproducibility | Release source code, pretrained models, dataset preprocessing scripts, random seeds, and configuration files whenever possible. |
| Benchmark | Domain | Typical Scenario | Common Metrics | Foundation Models | Main Limitations |
|---|---|---|---|---|---|
| Split CIFAR-100 | Vision | Class-IL | Accuracy, Forgetting | Limited | Small-scale dataset with limited semantic diversity. |
| Split Tiny-ImageNet | Vision | Class-IL | Accuracy, BWT, FWT | Moderate | Lower visual diversity than ImageNet and limited realism. |
| Split ImageNet | Vision | Class-IL | Accuracy, Forgetting | Excellent | High computational cost and long training time. |
| CORe50 | Vision | Domain-IL | Accuracy | Moderate | Limited object diversity and relatively small scale. |
| DomainNet | Vision | Domain-IL | Accuracy, Domain Generalization | Excellent | Domain imbalance and high computational requirements. |
| CLVision Benchmark | Vision | Multiple CL Settings | FAA, AF, BWT | Excellent | Protocols vary across benchmark configurations. |
| CLUE/NLP Benchmarks | NLP | Task-IL, Domain-IL | Accuracy, F1 | Excellent | Limited long-horizon continual adaptation. |
| LLM Continual Learning Benchmarks | LLMs | Instruction Continual Learning | Task Accuracy, Forgetting | Native | Lack of standardized protocols and high computational cost. |
| Method | Category | Benchmark | Final Avg. Accuracy (%) | Avg. Forgetting (%) |
|---|---|---|---|---|
| EWC [12] | Regularization | Split CIFAR-100 | ≈58 | ≈18 |
| iCaRL [66] | Replay | Split CIFAR-100 | ≈64 | ≈13 |
| DER++ [186] | Replay | Split CIFAR-100 | ≈74 | ≈7 |
| L2P [207] | Prompt-based | Split ImageNet-R | ≈81 | Low |
| DualPrompt [206] | Prompt-based | Split ImageNet-R | ≈84 | Very Low |
| CODA-Prompt [208] | Prompt-based | Split ImageNet-R | ≈87 | Very Low |
| DER++ [219] | Replay | CORe50 | ≈87 | Low |
| Method Category | Key Characteristics | Main Challenges | Replay Memory | Parameter Overhead | Computational Overhead | Typical Applications |
|---|---|---|---|---|---|---|
| Regularization-Based Methods | Constrain parameter updates to preserve previous knowledge; memory-efficient and easy to integrate | Limited performance under severe domain shifts and long task sequences | None | Low | Low | Task-Incremental Learning, resource-constrained systems, privacy-sensitive applications |
| Replay-Based Methods | Replay stored or generated samples to reinforce previous knowledge; strong retention performance | Replay buffer management, privacy concerns, and storage overhead | Required (fixed-size buffer or synthetic replay) | Low | Moderate | Class-incremental learning, reinforcement learning, streaming adaptation |
| Architecture-Based Methods | Allocate task-specific modules or expandable subnetworks to reduce interference | Poor scalability due to parameter growth and increasing model complexity | None | High (grows with tasks) | Moderate–High | Task-Incremental Learning with explicit task boundaries |
| Optimization-Based Methods | Modify gradient updates to balance stability and plasticity during training | High optimization complexity and gradient computation overhead | None | Low | High | Gradient-constrained continual adaptation and stability-focused learning |
| Representation-Learning Methods | Learn transferable and domain-invariant feature representations across tasks | Representation drift under highly heterogeneous task distributions | Optional | Low–Moderate | Moderate | Domain-incremental learning and self-supervised continual adaptation |
| Prompt-Based and PEFT Methods | Adapt pretrained foundation models using prompts, adapters, or low-rank updates | Prompt interference, adapter scalability, and long-term stability | None | Very Low (small trainable modules) | Low | Foundation models, multimodal systems, and large-scale deployment |
| Federated and Privacy-Aware CL | Enable continual learning across distributed clients without centralized data sharing | Client drift, communication overhead, and heterogeneous data distributions | Optional (local buffers) | Moderate | High (communication + synchronization) | Healthcare, finance, edge AI, and mobile systems |
| Method Category | Training Memory | Inference Memory | Long-Term Storage | Primary Source of Resource Consumption |
|---|---|---|---|---|
| Regularization-Based | Low–Moderate | Low | Low | Parameter importance statistics (e.g., Fisher information and importance weights). |
| Replay-Based | High | Low | High | Replay buffer or generative model used to preserve previous knowledge. |
| Architecture-Based | Moderate | High | Moderate | Additional task-specific subnetworks, adapters, or classifier heads. |
| Optimization-Based | Moderate–High | Low | Low | Gradient manipulation and optimization statistics. |
| Representation-Learning | Moderate | Low–Moderate | Low–Moderate | Feature representations and auxiliary embedding spaces. |
| Prompt-Based and PEFT | Low | Low–Moderate | Moderate | Task-specific prompts, adapters, or low-rank parameter updates. |
| Federated Continual Learning | Moderate | Moderate | Moderate | Local models, communication buffers, and client synchronization. |
| Application Area | Description | Key Benefits | Examples |
|---|---|---|---|
| Healthcare and Medical Imaging | Enables dynamic adaptation to evolving medical knowledge, diseases, and patient data over time. | Personalized diagnostics, improved adaptability, and long-term patient monitoring. | Radiology systems adapting to new imaging techniques or emerging diseases like novel cancer types. |
| Robotics and Autonomous Systems | Allows robots and autonomous systems to learn new tasks, adapt to dynamic environments, and retain prior knowledge. | Efficient task performance, knowledge transfer, and adaptability in real-world scenarios. | Household robots learning new cleaning techniques while retaining old capabilities like object recognition. |
| Natural Language Processing (NLP) | Helps models stay updated with evolving language patterns, domain-specific knowledge, and user preferences. | Better understanding of new language constructs, improved domain adaptation, and enhanced usability. | Chatbots adapting to new slang or technical jargon while maintaining general conversational abilities. |
| Recommender Systems | Adapts to changing user preferences and updates content or product catalogs dynamically. | Improved user engagement, personalized recommendations, and scalability for diverse user bases. | Streaming platforms suggesting trending shows based on current preferences without forgetting past ones. |
| Cybersecurity | Learns from new attack patterns and threat vectors while retaining the ability to recognize older threats. | Improved security, real-time threat detection, and reduced vulnerability to emerging cyberattacks. | Intrusion detection systems identifying novel malware while protecting against traditional viruses. |
| Application | Representative Benchmarks | Common Evaluation Metrics | Domain-Specific Continual Learning Challenge |
|---|---|---|---|
| Computer Vision | Split MNIST, Split CIFAR-100, Tiny ImageNet, CORe50 | Average Accuracy, Forgetting, BWT, FWT | Large class expansion, domain shifts, long task sequences |
| Healthcare | BraTS, ISIC, CAMUS, EchoNet-Dynamic, CheXpert | Dice, IoU, HD95, AUC, Sensitivity, Average Forgetting | Privacy constraints, scarce annotations, clinical reliability |
| Robotics | Meta-World, RoboSuite, ManiSkill, Habitat | Task Success Rate, Cumulative Reward, Adaptation Speed | Embodied interaction, non-stationary environments, safety |
| Natural Language Processing | CLINC150, AG News, Amazon Reviews, SuperGLUE | Accuracy, Macro-F1, BLEU, ROUGE, Perplexity | Vocabulary expansion, instruction drift, long-context retention |
| Autonomous Driving | BDD100K, nuScenes, Waymo Open Dataset | mAP, mIoU, Recall, Adaptation Accuracy | Weather variation, continual object discovery, safety-critical deployment |
| Challenge | Description | Impact | Examples | Potential Solutions |
|---|---|---|---|---|
| Catastrophic Forgetting | Overwriting of previous knowledge when learning new tasks. | Loss of performance on earlier tasks, limiting multi-task applications. | A model trained on new object classes forgets previously learned ones. | Replay methods, regularization techniques (e.g., EWC, SI), parameter isolation (e.g., PNNs, PackNet). |
| Scalability to Real-World Tasks | Difficulty in handling diverse, undefined, and open-ended tasks found in real-world environments. | Limits practical applications, especially in dynamic or multi- domain environments. | A robot operating in a dynamic home environment fails to generalize across diverse tasks. | Dynamic architectures (e.g., expandable networks), meta-learning, unsupervised task detection. |
| Memory Constraints | Storing data from previous tasks is often infeasible for large-scale or resource-limited applications. | Limits model ability to effectively retain and replay past information. | Replay-based methods requiring storage of vast datasets for continual adaptation. | Efficient memory management techniques, synthetic replay using generative models, data pruning. |
| Computational Overhead | Increased computational demands for training and inference due to replay, regularization, or parameter isolation techniques. | Hinders real-time applications on edge devices or systems with limited resources. | On-device CL in IoT systems is slowed by high computational requirements. | Lightweight models, parameter optimization, pruning, and efficient task-specific parameter allocation. |
| Bias Amplification | Sequential learning may reinforce biases present in earlier data or tasks. | Skewed model behavior, disproportionately affecting certain demographic groups. | A financial model favoring certain demographics due to biased historical data. | Fairness-aware training, regular bias audits, diversity-focused data augmentation. |
| Transparency and Explainability | Models evolving continuously can become opaque, making their decision-making hard to interpret. | Erodes trust, particularly in sensitive applications like healthcare or finance. | Difficulty auditing a continually adapting medical diagnostic system. | Explainability frameworks, interpretable architecture designs, and model debugging tools. |
| Privacy Concerns | Replay-based methods storing or processing user data may violate privacy regulations. | Non-compliance with privacy laws (e.g., GDPR, HIPAA), leading to legal and ethical implications. | Retaining user data for replay in recommendation systems could breach user consent. | Privacy-preserving methods like federated learning, data anonymization, and synthetic data generation. |
| Unintended Consequences | Autonomous learning systems may exhibit behaviors or decisions not aligned with human intentions or societal norms. | Potential safety risks, ethical conflicts, or misaligned system behavior in real-world scenarios. | A self-learning robot adopts unsafe behaviors while optimizing a task autonomously. | Strict behavioral constraints, ethical guidelines for autonomous systems, and robust oversight mechanisms during model deployment. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ullah, Z.; Hong, M.; Kim, J. Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions. Mathematics 2026, 14, 2774. https://doi.org/10.3390/math14152774
Ullah Z, Hong M, Kim J. Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions. Mathematics. 2026; 14(15):2774. https://doi.org/10.3390/math14152774
Chicago/Turabian StyleUllah, Zahid, Minki Hong, and Jihie Kim. 2026. "Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions" Mathematics 14, no. 15: 2774. https://doi.org/10.3390/math14152774
APA StyleUllah, Z., Hong, M., & Kim, J. (2026). Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions. Mathematics, 14(15), 2774. https://doi.org/10.3390/math14152774

