Next Article in Journal
Multi-Agent Reinforcement Learning Game Model for Market Economic Equilibrium Regulation
Previous Article in Journal
Artificial Intelligence Exposure, Perceived Job Replaceability, and Perceived Income Change: The Moderating Role of Task Codifiability—Evidence from the China General Social Survey
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions

Department of Computer Science and Artificial Intelligence, Dongguk University, Seoul 04620, Republic of Korea
*
Author to whom correspondence should be addressed.
Mathematics 2026, 14(15), 2774; https://doi.org/10.3390/math14152774
Submission received: 16 May 2026 / Revised: 17 July 2026 / Accepted: 24 July 2026 / Published: 3 August 2026

Abstract

Continual learning (CL) aims to develop intelligent systems capable of learning continuously from sequential data while retaining previously acquired knowledge. As AI systems are increasingly deployed in dynamic real-world environments, CL has become essential for enabling long-term adaptation without catastrophic forgetting. This review provides a structured overview of major CL paradigms, including task-incremental, domain-incremental, class-incremental, online, multimodal, and federated CL. We examine the theoretical foundations of CL, particularly the stability–plasticity dilemma, catastrophic forgetting, transfer dynamics, and representation learning. In addition, we analyze major methodological categories, including regularization-based, replay-based, architecture-based, optimization-based, representation-learning, and parameter-efficient approaches. Recent developments involving transformers, prompt learning, foundation models, and multimodal adaptation are also discussed as emerging directions in modern CL research. Furthermore, this review highlights important issues related to benchmark fragmentation, evaluation inconsistency, memory constraints, computational efficiency, scalability, and privacy-aware learning. We also summarize key application domains, including computer vision, natural language processing, robotics, healthcare, and medical imaging. Finally, we identify open research challenges and future directions toward scalable, reliable, and deployment-oriented lifelong learning systems capable of operating effectively in continuously evolving environments.

1. Introduction

Artificial intelligence (AI) enables computational systems to perform tasks that typically require human intelligence, including reasoning, learning, perception, language understanding, and decision-making [1,2,3]. AI systems can analyze data, recognize patterns, generate language, and support decision-making across a wide range of domains [4,5]. However, many conventional machine learning (ML) models are trained on fixed datasets and deployed under the assumption that the data distribution will remain stable [6]. This assumption is often unrealistic in dynamic environments, where new classes may emerge, data distributions may shift, and previously learned knowledge may become insufficient or outdated [7].
Continual learning (CL), also referred to as incremental or lifelong learning, addresses this limitation by enabling models to learn from sequential data while retaining previously acquired knowledge [8,9]. This capability is essential for AI systems operating in real-world environments, where models must adapt to changing conditions, incorporate new tasks, and maintain reliable performance over time [10,11]. Without effective CL mechanisms, neural networks are vulnerable to catastrophic forgetting, in which learning new information causes performance degradation on earlier tasks [10,12]. Therefore, the central goal of CL is to balance plasticity, the ability to acquire new knowledge, with stability, the ability to preserve prior knowledge.
The need for CL has become more urgent with the rapid growth of continuous data streams from sensor networks, financial systems, social media platforms, Internet of Things devices, and other real-time sources [13,14]. Unlike traditional batch learning, these settings require models that can update incrementally as new data arrive, without complete retraining from scratch [9,15]. For example, an autonomous driving model trained under clear weather may need to adapt to rain, snow, or fog, while a fraud detection system must continuously respond to new fraudulent behaviors. CL provides a framework for such adaptation by allowing models to refine their parameters in response to evolving data distributions while preserving useful prior knowledge [10,16].
Beyond adaptability, CL also supports computational efficiency and sustainability. Retraining large models whenever new data become available is often costly, especially in resource-constrained environments such as mobile devices, embedded systems, robotics, and edge AI platforms [9]. By updating models incrementally, CL can reduce computational overhead and support long-term deployment in dynamic environments. These properties make CL relevant to a wide range of applications, including autonomous systems, healthcare, personalized recommendation, NLP, robotics, and adaptive decision-support systems.
Despite substantial progress, CL research remains fragmented. Existing studies differ in learning scenarios, task construction, evaluation protocols, memory budgets, and assumptions about task identity. Moreover, recent advances in transformers, prompt learning, PEFT, foundation models, diffusion models, multimodal learning, and privacy-preserving learning have introduced new opportunities and challenges that are not fully synthesized in earlier surveys. This review therefore aims to provide an updated and structured overview of CL. To improve the transparency and reproducibility of the literature selection process, a structured search strategy guided by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 reporting framework was adopted for identifying, screening, and selecting the publications included in this review.
The main contributions of this review are as follows. First, it presents a structured organization of major CL settings, including task-incremental, domain-incremental, class-incremental, online, data-incremental, multimodal, and federated CL. Second, it summarizes major methodological categories, including regularization, replay, architecture-based, optimization-based, representation-learning, prompt-based, and parameter-efficient approaches. Third, it discusses evaluation protocols, benchmark fragmentation, memory constraints, and reproducibility issues that affect fair comparison across CL methods. Finally, it identifies emerging research directions related to foundation-model-based CL, multimodal continual adaptation, privacy-preserving CL, and real-world deployment.

1.1. Literature Search Methodology

This review follows the PRISMA 2020 reporting framework to ensure a transparent, systematic, and reproducible process for literature identification, screening, eligibility assessment, and study selection. The completed PRISMA 2020 Checklist is provided as Supplementary Material S1 and follows the PRISMA 2020 Statement [17]. The literature search was conducted using predefined search strategies, inclusion and exclusion criteria, and multiple bibliographic databases to provide comprehensive coverage of recent advances in CL while minimizing selection bias. Although no quantitative meta-analysis was performed because of the methodological heterogeneity of the included studies, the review employs a structured qualitative evidence synthesis to comprehensively analyze current developments, evaluation protocols, and future research directions in continual learning.
This systematic review was retrospectively registered with the Open Science Framework (OSF) during the peer-review process in response to the journal editor’s request to enhance transparency and reproducibility. The review protocol is available through the OSF at https://osf.io/kznrc/ (accessed on 20 July 2026). The registration is currently under embargo while the manuscript is under peer review.
The literature search was conducted using several major scientific databases, including IEEE Xplore, ACM Digital Library, SpringerLink, ScienceDirect, Web of Science, Scopus, and Google Scholar. These databases were selected because they collectively provide comprehensive coverage of high-quality journals and conference proceedings in AI, ML, computer vision, NLP, and related research fields. The search primarily covered publications from 2017 to January 2026, corresponding to the period during which CL experienced rapid development with the emergence of deep neural networks, vision transformers (ViTs), foundation models, prompt learning, PEFT, and multimodal learning. The literature search was conducted between 03 January 2026 and 30 January 2026. The final search was completed on 30 January 2026, and only publications available up to this date were considered for inclusion.
The search strategy employed combinations of keywords related to CL and its major research directions. Representative search expressions included “continual learning,” “lifelong learning,” “incremental learning,” “catastrophic forgetting,” “experience replay,” “knowledge distillation,” “regularization,” “vision transformer,” “foundation model,” “prompt learning,” “Parameter-Efficient Fine-Tuning,” “multimodal continual learning,” and “federated continual learning.” Boolean operators (AND/OR) were employed to combine these keywords according to the syntax supported by each digital library. Representative database-specific search queries are summarized in Table 1. The PRISMA 2020 Checklist is provided as Supplementary Material.

1.1.1. Database-Specific Search Strategy

Following the initial search, duplicate records were removed before title and abstract screening. Full-text articles were subsequently evaluated for relevance based on predefined inclusion and exclusion criteria. Studies were included if they presented novel CL algorithms, theoretical developments, benchmark datasets, evaluation methodologies, comprehensive surveys, or significant applications of CL across computer vision, NLP, robotics, healthcare, medical imaging, or related domains. Both journal articles and peer-reviewed conference papers published in leading venues were considered. Highly influential preprints were included only when they represented important emerging research directions that had not yet appeared in peer-reviewed publications. Figure 1 illustrates the complete literature identification, screening, eligibility assessment, and study selection process following the PRISMA 2020 reporting framework.
Studies were excluded if they focused exclusively on transfer learning, conventional domain adaptation, multi-task learning without continual adaptation, hardware implementation without methodological contributions, duplicate publications, or articles lacking sufficient technical detail. Non-English publications and studies whose primary contributions were outside the scope of CL were also excluded. During the full-text assessment, representative examples of studies excluded during the full-text eligibility assessment are summarized in Table 2. These studies were excluded because they did not satisfy the predefined eligibility criteria despite being retrieved during the initial database search.
After the screening and eligibility assessment, a total of 285 publications satisfied the inclusion criteria and were selected for detailed analysis.

1.1.2. Quality Assessment

To improve the reliability and scientific rigor of this review, all studies satisfying the predefined inclusion criteria underwent a structured qualitative assessment before being included in the final synthesis. This review follows the PRISMA 2020 reporting framework for systematic literature reviews; however, because of the substantial methodological heterogeneity among the included studies, including differences in CL scenarios, benchmark datasets, evaluation protocols, and performance metrics, a qualitative appraisal approach was adopted instead of a formal quantitative risk-of-bias assessment or meta-analysis. This approach enabled a comprehensive and systematic synthesis of the current state of continual learning research while maintaining methodological transparency and reproducibility.
Each full-text article was evaluated according to several criteria, including (i) relevance to CL, (ii) methodological novelty and technical contribution, (iii) experimental rigor and validation, (iv) clarity and reproducibility of the proposed methodology, and (v) publication quality. Preference was given to peer-reviewed journal articles and papers published in leading conferences in AI and ML, including NeurIPS, ICML, ICLR, CVPR, ICCV, ECCV, ACL, EMNLP, AAAI, IJCAI, TPAMI, IJCV, and JMLR.
Highly influential preprints available on arXiv were included only when they introduced important emerging research directions, such as foundation-model adaptation, prompt learning, parameter-efficient CL, or multimodal CL, and when no corresponding peer-reviewed publication was available at the time of the literature search. Whenever both a preprint and its subsequently published journal or conference version were available, only the peer-reviewed version was retained to avoid duplication.
The quality assessment was conducted manually by the authors through full-text examination. Studies with insufficient methodological descriptions, inadequate experimental validation, unclear evaluation protocols, or contributions outside the scope of CL were excluded during the eligibility assessment. This qualitative appraisal ensured that the final set of 285 publications represented technically sound, reproducible, and scientifically relevant contributions to modern CL research.

2. Existing Continual Learning Surveys and Remaining Gaps

Although CL has been extensively reviewed in prior surveys, most existing works either focus on foundational concepts and traditional methodologies or specialize in narrow subdomains such as class-incremental learning (CIL), online CL, or biologically inspired approaches. Earlier surveys primarily emphasized classical paradigms including replay, regularization, and architecture-based methods, with limited discussion of recent developments driven by foundation models, ViTs, prompt-based learning, and parameter-efficient adaptation techniques. Furthermore, many surveys provide descriptive summaries of methods without critically analyzing evaluation inconsistencies, benchmark fragmentation, memory–computation trade-offs, or the practical limitations affecting real-world deployment. Recent advances in multimodal learning, continual adaptation of large language models (LLMs), diffusion-based continual generation, and privacy-aware or federated CL remain insufficiently synthesized in the literature. In addition, there is still a lack of unified discussion regarding reproducibility, task-split design, replay memory constraints, and standardized evaluation protocols, all of which significantly influence reported performance and fair comparison across methods. Motivated by these gaps, this review provides an updated and structured synthesis of modern CL research, with particular emphasis on emerging trends, evaluation challenges, scalable adaptation strategies, and open deployment issues in dynamic real-world environments.
Unlike previous surveys that primarily focus on specific CL settings or methodological categories, this review makes four structural contributions. First, it introduces an expanded taxonomy of CL scenarios that unifies traditional and emerging settings, including online, multimodal, federated, and data-incremental learning. Second, it organizes modern CL methods into a unified methodological taxonomy covering regularization-, replay-, optimization-, representation-, architecture-, prompt-, and parameter-efficient learning approaches. Third, it provides a critical analysis of evaluation protocols, benchmark fragmentation, reproducibility, and computational reporting practices. Finally, it synthesizes recent advances in foundation-model-based CL while identifying open challenges and future research directions for scalable lifelong learning systems. Table 3 compares representative continual learning survey papers and highlights the scope, coverage, recent trends, and key contributions of the present review.

2.1. Continual Learning

CL is an ML paradigm where models are designed to learn continuously from a stream of data, integrating new information while retaining and utilizing previously acquired knowledge [22] as shown in Figure 2A. Unlike traditional static learning methods that require retraining from scratch with all available data, CL enables systems to adapt to evolving tasks and data distributions incrementally. One of the primary challenges is catastrophic forgetting [23,24], where learning from new data typically leads to a significant decline in the model’s ability to retain previously learned information. This challenge reflects the inherent trade-off between learning plasticity and memory stability. Too much plasticity can compromise memory retention, while excessive stability may hinder the integration of new knowledge. Rather than merely adjusting the balance between these two factors, an effective CL approach should also ensure strong generalization capabilities to handle variations both within and across tasks as shown in Figure 2B. This approach is inspired by the human ability to acquire knowledge progressively and apply it to diverse and dynamic contexts, making it essential for creating adaptive and intelligent AI systems capable of operating effectively in real-world, ever-changing environments. In recent years, numerous CL strategies have been developed to address different challenges within ML. These methods can be broadly categorized into five conceptual groups as shown in Figure 2C: (1) regularization-based approaches, which introduce penalty terms based on the previous model; (2) replay-based approaches, which aim to approximate and reconstruct past data distributions; (3) optimization-based approaches, which directly adjust the optimization process; (4) representation-based approaches, which focus on learning stable and transferable feature representations; and (5) architecture-based approaches, which design adaptable model components tailored to specific tasks.
This classification builds upon traditional taxonomies by incorporating recent advancements and offering more nuanced subcategories. We provide a detailed overview of how these techniques contribute to the goals of CL, analyzing both their theoretical underpinnings and practical implementations. Notably, these approaches often intersect; for instance, regularization and replay methods both influence gradient direction during optimization-and they can complement each other, such as by enhancing replay effectiveness through knowledge distillation from prior models. Real-world applications introduce unique challenges to CL, which can be broadly divided into scenario complexity and task-specific demands as shown in Figure 2D. Regarding scenario complexity, issues such as the absence of task identity during training or testing and the limited availability of data-sometimes arriving in small batches or even just once-pose significant hurdles. Additionally, due to the high cost and limited availability of labeled data, CL must perform well in few-shot, semi-supervised, and even unsupervised settings [25]. On the other hand, task specificity highlights that while most research has concentrated on visual classification, there is growing interest in other areas like object detection, semantic segmentation, conditional generation, reinforcement learning, NLP, and ethical considerations, each with their own distinct challenges [11]. This review outlines these domain-specific issues and explores how CL approaches have been adapted to address them.
CL, while sharing some similarities with transfer learning, multi-task learning, and online learning, is fundamentally distinct in its objectives and methodology. Each of these paradigms addresses specific challenges in ML, but their differences lie in how they approach task sequencing, knowledge transfer, and adaptation to dynamic environments, as shown in Table 4.

2.1.1. Continual Learning vs. Transfer Learning

Transfer learning focuses on leveraging knowledge from a source task or domain to improve performance on a related target task or domain. Typically, this involves training a model on a large dataset (e.g., ImageNet) and fine-tuning it for a specific application [26]. The transfer process is usually one-directional and occurs only once. In contrast, CL emphasizes the ongoing acquisition of knowledge from a sequence of tasks [22], with the critical goal of retaining prior knowledge while learning new tasks. Unlike transfer learning, CL explicitly addresses the problem of catastrophic forgetting, ensuring that performance on earlier tasks is not degraded as new tasks are introduced. While both paradigms reuse knowledge, CL operates in a dynamic and evolving environment, whereas transfer learning assumes static datasets and a single knowledge transfer event.

2.1.2. Continual Learning vs. Multi-Task Learning

Multi-task learning focuses on learning multiple tasks simultaneously by optimizing a shared representation across all tasks [27]. This approach assumes that all tasks and their data are available during training, which allows the model to generalize effectively across tasks. In contrast, CL deals with tasks arriving sequentially, where access to prior task data may be limited or unavailable [22,28]. The primary challenge in CL is to integrate new knowledge without overwriting or forgetting previously learned tasks, while multi-task learning does not encounter this issue since all tasks are trained together [29]. However, there is an overlap in the shared goal of improving performance across tasks, as CL can be viewed as a sequential extension of multi-task learning in dynamic environments.

2.1.3. Continual Learning vs. Online Learning

Online learning involves updating a model incrementally as new data arrives, typically in a single task or stationary data distribution setting. It focuses on optimizing for real-time updates and minimizing latency, often without addressing how to handle changes in task or data distribution over time [30]. CL, on the other hand, is designed to operate in non-stationary environments where new tasks or concepts emerge sequentially. Unlike online learning, CL emphasizes the retention of knowledge across tasks and adapts to evolving distributions, ensuring that performance remains robust over time. While both paradigms involve incremental updates, CL prioritizes long-term adaptability and knowledge integration.
CL distinguishes itself by addressing the unique challenges of dynamic, real-world environments where tasks and data evolve over time. While it shares elements of knowledge transfer with transfer learning, task generalization with multi-task learning, and incremental updates with online learning, its emphasis on lifelong learning without forgetting sets it apart. This makes CL a critical paradigm for building adaptive, resilient, and intelligent systems capable of operating effectively in ever-changing settings.
A fundamental challenge in CL is catastrophic forgetting as discussed in Section 4.2, a phenomenon where a model loses the ability to perform previously learned tasks when trained on new ones. This problem arises because traditional ML models are typically optimized for single-task scenarios, where the entire dataset is available at once. When these models are incrementally updated with new data, the weights and representations optimized for earlier tasks are overwritten, leading to a sharp decline in performance on prior tasks.
Addressing catastrophic forgetting is a core objective of CL. Strategies such as memory replay, parameter isolation, and regularization have been developed to mitigate this issue. These methods aim to preserve important parameters associated with past tasks, either by replaying prior data, selectively freezing certain parameters, or imposing constraints that minimize changes to previously learned representations. Despite significant progress, catastrophic forgetting remains a central obstacle in the development of robust CL systems, and overcoming it is crucial for advancing AI’s ability to function effectively in dynamic and evolving environments.

2.1.4. Scope and Contribution of This Review

The primary objective of this review paper is to provide a comprehensive and systematic exploration of CL, a rapidly emerging paradigm in AI [31,32,33] that enables models to learn and adapt continuously over time. As the need for adaptive systems grows in dynamic and evolving real-world environments, this paper aims to elucidate the foundational aspects of CL while offering critical insights into its current state and future prospects. One of the key goals is to categorize and analyze the different types of CL approaches, including task-based, class-based, and domain-based learning. By delineating these types, the review aims to provide clarity on the diverse ways in which CL addresses various challenges in dynamic data scenarios. Additionally, the paper seeks to delve into the theoretical underpinnings of CL, drawing connections to biological inspirations, mathematical frameworks, and its interplay with related fields such as multi-task and transfer learning.
Another important objective is to evaluate and synthesize the state-of-the-art methods in CL. This includes examining strategies such as memory-based replay, regularization techniques, parameter isolation, and hybrid approaches. Each method will be analyzed for its strengths, limitations, and suitability for specific types of tasks and data distributions. Alongside this, the review will identify and elaborate on key challenges faced by the field, particularly catastrophic forgetting, scalability, resource constraints, and the lack of standardized evaluation metrics. Furthermore, this paper will highlight the diverse applications of CL in areas such as robotics, healthcare, autonomous systems, and personalized recommendations. These examples will illustrate the transformative potential of CL in creating intelligent systems capable of adapting to changing conditions. Finally, the review will propose future directions for research, emphasizing opportunities to address existing limitations, integrate CL with complementary paradigms, and explore ethical and practical considerations in deploying such systems. By achieving these objectives, this paper aspires to serve as a comprehensive resource for researchers, practitioners, and policymakers, fostering innovation and collaboration in the advancement of CL.
As shown in Table 3, while some earlier surveys provide broad overviews of CL [34,35], they do not reflect the significant advancements made in recent years. Conversely, more recent surveys often focus on specific aspects of CL, such as its biological foundations [36,37], tailored approaches for visual classification [38,39], or particular applications in NLP [40,41] and reinforcement learning [42]. This paper provides an updated synthesis focused on modern CL trends, evaluation inconsistencies, foundation models, and open deployment challenges. Building on this foundation, we offer a detailed exploration of the field, discussing emerging trends, interdisciplinary prospects, methodologies, and open challenges.
Although numerous surveys have reviewed CL from theoretical, methodological, or application-oriented perspectives, most focus either on classical CL techniques or on specific subfields such as CIL, online CL, or biologically inspired approaches. Recent advances involving foundation models, prompt learning, PEFT, multimodal learning, diffusion models, and LLMs have substantially changed the CL landscape, yet these developments remain fragmented across the literature.
This review distinguishes itself from previous surveys in several important ways. First, we present an expanded taxonomy that unifies both traditional and emerging CL scenarios, including task-, domain-, class-, data-incremental, online, multimodal, and federated CL within a single conceptual framework. Second, we provide a unified methodological taxonomy covering regularization-, replay-, optimization-, representation-, architecture-, prompt-, and parameter-efficient approaches, enabling readers to understand the relationships among modern CL strategies.
Third, unlike earlier reviews that primarily summarize algorithms, we critically examine evaluation methodologies, benchmark fragmentation, reproducibility, computational efficiency, memory constraints, and standardized reporting practices, highlighting how inconsistent evaluation protocols often lead to misleading comparisons between methods. Fourth, we comprehensively synthesize recent developments in foundation-model adaptation, ViTs, prompt-based CL, PEFT, multimodal CL, diffusion models, and LLMs, topics that have rapidly emerged but have not been comprehensively integrated into previous surveys.
Finally, beyond summarizing existing research, this review identifies key open challenges and future research directions, including continual reasoning, continual alignment, continual reinforcement learning, lifelong multimodal agents, embodied AI, and safe and trustworthy CL, thereby providing a forward-looking perspective for next-generation CL systems.

2.2. Setup

In this section, we begin by outlining the fundamental framework of CL. Following that, we discuss common scenarios and other emerging paradigms in CL.

2.3. Basic Formulation of Continual Learning

CL focuses on adapting to evolving data distributions, where training samples from different distributions are introduced sequentially. A CL model, parameterized by θ , must effectively learn the given task(s) with limited or no access to prior training samples, while maintaining strong performance on their corresponding test sets. Formally, a batch of training samples for a task t is represented as D t , b = X t , b , Y t , b , where X t , b , is the input data, Y t , b , is the data label, t T = { 1 , , k } is the task identity and b B t is the batch index ( T and B t representing their space, respectively). Here we define a task by its training samples D t following the distribution D t : = p X t , Y t ( D t denotes the entire training set by omitting the batch index, likewise for X t and Y t ) , and assume that there is no difference in distribution between training and testing. Under realistic constraints, the data label Y t and the task identity t might not always be available. In CL, the training samples of each task can arrive incrementally in batches (i.e., { { D t , b } b B t } t T ) .
If not explicitly stated, it is generally assumed that each task has enough labeled training data available, which aligns with the principles of supervised CL. Based on the given X t and Y t in each D t , CL has evolved to encompass a wide range of learning scenarios, including zero-shot learning [43,44], few-shot learning [45], semi-supervised learning [46], open-world learning (involving the identification of unknown classes followed by label incorporation) [47,48], and unsupervised or self-supervised learning [49,50].

3. Types of Continual Learning

CL encompasses various scenarios as shown in Table 5, each tailored to specific types of data streams, tasks, and requirements. These scenarios, or settings, define how models interact with new information and retain prior knowledge. The three primary types of CL are task-based learning, class-based learning, and domain-based learning. The taxonomy illustrated in Figure 3 follows widely adopted continual learning scenarios described by Van de Ven et al. [18] and Wang et al. [11], while incorporating emerging settings such as open-world CL and task-agnostic learning.

3.1. Task-Incremental Learning

Task-incremental learning (TIL) is a paradigm within CL where models are trained sequentially on a series of distinct tasks [51,52,53]. A defining characteristic of TIL is that the identity of the task is explicitly provided during both training and inference, enabling the model to tailor its predictions based on the known task. This approach simplifies the problem of learning incrementally by focusing on task-specific outputs while preserving performance across previously encountered tasks. Table 6 summarizes the characteristics of TIL.
In TIL, each task is typically associated with a unique output head or module within the model. For example, in a classification scenario, the model might have separate output layers for each task, and the task identity directs the model to use the corresponding output layer during inference. This separation helps mitigate catastrophic forgetting, as the updates to the model parameters for one task are less likely to interfere with those of other tasks. However, this approach assumes that task boundaries are clearly defined, and the system knows which task it is addressing at any given time.
One of the primary benefits of TIL is its robustness to forgetting. Since the model is explicitly informed about the task identity, it does not need to generalize across tasks. This allows the use of task-specific components, such as output heads, which can independently optimize for the unique characteristics of each task [54]. Additionally, the modular nature of TIL enables efficient learning and evaluation of tasks without requiring retraining on the entire dataset. TIL is particularly advantageous in scenarios where task boundaries are clear and well-defined, such as robotics, where different tasks like object manipulation, navigation, and gesture recognition may need to be learned sequentially. By isolating these tasks, the model can adapt to new functionalities without compromising existing ones.

3.1.1. Applications of Task-Incremental Learning

TIL is commonly applied in domains where tasks are naturally discrete and identifiable. For example, in healthcare, models are trained to analyze different types of medical imaging data (e.g., X-rays, MRIs, CT scans) and can use TIL to maintain task-specific expertise [55]. Also, robots can learn individual tasks, such as grasping objects, path planning, and object categorization, without interference between tasks. Language models can incrementally learn new languages or domains while retaining the ability to process previously learned ones. By providing explicit task identities, TIL offers a structured approach to CL, making it a suitable choice for scenarios with clear task demarcation. However, as the need for more flexible and generalized models grows, hybrid approaches that incorporate elements of TIL while addressing its limitations are gaining traction.

3.1.2. Examples of Task-Incremental Learning

Sequential Image Classification Across Datasets: A common example of TIL involves training a model to classify data from different image datasets sequentially [54]. For instance, a model could first learn to classify handwritten digits in the MNIST dataset. Once this task is complete, the model may be trained to classify natural images in the CIFAR-10 dataset. In this scenario, TIL ensures that the model retains its ability to classify handwritten digits while learning to classify natural images. During inference, the task identity (e.g., “MNIST” or “CIFAR-10”) directs the model to use the relevant task-specific output head, ensuring accurate predictions for each dataset. This setup is particularly useful in academic and industrial settings where models need to be updated with new data or tasks without retraining from scratch.
Robotics: Learning Discrete Tasks: In robotics, TIL can be applied to teach a robot distinct skills sequentially [56]. For example, a robot may first learn to recognize and manipulate specific objects, followed by learning to navigate different environments. Each task is distinct, with its own set of parameters and objectives. The task identity ensures that the robot applies the appropriate skill set during operation, avoiding confusion between object manipulation and navigation.
Healthcare: Multimodal Diagnosis Systems: TIL is particularly beneficial in healthcare applications where systems need to process distinct datasets for different diagnostic tasks [55]. For instance, a diagnostic model could sequentially learn to analyze X-ray images for lung diseases, MRI scans for brain abnormalities, and CT scans for cardiovascular conditions. With TIL, the model can retain its diagnostic capabilities for earlier tasks while learning new ones, ensuring that expertise in analyzing X-rays is not lost when training on MRI or CT data.
In summary, each task is treated independently, reducing interference and preserving performance on previous tasks. In addition, models can optimize parameters for each task individually, enhancing performance. TIL is well suited for domains with clearly defined and non-overlapping tasks. By applying TIL in these examples, systems can incrementally learn new tasks while retaining the performance and accuracy of previously learned ones. However, the reliance on task identity highlights one of the key limitations of TIL; it is not suited for environments where tasks are ambiguous or where the task identity is unavailable during inference. Addressing these limitations requires hybrid or alternative approaches in CL.

3.1.3. Challenges

Despite its advantages, TIL has certain limitations. The reliance on task identity during inference limits its applicability in scenarios where the task is unknown or ambiguous. For instance, in real-world environments where tasks may blend or overlap, the assumption of clear task boundaries may not hold. Furthermore, as the number of tasks increases, the need for task-specific components can lead to scalability issues, including memory and computational overhead. Another challenge is the lack of knowledge transfer between tasks. Since TIL separates tasks into distinct components, it often misses opportunities to share learned representations, which could enhance performance on new tasks. Research efforts are exploring ways to balance task isolation with shared representations to overcome this limitation [55].
The primary challenge in TIL is not merely to prevent catastrophic forgetting but to develop effective methods for sharing learned representations across tasks. This includes optimizing the balance between performance and computational efficiency, as well as leveraging knowledge from one task to enhance performance on others-achieving positive forward or even backward transfer between tasks [57,58]. These remain unresolved challenges. Real-world examples of TIL include learning to play various sports or musical instruments, where it is generally clear which specific sport or instrument is being practiced.

3.2. Domain-Incremental Learning

Domain-incremental learning (DIL) is a CL paradigm designed to handle scenarios where the model encounters new data distributions from different domains over time. Unlike TIL, the underlying task in DIL remains consistent, but the data characteristics, such as the input distribution, environment, or context, change [59,60]. This requires the model to adapt to new domains without forgetting the knowledge learned from previous domains, making it particularly suitable for real-world applications where environmental or contextual variations are frequent. In DIL, the goal is to achieve domain adaptation while maintaining performance on previously seen domains. For example, a model trained to recognize objects may initially be trained on images captured in sunny weather and later exposed to images taken in rainy or snowy conditions. The task of object recognition remains unchanged, but the input data’s domain shifts due to variations in lighting, background, or environmental factors. Table 7 summarizes the characteristics of DIL.
Unlike TIL, the task identity is typically unavailable during inference in DIL. This means that the model must generalize across domains without explicit guidance about which domain the input belongs to. Achieving this requires robust learning strategies that can minimize domain-specific biases while preserving domain-agnostic features [61,62]. DIL is particularly challenging due to catastrophic forgetting, as updates to accommodate new domains can overwrite knowledge about previous domains. In DIL, preventing forgetting “by design” is not feasible, making the mitigation of catastrophic forgetting a critical and unresolved challenge. Examples of this scenario include incrementally learning to recognize objects under varying lighting conditions [63] (e.g., indoors versus outdoors) or adapting to driving in different weather conditions [60]. Strategies to address this issue often involve regularization techniques, memory-based approaches, or domain-specific adaptation mechanisms.

3.2.1. Examples of Domain-Incremental Learning

Object Recognition Across Weather Conditions: One of the most illustrative examples of DIL is in object recognition systems deployed in autonomous vehicles. A model trained to recognize traffic signs in clear weather may later need to adapt to recognizing the same signs in foggy, rainy, or snowy conditions. While the core task of traffic sign recognition remains unchanged, the model must accommodate shifts in the data distribution caused by weather changes. DIL enables the model to adapt to these conditions while retaining its ability to perform well in previously encountered weather scenarios.
Medical Imaging: Cross-Domain Adaptation: In healthcare, DIL can be applied to diagnostic systems that encounter data from different imaging devices or medical institutions [64]. For instance, a model trained to detect tumors in CT scans from one hospital may need to analyze scans from another hospital, where variations in imaging protocols, equipment, and patient demographics create domain shifts. Here, DIL ensures that the model adapts to the new domain while preserving its diagnostic capabilities for data from the original domain.
NLP: Sentiment Analysis Across Domains: In NLP, DIL can address challenges such as performing sentiment analysis on text from different sources [65]. For example, a model trained on movie reviews may need to adapt to analyzing customer reviews for products or services. While the task-sentiment classification remains the same, the language style, vocabulary, and context vary across domains. DIL allows the model to generalize across these domains without sacrificing performance on the original dataset. The main advantages in these examples are:
Consistent task objective: The core task remains the same, simplifying the learning objective compared to TIL.
Robustness to domain shifts: DIL enables models to handle variations in data distributions effectively, making them suitable for dynamic environments.
Scalability: By focusing on domain adaptation, DIL can be scaled to handle diverse contexts within a single task framework.

3.2.2. Challenges

While DIL offers a structured approach to handling domain shifts, it faces significant challenges, such as balancing domain-specific adaptation with generalization and ensuring efficiency in memory and computational resources. Future research may focus on hybrid methods that combine domain adaptation with task transfer, as well as techniques to reduce catastrophic forgetting without excessive reliance on domain-specific components. By applying DIL, systems can achieve greater flexibility and adaptability, enabling them to operate effectively in diverse and evolving environments while retaining knowledge from past domains.

3.3. Class-Incremental Learning

CIL is a demanding paradigm in CL where a model must sequentially learn new classes while retaining knowledge of previously learned ones [45,66]. Unlike TIL, the task identity is not provided during inference, making CIL particularly challenging as the model must classify inputs across all learned classes without prior information about the context or task [62,67,68]. This setting closely mirrors real-world scenarios, where systems are often required to expand their knowledge incrementally without forgetting prior capabilities.
Table 8 presents the overview of CIL. In CIL, the goal is to extend the model’s classification abilities incrementally by introducing new object categories or class labels over time [69]. For example, a model trained to recognize animals in a dataset containing cats and dogs may later need to classify additional animals, such as horses and birds. Unlike other CL paradigms, the model is evaluated across all classes (e.g., cats, dogs, horses, birds) simultaneously, requiring it to integrate new knowledge without degrading its performance on earlier classes. The lack of task identity during inference adds complexity, as the model must generalize across a growing set of classes without explicit guidance. This necessitates robust strategies to address catastrophic forgetting, where new learning overwrites the knowledge of earlier classes, and class imbalance, as new classes are often presented in smaller batches than the original dataset [70,71].
To mitigate these challenges, various techniques are employed, such as memory-based replay (storing examples from earlier classes) [72], KD (transferring knowledge from previous versions of the model) [73] and dynamic architectures (expanding model capacity as new classes are introduced) [54]. However, achieving high performance in CIL remains an open research problem due to the trade-offs between memory usage, computational cost, and knowledge retention.

3.3.1. Examples of Class-Incremental Learning

Extending Image Classification Models: A classic application of CIL is extending image classifiers with new object categories. For instance, a model initially trained to recognize household items like chairs and tables may later need to classify electronic gadgets such as laptops and smartphones. In this scenario, the model must integrate the new categories into its knowledge base while maintaining its ability to correctly classify the earlier categories. CIL enables this incremental expansion without requiring access to the entire dataset of past classes, which may be impractical due to storage or privacy constraints. In CL, particularly CIL for visual classification tasks, state-of-the-art methods primarily concentrate on image classification, often utilizing complex and large-scale datasets like ILSVRC2012 [74] and its variations. Additionally, numerous benchmarks exist for video classification [63], differing in scale and objectives.
Autonomous Driving: Incremental Object Detection: In autonomous vehicles, systems must continually learn to recognize new objects, such as novel road signs or vehicles, as they are introduced in different regions or regulations. For example, a car deployed in one country might later be used in another with entirely different road sign categories. CIL ensures that the system adapts to the new categories while retaining its ability to detect and classify previously learned objects.
As one of the earlier works, ILOD [75] introduced response distillation for old classes to prevent catastrophic forgetting in Fast R-CNN [76]. Building on this, RKT [77] further refined the approach by distilling co-occurrence relationships from selected proposals. Knowledge distillation was subsequently applied to other object detectors, including SID [78] on CenterNet [79], RILOD [80] on RetinaNet [81], ERD [82] on GFLV1 [83], CIFRCN [84], Faster ILOD [85], DMC [86], BNC [87], and IOD-ML [88] on Faster R-CNN [89], among others.
Some methods leverage unlabeled, in-the-wild data to integrate old and new models into a unified framework, addressing challenges such as non-co-occurrence (BNC [87]) and improving the stability–plasticity trade-off (DMC [86]). To minimize the adverse effects of knowledge distillation on learning plasticity, IOD-ML [88] employs meta-learning to adjust parameter gradients, achieving a balance between old and new classes.
Incremental object detection (IOD) is applicable beyond 2D images, extending to 3D images [90] and videos [91]. Additional related scenarios include incremental few-shot detection [92], where a pretrained object detector incorporates new classes with minimal annotated data, and open-world object detection [47], where the detector identifies unknown object instances and registers them upon receiving annotations.
Healthcare: Incremental Diagnosis Systems: In medical imaging, diagnostic models often need to be updated as new diseases or imaging modalities are introduced. For instance, a model trained to detect common skin conditions might need to learn to diagnose rare disorders over time. Using CIL, the model can integrate these new categories without losing its diagnostic accuracy for previously learned conditions [93].

3.3.2. Challenges in Class-Incremental Learning

Catastrophic Forgetting: Without access to prior task data, models are prone to forgetting earlier classes as new ones are introduced.
Scalability: As the number of classes grows, managing model capacity and computational efficiency becomes increasingly complex.
Class imbalance: New classes are often introduced with fewer examples, leading to an imbalance that can skew model performance.

3.4. Data-Incremental Learning

Data-incremental learning is a CL scenario where new data instances arrive over time, potentially from existing or new classes, without explicit task boundaries. Unlike CIL, where new classes are introduced sequentially and the model’s output space expands accordingly, data-incremental learning focuses on the ongoing arrival of data-sometimes from previously seen classes, sometimes from novel ones-mirroring the unpredictable and non-stationary nature of real-world data streams. Table 9 depicts the overview of data-incremental learning.
In data-incremental learning, the model must adapt continuously, updating its knowledge as each new data chunk or instance is observed, while maintaining performance on previously encountered information. There are no clear demarcations between tasks or phases; instead, the learning process is fluid, with the model encountering a mix of familiar and unfamiliar data points at any time. This setting is particularly relevant for applications such as autonomous vehicles, online recommendation systems, and adaptive control systems, where data flows in constantly and the system must respond in real time to both recurring and novel patterns.
A key challenge in data-incremental learning is ensuring that the model does not forget previously learned information-catastrophic forgetting, while efficiently integrating new data, especially when the underlying data distribution changes or when new classes appear unexpectedly. The absence of explicit task boundaries and the mixture of old and new data make data-incremental learning a demanding but highly practical and realistic setting for CL research and applications. The key features of data-incremental learning are:
  • Incremental data arrival: The model receives data sequentially, one instance or batch at a time, without knowledge of whether the data introduces new classes or extends existing ones.
  • No explicit task boundaries: Unlike task-based scenarios, data-incremental learning does not provide information about task transitions, requiring the model to infer patterns and adjust its learning dynamically.
  • Challenges of catastrophic forgetting: As new data arrives, the model’s parameters may be updated in ways that overwrite knowledge of previously learned classes, leading to catastrophic forgetting.
  • Adaptation and generalization: The model must generalize well to new instances and classes while preserving accuracy on old ones, requiring a balance between plasticity (learning new data) and stability (retaining old knowledge).

3.4.1. Examples of Data-Incremental Learning

Dynamic Object Classification: In real-world object recognition systems, such as those used in smart home devices, data from cameras and sensors often arrives incrementally. For example, a system might first be trained to recognize common household objects like chairs and tables but later encounter new objects like plants or appliances. The model must adapt to recognize these new categories without losing its ability to classify previously learned objects.
EVolving Customer Preferences in Recommendation Systems: Recommendation systems often deal with continuously changing user preferences and new items being added to the catalog. For instance, a music recommendation system may encounter new genres or artists over time while users’ preferences also evolve. The system must integrate these new data points dynamically to provide accurate recommendations without retraining from scratch.
Healthcare: Continuous Data From Medical Devices: In healthcare, wearable devices and monitoring systems generate continuous streams of data. For instance, a model might initially learn to detect anomalies in heart rate data but later need to incorporate additional signals such as oxygen levels or blood pressure. Data-incremental learning enables the system to integrate this evolving information while preserving its diagnostic accuracy for previously monitored parameters.

3.4.2. Challenges in Data-Incremental Learning

  • Unstructured data streams: The lack of clear task boundaries increases the difficulty of organizing and processing data effectively.
  • Memory constraints: Retaining past data or features for replay becomes resource-intensive as the volume of data grows.
  • Class imbalance: Incrementally arriving data may introduce imbalanced class distributions, skewing the model’s performance.
To handle these challenges, several techniques are employed, including:
  • Replay mechanisms: Retaining a subset of past data or using generative models to recreate previous data for rehearsal as discussed in Section 6.2 in detail.
  • Dynamic networks: Expanding model capacity incrementally to accommodate new data without overwriting existing knowledge.
  • Regularization methods: Penalizing changes to parameters critical for previously learned data to mitigate forgetting as discussed in Section 6.1.

3.5. Other Emerging Paradigms in Continual Learning

While traditional paradigms like task-incremental, class-incremental, and DIL dominate the field, emerging paradigms are gaining traction as researchers address increasingly complex and nuanced challenges [94]. These paradigms include few-shot CL, unsupervised CL, and meta-CL, among others. Each represents an innovative approach tailored to specific needs or constraints in dynamic learning environments. Table 10 presents the overview of emerging paradigms in CL.

3.5.1. Few-Shot Continual Learning

Few-shot CL combines the principles of CL and few-shot learning [45]. It focuses on enabling models to learn new tasks or classes with very limited labeled data while retaining knowledge of previously learned tasks. This paradigm is particularly useful in scenarios where collecting extensive labeled datasets is impractical or costly, such as in medical diagnostics or rare object recognition [95,96,97]. This technique adapts effectively with minimal data; however, it avoids catastrophic forgetting.To overcome this issue, meta-learning, episodic memory, and generative replay are commonly employed to enhance the model’s ability to generalize from a few examples.

3.5.2. Unsupervised Continual Learning

Unsupervised CL eliminates the need for labeled data during training, focusing instead on discovering patterns and structure within data streams [49,98]. This paradigm is motivated by the vast amounts of unlabeled data generated in real-world environments, such as video surveillance, social media feeds, or IoT sensors [50,99]. However, without explicit labels, models must balance representation learning for new data while maintaining consistency with previously learned patterns. Therefore, self-supervised learning, clustering-based methods, and contrastive learning approaches are often used to extract meaningful features from unlabeled data.

3.5.3. Meta-Continual Learning

Meta-CL involves training models to adapt quickly to new tasks in a CL setting [100]. This paradigm leverages meta-learning principles to prepare models for future tasks by training them on a sequence of tasks, encouraging rapid adaptation and minimizing forgetting. It aims to combine the benefits of CL with the adaptability of meta-learning. The main challenge in this technique is designing training algorithms that balance fast adaptation with stability across tasks. Gradient-based meta-learning and memory-augmented neural networks are frequently used to enhance task adaptability [101,102].
Javed and White [103] present an online aware meta-learning training strategy that updates sequential inputs online while minimizing interference, naturally producing sparse representations well suited for CL. Neuromodulated meta-Learning [104] builds on this concept by incorporating a meta-learned, context-dependent gating function to selectively activate neurons based on incremental tasks. Attentive Independent Mechanisms [105] further refines this approach by using a mixture of experts to make predictions with representations learned by online-aware meta-learning [103] or neuromodulated meta-learning [104], achieving greater sparsity at the architectural level.
Meta-learning can also complement experience replay, enhancing the utilization of both old and new training samples. For instance, Meta-Experience Replay [100] aligns their gradient directions, while incremental task-agnostic meta-learning [106] uses a meta-updating rule to balance these gradients. Look-Ahead MAML [107] combines experience replay with an online optimization of online-aware meta-learning [103] objectives, incorporating an adaptively modulated learning rate. OSAKA [108] introduces a hybrid objective focused on knowledge accumulation and rapid adaptation, achieved by meta-training for a robust initialization and integrating incremental task knowledge into it.
Meta-learning also facilitates the optimization of specialized architectures. MERLIN [109] learns a meta-distribution of model parameters for each task’s representations, enabling task-specific model sampling and ensemble inference. Similarly, Henning and Cervera [110] employ a Bayesian approach to learn task-specific posteriors from a shared meta-model. MetA Reusable Knowledge or MARK [111] maintains shared weights incrementally updated through meta-learning and selectively masked for task-specific applications. Anti-Retroactive Interference for lifelong learning or ARI [112] integrates adversarial attacks with experience replay to generate task-specific models, which are then merged using meta-training.

3.5.4. Federated Continual Learning

In federated CL, models are trained across distributed devices or nodes, with each node continually receiving new data [113]. This paradigm is particularly relevant for privacy-preserving applications, such as personalized healthcare or mobile device personalization. However, the challenges in this technique are balancing knowledge sharing across nodes without violating privacy, handling heterogeneous data distributions, and avoiding forgetting across nodes. The key techniques are decentralized learning algorithms, secure aggregation, and adaptive synchronization protocols.

3.5.5. Multi-Agent Continual Learning

Multi-agent CL explores scenarios where multiple models or agents learn and interact in a shared environment, continually adapting to new tasks or domains. This paradigm is particularly useful for collaborative robotics, multi-player gaming AI, and distributed sensor networks.
Challenges: Coordinating knowledge transfer between agents and managing inter-agent dependencies [114].
Key techniques: Communication protocols, shared memory systems, and ensemble learning approaches.
While addressing catastrophic forgetting in agents, it is also essential to reduce interference from other agents, leverage their knowledge effectively, preserve client privacy, and limit the accumulation and spread of errors. Additionally, CL algorithms should be readily adaptable to multi-agent systems. Given that multi-agent learning introduces additional computational overhead, particularly in terms of communication, and is often deployed on edge devices, CL methods should aim to be as computationally efficient as possible. For instance, Yoon et al. [114] propose a federated CL framework aimed at reducing both inter-client interference and communication overhead. In typical federated learning setups, a central server aggregates updates from multiple clients and distributes a global model. However, merging knowledge from clients trained on different data distributions can lead to catastrophic forgetting. To address this, the authors divide the model parameters into three categories: (1) a global dense base parameter set that captures shared knowledge across all clients, (2) a local base parameter set for client-specific general knowledge, and (3) a sparse task-adaptive parameter set tailored to each client’s current task. Clients selectively activate the task-adaptive parameters using attention masks.
In a related approach, an enhanced version of learning without forgetting is employed to retain learned information locally while enabling knowledge distillation from the central server [115]. Park et al. [116] further investigate the integration of rehearsal strategies into federated CL. Due to the critical importance of privacy in federated settings, they introduce variational embeddings to encode and transmit task-relevant data to the server. These embeddings are then used for server-side training, allowing the system to replay past knowledge and mitigate forgetting.

4. Theoretical Foundations of Continual Learning

CL aims to develop models that can acquire new knowledge from sequential data while preserving previously learned information. Its theoretical foundations are centered on four closely related issues: the stability–plasticity dilemma, catastrophic forgetting, forward and backward transfer, and representation learning. These concepts explain why learning continuously is difficult and why different CL methods attempt to regulate parameter updates, preserve useful representations, or reuse prior knowledge. Table 11 shows the major theoretical foundation of CL.

4.1. Stability–Plasticity Dilemma

Building upon the formulation introduced in Section 2.3, consider a model with parameters θ R | θ | trained on a sequence of k tasks. For each task t = 1 , , k , the training set is defined as
D t = X t , Y t = x t , n , y t , n n = 1 N t ,
where N t denotes the number of samples in task t. The objective is to learn from the sequence
D 1 : k : = { D 1 , , D k }
while maintaining strong performance across both previous and current tasks. Assuming conditional independence across tasks, the joint likelihood can be written as
p ( D 1 : k θ ) = t = 1 k p ( D t θ ) .
For discriminative models, the log-likelihood of task t is commonly expressed as
log p ( D t θ ) = n = 1 N t log p θ ( y t , n x t , n ) .
The main difficulty arises because, when learning a new task D k , data from previous tasks { D 1 , , D k 1 } may be unavailable or only partially accessible. The model must therefore adapt to new information while preserving knowledge learned from earlier tasks. This tension is known as the stability–plasticity dilemma. Stability refers to the ability to retain previous knowledge, whereas plasticity refers to the ability to learn new information [117,118,119,120,121]. Excessive plasticity may cause catastrophic forgetting, while excessive stability may prevent adaptation to new tasks.
As shown in Figure 4A, CL requires balancing stability and plasticity to prevent catastrophic forgetting while maintaining the ability to acquire new knowledge. Figure 4B summarizes the three representative categories of CL methods that address this challenge through replay, parameter regularization, and architectural adaptation.

Formal Definition of Catastrophic Forgetting

Although catastrophic forgetting is often described qualitatively as the degradation of previously acquired knowledge after learning new tasks, CL research has introduced several mathematical formulations that enable rigorous analysis of forgetting. Consider a CL setting consisting of T sequential tasks, where A i , j denotes the test performance (e.g., classification accuracy) on task i after the model has completed learning task j, with j i . Following the widely adopted formulation, the forgetting associated with task i after learning all tasks is defined as
F i = max k { i , , T 1 } A i , k A i , T ,
where the first term represents the highest performance achieved on task i during sequential learning, while A i , T denotes the final performance after learning all tasks. The average forgetting over the entire CL sequence is therefore computed as
F = 1 T 1 i = 1 T 1 F i .
These formulations quantify knowledge degradation independently of the absolute predictive performance and have therefore become one of the standard evaluation measures for CL algorithms. A smaller value of F indicates better preservation of previously acquired knowledge, whereas larger values imply more severe catastrophic forgetting. Unlike conventional accuracy measures, forgetting explicitly captures the sequential nature of CL and provides a direct assessment of long-term knowledge retention.
Recent theoretical studies have further investigated upper bounds on catastrophic forgetting. Bennani et al. [122] analyzed replay-based CL and showed that forgetting can be bounded under assumptions regarding task similarity, replay memory size, and hypothesis complexity. Their analysis demonstrates that increasing replay memory and reducing distributional discrepancy between consecutive tasks both contribute to reducing the upper bound on forgetting. Similarly, Doan et al. [123] studied CL from the perspective of sequential optimization and established convergence guarantees showing that appropriate regularization and replay mechanisms reduce optimization error accumulation across task sequences. These theoretical analyses provide important mathematical justification for why replay and regularization methods remain among the most successful approaches for mitigating catastrophic forgetting.
Although existing theoretical guarantees rely on assumptions such as bounded stochastic gradients, task similarity, smooth loss functions, or limited distribution shift, they represent important steps toward establishing a rigorous mathematical understanding of CL. Developing tighter forgetting bounds under realistic non-stationary environments, foundation models, and multimodal CL remains an important open research problem.

4.2. Catastrophic Forgetting

Catastrophic forgetting is the degradation of performance on previously learned tasks after training on new data [124,125]. In neural networks, this problem occurs because gradient-based optimization updates shared parameters, which may overwrite representations that were important for earlier tasks. Forgetting is especially severe when tasks share overlapping parameters but require different decision boundaries or feature representations.
Most CL methods can be interpreted as attempts to reduce destructive interference. Regularization-based methods penalize changes to parameters that are important for previous tasks. Replay-based methods preserve earlier knowledge by revisiting stored or generated samples from previous tasks. Architecture-based methods reduce interference by allocating task-specific parameters or expandable modules. Although these strategies differ technically, they share the same goal: preserving useful prior knowledge while enabling adaptation to new data.

4.3. Forward and Backward Transfer

CL is not only concerned with avoiding forgetting. A strong CL system should also support transfer across tasks. Forward transfer occurs when knowledge learned from previous tasks improves learning on future tasks. For example, representations learned for object recognition may accelerate learning of related categories. Backward transfer occurs when learning a new task improves performance on earlier tasks, indicating that the model has refined or reorganized previous knowledge in a useful way.
Positive transfer is desirable because it allows CL systems to reuse knowledge rather than treating each task independently. However, transfer can also be negative when learning one task interferes with another. Therefore, an effective CL method should maximize beneficial transfer while minimizing harmful interference.

4.4. Negative Transfer

While positive transfer enables previously acquired knowledge to accelerate or improve learning on new tasks, CL systems may also experience negative transfer, where knowledge learned from earlier tasks interferes with subsequent learning. Unlike catastrophic forgetting, which refers to the degradation of performance on previously learned tasks after learning new ones, negative transfer affects the acquisition of future tasks by biasing the optimization process toward outdated or task-specific representations.
Negative transfer commonly arises when sequential tasks exhibit weak semantic relatedness, conflicting feature distributions, or incompatible decision boundaries. In such cases, representations learned for earlier tasks may constrain the optimization landscape, causing slower convergence, reduced generalization, or lower final performance on subsequent tasks. For example, features learned for natural image classification may hinder adaptation to medical image analysis because the learned representations emphasize textures and object semantics that are less relevant for anatomical structures.
From an optimization perspective, negative transfer is often associated with conflicting gradients between tasks. Let
g i = θ L i , g j = θ L j ,
denote the gradients of two sequential tasks. When
g i g j < 0 ,
the gradients point in opposing directions, indicating that optimizing one task may increase the loss of the other. Such gradient conflicts can impede efficient knowledge accumulation and reduce CL performance.
Several recent CL methods attempt to mitigate negative transfer through gradient projection, orthogonal gradient optimization, adaptive parameter isolation, modular architectures, sparse routing, task-specific prompts, representation disentanglement, and similarity-aware replay. Foundation-model-based CL further reduces negative transfer by updating only a small subset of parameters through PEFT, thereby preserving general-purpose representations while allowing task-specific adaptation.
Consequently, modern CL aims not only to minimize catastrophic forgetting but also to maximize positive transfer while avoiding negative transfer across heterogeneous task sequences. Understanding the interaction between these competing transfer dynamics is essential for developing scalable lifelong learning systems capable of robust adaptation in real-world environments. Table 12 compares catastrophic forgetting and negative transfer.

4.5. Representation Learning

Representation learning plays a central role in CL because the quality of learned features strongly affects both forgetting and transfer. If a model learns task-specific representations that are highly dependent on one dataset or task distribution, it may struggle to generalize to future tasks. In contrast, stable and reusable representations can reduce interference and improve adaptation.
Recent CL research increasingly emphasizes task-agnostic, domain-invariant, and transferable representations. Self-supervised learning and contrastive learning are often used to learn features that remain useful across different tasks and domains. Such representations can improve forward transfer and reduce the need for large replay buffers, although they may still require additional mechanisms when task distributions are highly heterogeneous.

4.6. Gradient Interference and Optimization Geometry

A fundamental reason for catastrophic forgetting arises from gradient interference, where optimization steps that improve performance on a newly encountered task simultaneously deteriorate performance on previously learned tasks. Consider two sequential tasks, T i and T j , with corresponding loss functions L i ( θ ) and L j ( θ ) , where θ denotes the model parameters. Their gradients are given by
g i = θ L i ( θ ) , g j = θ L j ( θ ) .
The degree of compatibility between the two optimization directions can be quantified using the cosine similarity
cos ( g i , g j ) = g i g j g i g j .
When cos ( g i , g j ) > 0 , both tasks benefit from similar optimization directions, promoting positive knowledge transfer. Conversely, cos ( g i , g j ) < 0 indicates conflicting gradients, implying that optimizing the new task moves the parameters away from the optimum of the previous task, thereby increasing catastrophic forgetting. When the cosine similarity approaches zero, the tasks are approximately independent, and parameter updates have limited influence on one another. Consequently, CL algorithms increasingly seek optimization strategies that maximize gradient agreement while minimizing destructive interference.
From an optimization perspective, CL may therefore be interpreted as searching for parameter updates that jointly minimize the losses of previously learned and newly arriving tasks. Rather than optimizing only the current task objective, modern CL algorithms attempt to identify update directions that remain compatible across multiple tasks, thereby balancing stability and plasticity throughout sequential learning.
Recent theoretical studies further relate gradient interference to the geometry of the optimization landscape. As neural networks are trained sequentially, parameter updates move the solution from one local minimum to another. If the curvature of the loss surface differs substantially between consecutive tasks, gradient descent may leave the basin corresponding to previously learned solutions, causing severe forgetting. This observation motivates several optimization-based CL algorithms that explicitly constrain gradient directions, project conflicting gradients, or preserve flat minima that remain robust under subsequent parameter updates.
Another theoretical perspective is provided by the Neural Tangent Kernel (NTK), which approximates deep neural network training by linearizing the network around its initialization. Under the NTK regime, catastrophic forgetting can be interpreted as interference between task-specific kernel updates. If two sequential tasks induce highly correlated kernel representations, continual adaptation becomes easier because parameter updates remain compatible. In contrast, weakly correlated or conflicting kernel representations produce larger optimization conflicts and greater forgetting. Although practical deep networks often operate beyond the strict NTK assumptions, this perspective provides valuable theoretical insight into why representation similarity strongly influences CL performance.
These optimization-geometric viewpoints provide a unifying explanation for numerous CL strategies. Regularization-based methods attempt to restrict movement in sensitive parameter directions, replay-based approaches approximate previous optimization objectives using stored samples, optimization-based methods explicitly modify conflicting gradients, while representation-learning approaches seek latent spaces in which gradients from different tasks become more compatible. Consequently, gradient interference provides a common theoretical framework for understanding catastrophic forgetting across diverse CL methodologies.

4.7. Neuroscientific Motivation

CL is partly inspired by human and biological learning systems, which can acquire new knowledge without rapidly erasing prior memories [36,37,126]. Several biological mechanisms have influenced CL research. Synaptic consolidation preserves important memories by stabilizing relevant neural connections, inspiring methods such as Elastic Weight Consolidation (EWC), which penalizes changes to parameters important for earlier tasks [127,128,129,130]. Neurogenesis motivates expandable architectures that allocate new capacity for new tasks. Rehearsal-based memory mechanisms inspire replay methods, where previous samples or generated approximations are revisited during training.
These biological analogies should not be treated as direct implementations of human learning. Rather, they provide useful conceptual guidance for designing artificial systems that balance retention, adaptation, and efficient memory use.
In summary, the theoretical foundations of CL revolve around the need to balance stability and plasticity, reduce catastrophic forgetting, promote useful knowledge transfer, and learn reusable representations. These principles provide the basis for the major methodological families discussed in later sections, including regularization, replay, architecture-based adaptation, representation learning, and parameter-efficient continual adaptation.

4.8. Mathematical Frameworks

The theoretical understanding of CL is further supported by formal mathematical models that offer precise descriptions of its mechanisms and challenges.

4.8.1. Regularization-Based Models

These models introduce constraints during the optimization process to prevent significant changes to parameters critical for previous tasks. For instance, EWC adds a quadratic penalty to the loss function for parameters identified as important for earlier tasks, preserving stability without hindering plasticity.

4.8.2. Replay and Memory Models

Replay-based methods integrate past and current data to retain earlier knowledge while learning new tasks. For example, experience replay minimizes the combined loss from new task data and stored samples from previous tasks, balancing plasticity and stability. Synthetic replay techniques recreate data distributions from previous tasks using generative models.

4.8.3. Dynamic Architectures

Dynamic approaches adapt the model’s structure to accommodate new tasks without overwriting prior knowledge. For example, progressive neural networks add new parameters for each task while maintaining frozen connections to prior tasks, ensuring stability.

4.8.4. Bayesian Models

Bayesian frameworks incorporate uncertainty into parameter updates, balancing the importance of old and new knowledge. For example, variational CL uses Bayesian inference to estimate parameter importance and preserve previously learned distributions.

4.8.5. Information-Theoretic Models

These models quantify the trade-off between stability and plasticity using information theory: for instance, methods based on mutual information optimize how much knowledge from prior tasks is retained while maximizing adaptability to new tasks.
Overall, the theoretical foundations of CL integrate key concepts, biological inspirations, and mathematical models to address the dual challenge of retaining past knowledge and acquiring new knowledge. By leveraging insights from neuroscience, cognitive science, and formal mathematical frameworks, researchers continue to develop algorithms that strike an optimal balance between stability and plasticity, paving the way for robust and adaptable AI systems. These principles form the cornerstone of ongoing advancements in the field of CL.

4.9. Unique Challenges in Foundation-Model Continual Learning

Although foundation models have significantly improved CL through pretrained representations, prompt learning, and PEFT, they also introduce challenges that differ fundamentally from those encountered in conventional CL. Unlike traditional convolutional neural networks, foundation models contain billions of pretrained parameters whose knowledge is distributed across highly interconnected representations. Consequently, preserving previously acquired knowledge while continuously adapting to new tasks becomes considerably more complex. Table 13 shows challenges in foundation-model CL.
One important challenge is prompt interference. Prompt-based CL assumes that task-specific prompts can be incrementally optimized while the backbone model remains largely frozen. However, as the number of sequential tasks increases, newly learned prompts may gradually interfere with previously optimized prompts, resulting in degraded task representations and reduced knowledge retention. This phenomenon becomes increasingly pronounced in long CL sequences where hundreds of tasks must coexist within a limited prompt space. Developing scalable prompt management, prompt routing, and prompt composition strategies therefore remains an important open research problem.
Another emerging challenge is multimodal alignment collapse. Modern foundation models frequently learn joint representations across multiple modalities such as images, text, audio, and video. During continual adaptation, optimizing one modality may gradually distort the shared embedding space established during large-scale pretraining. As a result, semantic correspondence between modalities can deteriorate over time, reducing cross-modal retrieval accuracy, multimodal reasoning capability, and overall downstream performance. Preserving stable multimodal representations while accommodating continual domain adaptation remains largely unexplored and represents an important direction for future research.
Foundation models also present unique scalability challenges. Although PEFT techniques substantially reduce the number of trainable parameters, storing task-specific adapters, prompts, or low-rank modules for numerous sequential tasks may eventually lead to considerable storage requirements and increasingly complex inference pipelines. Furthermore, balancing parameter efficiency with positive knowledge transfer across related tasks remains an open challenge, particularly for very LLMs and vision–language foundation models.
Finally, evaluating foundation-model CL requires benchmarks that extend beyond traditional image classification. Future evaluation protocols should measure long-term prompt stability, multimodal representation preservation, instruction-following capability, reasoning consistency, computational efficiency, and robustness across extended CL sequences. Addressing these challenges will be essential for developing scalable lifelong foundation models capable of reliable deployment in dynamic real-world environments.

4.10. Recent Theoretical Developments in Continual Learning

Although CL has traditionally been studied from the perspective of catastrophic forgetting and the stability–plasticity dilemma, recent theoretical research has increasingly focused on understanding the optimization behavior, convergence properties, and generalization capabilities of CL algorithms. These theoretical developments seek to explain why certain CL strategies remain stable over long task sequences while others exhibit severe performance degradation. Rather than viewing CL solely as a memory preservation problem, recent studies formulate it as a constrained sequential optimization problem in which learning new tasks should improve the current objective while minimizing interference with previously acquired knowledge.
Most CL algorithms can be interpreted as solving a constrained optimization problem of the form
min θ t = 1 T L t ( θ ) + λ Ω ( θ ) ,
where L t denotes the empirical loss for task t, θ represents the model parameters, and Ω ( θ ) is a regularization or replay objective that preserves knowledge acquired from previous tasks. Different CL paradigms mainly differ in how this constraint is implemented. Regularization-based approaches estimate parameter importance, replay-based methods approximate previous data distributions, architecture-based methods isolate subsets of parameters, and parameter-efficient approaches constrain updates to a small fraction of trainable parameters.
Recent theoretical analyses have also investigated the convergence properties of CL algorithms. Under assumptions such as smooth objective functions, bounded stochastic gradients, and appropriately decreasing learning rates, several optimization-based CL methods can be shown to converge toward stationary solutions similarly to stochastic gradient descent. Nevertheless, CL introduces additional non-stationarity because the optimization objective changes whenever new tasks arrive. Consequently, convergence guarantees generally depend on assumptions regarding task similarity, bounded gradient interference, replay memory size, or regularization strength, making rigorous theoretical analysis considerably more challenging than conventional supervised learning.
Another important research direction concerns the optimization landscape of CL. Sequential task learning often introduces conflicting gradients between old and new tasks, producing optimization trajectories that deviate from those obtained through joint training. Recent optimization-based methods therefore attempt to reduce gradient interference by projecting conflicting gradients, balancing task-specific objectives, or identifying flatter optimization minima that improve long-term knowledge retention. These observations suggest that catastrophic forgetting is closely related not only to parameter overwriting but also to the geometric properties of the optimization landscape encountered during sequential training.
Generalization theory has likewise received increasing attention in recent CL research. Existing studies investigate how replay strategies, parameter regularization, and representation learning influence generalization across both previously learned and unseen tasks. Although comprehensive theoretical guarantees remain limited, recent analyses indicate that stable feature representations, bounded model complexity, and reduced gradient interference can improve both knowledge retention and forward transfer. Establishing tighter generalization bounds for sequential learning under evolving data distributions remains an important open research problem, particularly for foundation models, multimodal CL, and LLMs.
Despite these encouraging advances, a unified mathematical theory of CL has not yet emerged. Developing rigorous convergence guarantees, optimization analyses under non-stationary task distributions, and generalization bounds for large-scale foundation models therefore represents one of the most important directions for future theoretical research. Such developments will be essential for bridging the gap between empirical performance and mathematically grounded CL algorithms.

4.11. Worked Mathematical Example: Stability–Plasticity Trade-Off

To illustrate the stability–plasticity trade-off mathematically, consider a simple one-dimensional CL problem with two sequential tasks. Let the model be parameterized by a scalar parameter θ R . The first task has the quadratic loss
L 1 ( θ ) = ( θ 1 ) 2 .
The minimizer of the first task is obtained by setting the derivative to zero,
d L 1 d θ = 2 ( θ 1 ) = 0 ,
which gives
θ 1 * = 1 .
Now suppose that the second task arrives with the loss
L 2 ( θ ) = ( θ + 1 ) 2 .
If the model is trained only on the second task, the optimal parameter is
d L 2 d θ = 2 ( θ + 1 ) = 0 ,
θ 2 * = 1 .
This solution perfectly optimizes the second task but completely moves the parameter away from the optimum of the first task. The loss on the first task after learning the second task becomes
L 1 ( θ 2 * ) = ( 1 1 ) 2 = 4 ,
whereas the original first-task loss was
L 1 ( θ 1 * ) = 0 .
Thus, the increase in the previous-task loss is
Δ L 1 = L 1 ( θ 2 * ) L 1 ( θ 1 * ) = 4 ,
which represents catastrophic forgetting in this simplified setting.
A regularization-based CL strategy can reduce this forgetting by penalizing movement away from the previous optimum. For example, an EWC-like objective for learning the second task can be written as
J ( θ ) = L 2 ( θ ) + λ ( θ θ 1 * ) 2 ,
where λ 0 controls the strength of the stability constraint. Substituting θ 1 * = 1 gives
J ( θ ) = ( θ + 1 ) 2 + λ ( θ 1 ) 2 .
The minimizer is obtained by differentiating with respect to θ ,
d J d θ = 2 ( θ + 1 ) + 2 λ ( θ 1 ) .
Setting the derivative equal to zero yields
2 ( θ + 1 ) + 2 λ ( θ 1 ) = 0 .
Dividing by two and rearranging gives
θ + 1 + λ θ λ = 0 ,
( 1 + λ ) θ + 1 λ = 0 ,
θ λ * = λ 1 λ + 1 .
Equation (21) explicitly shows how the regularization parameter controls the stability–plasticity trade-off. When λ = 0 , the solution becomes
θ 0 * = 1 ,
which corresponds to full plasticity and complete adaptation to the second task, but also maximum forgetting of the first task. Conversely, as λ ,
lim λ θ λ * = 1 ,
which preserves the first task but prevents effective learning of the second task. For intermediate values of λ , the solution lies between the two task optima, providing a compromise between preserving old knowledge and learning new information.
This example demonstrates that CL can be viewed as a constrained optimization problem in which stability is enforced by limiting movement away from parameters important to previous tasks, while plasticity is achieved by allowing sufficient adaptation to new tasks. Although real neural networks involve high-dimensional, non-convex loss landscapes, the same principle underlies many regularization-based CL methods, including EWC, SI, and Memory-Aware Synapses.

5. The Catastrophic Forgetting Problem

Catastrophic forgetting, also known as catastrophic interference, is a significant challenge in neural networks, particularly in CL scenarios [124,131].
It refers to the dramatic loss of performance on previously learned tasks when the model is trained on new tasks. This problem arises because of the way neural networks learn and update their parameters, creating conflicts between the objectives of retaining past knowledge and learning new information. Table 14 presents the overview of catastrophic forgetting.

5.1. Why Neural Networks Forget

Neural networks are typically trained using gradient-based optimization techniques, such as stochastic gradient descent. During training, the network adjusts its weights and biases to minimize a loss function that quantifies the model’s error on the current task. These updates are global, meaning they affect all parameters of the network, regardless of their relevance to previously learned tasks. When a new task is introduced, the model updates its weights to optimize performance on the new task. However, this optimization does not explicitly account for the importance of certain weights to earlier tasks [132]. Consequently, the new task’s learning process can overwrite or “interfere” with the representations encoded in the network for prior tasks. This phenomenon is often referred to as parameter drift, where the values of critical parameters for earlier tasks shift during the learning of new tasks.

5.2. Weight Updates and Parameter Drift

The neural network’s capacity to learn and retain knowledge is directly tied to its parameters (weights and biases). For each task, certain parameters become more significant in capturing its patterns and features. For instance, Task A might heavily rely on a subset of parameters to classify images of animals. Task B, introduced later, might require modifying the same subset of parameters to classify images of vehicles. In the absence of mechanisms to preserve the importance of parameters for Task A, updates during Task B’s training overwrite the information stored in those parameters. This leads to a degradation in the network’s ability to perform Task A, which manifests as catastrophic forgetting. Parameter drift occurs because standard gradient-based optimization lacks constraints to differentiate between parameters that are crucial for previously learned tasks and those that can be safely modified for new learning. Without explicit mechanisms to mitigate this drift, the network effectively “forgets” earlier tasks as it learns new ones.

5.3. Factors Exacerbating Catastrophic Forgetting

Overlapping Representations: Neural networks often use overlapping representations for different tasks, meaning that the same subset of neurons and parameters is reused across tasks. While this enables compact and efficient learning, it also increases the likelihood of interference, as updates for one task can disrupt representations needed for another [133].
Sequential Data Access: In CL, data from previous tasks is typically unavailable during the training of new tasks due to memory constraints or privacy concerns. This sequential access to data makes it difficult for the model to revisit and reinforce earlier learning, exacerbating forgetting [133].
Lack of Task Awareness: In class-incremental and DIL, the model is not explicitly informed about the task identity during inference. This forces the network to integrate knowledge across tasks, increasing the risk of overwriting previously learned information [9].

5.4. Mitigation Strategies

Researchers have developed various strategies to address catastrophic forgetting. Table 15 presents the summary of mitigation strategies.

6. Method Taxonomy in CL

CL methods are designed to mitigate catastrophic forgetting while enabling models to acquire new knowledge from sequential tasks. Existing approaches address the stability–plasticity trade-off through different mechanisms, including constraining parameter updates, replaying past information, expanding or isolating model components, modifying optimization dynamics, learning transferable representations, and adapting large pretrained models through parameter-efficient mechanisms. These methods are not mutually exclusive; many recent approaches combine replay, regularization, distillation, representation learning, and architectural adaptation to improve performance across different CL settings.

6.1. Regularization-Based Methods

Regularization-based methods mitigate catastrophic forgetting by constraining changes to parameters or model outputs that are important for previously learned tasks [16,108,134]. These methods are attractive because they usually do not require storing large replay buffers, making them suitable for memory-limited or privacy-sensitive settings. They can be broadly divided into weight regularization and function regularization.
Weight regularization constrains changes in network parameters. A common strategy is to add a penalty term to the loss function that discourages updates to parameters estimated to be important for previous tasks. EWC uses the Fisher Information Matrix (FIM) to estimate parameter importance [12], while later variants improve scalability or approximation quality [135,136]. Function regularization, in contrast, preserves the behavior of the previous model by aligning intermediate features or output distributions. This is commonly implemented through knowledge distillation, where the previous model acts as a teacher and the current model is trained to retain its functional behavior [137]. Because previous data are often unavailable in CL, distillation may rely on current samples [138,139,140], limited old samples [141,142,143], external unlabeled data, or synthetic samples [144]. Figure 5 provides an overview of regularization-based continual learning, highlighting weight and function regularization strategies used to reduce catastrophic forgetting.

6.1.1. Elastic Weight Consolidation

EWC is inspired by synaptic consolidation and penalizes changes to parameters that are important for previous tasks [135,136,145]. It estimates parameter importance using the FIM [146]. The EWC objective is expressed as
L EWC = L current + λ i F i θ i θ i * 2 ,
where L current is the loss on the current task, θ i is the current parameter value, θ i * is the parameter value learned from previous tasks, F i is the Fisher importance estimate, and λ controls the strength of regularization.
EWC is simple and effective when task boundaries are known, but its performance can degrade as the number of tasks increases. Storing task-specific Fisher estimates may also become costly. Several extensions address these limitations. Liu et al. [145] reduce errors from diagonal FIM approximations through parameter-space rotation. Ritter et al. [135] use a Kronecker-factored block-diagonal Hessian approximation, while online EWC maintains a single running penalty instead of storing separate Fisher matrices for each task [136]. Incremental Moment Matching follows a related Bayesian view by using the posterior of previous tasks as the prior for new tasks [147].

6.1.2. Information-Geometric Interpretation of Elastic Weight Consolidation

Although EWC is commonly interpreted as a regularization-based CL method, it can also be understood from the perspective of information geometry. Rather than simply penalizing changes in model parameters, EWC attempts to preserve parameters that contain a large amount of information about previously learned tasks. This interpretation establishes a theoretical connection between CL and statistical estimation theory through the FIM, which locally characterizes the geometry of the parameter space.
Assume that a neural network with parameters θ has been optimized for a previous task, resulting in the parameter estimate θ * . The objective of learning a new task while preserving previously acquired knowledge can be formulated as
L ( θ ) = L new ( θ ) + λ 2 ( θ θ * ) T F ( θ θ * ) ,
where L new ( θ ) denotes the loss associated with the current task, λ controls the regularization strength, and F represents the FIM computed from the previous task.
The FIM is defined as
F = E ( x , y ) D θ log p ( y | x ; θ ) θ log p ( y | x ; θ ) T .
The expectation is taken over the data distribution of the previously learned task. Intuitively, each element of F measures the sensitivity of the likelihood function with respect to small perturbations of the model parameters. Parameters associated with larger Fisher values contribute more significantly to the predictive distribution and therefore contain more information about the previous task. Consequently, EWC discourages large updates along directions with high Fisher curvature while allowing greater flexibility in directions that contribute relatively little to previously acquired knowledge.
From an information-geometric perspective, the FIM defines a Riemannian metric on the statistical manifold of neural network parameters. Instead of measuring parameter displacement using the conventional Euclidean distance, EWC measures movement according to the local geometry induced by the probability distribution represented by the model. Therefore, two parameter vectors that appear close in Euclidean space may correspond to substantially different predictive distributions, whereas larger Euclidean changes along flat directions may have negligible influence on model behavior.
Under a second-order Taylor approximation of the previous task objective around the optimum θ * , the loss can be approximated as
L old ( θ ) L old ( θ * ) + 1 2 ( θ θ * ) T H ( θ θ * ) ,
where H denotes the Hessian matrix evaluated at the optimum. Under common regularity assumptions for maximum likelihood estimation, the Hessian can be approximated by the FIM,
H F .
This approximation provides the theoretical justification for replacing the computationally expensive Hessian with the FIM in EWC, substantially reducing computational complexity while preserving second-order information about parameter importance.
In practice, computing the complete FIM is computationally infeasible for modern deep neural networks because it requires storing and inverting an extremely large dense matrix. Consequently, EWC employs a diagonal approximation
F diag ( F 1 , F 2 , , F n ) ,
which assumes that parameter correlations are negligible. This simplification leads to the widely used diagonal EWC objective
L ( θ ) = L new ( θ ) + λ 2 i F i ( θ i θ i * ) 2 .
Although this approximation greatly improves computational efficiency, it introduces several important limitations. The diagonal Fisher neglects correlations among parameters, which can become substantial in highly over-parameterized neural networks. Consequently, parameter importance may be either overestimated or underestimated, particularly for deep transformer architectures and foundation models containing billions of parameters. Recent studies have therefore explored richer approximations, including block-diagonal Fisher matrices, Kronecker-factored approximations, and low-rank second-order methods, to better capture parameter dependencies while maintaining computational tractability.
This information-geometric interpretation highlights that EWC is fundamentally more than a regularization technique. It performs CL by constraining optimization within regions of the parameter manifold that preserve the predictive distribution of previously learned tasks. This viewpoint establishes a strong theoretical connection between CL, Bayesian inference, and information geometry, and has motivated numerous subsequent CL algorithms that estimate parameter importance using alternative geometric or probabilistic criteria.

6.1.3. Synaptic Intelligence and Related Methods

SI estimates parameter importance online by measuring how much each parameter contributes to reducing the training loss during learning. Unlike EWC, SI does not require computing the FIM. Instead, it accumulates the contribution of each parameter along the optimization trajectory of a task.
Let θ i ( s ) denote the value of parameter i at optimization step s, and let Δ θ i ( s ) = θ i ( s + 1 ) θ i ( s ) denote its update between two consecutive optimization steps. Let g i ( s ) = L ( s ) / θ i denote the gradient of the loss with respect to parameter θ i at step s. During training on a task, SI accumulates the contribution of parameter i as
ω i = s = 1 S g i ( s ) Δ θ i ( s ) ,
where S is the number of optimization steps for the current task. The negative sign appears because a parameter update that moves in the direction of loss reduction should contribute positively to the importance estimate.
After completing the task, the normalized importance of parameter i is computed as
Ω i = ω i θ i θ i * 2 + ϵ ,
where θ i * denotes the parameter value at the end of the previous task, θ i denotes the parameter value after learning the current task, and ϵ is a small damping constant used to avoid division by zero. The SI regularization objective for a new task is then written as
L S I = L c u r r e n t + λ i Ω i θ i θ i * 2 ,
where λ controls the strength of the regularization penalty.

6.2. Replay-Based Methods

Replay-based methods mitigate forgetting by revisiting information from previous tasks during training. They are among the most effective CL approaches, particularly in CIL. Replay can be implemented by storing raw samples, storing compressed representations, replaying features, or generating synthetic data that approximates previous task distributions. Figure 6 provides an overview of replay-based continual learning, where representative samples, generated data, or feature representations from previous tasks are replayed to preserve prior knowledge while learning new tasks.

6.2.1. Experience Replay

Experience replay stores a subset of previous samples in a memory buffer and mixes them with current task data during training. The replay objective can be written as
L Replay = α L current + ( 1 α ) L replay ,
where L current is the loss on new task data, L replay is the loss on replayed data, and α balances current and past information.
The main advantage of experience replay is that it directly preserves representative samples from previous tasks. However, its effectiveness depends strongly on memory size, sample selection, and replay scheduling. Storing raw data may also be infeasible in privacy-sensitive applications.
Early memory selection strategies include Reservoir Sampling [100,148,149], Ring Buffer [57], and mean-of-feature selection as used in iCaRL [66]. Other strategies use clustering, plane distance, or entropy-based selection [100,148]. More advanced methods select samples based on gradient diversity or optimization objectives, including GSS [134], CCBO [150], OCS [151], ASER [152], Rainbow Memory [153], and GCR [154].
Several works improve replay efficiency by compressing, augmenting, or editing memory samples. Adaptive Quantization Modules use vector-quantized compression for memory-efficient replay [108,155]. Memory replay with data compression models storage allocation using determinantal point processes [156,157]. Rainbow Memory increases diversity through augmentation [153], while Retrospective Adversarial Replay generates challenging samples near forgetting boundaries and applies MixUp [158,159]. Other approaches store auxiliary information such as dual memory statistics [160] or attention maps [161]. Memory samples can also be updated to become more representative or more challenging, as in Mnemonics [162] and Gradient-based Memory Editing [163].
Replay is also frequently combined with constrained optimization and knowledge distillation. Gradient Episodic Memory (GEM) constrains updates so that losses on stored samples do not increase [57], while Averaged GEM (A-GEM) improves efficiency by replacing task-specific constraints with a global replay constraint [164]. Meta-Experience Replay encourages gradient alignment between old and new samples [100], and later works explore task-gradient decomposition, saddle-point optimization, Pareto balancing, and selective replay [77,165,166,167,168].
In CIL, replay is often paired with distillation. iCaRL [66] and EEIL [141] combine exemplars with knowledge distillation. LUCIR improves feature consistency and reduces classifier bias [143], while BiC [169], WA [170], and SS-IL [171] address class imbalance and bias. PODNet preserves spatial representations through distillation [142], Co2L uses self-supervised distillation [172], GeoDL aligns old and new feature spaces [173], and ELI uses energy-based alignment [174]. Other methods enhance distillation through uncertainty, feature-space structure, task attention, or dynamic expansion [175,176,177,178,179,180]. Weight regularization can also be combined with replay to improve stability [181,182].
Despite its strength, experience replay may overfit to the limited stored samples [183]. LiDER addresses this by enforcing Lipschitz continuity [184], while MOCA increases representation variability to prevent feature contraction [185]. Strong simple baselines such as DER/DER++ [186], X-DER [187], and GDumb [188] show that replay design remains a critical factor in fair CL evaluation.

6.2.2. Generative Replay

Generative replay reduces the need to store raw samples by training a generative model to synthesize data from previous tasks [67,68]. Synthetic samples are combined with current task data during training, allowing the model to rehearse earlier distributions without explicit access to old data.
Generative replay is appealing for privacy-sensitive and memory-constrained settings, but it introduces additional computational cost and depends heavily on the quality of generated samples. Poor generative models may produce biased or low-diversity samples, which can weaken retention. GAN-based methods often generate high-quality samples but may suffer from label inconsistency or mode collapse [189,190]. Autoencoder-based methods offer more explicit label control but may produce less detailed samples, as seen in FearNet [191], SRM [192], CLEER [193], EEC [189], GMR [194], and Flashcards [195]. Hybrid approaches such as L-VAEGAN combine generative quality with more precise inference [196].
DGR provides a foundational generative replay framework by replaying samples from a previous generator while learning new tasks [67]. MeRGAN improves consistency through replay alignment [144]. Generative replay can also be combined with weight regularization [46,197,198], experience replay [46,199], masking and expandable architectures [190], and pretrained feature statistics [191,200,201]. Because full data generation remains expensive, feature replay has emerged as a lighter alternative. GFR replays generated features after the feature extractor [202], and BI-R replays internal representations using context-modulated feedback connections [68]. Large-scale pretraining can further stabilize feature representations for downstream CL [203]. As shown in Table 16, experience replay and generative replay exhibit distinct trade-offs with respect to memory requirements, privacy, computational complexity, and replay fidelity.

6.3. Architecture-Based Methods

Architecture-based methods reduce forgetting by modifying model structure. They allocate task-specific parameters, subnetworks, masks, or expandable modules to reduce interference between tasks. These approaches are effective when task boundaries are known because the model can activate task-specific components during training and inference.
Progressive neural networks, dynamically expandable networks, PackNet-like pruning strategies, and modular expert-based models are representative examples. Their major strength is strong knowledge preservation through parameter isolation. However, their main weakness is scalability: as the number of tasks increases, model size and computational cost may grow substantially. Therefore, architecture-based methods are best suited for task-incremental settings or applications where a moderate number of clearly defined tasks is expected.

6.4. Optimization-Based Methods

Optimization-based methods directly control gradient updates to reduce interference between old and new tasks. Rather than storing knowledge only through parameters or samples, these methods modify the optimization trajectory so that learning new tasks does not substantially harm previous tasks. GEM and A-GEM are common examples because they project or constrain gradients using replay memory [57,164]. Other methods encourage gradient alignment, reduce conflicting gradients, or balance stability and plasticity through multi-objective optimization [77,100,166].
These methods provide principled mechanisms for reducing destructive interference, but they may introduce computational overhead due to gradient storage, projection, or constraint solving. Their effectiveness also depends on the quality and representativeness of replay samples.

6.5. Representation-Learning Methods

Representation-learning methods aim to learn features that remain stable and transferable across tasks. Instead of only protecting parameters or replaying data, these approaches improve the quality of the feature space so that future tasks can be learned with less interference. Self-supervised learning, contrastive learning, feature disentanglement, and pretrained representations are increasingly used for this purpose [26,49,99].
Strong representations can improve forward transfer and reduce the need for large replay buffers. However, they are not sufficient by themselves when task distributions are highly heterogeneous or when new tasks require substantially different decision boundaries. Therefore, representation learning is often combined with replay, distillation, or regularization.

7. Foundation-Model Adaptation Through Prompt Learning and Parameter-Efficient Fine-Tuning

The remarkable success of foundation models has significantly transformed CL research. Unlike conventional CL methods that primarily focus on training task-specific neural networks from scratch, modern approaches increasingly rely on large pretrained models, including ViTs, LLMs, and multimodal foundation models, which possess rich transferable representations acquired from massive pretraining corpora. Rather than modifying all model parameters during continual adaptation, recent research has shifted toward parameter-efficient learning strategies that preserve pretrained knowledge while introducing only a small number of trainable parameters. This paradigm substantially reduces computational cost, lowers memory consumption, and mitigates catastrophic forgetting, making continual adaptation of large-scale models more practical for real-world deployment.
Among these approaches, prompt-based CL has emerged as one of the most influential research directions. Instead of updating the backbone network, prompt-based methods learn a set of task-adaptive prompts that guide the pretrained model toward new tasks while keeping the majority of model parameters frozen. Representative methods such as Learn-to-Prompt (L2P), DualPrompt, CODA-Prompt, S-Prompts, and HiDe-Prompt have demonstrated that carefully designed prompt representations can effectively balance knowledge retention and adaptation without requiring replay buffers or extensive parameter updates. These methods have become particularly attractive for continual adaptation of large ViTs because they significantly reduce both computational complexity and storage requirements compared with full fine-tuning.
Another important direction is PEFT, which adapts foundation models by introducing lightweight trainable modules while freezing most pretrained parameters. Representative PEFT techniques include adapter layers, prefix tuning, prompt tuning, and Low-Rank Adaptation (LoRA), all of which aim to minimize the number of trainable parameters while preserving the general knowledge encoded in large pretrained models. More recent approaches further improve scalability by dynamically selecting or composing task-specific adaptation modules, allowing CL systems to support long task sequences with substantially lower computational overhead than conventional fine-tuning.
Recent studies have also explored replay-free adaptation strategies specifically designed for foundation models. Instead of storing previous samples or explicitly estimating parameter importance, these approaches exploit pretrained feature representations together with efficient classification or adaptation mechanisms to achieve competitive CL performance while satisfying strict memory and privacy constraints. Such methods are particularly appealing in applications where replay is prohibited because of privacy regulations or storage limitations.
Although prompt learning and parameter-efficient adaptation have significantly advanced CL for foundation models, several important challenges remain. Prompt interference may accumulate as the number of sequential tasks increases, reducing the effectiveness of prompt selection and retrieval. Similarly, continual addition of adapters or low-rank modules may increase inference latency and long-term storage requirements. Furthermore, adapting LLMs introduces additional challenges, including continual instruction tuning, reasoning preservation, alignment maintenance, and knowledge updating, which are substantially different from traditional class-incremental visual recognition problems. Consequently, recent research has increasingly focused on scalable prompt optimization, adaptive parameter-efficient learning, and continual instruction tuning to enable lifelong adaptation of increasingly capable foundation models.

7.1. Parameter-Efficient and Prompt-Based Continual Learning

With the increasing use of large pretrained models, parameter-efficient and prompt-based CL methods have become important. Instead of updating the entire model, these methods adapt only a small number of parameters, such as prompts, adapters, prefixes, or low-rank modules. This reduces computational cost and limits interference with pretrained representations.
Prompt-based methods such as L2P, DualPrompt, and related approaches use learnable prompts to guide task adaptation while keeping most backbone parameters fixed [204,205,206]. Parameter-efficient transfer learning methods from NLP also provide useful tools for CL because they allow large models to incorporate new knowledge without full fine-tuning. These approaches are especially promising for foundation models, multimodal systems, and large-scale deployment scenarios. However, prompt selection, prompt interference, adapter growth, and task identity uncertainty remain open problems.
Prompt-based CL has recently emerged as one of the most successful strategies for adapting large pretrained ViTs and foundation models without modifying the backbone parameters. Instead of updating millions or even billions of network parameters during CL, prompt-based methods learn a relatively small collection of trainable prompt embeddings that guide the pretrained model toward new tasks while preserving previously acquired knowledge. Since the backbone remains frozen throughout training, these approaches substantially reduce computational cost, alleviate catastrophic forgetting, and eliminate the need for replay buffers in many CL scenarios. Consequently, prompt learning has become one of the dominant paradigms for rehearsal-free CL using foundation models.
One of the earliest prompt-based CL approaches is L2P [207]. Rather than relying on explicit task identifiers or replay memory, L2P maintains a pool of learnable prompts and dynamically retrieves the most relevant prompts for each input using feature similarity. During continual adaptation, only the selected prompts are optimized, while the pretrained ViT remains frozen. This design significantly reduces catastrophic forgetting and enables efficient continual adaptation without modifying the backbone network. Nevertheless, because prompts are retrieved independently for each task, prompt competition may increase as the number of tasks grows, leading to reduced prompt specialization and degraded long-term performance.
DualPrompt [206] extends L2P by introducing two complementary prompt types. General prompts capture knowledge shared across multiple tasks, whereas expert prompts specialize in task-specific representations. During inference, both prompt types are jointly utilized to balance knowledge transfer and task discrimination. Compared with L2P, this hierarchical prompt decomposition improves knowledge retention and reduces interference among related tasks. However, DualPrompt still assumes that appropriate prompt selection can effectively separate old and new knowledge, an assumption that becomes increasingly challenging under long CL sequences involving diverse task distributions.
To further alleviate prompt interference, CODA-Prompt [208] introduces an attention-based prompt composition mechanism that decomposes prompts into multiple learnable components rather than assigning a fixed prompt to each task. Instead of retrieving prompts independently, CODA-Prompt dynamically composes prompt representations through attention, allowing different prompt components to be shared across tasks according to their semantic relevance. This decomposition improves prompt utilization, increases parameter sharing, and substantially reduces prompt competition during continual adaptation. Experimental results demonstrate improved performance over earlier prompt-based methods across several CIL benchmarks, particularly for long task sequences where prompt interference becomes increasingly severe.
While most prompt-learning approaches focus primarily on CIL, S-Prompts extend prompt learning toward domain-incremental CL by explicitly separating shared and domain-specific prompt representations. Shared prompts encode transferable knowledge that remains useful across multiple domains, whereas domain-specific prompts capture distribution-dependent characteristics unique to individual domains. This decomposition improves domain generalization while reducing catastrophic forgetting caused by distribution shifts. Consequently, S-Prompts demonstrate that prompt decomposition can be beneficial not only across tasks but also across heterogeneous domains, making them particularly suitable for continual domain adaptation scenarios.
HiDe-Prompt further advances prompt learning by introducing a hierarchical prompt decomposition strategy. Instead of representing all continual knowledge using a single prompt pool, HiDe-Prompt organizes prompts across multiple semantic levels, enabling the model to capture both global knowledge shared among tasks and fine-grained task-specific information. The hierarchical prompt organization facilitates more effective knowledge reuse, improves scalability as task sequences become longer, and reduces prompt interference compared with flat prompt pools. Such hierarchical representations are particularly attractive for large ViTs because they better exploit the multi-level feature hierarchies learned during large-scale pretraining.
Despite their impressive performance, prompt-based CL methods remain subject to several important limitations. First, prompt interference continues to accumulate as the number of sequential tasks increases, making prompt retrieval increasingly ambiguous. Second, most existing methods assume relatively stable feature representations learned during pretraining, which may not hold under substantial domain shifts or multimodal continual adaptation. Third, prompt pools and expert prompts typically grow with the number of tasks, increasing storage requirements and inference complexity for long-term CL. Finally, most current prompt-learning methods have been developed primarily for ViTs, while extending similar mechanisms to large multimodal foundation models and LLMs remains an active research direction.
Overall, prompt-based CL has rapidly evolved from simple prompt retrieval strategies such as L2P to increasingly sophisticated prompt composition and hierarchical decomposition approaches, including DualPrompt, CODA-Prompt, S-Prompts, and HiDe-Prompt. Collectively, these methods demonstrate that parameter-efficient prompt optimization can substantially reduce catastrophic forgetting while preserving the powerful representations learned by foundation models. Nevertheless, achieving scalable prompt management, minimizing prompt interference, and supporting continual adaptation across vision, language, and multimodal foundation models remain important open research challenges.
In summary, CL methods differ in how they preserve previous knowledge and support adaptation. Regularization-based methods constrain parameter or function changes, replay-based methods revisit previous information, architecture-based methods isolate or expand capacity, optimization-based methods control gradient interference, representation-learning methods improve feature transferability, and parameter-efficient methods adapt large pretrained models with limited updates. In practice, the strongest CL systems often combine several of these strategies to balance accuracy, memory efficiency, computational cost, scalability, and robustness.

7.1.1. Replay-Free Adaptation of Pretrained Vision Models

Although prompt-based CL has substantially reduced the need for replay buffers, recent research has explored even simpler replay-free adaptation strategies that exploit the strong feature representations learned by pretrained foundation models. These approaches operate under the observation that large ViTs pretrained on massive datasets already possess highly transferable semantic representations, thereby reducing the need for continual optimization of the backbone network. Instead, continual adaptation can be achieved by learning lightweight classification or adaptation modules on top of frozen pretrained features. This strategy not only minimizes catastrophic forgetting but also significantly reduces computational cost, storage requirements, and privacy concerns associated with maintaining replay memories.
A representative example of this paradigm is RanPAC (Random Projection and Pretrained Representations for CL) [209]. Rather than learning task-specific prompts or replaying previous training samples, RanPAC keeps the pretrained ViT completely frozen throughout CL and performs adaptation only within a lightweight classifier. Specifically, pretrained feature representations are first projected into a higher-dimensional random feature space using a fixed random projection matrix. The projected features are subsequently used to construct an efficient prototype-based classifier that can be updated incrementally as new classes become available. Since neither the backbone parameters nor previously observed samples require modification, RanPAC achieves efficient continual adaptation while completely avoiding catastrophic forgetting caused by repeated optimization of deep neural network parameters.
An important advantage of RanPAC is that its computational complexity remains substantially lower than conventional replay-based CL algorithms. Because feature extraction is performed only once by the frozen foundation model, CL primarily involves updating the lightweight classifier rather than repeatedly optimizing millions of network parameters. Consequently, training becomes significantly faster, GPU memory consumption is considerably reduced, and privacy concerns associated with replay buffers are eliminated. These characteristics make RanPAC particularly attractive for resource-constrained environments, edge computing platforms, and privacy-sensitive applications where storing previous training data is either undesirable or prohibited.
Recent experimental studies have demonstrated that RanPAC achieves competitive performance on several standard CL benchmarks, including CIL scenarios based on CIFAR-100 and ImageNet. Despite requiring substantially fewer trainable parameters than conventional replay-based methods, RanPAC often approaches or even surpasses the performance of more computationally intensive algorithms. These results suggest that the rich semantic representations learned during large-scale pretraining can significantly reduce the reliance on replay memory when appropriate lightweight adaptation mechanisms are employed.
Nevertheless, RanPAC also exhibits several important limitations. First, because the pretrained backbone remains completely frozen, its ability to learn genuinely novel visual concepts is restricted when newly arriving tasks differ substantially from the original pretraining distribution. Second, random projection assumes that the projected feature space preserves sufficient discriminative information for continual classification, an assumption that may become less reliable under severe domain shifts or highly heterogeneous task sequences. Third, although replay-free adaptation reduces storage requirements, maintaining increasingly large prototype sets may still introduce computational overhead during long CL sequences involving thousands of classes. Finally, extending RanPAC beyond visual recognition toward multimodal foundation models and LLMs remains an important open research problem.
Overall, replay-free adaptation methods such as RanPAC represent an important shift in CL research. Instead of combating catastrophic forgetting through replay, regularization, or prompt optimization, these approaches leverage the strong representational capability of pretrained foundation models and perform continual adaptation using lightweight classifiers or feature transformations. Their excellent computational efficiency, low memory requirements, and privacy-preserving characteristics make replay-free adaptation an increasingly promising direction for scalable CL, particularly as foundation models continue to increase in size and capability.

7.1.2. Adapter-Based and Parameter-Efficient Continual Learning

PEFT has become one of the most important paradigms for continual adaptation of foundation models. Unlike conventional fine-tuning, which updates all model parameters during learning, PEFT modifies only a small subset of trainable parameters while keeping the pretrained backbone frozen. This strategy substantially reduces computational complexity, memory consumption, and optimization instability, making CL feasible for modern ViTs, LLMs, and multimodal foundation models containing billions of parameters. Furthermore, because the majority of pretrained parameters remain unchanged, PEFT naturally alleviates catastrophic forgetting by preserving previously acquired representations while enabling efficient adaptation to new tasks.
Among the numerous PEFT techniques, LoRA has become one of the most widely adopted approaches. LoRA [210] assumes that task-specific parameter updates can be represented using a low-rank decomposition instead of learning a full weight matrix. For a pretrained weight matrix W, LoRA approximates the update as
Δ W = B A ,
where B R d × r and A R r × k are trainable low-rank matrices with rank r min ( d , k ) . Consequently, the adapted weight becomes
W = W + B A .
Since only the low-rank matrices are optimized during CL, the number of trainable parameters is dramatically reduced compared with full fine-tuning. This significantly decreases GPU memory requirements and training time while maintaining competitive performance across numerous downstream tasks. Owing to these advantages, LoRA has become one of the standard adaptation techniques for CL with LLMs, ViTs, and multimodal foundation models. Nevertheless, repeated accumulation of LoRA modules for long task sequences may gradually increase storage requirements and complicate inference when numerous task-specific adaptations must be maintained simultaneously.
Another widely adopted PEFT strategy is adapter-based learning, in which lightweight bottleneck modules are inserted between transformer layers while the pretrained backbone remains frozen. During CL, only these adapter parameters are updated, allowing new knowledge to be incorporated without modifying the original foundation model. Compared with full fine-tuning, adapters provide greater modularity because different tasks may utilize independent adapter modules while sharing the same pretrained backbone. This modular design facilitates continual adaptation, improves parameter reuse, and simplifies deployment in multi-task environments where several adaptation modules can coexist without requiring multiple copies of the complete model.
More recent research has focused on adaptive parameter-efficient learning strategies that dynamically allocate adaptation modules according to the characteristics of incoming tasks. A representative example is Elastic Adapter Selection (EASE), which introduces an adaptive mechanism for selecting or composing task-specific adapter modules rather than statically assigning new adapters to every task. By encouraging parameter sharing among related tasks while preserving specialized adapters for task-specific knowledge, EASE reduces redundant parameter growth and improves scalability over long CL sequences. Such adaptive routing mechanisms represent an important step toward continual foundation models capable of supporting hundreds or thousands of sequential tasks without proportional increases in model size.
Despite the remarkable success of PEFT techniques, several important challenges remain. First, although LoRA and adapter modules greatly reduce the number of trainable parameters, continual accumulation of adaptation modules inevitably increases long-term storage requirements as additional tasks are encountered. Second, efficient routing or selection of appropriate adaptation modules becomes increasingly difficult as the number of stored adapters grows, potentially increasing inference latency. Third, most existing PEFT methods assume relatively stable pretrained feature representations, whereas substantial domain shifts may require updating backbone representations beyond lightweight parameter adaptation. Finally, extending parameter-efficient CL from ViTs to multimodal foundation models and LLMs introduces additional challenges related to instruction following, reasoning consistency, safety alignment, and cross-modal knowledge transfer.
Overall, PEFT has fundamentally changed the landscape of CL for foundation models. Rather than continually optimizing billions of parameters, modern PEFT approaches achieve continual adaptation by learning compact task-specific modules that preserve pretrained knowledge while maintaining high computational efficiency. Techniques such as LoRA, adapter-based learning, and adaptive frameworks including EASE demonstrate that scalable CL can be achieved through efficient parameter reuse instead of repeated full-model optimization. Future research is expected to further improve adaptive parameter routing, reduce long-term storage overhead, and develop unified PEFT frameworks capable of continual adaptation across vision, language, and multimodal foundation models.

7.1.3. Continual Instruction Tuning for Large Language Models

The rapid emergence of LLMs has significantly broadened the scope of CL beyond traditional vision tasks. Unlike conventional CL, which primarily focuses on sequential visual classification or representation learning, continual adaptation of LLMs aims to update instruction-following capabilities, factual knowledge, reasoning ability, and conversational behavior while preserving previously acquired competencies. Since modern LLMs are typically pretrained on massive corpora containing billions of tokens, retraining these models whenever new knowledge becomes available is computationally impractical. Consequently, continual instruction tuning has emerged as an efficient paradigm for incrementally adapting pretrained LLMs to evolving user requirements without performing full model retraining.
Continual instruction tuning extends supervised instruction tuning to sequential learning environments, where new instruction datasets arrive over time. Instead of independently fine-tuning a model for each collection of instructions, the objective is to continuously incorporate new capabilities while preserving performance on previously learned instructions. Typical instruction datasets contain diverse reasoning, question answering, summarization, translation, code generation, and dialog tasks, requiring CL algorithms to balance stability and plasticity across heterogeneous instruction distributions. PEFT methods such as LoRA, adapters, and prompt tuning are frequently employed to reduce computational cost during continual instruction tuning while minimizing catastrophic forgetting.
To facilitate systematic evaluation of continual instruction tuning, recent studies have introduced benchmark suites such as TRACE, which evaluates the ability of LLMs to incrementally acquire new instruction-following skills across sequential tasks. TRACE measures not only predictive performance on newly introduced instructions but also knowledge retention, forward transfer, backward transfer, and forgetting throughout continual adaptation. By organizing diverse instruction datasets into sequential learning scenarios, TRACE enables more comprehensive evaluation of lifelong instruction following than traditional single-task fine-tuning benchmarks. Such evaluation protocols provide valuable insights into how effectively LLMs preserve previously acquired capabilities while learning new instruction distributions.
Another important benchmark is ConTinTin, which focuses specifically on continual instruction tuning of LLMs under realistic lifelong learning settings. Rather than evaluating isolated downstream tasks, ConTinTin considers sequential instruction adaptation across multiple domains while measuring knowledge retention, reasoning consistency, instruction-following accuracy, and continual generalization. The benchmark also emphasizes long-term evaluation across extended task sequences, thereby exposing limitations that are often hidden by conventional short-term CL benchmarks. Such benchmark suites represent an important step toward standardized evaluation protocols for continual adaptation of foundation language models.
Despite substantial progress, catastrophic forgetting remains a major challenge for continual adaptation of LLMs. Unlike conventional image classification, forgetting in language models manifests not only as degraded predictive accuracy but also as deterioration of factual knowledge, reasoning capability, instruction-following behavior, dialog quality, safety alignment, and multilingual competence. Sequential fine-tuning on new instruction datasets may unintentionally overwrite previously learned knowledge, resulting in inconsistent responses, hallucinations, or degraded reasoning performance on earlier tasks. Furthermore, preserving alignment with human preferences during continual adaptation remains particularly challenging because updating instruction-following behavior may unintentionally alter previously aligned safety constraints or conversational characteristics.
Several strategies have been proposed to mitigate catastrophic forgetting during continual instruction tuning, including replay-based learning using previously observed instructions, regularization methods that preserve important parameters, PEFT, adapter composition, prompt routing, and hybrid replay-distillation frameworks. More recently, retrieval-augmented generation and external memory mechanisms have also been investigated as complementary approaches for updating factual knowledge without permanently modifying model parameters. Although these techniques substantially improve continual adaptation, achieving scalable lifelong learning for billion-parameter language models remains an open research problem because computational cost, memory efficiency, long-term knowledge retention, and alignment preservation become increasingly difficult as the number of sequential tasks grows.
Overall, continual instruction tuning represents one of the most rapidly developing areas of CL research. Benchmarks such as TRACE and ConTinTin provide standardized evaluation protocols for measuring long-term adaptation of instruction-following capabilities, while PEFT methods enable practical deployment on increasingly large foundation models. Nevertheless, developing CL algorithms capable of preserving reasoning ability, factual knowledge, safety alignment, and multimodal understanding throughout prolonged sequential adaptation remains a fundamental challenge for next-generation foundation models.

7.1.4. Critical Discussion and Open Challenges

Despite the remarkable progress achieved by prompt learning, replay-free adaptation, and PEFT, several fundamental challenges remain before foundation models can support robust lifelong learning in real-world environments. Although recent methods substantially reduce catastrophic forgetting while maintaining computational efficiency, they often rely on assumptions that become increasingly difficult to satisfy as the number of sequential tasks, model size, and data diversity continue to grow. Consequently, improving the scalability, robustness, and deployment readiness of continual foundation models remains an important research direction.
One of the most widely recognized challenges is prompt interference. Prompt-based methods assume that task-specific prompts can effectively capture new knowledge while preserving previously learned representations. However, as the number of sequential tasks increases, different prompts frequently compete for similar feature representations, making prompt retrieval increasingly ambiguous. This phenomenon may reduce prompt specialization, introduce negative transfer among related tasks, and ultimately degrade CL performance. Although recent approaches such as CODA-Prompt and HiDe-Prompt attempt to alleviate prompt interference through prompt decomposition and hierarchical organization, designing scalable prompt management strategies for long CL sequences remains an open problem.
Another important limitation is adapter accumulation in PEFT methods. While LoRA, adapters, and related PEFT techniques dramatically reduce the number of trainable parameters for individual tasks, CL often requires introducing new adaptation modules as additional tasks arrive. Over long task sequences, the cumulative number of adapters or low-rank modules may substantially increase storage requirements, inference latency, and model management complexity. Developing adaptive parameter-sharing mechanisms capable of reusing previously learned adapters without sacrificing task-specific performance therefore represents an important direction for future research.
Closely related to adapter accumulation is the challenge of parameter isolation. Many CL algorithms attempt to prevent catastrophic forgetting by allocating independent parameters, prompts, or adapters for different tasks. Although parameter isolation effectively reduces destructive interference, excessive isolation limits knowledge transfer between related tasks and gradually increases model complexity. Conversely, excessive parameter sharing may improve transfer but increase catastrophic forgetting. Designing adaptive mechanisms that dynamically balance parameter sharing and parameter isolation according to task similarity remains one of the central challenges in CL theory and practice.
The continual adaptation of multimodal foundation models introduces additional challenges associated with multimodal alignment collapse. Modern vision–language models rely on carefully aligned visual and textual representations acquired during large-scale pretraining. Sequential adaptation using only one modality or a limited set of multimodal tasks may gradually distort this alignment, reducing cross-modal retrieval accuracy, visual grounding capability, and multimodal reasoning performance. Preserving semantic consistency across multiple modalities while continuously incorporating new knowledge remains considerably more challenging than continual adaptation of unimodal vision or language models. Consequently, future research should investigate continual alignment strategies capable of maintaining robust cross-modal representations throughout long-term sequential learning.
Scalability also remains a major concern as foundation models continue to increase in size. Although prompt learning and PEFT substantially reduce the number of trainable parameters, training and deploying billion-parameter models still require considerable computational resources, GPU memory, and energy consumption. Furthermore, evaluation protocols have not yet kept pace with recent advances in foundation models, making it difficult to compare algorithms fairly across different architectures, parameter budgets, and CL scenarios. Developing scalable optimization algorithms together with standardized benchmarks for large foundation models therefore represents an essential research direction.
Finally, several practical deployment challenges remain largely unexplored. Real-world CL systems must operate under limited computational resources, strict latency requirements, privacy regulations, non-stationary data distributions, and evolving user preferences. In many applications, replay buffers cannot be maintained because of privacy concerns, while continual adaptation must occur on resource-constrained edge devices or distributed systems. Furthermore, LLMs require continual updates without compromising factual accuracy, reasoning capability, or safety alignment. Addressing these practical constraints requires integrating CL with efficient optimization, privacy-preserving learning, retrieval-augmented generation, distributed training, and trustworthy AI principles.
Overall, foundation-model CL has rapidly evolved from replay-based optimization toward increasingly parameter-efficient adaptation strategies based on prompts, adapters, and lightweight task-specific modules. Nevertheless, prompt interference, adapter accumulation, parameter isolation, multimodal alignment preservation, scalability, and deployment constraints continue to limit practical lifelong learning. Future research should therefore focus on developing unified CL frameworks capable of supporting scalable, trustworthy, and computationally efficient adaptation across vision, language, and multimodal foundation models while maintaining strong long-term knowledge retention and generalization.

7.2. Practical Comparison of Continual Learning Paradigms

The previous sections reviewed the major categories of CL methods from methodological and theoretical perspectives. However, selecting an appropriate CL algorithm for practical deployment requires considering multiple factors beyond predictive performance alone. Different paradigms exhibit distinct trade-offs in computational complexity, memory consumption, scalability, robustness, and deployment feasibility, making them suitable for different application scenarios. Consequently, a comprehensive comparison should jointly consider algorithmic performance, resource requirements, and compatibility with modern learning paradigms rather than relying solely on classification accuracy or forgetting metrics.
Table 17 presents a unified comparison of the major CL paradigms from a deployment-oriented perspective. Specifically, the table compares representative method categories with respect to memory usage, computational complexity, scalability, representative benchmark performance, suitability for different CL scenarios, and applicability to foundation models. These criteria collectively provide a more holistic assessment of practical strengths and limitations than isolated performance metrics. For example, replay-based methods generally achieve superior knowledge retention but require substantial replay memory and raise privacy concerns, whereas regularization-based approaches maintain low memory requirements but often experience increased forgetting under long task sequences and severe domain shifts. Architecture-based methods effectively mitigate catastrophic forgetting through parameter isolation but exhibit reduced scalability due to continual model expansion. More recently, prompt-based and PEFT approaches have demonstrated strong scalability and competitive benchmark performance while updating only a small fraction of model parameters, making them particularly attractive for adapting large foundation models. Nevertheless, prompt interference, adapter accumulation, and long-term parameter management remain important challenges that require further investigation.
It is important to emphasize that the qualitative ratings summarized in Table 17 represent general trends consistently reported across the recent CL literature rather than absolute rankings. Since individual studies employ different backbone architectures, benchmark datasets, replay budgets, evaluation protocols, and experimental settings, direct numerical comparison across all methods is generally inappropriate. Instead, the table is intended to provide readers with a concise overview of the practical trade-offs associated with each methodological category and to facilitate informed selection of CL approaches for diverse real-world applications.

8. Evaluation Protocols, Benchmarks, and Metrics in CL

A major challenge in CL research is the lack of standardized evaluation protocols and benchmark settings. Although numerous methods have been proposed to mitigate catastrophic forgetting, fair comparison across studies remains difficult due to variations in task construction, dataset partitioning, memory budgets, evaluation metrics, and training protocols. As a result, reported performance often depends not only on the effectiveness of the proposed method but also on the experimental setup itself. Establishing consistent and reproducible evaluation practices is therefore essential for accurately assessing CL systems.

8.1. Benchmark Datasets

Benchmark datasets play a central role in evaluating CL methods. Existing benchmarks can generally be categorized into image classification, object detection, video understanding, reinforcement learning, NLP, and medical imaging benchmarks. In image classification, commonly used benchmarks include Split-MNIST, Permuted-MNIST, Split CIFAR-10, Split CIFAR-100, TinyImageNet, and ImageNet-based incremental benchmarks. These datasets are frequently used in task-incremental, domain-incremental, and CIL scenarios due to their controllable task construction and broad adoption in prior studies.
More challenging and realistic benchmarks have also emerged in recent years. CORe50 introduces continuous object recognition under varying environmental conditions, while Stream-51 focuses on streaming video-based CL. CLEAR (CL on real-world imagery) was proposed to evaluate CL under more realistic temporal and environmental variations. Such benchmarks aim to move beyond artificially segmented tasks toward real-world continual adaptation scenarios.
In medical imaging, CL benchmarks remain relatively limited despite increasing interest in lifelong medical AI systems. Existing studies commonly utilize datasets from segmentation and classification tasks, including retinal imaging, brain tumor segmentation, cardiac ultrasound, and chest X-ray analysis. However, evaluation protocols in medical CL are often inconsistent due to domain shifts across hospitals, imaging devices, and annotation standards.

8.2. Continual Learning Evaluation Settings

CL evaluation protocols typically differ according to the underlying learning scenario. The three most widely studied settings are TIL, domain-incremental learning (DIL), and CIL. In TIL, task identity is available during inference, allowing the model to utilize task-specific components or output heads. This setting is generally considered less challenging because the model is informed about the current task context.
In DIL, the task objective remains unchanged while the input distribution changes over time. The model must adapt to domain shifts without explicit task identity information during inference. CIL represents the most challenging and practically relevant setting. In this scenario, new classes are introduced sequentially, and the model must classify samples across all previously encountered classes without access to task identity during inference. This setting closely resembles real-world deployment conditions where task boundaries are often unavailable.
Recent research has also explored online CL, streaming CL, multimodal CL, and federated CL settings. These protocols introduce additional constraints such as single-pass training, privacy preservation, communication efficiency, and cross-modal adaptation.

8.3. Task Construction and Data Splits

Task construction significantly influences CL performance. Different studies often use distinct task partitioning strategies even when evaluating on the same dataset, making direct comparison difficult. For example, CIFAR-100 may be divided into 10 tasks with 10 classes each, 20 tasks with five classes each, or arbitrary task groupings depending on the experimental design. Similarly, the order of tasks can substantially affect forgetting behavior and knowledge transfer. Some studies use fixed task orders, whereas others employ random task permutations across multiple runs. Unfortunately, many works report results using only a single task ordering, which may introduce evaluation bias. Another important issue involves base initialization protocols. Some methods begin with a large base task followed by incremental updates, whereas others use fully balanced sequential tasks. These choices directly influence feature stability, representation quality, and replay effectiveness. To improve reproducibility, recent works increasingly recommend reporting task construction details explicitly, evaluating across multiple random seeds, testing multiple task orders, and using consistent train–validation–test splits.

8.4. Evaluation Metrics

Several evaluation metrics have been proposed to measure CL performance. However, no single metric fully captures all aspects of continual adaptation.

8.4.1. Average Accuracy

Average accuracy is one of the most widely used metrics in CL. After learning task T, the average accuracy is computed as
A T = 1 T i = 1 T a T , i
where a T , i denotes the test accuracy on task i after learning task T.
Although average accuracy provides a general measure of overall performance, it does not explicitly quantify forgetting or transfer behavior.

8.4.2. Forgetting Measure

Forgetting evaluates how much performance degrades on previous tasks after learning new tasks. A commonly used formulation is
F T = 1 T 1 i = 1 T 1 max l { 1 , , T 1 } a l , i a T , i
Lower forgetting values indicate better knowledge retention.

8.4.3. Forward Transfer

Forward transfer measures whether previously acquired knowledge improves learning efficiency on future tasks. Positive forward transfer indicates that earlier representations facilitate adaptation to new tasks.

8.4.4. Backward Transfer

Backward transfer evaluates whether learning new tasks improves performance on earlier tasks. Positive backward transfer reflects beneficial knowledge integration across tasks, whereas negative backward transfer indicates forgetting.

8.4.5. Memory and Computational Efficiency

In replay-based methods, memory budget is a critical evaluation factor. Some methods achieve high performance by storing large replay buffers, making comparisons unfair when memory constraints differ substantially across studies. Computational efficiency is also increasingly important, particularly for large-scale transformer-based and foundation-model CL systems. Metrics such as training time, parameter growth, inference latency, and energy consumption are becoming relevant for real-world deployment.

8.5. Challenges in Continual Learning Evaluation

Despite substantial progress, CL evaluation remains fragmented and inconsistent. Several key issues continue to hinder fair comparison across methods. These include different task splits and benchmark configurations, inconsistent replay memory budgets, variation in model initialization strategies, different validation protocols and hyperparameter tuning approaches, limited reporting of statistical variance across multiple runs, and heavy reliance on simplified academic benchmarks. Furthermore, many studies evaluate methods primarily on image classification datasets while neglecting more realistic settings such as multimodal learning, streaming environments, dense prediction tasks, and large-scale foundation-model adaptation. Another emerging challenge involves evaluating CL under realistic deployment constraints, including privacy preservation, domain shift, continual annotation updates, and resource-limited edge devices. Although the previous sections qualitatively compare the major CL paradigms, qualitative descriptions alone are often insufficient for understanding the practical trade-offs among different approaches. To provide a more comprehensive synthesis of the existing literature, Table 18 and Table 19 summarize representative quantitative comparisons reported across recent CL studies. Table 18 provides a high-level comparison of the major CL paradigms in terms of final performance, forgetting behavior, memory requirements, computational cost, and scalability, highlighting the inherent trade-offs among different methodological categories. Complementarily, Table 19 summarizes representative performance ranges reported on widely used CL benchmarks, illustrating the relative strengths and limitations of representative methods under commonly adopted evaluation protocols. Since these results originate from different studies employing diverse backbone architectures, replay budgets, task sequences, and experimental settings, the reported values should be interpreted as representative trends rather than direct head-to-head comparisons. Together, these quantitative summaries complement the qualitative analysis presented throughout this review and provide readers with a clearer understanding of the practical performance characteristics of modern CL approaches.
Overall, replay-based methods consistently achieve higher final accuracy and lower forgetting than regularization-based approaches, although they require additional memory for sample storage. Architecture-based methods provide excellent knowledge retention but exhibit limited scalability because model size grows with the number of tasks. Recent prompt-learning and PEFT approaches demonstrate competitive performance while substantially reducing the number of trainable parameters, making them attractive for adapting large foundation models. Nevertheless, quantitative comparison across studies remains challenging because evaluation protocols, benchmark construction, replay budgets, backbone architectures, and reporting metrics differ considerably. These inconsistencies highlight the need for standardized CL benchmarks and reporting practices.

8.6. Toward a Standardized Evaluation Protocol for Continual Learning

Evaluation inconsistency remains one of the major barriers to fair and reproducible comparison of CL methods. Existing studies differ considerably in benchmark selection, task construction, memory allocation, evaluation metrics, computational reporting, and experimental settings. Consequently, reported improvements are often difficult to interpret, reproduce, or compare across studies, even when evaluating similar algorithms. Benchmark fragmentation further complicates objective assessment, as performance can vary substantially depending on task order, memory budget, and evaluation protocol.
To improve reproducibility and facilitate more meaningful comparisons, we propose a standardized evaluation protocol that serves as a practical reporting guideline for future CL research. Rather than prescribing a specific benchmark or algorithm, the proposed protocol defines the essential experimental information that should be consistently reported across studies. These recommendations emphasize transparent benchmark construction, clearly specified CL scenarios, fixed replay memory constraints, standardized evaluation metrics, computational efficiency analysis, statistical validation over multiple random seeds, and comprehensive implementation details. Furthermore, the protocol encourages evaluation on realistic settings involving foundation models, multimodal data streams, domain shifts, and streaming environments to better reflect real-world deployment conditions. The proposed evaluation checklist, summarized in Table 20, is intended to complement rather than replace existing benchmark protocols. Its adoption would improve experimental transparency, strengthen reproducibility, reduce benchmark fragmentation, and facilitate fair comparison among future CL methods developed under diverse experimental settings.

8.7. Representative Continual Learning Benchmarks

Benchmark selection plays a fundamental role in the evaluation of CL algorithms because different benchmarks emphasize different aspects of lifelong adaptation, including catastrophic forgetting, domain adaptation, task transfer, scalability, and computational efficiency. Consequently, the choice of benchmark significantly influences the reported performance of CK methods and often contributes to the inconsistency observed across the literature. Table 21 summarizes several representative benchmarks that are widely adopted for evaluating CL algorithms across computer vision, NLP, and foundation-model adaptation.
Among vision benchmarks, Split CIFAR-100 remains one of the most frequently adopted datasets because of its simplicity, computational efficiency, and standardized experimental protocols. It is widely used for evaluating CIL methods and measuring catastrophic forgetting through metrics such as final average accuracy and average forgetting. However, its relatively small image resolution and limited semantic diversity make it insufficient for evaluating large foundation models or realistic deployment scenarios. Split Tiny-ImageNet provides greater visual diversity and increased task complexity, although it still represents a simplified approximation of real-world CL environments.
Split ImageNet has become an increasingly important benchmark because it contains substantially larger visual diversity and more realistic object categories. Many recent CL studies employ Split ImageNet to evaluate scalability and foundation-model adaptation under long task sequences. Nevertheless, its large computational requirements considerably increase training time, GPU memory consumption, and experimental cost, making comprehensive evaluation difficult for many researchers. Similarly, DomainNet introduces multiple visual domains that enable realistic domain-incremental evaluation under significant distribution shifts. Although DomainNet better reflects practical deployment environments than conventional image classification benchmarks, its heterogeneous domain distributions and computational complexity create additional experimental challenges.
Several benchmark suites have also been developed to improve evaluation consistency across CL scenarios. For example, the CL Vision benchmark provides standardized evaluation protocols covering multiple CL settings, datasets, and evaluation metrics. Such benchmark suites improve reproducibility and facilitate fairer comparison among competing methods by reducing variations in experimental protocols. Nevertheless, differences in backbone architectures, replay memory budgets, task sequences, and implementation details continue to affect reported performance, highlighting the ongoing need for standardized evaluation methodologies.
Beyond computer vision, CL has attracted increasing attention in NLP and LLMs. Continual NLP benchmarks typically evaluate sequential adaptation across multiple domains or language understanding tasks using metrics such as classification accuracy and F1-score. More recently, specialized benchmarks have emerged for evaluating continual adaptation of foundation models and LLMs under instruction tuning, prompt learning, PEFT, and knowledge editing scenarios. These benchmarks better reflect the practical challenges of maintaining instruction-following ability, preserving previously acquired knowledge, and mitigating catastrophic forgetting during long-term adaptation of large pretrained models.
Despite the substantial progress achieved by existing benchmark datasets, several important limitations remain. Many traditional vision benchmarks, including Split CIFAR-100 and Split Tiny-ImageNet, rely on predefined task boundaries, balanced class distributions, relatively small image collections, and short task sequences that fail to capture the complexity of real-world lifelong learning. Although larger benchmarks such as Split ImageNet and DomainNet provide greater semantic diversity and more realistic domain shifts, they require substantially higher computational resources, making comprehensive evaluation more expensive and less accessible. Furthermore, benchmark fragmentation has resulted in considerable variation in evaluation metrics, replay memory constraints, backbone architectures, and experimental protocols, making direct comparison across published studies difficult. In NLP and LLMs, benchmark standardization remains relatively limited, with considerable differences in task construction, evaluation protocols, and continual adaptation objectives. Table 21 summarizes representative CL benchmarks across computer vision, NLP, and foundation-model adaptation by comparing their application domains, typical CL scenarios, commonly reported evaluation metrics, applicability to foundation models, and principal limitations. Overall, future benchmark development should emphasize standardized evaluation protocols, multimodal CL, realistic domain shifts, long-term task sequences, foundation-model adaptation, and deployment-oriented evaluation under practical computational and memory constraints.

8.8. Quantitative Comparison of Representative Continual Learning Methods

Although qualitative comparisons provide valuable insights into the characteristics of different CL paradigms, quantitative evidence is equally important for understanding their practical performance. Table 22 summarizes representative results reported by several influential CL methods on widely adopted benchmarks, including Split CIFAR-100, Split ImageNet, and CORe50. Since these results originate from different studies employing different backbone networks, replay memory sizes, training protocols, and evaluation settings, they should not be interpreted as direct head-to-head comparisons. Instead, they provide a representative overview of current performance trends across major methodological categories.

9. Comparative Analysis of CL Methods

CL methods have evolved substantially over the past decade, leading to a diverse set of strategies designed to mitigate catastrophic forgetting while enabling adaptation to new tasks and domains. The representative quantitative results summarized in Table 22 further illustrate that no single CL method consistently achieves the highest performance across all benchmark datasets and CL scenarios. Replay-based methods generally exhibit strong performance on classical visual benchmarks, whereas prompt-based and parameter-efficient approaches demonstrate superior scalability for foundation models. Consequently, the most appropriate CL strategy remains highly dependent on the application scenario, computational constraints, memory budget, and model architecture. Different approaches exhibit distinct strengths and limitations depending on the task setting, memory constraints, computational budget, and availability of task identity information. This section provides a comparative analysis of major CL method categories, focusing on their underlying principles, advantages, limitations, scalability, and suitability for practical deployment. Table 23 analyses the major CL method categories.

9.1. Overview of CL Method Categories

Existing CL methods can generally be categorized into regularization-based methods, replay-based methods, architecture-based methods, optimization-based methods, representation-learning methods, and more recent parameter-efficient and prompt-based adaptation approaches. Although these categories are conceptually distinct, many modern approaches combine multiple strategies to improve stability and adaptability.
Regularization-based methods aim to preserve previously learned knowledge by constraining parameter updates during new task learning. Replay-based methods retain or reconstruct previous data distributions to reinforce earlier knowledge. Architecture-based methods dynamically expand or isolate network components to reduce interference between tasks. Optimization-based methods directly manipulate gradient updates to minimize forgetting. Representation-learning approaches focus on learning transferable and stable features across tasks. More recently, parameter-efficient and prompt-based methods have emerged as promising solutions for adapting large-scale foundation models under CL settings.
The computational characteristics summarized in Table 23 are intended to provide a qualitative comparison across major CL paradigms rather than absolute quantitative measurements. Exact memory consumption, parameter growth, and computational complexity depend on several factors, including the backbone architecture, replay buffer size, dataset characteristics, model scale, and hardware configuration. To improve reproducibility and facilitate fair comparison, future CL studies should consistently report the number of trainable parameters, total model size, replay memory footprint, computational complexity (FLOPs), GPU memory consumption, and training and inference times.

9.2. Regularization-Based Methods

Regularization-based methods are among the earliest and most widely studied CL approaches. These methods attempt to preserve previously acquired knowledge by penalizing changes to parameters that are considered important for earlier tasks. Representative approaches include EWC, SI, and Memory-Aware Synapses. A major advantage of regularization-based methods is their relatively low memory overhead since they do not require storing large replay buffers. These methods are computationally efficient and relatively easy to integrate into existing training pipelines. Furthermore, they are well suited for privacy-sensitive applications where retaining previous training data may be infeasible.
However, regularization-based methods often struggle in highly non-stationary environments or long task sequences where task distributions differ substantially. Since they rely primarily on parameter constraints, their ability to preserve fine-grained task representations decreases as the number of tasks grows. Consequently, forgetting may accumulate gradually over time. Regularization approaches are generally more effective in task-incremental settings where task boundaries are relatively distinct and interference between tasks is moderate.

9.3. Replay-Based Methods

Replay-based methods mitigate forgetting by revisiting previously learned data during training. These methods either store a subset of past samples in memory buffers or generate synthetic replay samples using generative models. Representative approaches include experience replay (ER), GEM, iCaRL, dark experience replay (DER), and generative replay methods. Replay-based strategies have demonstrated strong empirical performance across various CL benchmarks, particularly in CIL scenarios. By repeatedly exposing the model to previous data distributions, replay methods effectively stabilize feature representations and reduce catastrophic forgetting.
Despite their effectiveness, replay-based methods introduce several challenges. Storing previous samples increases memory requirements, which can become problematic in large-scale or privacy-sensitive applications. Moreover, replay performance depends heavily on memory selection strategies, replay buffer size, and sample diversity. Generative replay methods alleviate explicit storage requirements but often suffer from imperfect sample generation and increased computational complexity. Replay-based methods remain among the most effective approaches for CIL, especially when moderate memory budgets are available.

9.4. Architecture-Based Methods

Architecture-based methods address catastrophic forgetting by modifying the network structure itself. These approaches allocate dedicated parameters, subnetworks, or modules for different tasks, thereby reducing interference between old and new knowledge. Representative methods include progressive neural networks, dynamically expandable networks (DENs), PathNet, and expert-based modular architectures. One of the primary strengths of architecture-based methods is their ability to preserve previously learned representations with minimal forgetting. Since task-specific parameters are isolated, interference between tasks is significantly reduced. However, these methods often suffer from scalability limitations. As the number of tasks increases, network size and computational complexity may grow substantially. Furthermore, architecture expansion can become inefficient for long CL sequences or resource-constrained deployment environments. Architecture-based approaches are particularly effective in task-incremental settings where task identity is available during inference.

9.5. Optimization-Based Methods

Optimization-based methods attempt to reduce forgetting by directly controlling gradient updates during training. These approaches aim to prevent destructive interference between tasks by modifying optimization trajectories. Representative methods include GEM, A-GEM, Orthogonal Gradient Descent (OGD), and related gradient projection strategies. These methods provide a principled framework for balancing stability and plasticity during continual adaptation. By constraining gradient directions, optimization-based methods attempt to preserve previously learned knowledge while enabling efficient learning of new tasks.
Nevertheless, optimization-based approaches often incur substantial computational overhead due to gradient storage, projection operations, or optimization constraints. Their performance may also depend heavily on replay memory quality and gradient approximation strategies. Optimization-based methods are especially useful in scenarios where preserving task-specific gradient information is critical.

9.6. Representation Learning Approaches

Representation-learning methods focus on learning stable and transferable feature representations across tasks. Instead of solely preserving parameters or replaying data, these approaches attempt to disentangle task-specific and task-invariant representations. Recent advances in self-supervised learning and contrastive learning have significantly influenced representation-based CL. These methods aim to improve feature generalization while reducing sensitivity to task-specific distribution shifts. A key advantage of representation-learning approaches is their ability to improve forward transfer and domain generalization. Robust representations can facilitate adaptation to future tasks and reduce forgetting under moderate domain shifts. However, learning universally transferable representations remains challenging, particularly in highly heterogeneous task sequences. Representation-learning methods may still require replay mechanisms or auxiliary regularization strategies for long-term stability.

9.7. Prompt-Based and Parameter-Efficient CL

Recent years have witnessed increasing interest in applying CL to large-scale transformers and foundation models. In this context, PEFT and prompt-based adaptation strategies have emerged as scalable alternatives to full model retraining. Representative methods include Learning to Prompt (L2P), DualPrompt, CODA-Prompt, adapter-based tuning, prefix tuning, and LoRA. These approaches update only small subsets of model parameters while preserving the majority of pretrained weights.
Prompt-based CL methods offer several advantages. They reduce computational cost, improve scalability for large foundation models, and mitigate catastrophic forgetting by minimizing parameter interference. Furthermore, these methods are particularly suitable for multimodal and LLM adaptation scenarios. Despite their promise, prompt-based methods remain relatively underexplored in realistic continual deployment settings. Several open challenges remain, including prompt interference, prompt scalability, long-term adaptation stability, and efficient prompt selection strategies.

9.8. Comparison Across CL Settings

Different CL methods exhibit varying levels of effectiveness depending on the underlying CL scenario. In TIL, architecture-based and regularization-based methods often perform well because task identity information reduces ambiguity during inference. In DIL, representation-learning approaches and replay-based methods tend to provide stronger robustness against domain shifts. In CIL, replay-based methods generally outperform other categories because the model must simultaneously distinguish between old and new classes without access to task identity information. Online and streaming CL scenarios introduce additional constraints, including single-pass learning, limited memory, and real-time adaptation requirements. Under such settings, lightweight replay strategies and parameter-efficient adaptation methods become increasingly important.

9.9. Memory, Scalability, and Computational Trade-Offs

One of the central challenges in CL involves balancing performance with memory and computational efficiency. Replay-based methods often achieve strong retention performance but require explicit memory buffers. Architecture-based methods reduce forgetting effectively but may scale poorly due to parameter growth. Regularization-based approaches remain memory-efficient but may struggle in highly dynamic environments. Prompt-based and PEFT approaches represent a promising compromise for large-scale continual adaptation because they reduce computational overhead while preserving pretrained representations. However, their long-term scalability and robustness under severe distribution shifts remain active research topics. Consequently, selecting an appropriate CL strategy depends heavily on the target application, deployment constraints, and task characteristics.
The comparative analysis reveals that CL remains fundamentally a trade-off optimization problem involving stability, plasticity, scalability, memory efficiency, and computational cost. Existing methods often excel in specific scenarios but fail to generalize universally across diverse CL environments. Future research should increasingly focus on unified CL frameworks capable of handling realistic streaming conditions, multimodal inputs, large-scale foundation models, domain shifts, and deployment-oriented constraints. In addition, greater emphasis should be placed on standardized evaluation protocols, reproducibility, and practical scalability to bridge the gap between academic benchmarks and real-world continual adaptation systems.

9.10. Training Memory, Inference Memory, and Long-Term Storage

The computational feasibility of CL algorithms cannot be assessed solely by reporting overall memory consumption. Different CL methods exhibit distinct resource requirements during training, deployment, and long-term operation. Therefore, it is important to distinguish three complementary aspects of memory usage: training memory, inference memory, and long-term storage.
Training memory refers to the GPU or system memory required during model optimization. It includes memory used to store activations, gradients, optimizer states, replay samples, auxiliary statistics (e.g., FIM or parameter importance estimates), and temporary computational buffers. Training memory directly affects hardware requirements and determines whether a CL algorithm can be trained efficiently on resource-constrained platforms.
Inference memory denotes the memory required after training has been completed. It depends primarily on the size of the deployed model and any additional modules required during inference, such as task-specific adapters, prompts, expandable subnetworks, or multiple classifier heads. Architecture-based CL methods often exhibit increasing inference memory because the model grows as new tasks are introduced, whereas most regularization-based approaches maintain nearly constant inference memory.
Long-term storage represents the persistent storage required throughout the CL process. This includes replay buffers, synthetic generators, task-specific adapters, prompt libraries, parameter snapshots, or archived feature representations that must be retained across tasks. Replay-based methods generally require the largest long-term storage because representative samples or generated data must be preserved, whereas regularization-based methods usually store only lightweight importance statistics.
These three resource requirements influence deployment in different ways. Edge devices are often constrained by training and inference memory, cloud-based systems may primarily be limited by computational cost, and long-duration CL applications are frequently dominated by persistent storage requirements. Consequently, future CL studies should report these resource characteristics separately rather than presenting a single aggregate memory measurement. As shown in Table 24, different continual learning paradigms exhibit distinct trade-offs in training memory, inference memory, storage requirements, and computational resource consumption.

10. Applications of CL

CL has vast potential across various domains where systems need to learn and adapt to new information over time without forgetting prior knowledge. Below is a detailed explanation of its applications in healthcare and medical imaging, robotics and autonomous systems, NLP, recommender systems, and cybersecurity. Table 25 presents the summary of key applications of CL.

10.1. Applications in Healthcare and Medical Imaging

The healthcare domain stands to benefit immensely from the application of CL, particularly in medical imaging, where the technology can deal with the challenges of data heterogeneity, evolving diagnostic criteria, and the need for personalized treatment approaches [220]. Medical imaging datasets are often characterized by variability in image acquisition protocols, scanner types, and patient demographics, which can lead to performance degradation in traditional ML models. CL techniques can enable medical imaging systems to adapt to these variations by sequentially learning from new datasets without forgetting previously acquired knowledge [221]. One of the pivotal challenges in the deep learning field is data integration from various sources acquired using different hardware vendors, diverse acquisition protocols, experimental setups, and even inter-operator variabilities [222]. This leads to heterogeneous datasets, requiring careful harmonization before being usable to train AI algorithms. Moreover, CL can facilitate the integration of new imaging modalities or biomarkers into existing diagnostic workflows, enhancing the comprehensiveness and accuracy of clinical decision-making. Deep learning techniques have the capability to enhance diagnostic accuracy, streamline workflows, reduce interpretation time, and ultimately improve patient outcomes [223]. Deep learning algorithms integrated with NLP and computer vision can foster multimodal medical data analysis and clinical decision-support systems, leading to improvements in patient care [223]. CL can also play a crucial role in the development of personalized medicine approaches by enabling models to adapt to individual patient characteristics and treatment responses over time. This ensures the AI models remain relevant and effective as new data become available [223,224].
Moreover, the static nature of conventional AI models poses a significant challenge in dynamic medical environments where conditions and protocols are ever-evolving [225]. CL enables models to adapt seamlessly to new data distributions, accommodating the introduction of novel diseases, updated diagnostic criteria, and evolving treatment modalities, thereby bolstering the long-term reliability and effectiveness of AI-driven medical solutions [223]. Deep learning algorithms, trained on extensive datasets, possess the capability to recognize intricate patterns and features that may elude the human eye, offering new insights to enhance decision-making. CL facilitates the integration of new imaging modalities or biomarkers into existing diagnostic workflows, enhancing the comprehensiveness and accuracy of clinical decision-making [226]. Despite the documented efficacy of ML and deep learning models in improving the accuracy of breast cancer diagnostics, challenges persist regarding the generalizability and robustness of these models across diverse medical imaging modalities [227]. CL addresses the critical need for models to adapt to new clinical guidelines and emerging medical knowledge, ensuring that diagnostic tools remain current and aligned with best practices. Representative Benchmarks and Evaluation: In healthcare and medical imaging, CL is commonly evaluated using benchmark datasets such as BraTS for brain tumor segmentation, ISIC for skin lesion classification, CheXpert and ChestX-ray14 for thoracic disease diagnosis, CAMUS and EchoNet-Dynamic for echocardiographic analysis, and retinal image datasets including APTOS and EyePACS for diabetic retinopathy assessment. Depending on the application, evaluation metrics typically include classification accuracy, Dice similarity coefficient, Intersection over Union (IoU), sensitivity, specificity, area under the ROC curve (AUC), and the 95th percentile Hausdorff Distance (HD95) for segmentation tasks. CL studies additionally report CL metrics such as average accuracy, average forgetting, forward transfer, and backward transfer to quantify knowledge retention and adaptation throughout sequential learning. Table 26 depicts the CL benchmarks and evaluation metrics across major applications. Recent studies further demonstrate the practical impact of CL in medical imaging and healthcare applications. Ref. [228] investigated continual adaptation strategies for medical image analysis, showing that lifelong learning can effectively reduce catastrophic forgetting across sequential diagnostic tasks. Derakhshani et al. [229] proposed CL techniques for medical image segmentation under evolving clinical data distributions, highlighting the importance of maintaining previously acquired anatomical knowledge while adapting to new imaging modalities. Furthermore, Srivastava et al. [230] introduced the LifeLonger benchmark, providing a standardized evaluation framework for lifelong medical image analysis and facilitating reproducible comparison of CL algorithms in healthcare. These studies demonstrate that CL has become increasingly important for clinical decision-support systems, where data distributions evolve continuously due to new diseases, imaging devices, and patient populations.

10.2. Applications in Robotics and Autonomous Systems

Robotics and autonomous systems represent another promising area for the application of CL, with the potential to enable robots to adapt to changing environments, learn new skills, and improve their performance over time. Robots operating in real-world environments encounter a myriad of challenges, including dynamic environments, unexpected obstacles, and the need to interact with humans in a safe and intuitive manner. Catastrophic forgetting can be mitigated in robots by leveraging the locality of splines. CL algorithms can enable robots to overcome these challenges by continuously learning from their experiences and adapting their behavior accordingly. For instance, a robot deployed in a warehouse environment can learn to navigate new layouts, identify new objects, and optimize its path planning strategies through CL. In the context of autonomous driving, CL can be used to improve the perception, decision-making, and control capabilities of self-driving cars [231]. This locality ensures that the coefficients that are in far-away regions will have information about the data that needs to be preserved. Moreover, CL can enable robots to learn new skills and tasks through imitation learning or RL. CL addresses the challenge of robots operating in dynamic and unpredictable environments by enabling them to continuously adapt to new situations and tasks. Robots can also learn from human demonstrations or feedback, allowing them to acquire new skills more efficiently [232]. CL can address the problem of catastrophic forgetting, where robots lose previously learned skills when learning new ones [10]. CL offers a solution to the challenge of acquiring vast datasets for robotic learning by enabling robots to learn incrementally from limited data, thereby reducing the burden of data collection and annotation [232]. Representative Benchmarks and Evaluation: Robotics research frequently evaluates CL using simulation environments such as Meta-World, RoboSuite, ManiSkill, Habitat, and OpenAI Gym, together with continual reinforcement learning benchmarks including Procgen and Atari. Performance is commonly measured using task success rate, cumulative reward, navigation accuracy, sample efficiency, adaptation speed, and catastrophic forgetting after sequential task acquisition. These benchmarks better reflect the long-term adaptation requirements encountered in autonomous robots operating in dynamic environments.

10.3. Application in Natural Language Processing

NLP stands to gain significantly from CL, particularly in scenarios involving evolving language patterns, emerging topics, and the need to adapt to diverse user preferences. In domains such as social media monitoring and customer service, the language used by individuals is constantly evolving, necessitating models that can adapt to new words, phrases, and sentiments without forgetting previously learned information [10]. CL enables NLP models to stay current with the latest trends and maintain their accuracy over time. For instance, CL can be used to train sentiment analysis models that can accurately classify the sentiment of tweets or customer reviews, even as new slang terms and emojis emerge. As we progress toward an agentic world where LLM-based agents autonomously handle specialized tasks, it becomes crucial for these models to adapt to new tasks without forgetting previously learned information [233].
Moreover, CL can be applied to improve the personalization of NLP models, allowing them to adapt to the specific language preferences and communication styles of individual users. This can enhance the user experience in applications such as chatbots, virtual assistants, and personalized content recommendation systems. CL facilitates the dynamic adaptation of language models to individual learners, providing context-aware and precise responses that cater to their unique needs [234]. CL enables NLP models to effectively address the challenges posed by evolving language patterns, emerging topics, and diverse user preferences. Various CL scenarios have been explored in NLP, including DIL, TIL, CIL, online CL, and continual pretraining [41]. Numerous methods have been adapted to these scenarios and have demonstrated effectiveness, such as: Weight Regularization: RMR-DSE [235] and SRC [236]. KD: ExtendNER [237], CFID [238], CID [239], PAGeR [240], LFPT5 [241], DnR [242], CL-NMT [243], and COKD [244]. Experience Replay: CFID [238], CID [239], ELLE [245], IDBR [246], MBPA++ [247], MetaMBPA++ [248], EMAR [249], DnR [242], ARPER [250], and Total Recall [251]. Generative Replay: PAGeR [240], LAMOL [252], ACM [253], and NER [254]. Parameter Allocation: TPEM [255]. Modular Networks: ProgModel [256]. Meta-Learning: MetaMBPA++ [248], MeLL [257], and CML [258].
CL in NLP is also characterized by the widespread use of pretrained transformer architectures, leading to the development of PEFT techniques. These techniques enable transformers to adapt to new tasks by learning a small number of task-specific parameters. Examples include:
  • Adaptor Tuning: Inserting fully connected layers (CPT [41], CLIF [97], AdapterCL [259], ACM [253], ADA [260]).
  • Prompt Tuning: Using trainable prompt tokens (C-PT [261], LFPT5 [241], EMP [262]).
  • Instruction-Based Approaches: Adding short descriptive text for each task (PAGeR [240], ConTinTin [263], ENTAILMENT [264]).
Given the success of pretrained foundation models, these techniques are increasingly being applied to CL in visual domains [204,205,206,207,260]. Applications of NLP in CL span diverse tasks, creating unique opportunities for further research. Key areas include dialog systems [239,250,255,259]; text classification [246,247,264]; sentence generation [235,250,253]; relation learning [249,258]; neural machine translation [243,244,265]; and named entity recognition [237,241,254]. Additionally, some studies address the integration of vision and language, focusing on continual pretraining [98,266] or downstream tasks [98,267].
Representative Benchmarks and Evaluation: CL for NLP is commonly evaluated using datasets such as CLINC150, Amazon Reviews, AG News, SuperGLUE, and sequential versions of GLUE and XTREME benchmarks. Typical evaluation metrics include classification accuracy, macro F1-score, BLEU, ROUGE, perplexity, and average forgetting across sequential tasks. Recent studies involving LLMs additionally evaluate instruction-following ability, reasoning consistency, and long-context retention after continual adaptation.
Representative Benchmarks and Evaluation: Computer vision remains the most extensively studied application of CL. Representative benchmarks include Split MNIST, Permuted MNIST, Split CIFAR-10, Split CIFAR-100, Tiny ImageNet, ImageNet-100, CORe50, and CORe50 NICv2. Evaluation commonly reports average classification accuracy, forgetting rate, forward transfer, backward transfer, memory consumption, computational cost, and inference latency to assess both learning performance and deployment efficiency.

10.4. Recommender Systems

Recommender systems, which are ubiquitous in e-commerce, entertainment, and other online platforms, can also benefit significantly from CL [268]. Recommender systems use ML algorithms to predict the items or content that a user is most likely to be interested in, based on their past behavior and preferences. Recommender systems traditionally capture user interests by encoding their historical activities on the platforms [269]. However, as user preferences and interests evolve over time, recommender systems need to adapt to these changes in order to maintain their accuracy and relevance.
CL can be used to update recommender systems with new data and user feedback, without forgetting previously learned preferences [270]. Recommender systems may leverage CL to enhance their performance by personalizing recommendations, adapting to new items, and mitigating popularity bias. For example, a movie recommender system can use CL to incorporate new movie releases and user ratings, ensuring that its recommendations remain up to date and aligned with current trends. CL can enable recommender systems to adapt to changing user preferences, new items, and evolving trends, leading to more accurate and personalized recommendations. Adaptive e-learning scenarios powered by CL not only keep learners engaged, but also broaden their awareness of relevant courses [271]. Personalized services that cater to learner preferences can enhance the learning experience [272].

10.5. Cybersecurity

The field of cybersecurity faces a continuous stream of novel threats and attack vectors, requiring security systems to constantly adapt and learn in order to stay ahead of malicious actors. CL offers a promising approach to address this challenge by enabling security systems to learn from new attack patterns and vulnerabilities without forgetting previously learned knowledge. For instance, CL can be used to train intrusion detection systems that can identify new types of malware and network attacks, even if they differ significantly from previously seen threats. By leveraging the locality of splines, Kolmogorov–Arnold Networks [273] can avoid catastrophic forgetting. In cybersecurity, this is important as new threats emerge constantly. Furthermore, CL can be applied to improve the accuracy and efficiency of spam filters, fraud detection systems, and other security applications. CL algorithms, such as Knowledge-augmented neural networks, demonstrate the potential to mitigate catastrophic forgetting in neural networks.
The ability to learn continuously is critical for adapting to evolving adaptation spaces [221]. The B-splines, which are piecewise polynomial functions, offer local control. By continually updating models with new data and feedback, security systems can improve their ability to detect and prevent cyberattacks, protecting sensitive data and infrastructure. Cybersecurity systems that use CL can dynamically adjust their defense mechanisms in response to emerging threats, offering a more robust and adaptive security posture [274]. Traditional security techniques often struggle to adapt to new threats [275]. Intrusion detection systems can use ML and deep learning to proactively prevent persistent and complex external attacks [276].
Standard deep learning methods lose their ability to learn with extended training on new data, a phenomenon known as loss of plasticity [277]. ML has the potential to significantly improve the speed and accuracy of threat detection, making it a powerful tool in the fight against cybercrime [278]. As commercial and open-source software developers improve the security of their products and organizations implement sophisticated threat detection systems, attackers are expected to use increasingly sophisticated methods to infiltrate networks [279,280]. ML-based systems have been shown to outperform traditional, human-based security monitoring systems, especially with the increasing demand for security [281].
In summary, CL enables systems in healthcare, robotics, NLP, recommender systems, and cybersecurity to adapt to new information, evolving environments, and user needs. By overcoming the limitations of traditional static models, CL enhances the efficiency, relevance, and resilience of intelligent systems, paving the way for their integration into real-world applications.

11. Open Challenges and Future Directions

Despite substantial progress, CL remains far from a mature solution for real-world adaptive intelligence. Many existing methods perform well under controlled benchmark settings but struggle with long task sequences, severe domain shifts, limited memory, privacy restrictions, unclear task boundaries, and large-scale deployment constraints. Future CL research should therefore move beyond simplified academic benchmarks and focus on scalable, reproducible, and deployment-oriented learning systems.

11.1. Catastrophic Forgetting and Long-Term Knowledge Retention

Catastrophic forgetting remains the central challenge in CL. It occurs when learning new information overwrites knowledge acquired from previous tasks. Although replay, regularization, knowledge distillation, and parameter-isolation methods have reduced forgetting in many settings, their effectiveness often decreases as task sequences become longer and more heterogeneous. Future work should focus on long-term retention under realistic task streams, where task boundaries are unclear, data distributions evolve gradually, and previous data may be unavailable due to privacy or storage constraints.

11.2. Scalability and Realistic Streaming Benchmarks

Many CL studies still rely on artificial task splits created from static datasets. While useful for controlled comparison, these settings do not fully represent real-world data streams, where new classes, domains, and concepts may emerge continuously. Future benchmarks should include realistic streaming conditions, temporal distribution shifts, noisy labels, open-set categories, class imbalance, and limited annotation availability. Evaluation protocols should clearly report task order, memory budget, validation strategy, computational cost, and variance across multiple runs. Without such standardized reporting, comparisons across CL methods remain unreliable.

11.3. Memory, Computation, and Deployment Constraints

Practical CL systems must balance knowledge retention with memory and computational efficiency. Replay-based methods often provide strong performance, but storing past samples can be infeasible in privacy-sensitive or resource-limited environments. Architecture-based methods reduce interference but may suffer from uncontrolled parameter growth. Regularization-based methods are memory-efficient but can struggle under severe distribution shifts. Future work should prioritize lightweight CL methods that support efficient training, low-latency inference, energy-aware deployment, and stable performance on edge devices, robotics platforms, wearable systems, and embedded AI applications.

11.4. Continual Learning for Foundation Models

Foundation models, including LLMs, vision–language models, and large-scale viTs, have changed the landscape of modern AI. However, most CL research still focuses on smaller models and controlled datasets. Applying CL to foundation models introduces new challenges because full fine-tuning is computationally expensive and may damage pretrained general knowledge, while freezing the model limits adaptation. Future research should investigate how foundation models can continuously acquire new knowledge while preserving their broad generalization ability, reasoning capacity, and cross-task transfer performance. In foundation-model-based CL, two important adaptation settings are continual pretraining and continual instruction tuning [218]. Continual pretraining updates a pretrained model using newly available unlabeled corpora from emerging domains, time periods, or specialized fields such as medicine, law, finance, and scientific literature [210,218]. The goal is to incorporate new knowledge without erasing the broad linguistic, visual, or multimodal representations acquired during original pretraining [209,218]. Continual instruction tuning further adapts LLMs to new instruction-following datasets, requiring the model to preserve earlier instruction-following, reasoning, dialog, and task-completion capabilities while learning new user intents and task formats [218]. These settings differ from conventional classification-based CL because forgetting may appear not only as reduced accuracy, but also as degradation in factual recall, reasoning consistency, safety alignment, multilingual ability, code generation, and instruction-following behavior.

11.5. Parameter-Efficient Continual Adaptation

PEFT methods, including adapters, LoRA, prefix tuning, and prompt-based learning, offer a promising direction for scalable CL. These methods update only a small subset of parameters and can reduce interference with pretrained knowledge. However, their long-term behavior in continual settings remains insufficiently understood. Important questions include how to select, expand, merge, or prune task-specific prompts and adapters over long task sequences. Future studies should also examine whether PEFT-based CL remains stable when task identity is unavailable during inference or when tasks overlap substantially. Parameter isolation techniques are particularly important for continual adaptation of foundation models because they avoid modifying the full pretrained backbone. Instead, task-specific components such as adapters, prompts, prefixes, low-rank matrices, or sparse trainable modules are introduced while most pretrained parameters remain frozen [206,207]. LoRA and its variants have become especially influential in this setting because they reduce trainable parameters by learning low-rank updates to selected weight matrices [210]. Recent variants such as AdaLoRA [282], QLoRA [283], DoRA [284], and LongLoRA [285] further improve adaptation efficiency by adjusting rank allocation, reducing memory usage, improving low-rank representation quality, or supporting long-context adaptation. However, their CL behavior remains unresolved, since sequentially adding adapters or LoRA modules may increase long-term storage, complicate inference routing, and introduce interference among task-specific modules.

11.6. Multimodal Continual Learning

Most existing CL studies focus on single-modality data, especially image classification. However, real-world intelligent systems often process multiple modalities, including images, text, audio, video, sensor signals, and clinical records. Multimodal CL introduces additional complexity because forgetting may occur within individual modalities, across modalities, or in the alignment between modalities. Future research should investigate how to preserve cross-modal representations while allowing each modality to adapt to new distributions. This direction is particularly important for vision–language models, embodied AI, medical AI, and human-centered intelligent systems. The ongoing development of multimodal foundation models further expands the scope of CL. Models such as CLIP-style vision–language systems, visual instruction-following models, and large multimodal assistants must preserve alignment between modalities while continuously adapting to new visual concepts, textual instructions, domains, and downstream tasks [286]. A major challenge is multimodal alignment collapse, where updates in one modality distort the shared representation space and weaken image-text retrieval, visual reasoning, grounding, or cross-modal generation [209,218]. Future multimodal CL should therefore investigate modality-specific adapters, cross-modal replay, alignment-preserving regularization, modular routing, and parameter-efficient adaptation strategies that maintain stable multimodal representations over long task sequences [210].

11.7. Privacy-Preserving and Federated Continual Learning

In many practical applications, previous data cannot be stored or replayed because of privacy, legal, or institutional restrictions. This is especially important in healthcare, finance, mobile devices, and distributed edge systems. Federated CL offers one possible solution by allowing models to learn across distributed clients without centralizing raw data. However, it introduces additional challenges, including client drift, heterogeneous data distributions, communication cost, and forgetting across clients. Future work should develop privacy-preserving CL methods that jointly address knowledge retention, data protection, fairness, and communication efficiency.

11.8. CL in Medical Imaging Under Domain Shift

Medical imaging is a critical but difficult application area for CL. Models deployed in clinical environments must adapt to new scanners, imaging protocols, patient populations, annotation styles, and disease categories. However, medical CL is constrained by limited annotations, strict privacy regulations, class imbalance, and high safety requirements. Future studies should develop domain-aware CL protocols for segmentation, detection, diagnosis, and prognosis tasks. More attention is also needed on clinically meaningful evaluation, including cross-institution robustness, uncertainty estimation, failure analysis, and retention of performance on rare but clinically important cases.

11.9. Ethical, Fairness, and Safety Considerations

As CL systems continuously adapt, they may amplify biases, become harder to audit, or develop unintended behavior after deployment. These risks are especially serious in sensitive domains such as healthcare, finance, autonomous driving, and public services. Future CL systems should include mechanisms for bias monitoring, privacy protection, interpretability, uncertainty estimation, and human oversight. Responsible CL should not only optimize accuracy and forgetting metrics but also ensure fairness, transparency, robustness, and safety during long-term adaptation.

11.10. Toward Robust Lifelong AI

The long-term goal of CL is to support lifelong AI systems that can acquire, organize, and refine knowledge over extended periods. Achieving this goal requires methods that integrate stable memory, flexible adaptation, efficient resource use, and reliable decision-making. Future research should combine insights from ML, neuroscience, cognitive science, ethics, and human-centered AI. Progress will depend not only on stronger algorithms but also on better benchmarks, clearer evaluation standards, and more realistic deployment studies.

11.11. Critical Analysis of Current Continual Learning Approaches

Despite remarkable progress in CL, current methods remain far from achieving robust lifelong learning comparable to human intelligence. Most existing approaches demonstrate strong performance on controlled benchmark datasets but often exhibit limited scalability, high computational cost, and reduced robustness when deployed in realistic environments. Several fundamental limitations remain unresolved.
First, although PEFT methods substantially reduce the number of trainable parameters compared with full model fine-tuning, their scalability over long CL sequences remains insufficiently understood. As the number of sequential tasks increases, task-specific prompts, adapters, or low-rank parameter updates continue to accumulate, resulting in increasing storage requirements and growing management complexity. Furthermore, prompt interference may occur when prompts learned for earlier tasks conflict with those required for subsequent tasks, reducing both adaptation efficiency and knowledge transfer. These issues become increasingly significant for foundation models containing billions of parameters.
Second, even the most competitive CL methods remain susceptible to catastrophic forgetting and negative transfer during long task sequences. Replay-based methods generally achieve superior knowledge retention but require substantial replay memory and may violate privacy constraints. Regularization-based approaches are computationally efficient but often struggle under severe domain shifts. Architecture-based methods reduce forgetting by allocating task-specific parameters; however, continuous model expansion increases inference latency, parameter count, and deployment cost. Consequently, no existing method provides an ideal balance among stability, plasticity, scalability, and computational efficiency.
Third, robustness under domain shift remains a significant challenge. Many CL algorithms are evaluated on carefully designed academic benchmarks with relatively moderate distribution changes. In practical applications, however, evolving environments may involve simultaneous domain shifts, class expansion, sensor degradation, changing data quality, and previously unseen concepts. Such complex distribution shifts frequently degrade representation quality and increase catastrophic forgetting, highlighting the need for more realistic CL benchmarks.
Fourth, computational complexity remains an important barrier to large-scale deployment. Replay mechanisms require repeated optimization using both historical and current samples, architecture-based methods increase parameter counts over time, and foundation-model adaptation introduces substantial computational overhead despite parameter-efficient optimization. Future studies should therefore report computational complexity, training time, inference latency, GPU memory consumption, and energy efficiency alongside predictive performance to facilitate fair comparison among CL methods.
Finally, practical deployment introduces additional constraints that remain insufficiently addressed in the existing literature. Real-world CL systems must operate under limited memory, strict latency requirements, privacy regulations, unreliable network connectivity, and evolving data distributions. Healthcare, robotics, autonomous driving, and edge intelligence applications require continual adaptation while maintaining reliability, interpretability, and safety. Bridging the gap between controlled benchmark evaluation and deployment in continuously evolving environments therefore represents one of the most important research directions for future CL systems. Figure 7 illustrates representative continual learning strategies that preserve knowledge at the data, feature, and label levels through replay and knowledge distillation.

11.12. Summary

Future progress in CL will depend on shifting from simplified benchmark performance toward robust lifelong adaptation. The most important directions include long-term knowledge retention, realistic streaming evaluation, foundation-model-based CL, parameter-efficient adaptation, multimodal learning, privacy-preserving methods, and medical CL under domain shift. Addressing these challenges is essential for developing CL systems that are accurate, scalable, safe, and reliable in real-world environments. Table 27 presents a comparative analysis of the principal challenges in continual learning, highlighting their practical impacts and commonly adopted solution strategies.

12. Conclusions

This review provides a comprehensive synthesis of modern CL, covering theoretical foundations, learning scenarios, methodological paradigms, evaluation protocols, emerging foundation-model adaptation techniques, and diverse real-world applications. Rather than viewing CL as a single solution to catastrophic forgetting, our analysis demonstrates that it has evolved into a broad research paradigm addressing the challenges of lifelong adaptation in increasingly dynamic AI systems. Several key insights emerge from this survey. First, no single CL strategy consistently outperforms all others across different tasks, datasets, and evaluation settings. The effectiveness of replay-based, regularization-based, optimization-based, architecture-based, and parameter-efficient methods depends strongly on the learning scenario, computational constraints, memory budget, and application requirements. Consequently, selecting an appropriate CL strategy requires careful consideration of the deployment environment rather than relying solely on benchmark accuracy.
Second, the lack of standardized evaluation protocols remains one of the major obstacles to fair comparison among CL methods. Differences in benchmark construction, task ordering, replay memory budgets, backbone architectures, hyperparameter tuning, and evaluation metrics often make reported results difficult to compare directly. Establishing unified benchmarking practices, standardized reporting guidelines, and reproducible experimental protocols will therefore be essential for accelerating progress in the field. Third, foundation models have fundamentally changed the direction of CL research. Prompt learning, PEFT, adapters, LoRA-based methods, and replay-free adaptation strategies demonstrate that continual adaptation can increasingly be achieved without updating the entire model. However, new challenges have emerged, including prompt interference, adapter accumulation, continual alignment, long-term storage growth, and efficient adaptation of large multimodal models. Fourth, CL is rapidly expanding beyond traditional image classification toward practical deployment in healthcare, robotics, autonomous systems, cybersecurity, recommender systems, NLP, scientific discovery, and multimodal foundation models. Future CL systems must therefore optimize not only predictive performance but also computational efficiency, memory consumption, robustness, privacy preservation, interpretability, and deployment scalability.
Finally, several important research directions remain open. These include continual reasoning for LLMs, continual instruction tuning, lifelong multimodal representation learning, adaptive memory management, privacy-preserving CL, standardized evaluation frameworks, trustworthy continual AI, and scalable lifelong foundation models capable of continuous adaptation over extended time horizons. Addressing these challenges will be critical for developing intelligent systems that can learn continuously while maintaining reliable performance in real-world environments.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/math14152774/s1, Supplementary Material S1: PRISMA 2020 Checklist.

Author Contributions

Conceptualization, Z.U. and J.K.; Methodology, Z.U.; Formal analysis, Z.U. and J.K.; Investigation, Z.U. and J.K.; Writing—original draft, Z.U.; Writing—review and editing, Z.U.; Visualization, M.H.; Funding acquisition, J.K.; Supervision, J.K.; Project administration, J.K. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the MSIT (Ministry of Science and ICT), Korea, under the Artificial Intelligence Convergence Innovation Human Resources Development (IITP-2026-RS-2023-00254592), and the Global Research Support Program in the Digital Field (RS-2024-00426860), supervised by the IITP (Institute for Information & Communications Technology Planning & Evaluation).

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Russell, S.; Norvig, P. Artificial Intelligence: A Modern Approach, 4th ed.; Pearson: Hoboken, NJ, USA, 2020. [Google Scholar]
  2. Annan, R.; Qingge, L. Artificial intelligence in COVID-19 research: A comprehensive survey of innovations, challenges, and future directions. Comput. Sci. Rev. 2025, 57, 100751. [Google Scholar] [CrossRef] [Scilit]
  3. Herrera, F. Reflections and attentiveness on eXplainable Artificial Intelligence (XAI). The journey ahead from criticisms to human–AI collaboration. Inf. Fusion 2025, 121, 103133. [Google Scholar] [CrossRef] [Scilit]
  4. Utomo, S.; Pratap, A.; Karthikeyan, P.; Ayeelyan, J.; Hsu, H.C.; Hsiung, P.A. When explainable artificial intelligence meets data governance: Enhancing trustworthiness in multimodal gas classification. Inf. Fusion 2025, 125, 103440. [Google Scholar]
  5. Longo, L.; Brcic, M.; Cabitza, F.; Choi, J.; Confalonieri, R.; Del Ser, J.; Guidotti, R.; Hayashi, Y.; Herrera, F.; Holzinger, A.; et al. Explainable Artificial Intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions. Inf. Fusion 2024, 106, 102301. [Google Scholar] [CrossRef] [Scilit]
  6. Naser, M. From failure to fusion: A survey on learning from bad machine learning models. Inf. Fusion 2025, 120, 103122. [Google Scholar] [CrossRef] [Scilit]
  7. Escovedo, T.; Koshiyama, A.; da Cruz, A.A.; Vellasco, M. Neuroevolutionary learning in nonstationary environments. Appl. Intell. 2020, 50, 1590–1608. [Google Scholar] [CrossRef] [Scilit]
  8. Criado, M.F.; Casado, F.E.; Iglesias, R.; Regueiro, C.V.; Barro, S. Non-iid data and continual learning processes in federated learning: A long road ahead. Inf. Fusion 2022, 88, 263–280. [Google Scholar] [CrossRef] [Scilit]
  9. Nguyen, C.V.; Achille, A.; Lam, M.; Hassner, T.; Mahadevan, V.; Soatto, S. Toward understanding catastrophic forgetting in continual learning. arXiv 2019, arXiv:1908.01091. [Google Scholar]
  10. Parisi, G.; Kemker, R.; Part, J.; Kanan, C.; Wermter, S. Continual lifelong learning with neural networks: A review. Neural Netw. Off. J. Int. Neural Netw. Soc. 2019, 113, 54–71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Wang, L.; Zhang, X.; Su, H.; Zhu, J. A comprehensive survey of continual learning: Theory, method and application. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 5362–5383. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Kirkpatrick, J.; Pascanu, R.; Rabinowitz, N.; Veness, J.; Desjardins, G.; Rusu, A.A.; Milan, K.; Quan, J.; Ramalho, T.; Grabska-Barwinska, A.; et al. Overcoming catastrophic forgetting in neural networks. Proc. Natl. Acad. Sci. USA 2017, 114, 3521–3526. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Xu, X.; Chen, J.; Thakur, D.; Hong, D. Multi-modal disease segmentation with continual learning and adaptive decision fusion. Inf. Fusion 2025, 118, 102962. [Google Scholar] [CrossRef] [Scilit]
  14. Wu, Y.; Li, Z.; Gao, Y.; Chiclana, F.; Chen, X.; Dong, Y. An endogenous and continual learning approach to personalize individual semantics to support linguistic consensus reaching. Inf. Fusion 2025, 114, 102640. [Google Scholar] [CrossRef] [Scilit]
  15. Shahrivari, S. Beyond batch processing: Towards real-time and streaming big data. Computers 2014, 3, 117–129. [Google Scholar] [CrossRef] [Scilit]
  16. Parisi, G.I.; Lomonaco, V. Online continual learning on sequences. In Proceedings of the Recent Trends in Learning From Data: Tutorials from the INNS Big Data and Deep Learning Conference (INNSBDDL2019); Springer: Berlin/Heidelberg, Germany, 2020; pp. 197–221. [Google Scholar]
  17. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. van de Ven, G.M.; Tuytelaars, T.; Tolias, A.S. Three types of incremental learning. Nat. Mach. Intell. 2022, 4, 1185–1197. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Bidaki, S.A.; Mohammadkhah, A.; Rezaee, K.; Hassani, F.; Eskandari, S.; Salahi, M.; Ghassemi, M.M. Online continual learning: A systematic literature review of approaches, challenges, and benchmarks. arXiv 2025, arXiv:2501.04897. [Google Scholar]
  20. Zhou, D.W.; Wang, Q.W.; Qi, Z.H.; Ye, H.J.; Zhan, D.C.; Liu, Z. Class-incremental learning: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 9851–9873. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Wickramasinghe, B.; Saha, G.; Roy, K. Continual learning: A review of techniques, challenges, and future directions. IEEE Trans. Artif. Intell. 2023, 5, 2526–2546. [Google Scholar]
  22. Thrun, S.; Mitchell, T.M. Lifelong robot learning. Robot. Auton. Syst. 1995, 15, 25–46. [Google Scholar] [CrossRef] [Scilit]
  23. Tan, A.; Wang, Y.; Wu, W.Z.; Ding, W.; Liang, J. Multi-View Fusion Graph Attention Network for multilabel class incremental learning. Inf. Fusion 2025, 123, 103309. [Google Scholar] [CrossRef] [Scilit]
  24. Li, D.; Wang, T.; Chen, J.; Kawaguchi, K.; Lian, C.; Zeng, Z. Multi-view class incremental learning. Inf. Fusion 2024, 102, 102021. [Google Scholar] [CrossRef] [Scilit]
  25. Zheng, Y.; Zhang, X.; Tian, Z.; Du, S. Enhancing few-shot lifelong learning through fusion of cross-domain knowledge. Inf. Fusion 2025, 115, 102730. [Google Scholar] [CrossRef] [Scilit]
  26. Mehta, S.V.; Patil, D.; Chandar, S.; Strubell, E. An empirical investigation of the role of pre-training in lifelong learning. J. Mach. Learn. Res. 2023, 24, 1–50. [Google Scholar] [CrossRef] [Scilit]
  27. Kanakis, M.; Bruggemann, D.; Saha, S.; Georgoulis, S.; Obukhov, A.; Van Gool, L. Reparameterizing convolutions for incremental multi-task learning without task interference. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part XX 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 689–707. [Google Scholar]
  28. Vödisch, N.; Cattaneo, D.; Burgard, W.; Valada, A. Covio: Online continual learning for visual-inertial odometry. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 2464–2473. [Google Scholar]
  29. Ullah, Z.; Usman, M.; Gwak, J. MTSS-AAE: Multi-task semi-supervised adversarial autoencoding for COVID-19 detection based on chest X-ray images. Expert Syst. Appl. 2023, 216, 119475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Bonicelli, L.; Boschini, M.; Frascaroli, E.; Porrello, A.; Pennisi, M.; Bellitto, G.; Palazzo, S.; Spampinato, C.; Calderara, S. On the effectiveness of equivariant regularization for robust online continual learning. arXiv 2023, arXiv:2305.03648. [Google Scholar]
  31. Ali, S.; Abuhmed, T.; El-Sappagh, S.; Muhammad, K.; Alonso-Moral, J.M.; Confalonieri, R.; Guidotti, R.; Del Ser, J.; Díaz-Rodríguez, N.; Herrera, F. Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence. Inf. Fusion 2023, 99, 101805. [Google Scholar] [CrossRef] [Scilit]
  32. Abbass, H. What is artificial intelligence? IEEE Trans. Artif. Intell. 2021, 2, 94–95. [Google Scholar] [CrossRef] [Scilit]
  33. Smith, P.D. Hands-On Artificial Intelligence for Beginners: An Introduction to AI Concepts, Algorithms, and Their Implementation; Packt Publishing Ltd.: Birmingham, UK, 2018. [Google Scholar]
  34. Chen, Z.; Liu, B. Lifelong Machine Learning; Morgan & Claypool Publishers: San Rafael, CA, USA, 2018. [Google Scholar]
  35. Yang, Y.; Zhou, J.; Ding, X.; Huai, T.; Liu, S.; Chen, Q.; Xie, Y.; He, L. Recent advances of foundation language models-based continual learning: A survey. ACM Comput. Surv. 2025, 57, 1–38. [Google Scholar] [CrossRef] [Scilit]
  36. Kudithipudi, D.; Aguilar-Simon, M.; Babb, J.; Bazhenov, M.; Blackiston, D.; Bongard, J.; Brna, A.P.; Chakravarthi Raja, S.; Cheney, N.; Clune, J.; et al. Biological underpinnings for lifelong learning machines. Nat. Mach. Intell. 2022, 4, 196–210. [Google Scholar] [CrossRef] [Scilit]
  37. Hadsell, R.; Rao, D.; Rusu, A.A.; Pascanu, R. Embracing change: Continual learning in deep neural networks. Trends Cogn. Sci. 2020, 24, 1028–1040. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  38. Qu, H.; Rahmani, H.; Xu, L.; Williams, B.; Liu, J. Recent advances of continual learning in computer vision: An overview. IET Comput. Vis. 2025, 19, e70013. [Google Scholar] [CrossRef] [Scilit]
  39. Masana, M.; Twardowski, B.; Van de Weijer, J. On class orderings for incremental learning. arXiv 2020, arXiv:2007.02145. [Google Scholar]
  40. Biesialska, M.; Biesialska, K.; Costa-Jussa, M.R. Continual lifelong learning in natural language processing: A survey. In Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain, 8–13 December 2020; pp. 6523–6541. [Google Scholar]
  41. Ke, Z.; Liu, B. Continual learning of natural language processing tasks: A survey. arXiv 2022, arXiv:2211.12701. [Google Scholar]
  42. Khetarpal, K.; Riemer, M.; Rish, I.; Precup, D. Towards continual reinforcement learning: A review and perspectives. J. Artif. Intell. Res. 2022, 75, 1401–1476. [Google Scholar] [CrossRef] [Scilit]
  43. Ghosh, S. Dynamic vaes with generative replay for continual zero-shot learning. arXiv 2021, arXiv:2104.12468. [Google Scholar]
  44. Singh, P.; Mazumder, P.; Rai, P.; Namboodiri, V.P. Rectification-based knowledge retention for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 15282–15291. [Google Scholar]
  45. Tao, X.; Hong, X.; Chang, X.; Dong, S.; Wei, X.; Gong, Y. Few-shot class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 12183–12192. [Google Scholar]
  46. Wang, L.; Yang, K.; Li, C.; Hong, L.; Li, Z.; Zhu, J. Ordisco: Effective and efficient usage of incremental unlabeled data for semi-supervised continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 5383–5392. [Google Scholar]
  47. Joseph, K.; Khan, S.; Khan, F.S.; Balasubramanian, V.N. Towards open world object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 5830–5840. [Google Scholar]
  48. Wang, Q.F.; Geng, X.; Lin, S.X.; Xia, S.Y.; Qi, L.; Xu, N. Learngene: From open-world to your learning task. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2022; Volume 36, pp. 8557–8565. [Google Scholar]
  49. Hu, D.; Yan, S.; Lu, Q.; Hong, L.; Hu, H.; Zhang, Y.; Li, Z.; Wang, X.; Feng, J. How well does self-supervised pre-training perform with streaming data? arXiv 2021, arXiv:2104.12081. [Google Scholar]
  50. Rao, D.; Visin, F.; Rusu, A.; Pascanu, R.; Teh, Y.W.; Hadsell, R. Continual unsupervised representation learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  51. Ruvolo, P.; Eaton, E. ELLA: An efficient lifelong learning algorithm. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2013; pp. 507–515. [Google Scholar]
  52. Masse, N.Y.; Grant, G.D.; Freedman, D.J. Alleviating catastrophic forgetting using context-dependent gating and synaptic stabilization. Proc. Natl. Acad. Sci. USA 2018, 115, E10467–E10475. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  53. Ramesh, R.; Chaudhari, P. Model zoo: A growing “brain” that learns continually. arXiv 2021, arXiv:2106.03027. [Google Scholar]
  54. PourKeshavarzi, M.; Zhao, G.; Sabokrou, M. Looking back on learned experiences for class/task incremental learning. In Proceedings of the International Conference on Learning Representations; OpenReview.net: Newton Highlands, MA, USA, 2021. [Google Scholar]
  55. Xie, X.; Xu, J.; Hu, P.; Zhang, W.; Huang, Y.; Zheng, W.; Wang, R. Task-incremental medical image classification with task-specific batch normalization. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision (PRCV); Springer: Berlin/Heidelberg, Germany, 2023; pp. 309–320. [Google Scholar]
  56. Feng, F.; Chan, R.H.; Shi, X.; Zhang, Y.; She, Q. Challenges in task incremental learning for assistive robotics. IEEE Access 2019, 8, 3434–3441. [Google Scholar]
  57. Lopez-Paz, D.; Ranzato, M. Gradient episodic memory for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  58. Vogelstein, J.T.; Dey, J.; Helm, H.S.; LeVine, W.; Mehta, R.D.; Tomita, T.M.; Xu, H.; Geisa, A.; Wang, Q.; van de Ven, G.M.; et al. A Simple Lifelong Learning Approach. arXiv 2020, arXiv:2004.12908. [Google Scholar]
  59. Ke, Z.; Liu, B.; Xu, H.; Shu, L. CLASSIC: Continual and contrastive learning of aspect sentiment classification tasks. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 6871–6883. [Google Scholar]
  60. Mirza, M.J.; Masana, M.; Possegger, H.; Bischof, H. An efficient domain-incremental learning approach to drive in all weather conditions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 3001–3011. [Google Scholar]
  61. Aljundi, R.; Chakravarty, P.; Tuytelaars, T. Expert gate: Lifelong learning with a network of experts. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 3366–3375. [Google Scholar]
  62. Von Oswald, J.; Henning, C.; Grewe, B.F.; Sacramento, J. Continual learning with hypernetworks. arXiv 2019, arXiv:1906.00695. [Google Scholar]
  63. Lomonaco, V.; Maltoni, D. Core50: A new dataset and benchmark for continuous object recognition. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2017; pp. 17–26. [Google Scholar]
  64. Garg, P.; Saluja, R.; Balasubramanian, V.N.; Arora, C.; Subramanian, A.; Jawahar, C. Multi-domain incremental learning for semantic segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2022; pp. 761–771. [Google Scholar]
  65. Capuano, N.; Greco, L.; Ritrovato, P.; Vento, M. Sentiment analysis for customer relationship management: An incremental learning approach. Appl. Intell. 2021, 51, 3339–3352. [Google Scholar]
  66. Rebuffi, S.A.; Kolesnikov, A.; Sperl, G.; Lampert, C.H. icarl: Incremental classifier and representation learning. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 2001–2010. [Google Scholar]
  67. Shin, H.; Lee, J.K.; Kim, J.; Kim, J. Continual learning with deep generative replay. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  68. van de Ven, G.M.; Siegelmann, H.T.; Tolias, A.S. Brain-inspired replay for continual learning with artificial neural networks. Nat. Commun. 2020, 11, 4069. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Zhou, D.W.; Yang, Y.; Zhan, D.C. Learning to classify with incremental new class. IEEE Trans. Neural Netw. Learn. Syst. 2021, 33, 2429–2443. [Google Scholar]
  70. Belouadah, E.; Popescu, A.; Kanellos, I. A comprehensive study of class incremental learning algorithms for visual tasks. Neural Netw. 2021, 135, 38–54. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Masana, M.; Liu, X.; Twardowski, B.; Menta, M.; Bagdanov, A.D.; Van De Weijer, J. Class-incremental learning: Survey and performance evaluation on image classification. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 5513–5533. [Google Scholar]
  72. Channappayya, S.; Tamma, B.R. Augmented memory replay-based continual learning approaches for network intrusion detection. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2023; Volume 36, pp. 17156–17169. [Google Scholar]
  73. Li, X.; Wang, S.; Sun, J.; Xu, Z. Variational data-free knowledge distillation for continual learning. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 12618–12634. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. Krizhevsky, A.; Sutskever, I.; Hinton, G.E. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2012; Volume 25. [Google Scholar]
  75. Shmelkov, K.; Schmid, C.; Alahari, K. Incremental learning of object detectors without catastrophic forgetting. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 3400–3409. [Google Scholar]
  76. Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2015; pp. 1440–1448. [Google Scholar]
  77. Ramakrishnan, K.; Panda, R.; Fan, Q.; Henning, J.; Oliva, A.; Feris, R. Relationship matters: Relation guided knowledge transfer for incremental learning of object detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2020; pp. 250–251. [Google Scholar]
  78. Paik, I.; Oh, S.; Kwak, T.; Kim, I. Overcoming catastrophic forgetting by neuron-level plasticity control. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2020; Volume 34, pp. 5339–5346. [Google Scholar]
  79. Zhou, X.; Wang, D.; Krähenbühl, P. Objects as points. arXiv 2019, arXiv:1904.07850. [Google Scholar]
  80. Li, D.; Tasci, S.; Ghosh, S.; Zhu, J.; Zhang, J.; Heck, L. RILOD: Near real-time incremental learning for object detection at the edge. In Proceedings of the 4th ACM/IEEE Symposium on Edge Computing; ACM: New York, NY, USA, 2019; pp. 113–126. [Google Scholar]
  81. Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
  82. Feng, T.; Wang, M.; Yuan, H. Overcoming catastrophic forgetting in incremental object detection via elastic response distillation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 9427–9436. [Google Scholar]
  83. Li, X.; Wang, W.; Wu, L.; Chen, S.; Hu, X.; Li, J.; Tang, J.; Yang, J. Generalized focal loss: Learning qualified and distributed bounding boxes for dense object detection. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 21002–21012. [Google Scholar]
  84. Hao, Y.; Fu, Y.; Jiang, Y.G.; Tian, Q. An end-to-end architecture for class-incremental object detection with knowledge distillation. In Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME); IEEE: New York, NY, USA, 2019; pp. 1–6. [Google Scholar]
  85. Peng, C.; Zhao, K.; Lovell, B.C. Faster ilod: Incremental learning for object detectors based on faster rcnn. Pattern Recognit. Lett. 2020, 140, 109–115. [Google Scholar] [CrossRef] [Scilit]
  86. Zhang, J.; Zhang, J.; Ghosh, S.; Li, D.; Tasci, S.; Heck, L.; Zhang, H.; Kuo, C.C.J. Class-incremental learning via deep model consolidation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2020; pp. 1131–1140. [Google Scholar]
  87. Dong, N.; Zhang, Y.; Ding, M.; Lee, G.H. Bridging non co-occurrence with unlabeled in-the-wild data for incremental object detection. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 30492–30503. [Google Scholar]
  88. Joseph, K.; Rajasegaran, J.; Khan, S.; Khan, F.S.; Balasubramanian, V.N. Incremental object detection via meta-learning. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 9209–9216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. Ren, S.; He, K.; Girshick, R.; Sun, J. Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2015; Volume 28. [Google Scholar]
  90. Zhao, N.; Lee, G.H. Static-dynamic co-teaching for class-incremental 3d object detection. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2022; Volume 36, pp. 3436–3445. [Google Scholar]
  91. Wang, J.; Wang, X.; Shang-Guan, Y.; Gupta, A. Wanderlust: Online continual object detection in the real world. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 10829–10838. [Google Scholar]
  92. Perez-Rua, J.M.; Zhu, X.; Hospedales, T.M.; Xiang, T. Incremental few-shot object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13846–13855. [Google Scholar]
  93. Feng, J.; Phillips, R.V.; Malenica, I.; Bishara, A.; Hubbard, A.E.; Celi, L.A.; Pirracchio, R. Clinical artificial intelligence quality improvement: Towards continual monitoring and updating of AI algorithms in healthcare. npj Digit. Med. 2022, 5, 66. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Franklin, S. Autonomous agents as embodied AI. Cybern. Syst. 1997, 28, 499–520. [Google Scholar] [CrossRef] [Scilit]
  95. Shi, G.; Wu, Y.; Liu, J.; Wan, S.; Wang, W.; Lu, T. Incremental few-shot semantic segmentation via embedding adaptive-update and hyper-class representation. In Proceedings of the 30th ACM International Conference on Multimedia; ACM: New York, NY, USA, 2022; pp. 5547–5556. [Google Scholar]
  96. Ganea, D.A.; Boom, B.; Poppe, R. Incremental few-shot instance segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 1185–1194. [Google Scholar]
  97. Jin, X.; Lin, B.Y.; Rostami, M.; Ren, X. Learn continually, generalize rapidly: Lifelong knowledge accumulation for few-shot learning. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2021; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 714–729. [Google Scholar]
  98. Cossu, A.; Carta, A.; Passaro, L.; Lomonaco, V.; Tuytelaars, T.; Bacciu, D. Continual pre-training mitigates forgetting in language and vision. Neural Netw. 2024, 179, 106492. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Madaan, D.; Yoon, J.; Li, Y.; Liu, Y.; Hwang, S.J. Representational continuity for unsupervised continual learning. arXiv 2021, arXiv:2110.06976. [Google Scholar]
  100. Riemer, M.; Cases, I.; Ajemian, R.; Liu, M.; Rish, I.; Tu, Y.; Tesauro, G. Learning to learn without forgetting by maximizing transfer and minimizing interference. arXiv 2018, arXiv:1810.11910. [Google Scholar]
  101. Guo, Q.; Zhao, W.; Lyu, Z.; Zhao, T. A GAN enhanced meta-deep reinforcement learning approach for DCN routing optimization. Inf. Fusion 2025, 121, 103160. [Google Scholar] [CrossRef] [Scilit]
  102. Zhao, Y.; Zhong, Z.; Yang, F.; Luo, Z.; Lin, Y.; Li, S.; Sebe, N. Learning to generalize unseen domains via memory-based multi-source meta-learning for person re-identification. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 6277–6286. [Google Scholar]
  103. Javed, K.; White, M. Meta-learning representations for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  104. Beaulieu, S.; Frati, L.; Miconi, T.; Lehman, J.; Stanley, K.O.; Clune, J.; Cheney, N. Learning to continually learn. In ECAI 2020; IOS Press: Amsterdam, The Netherlands, 2020; pp. 992–1001. [Google Scholar]
  105. Lee, E.; Huang, C.H.; Lee, C.Y. Few-shot and continual learning with attentive independent mechanisms. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 9455–9464. [Google Scholar]
  106. Rajasegaran, J.; Khan, S.; Hayat, M.; Khan, F.S.; Shah, M. itaml: An incremental task-agnostic meta-learning approach. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13588–13597. [Google Scholar]
  107. Gupta, G.; Yadav, K.; Paull, L. Look-ahead meta learning for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 11588–11598. [Google Scholar]
  108. Caccia, L.; Belilovsky, E.; Caccia, M.; Pineau, J. Online learned continual compression with adaptive quantization modules. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2020; pp. 1240–1250. [Google Scholar]
  109. KJ, J.; N Balasubramanian, V. Meta-consolidation for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 14374–14386. [Google Scholar]
  110. Henning, C.; Cervera, M.; D’Angelo, F.; Von Oswald, J.; Traber, R.; Ehret, B.; Kobayashi, S.; Grewe, B.F.; Sacramento, J. Posterior meta-replay for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 14135–14149. [Google Scholar]
  111. Hurtado, J.; Raymond, A.; Soto, A. Optimizing reusable knowledge for continual learning via metalearning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 14150–14162. [Google Scholar]
  112. Wang, R.; Bao, Y.; Zhang, B.; Liu, J.; Zhu, W.; Guo, G. Anti-retroactive interference for lifelong learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 163–178. [Google Scholar]
  113. McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; y Arcas, B.A. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2017; pp. 1273–1282. [Google Scholar]
  114. Yoon, J.; Jeong, W.; Lee, G.; Yang, E.; Hwang, S.J. Federated continual learning with weighted inter-client transfer. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 12073–12086. [Google Scholar]
  115. Usmanova, A.; Portet, F.; Lalanda, P.; Vega, G. A distillation-based approach integrating continual learning and federated learning for pervasive services. arXiv 2021, arXiv:2109.04197. [Google Scholar]
  116. Park, T.J.; Kumatani, K.; Dimitriadis, D. Tackling dynamics in federated incremental learning with variational embedding rehearsal. arXiv 2021, arXiv:2110.09695. [Google Scholar]
  117. Mermillod, M.; Bugaiska, A.; Bonin, P. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects. Front. Psychol. 2013, 4, 504. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  118. Grossberg, S. Adaptive Resonance Theory: How a brain learns to consciously attend, learn, and recognize a changing world. Neural Netw. 2013, 37, 1–47. [Google Scholar] [PubMed]
  119. Abraham, W.C.; Robins, A. Memory retention–the synaptic stability versus plasticity dilemma. Trends Neurosci. 2005, 28, 73–78. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  120. Hebb, D.O. The Organization of Behavior: A Neuropsychological Theory; Psychology Press: Hove, UK, 2005. [Google Scholar]
  121. Power, J.D.; Schlaggar, B.L. Neural plasticity across the lifespan. Wiley Interdiscip. Rev. Dev. Biol. 2017, 6, e216. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  122. Bennani, M.A.; Doan, T.; Sugiyama, M. Generalisation guarantees for continual learning with orthogonal gradient descent. arXiv 2020, arXiv:2006.11942. [Google Scholar]
  123. Doan, T.; Bennani, M.A.; Mazoure, B.; Rabusseau, G.; Alquier, P. A theoretical analysis of catastrophic forgetting through the ntk overlap matrix. In Proceedings of the International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2021; pp. 1072–1080. [Google Scholar]
  124. McCloskey, M.; Cohen, N.J. Catastrophic interference in connectionist networks: The sequential learning problem. In Psychology of Learning and Motivation; Elsevier: Amsterdam, The Netherlands, 1989; Volume 24, pp. 109–165. [Google Scholar]
  125. Ratcliff, R. Connectionist models of recognition memory: Constraints imposed by learning and forgetting functions. Psychol. Rev. 1990, 97, 285. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  126. Zhao, J.; Zhang, X.; Zhao, B.; Hu, W.; Diao, T.; Wang, L.; Zhong, Y.; Li, Q. Genetic dissection of mutual interference between two consecutive learning tasks in Drosophila. eLife 2023, 12, e83516. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  127. Hayashi-Takagi, A.; Yagishita, S.; Nakamura, M.; Shirai, F.; Wu, Y.I.; Loshbaugh, A.L.; Kuhlman, B.; Hahn, K.M.; Kasai, H. Labelling and optical erasure of synaptic memory traces in the motor cortex. Nature 2015, 525, 333–338. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  128. Yang, G.; Pan, F.; Gan, W.B. Stably maintained dendritic spines are associated with lifelong memories. Nature 2009, 462, 920–924. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  129. Zhang, X.; Li, Q.; Wang, L.; Liu, Z.J.; Zhong, Y. Active protection: Learning-activated Raf/MAPK activity protects labile memory from Rac1-independent forgetting. Neuron 2018, 98, 142–155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  130. Huszár, F. Note on the quadratic penalties in elastic weight consolidation. Proc. Natl. Acad. Sci. USA 2018, 115, E2496–E2497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  131. McNaughton, B.L.; O’Reilly, R.C. Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of. Psychol. Rev. 1995, 102, 419–457. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  132. Graves, L.; Nagisetty, V.; Ganesh, V. Does AI remember? Neural Networks and the Right to be Forgotten. Master’s Thesis University of Waterloo, Waterloo, ON, Canada, 2020. [Google Scholar]
  133. Ding, M.; Ji, K.; Wang, D.; Xu, J. Understanding forgetting in continual learning with linear regression. arXiv 2024, arXiv:2405.17583. [Google Scholar]
  134. Aljundi, R.; Lin, M.; Goujaud, B.; Bengio, Y. Gradient based sample selection for online continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  135. Ritter, H.; Botev, A.; Barber, D. Online structured laplace approximations for overcoming catastrophic forgetting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
  136. Schwarz, J.; Czarnecki, W.; Luketina, J.; Grabska-Barwinska, A.; Teh, Y.W.; Pascanu, R.; Hadsell, R. Progress & compress: A scalable framework for continual learning. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; pp. 4528–4537. [Google Scholar]
  137. Gou, J.; Yu, B.; Maybank, S.J.; Tao, D. Knowledge distillation: A survey. Int. J. Comput. Vis. 2021, 129, 1789–1819. [Google Scholar] [CrossRef] [Scilit]
  138. Dhar, P.; Singh, R.V.; Peng, K.C.; Wu, Z.; Chellappa, R. Learning without memorizing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 5138–5146. [Google Scholar]
  139. Iscen, A.; Zhang, J.; Lazebnik, S.; Schmid, C. Memory-efficient incremental learning through feature adaptation. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 699–715. [Google Scholar]
  140. Li, Z.; Hoiem, D. Learning without forgetting. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 40, 2935–2947. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  141. Castro, F.M.; Marín-Jiménez, M.J.; Guil, N.; Schmid, C.; Alahari, K. End-to-end incremental learning. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 233–248. [Google Scholar]
  142. Douillard, A.; Cord, M.; Ollion, C.; Robert, T.; Valle, E. Podnet: Pooled outputs distillation for small-tasks incremental learning. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part XX 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 86–102. [Google Scholar]
  143. Hou, S.; Pan, X.; Loy, C.C.; Wang, Z.; Lin, D. Learning a unified classifier incrementally via rebalancing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 831–839. [Google Scholar]
  144. Wu, C.; Herranz, L.; Liu, X.; Van De Weijer, J.; Raducanu, B. Memory replay gans: Learning to generate new categories without forgetting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2018; Volume 31. [Google Scholar]
  145. Liu, X.; Masana, M.; Herranz, L.; Van de Weijer, J.; Lopez, A.M.; Bagdanov, A.D. Rotate your networks: Better weight consolidation and less catastrophic forgetting. In Proceedings of the 2018 24th International Conference on Pattern Recognition (ICPR); IEEE: New York, NY, USA, 2018; pp. 2262–2268. [Google Scholar]
  146. Benzing, F. Unifying importance based regularisation methods for continual learning. In Proceedings of the International Conference on Artificial Intelligence and Statistics; PMLR: Cambridge, MA, USA, 2022; pp. 2372–2396. [Google Scholar]
  147. Lee, S.W.; Kim, J.H.; Jun, J.; Ha, J.W.; Zhang, B.T. Overcoming catastrophic forgetting by incremental moment matching. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  148. Chaudhry, A.; Rohrbach, M.; Elhoseiny, M.; Ajanthan, T.; Dokania, P.K.; Torr, P.H.; Ranzato, M. On tiny episodic memories in continual learning. arXiv 2019, arXiv:1902.10486. [Google Scholar]
  149. Vitter, J.S. Random sampling with a reservoir. ACM Trans. Math. Softw. (TOMS) 1985, 11, 37–57. [Google Scholar] [CrossRef] [Scilit]
  150. Borsos, Z.; Mutny, M.; Krause, A. Coresets via bilevel optimization for continual learning and streaming. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 14879–14890. [Google Scholar]
  151. Yoon, J.; Madaan, D.; Yang, E.; Hwang, S.J. Online coreset selection for rehearsal-based continual learning. arXiv 2021, arXiv:2106.01085. [Google Scholar]
  152. Shim, D.; Mai, Z.; Jeong, J.; Sanner, S.; Kim, H.; Jang, J. Online class-incremental continual learning with adversarial shapley value. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 9630–9638. [Google Scholar]
  153. Bang, J.; Kim, H.; Yoo, Y.; Ha, J.W.; Choi, J. Rainbow memory: Continual learning with a memory of diverse samples. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 8218–8227. [Google Scholar]
  154. Tiwari, R.; Killamsetty, K.; Iyer, R.; Shenoy, P. Gcr: Gradient coreset based replay buffer selection for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 99–108. [Google Scholar]
  155. Van Den Oord, A.; Vinyals, O. Neural discrete representation learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar]
  156. Wang, L.; Zhang, X.; Yang, K.; Yu, L.; Li, C.; Hong, L.; Zhang, S.; Li, Z.; Zhong, Y.; Zhu, J. Memory replay with data compression for continual learning. arXiv 2022, arXiv:2202.06592. [Google Scholar]
  157. Kulesza, A.; Taskar, B. Determinantal point processes for machine learning. Found. Trends® Mach. Learn. 2012, 5, 123–286. [Google Scholar] [CrossRef] [Scilit]
  158. Kumari, L.; Wang, S.; Zhou, T.; Bilmes, J.A. Retrospective adversarial replay for continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 28530–28544. [Google Scholar]
  159. Zhang, H.; Cisse, M.; Dauphin, Y.N.; Lopez-Paz, D. mixup: Beyond empirical risk minimization. arXiv 2017, arXiv:1710.09412. [Google Scholar]
  160. Belouadah, E.; Popescu, A. Il2m: Class incremental learning with dual memory. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 583–592. [Google Scholar]
  161. Ebrahimi, S.; Petryk, S.; Gokul, A.; Gan, W.; Gonzalez, J.E.; Rohrbach, M.; Darrell, T. Remembering for the right reasons: Explanations reduce catastrophic forgetting. Appl. AI Lett. 2021, 2, e44. [Google Scholar] [CrossRef] [Scilit]
  162. Liu, Y.; Su, Y.; Liu, A.A.; Schiele, B.; Sun, Q. Mnemonics training: Multi-class incremental learning without forgetting. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 12245–12254. [Google Scholar]
  163. Jin, X.; Sadhu, A.; Du, J.; Ren, X. Gradient-based editing of memory examples for online task-free continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 29193–29205. [Google Scholar]
  164. Chaudhry, A.; Ranzato, M.; Rohrbach, M.; Elhoseiny, M. Efficient Lifelong Learning with A-GEM. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2019. [Google Scholar]
  165. Tang, S.; Chen, D.; Zhu, J.; Yu, S.; Ouyang, W. Layerwise optimization by gradient decomposition for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 9634–9643. [Google Scholar]
  166. Sun, Q.; Lyu, F.; Shang, F.; Feng, W.; Wan, L. Exploring example influence in continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 27075–27086. [Google Scholar]
  167. Aljundi, R.; Belilovsky, E.; Tuytelaars, T.; Charlin, L.; Caccia, M.; Lin, M.; Page-Caccia, L. Online continual learning with maximal interfered retrieval. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  168. Chaudhry, A.; Gordo, A.; Dokania, P.; Torr, P.; Lopez-Paz, D. Using hindsight to anchor past knowledge in continual learning. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 6993–7001. [Google Scholar]
  169. Wu, Y.; Chen, Y.; Wang, L.; Ye, Y.; Liu, Z.; Guo, Y.; Fu, Y. Large scale incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 374–382. [Google Scholar]
  170. Zhao, B.; Xiao, X.; Gan, G.; Zhang, B.; Xia, S.T. Maintaining discrimination and fairness in class incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13208–13217. [Google Scholar]
  171. Ahn, H.; Kwak, J.; Lim, S.; Bang, H.; Kim, H.; Moon, T. Ss-il: Separated softmax for incremental learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 844–853. [Google Scholar]
  172. Cha, H.; Lee, J.; Shin, J. Co2l: Contrastive continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 9516–9525. [Google Scholar]
  173. Simon, C.; Koniusz, P.; Harandi, M. On learning the geodesic path for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 1591–1600. [Google Scholar]
  174. Joseph, K.; Khan, S.; Khan, F.S.; Anwer, R.M.; Balasubramanian, V.N. Energy-based latent aligner for incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 7452–7461. [Google Scholar]
  175. Kurmi, V.K.; Patro, B.N.; Subramanian, V.K.; Namboodiri, V.P. Do not forget to attend to uncertainty while mitigating catastrophic forgetting. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2021; pp. 736–745. [Google Scholar]
  176. Ashok, A.; Joseph, K.; Balasubramanian, V.N. Class-incremental learning with cross-space clustering and controlled transfer. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 105–122. [Google Scholar]
  177. Hu, X.; Tang, K.; Miao, C.; Hua, X.S.; Zhang, H. Distilling causal effect of data in class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 3957–3966. [Google Scholar]
  178. Bhat, P.; Zonooz, B.; Arani, E. Task-aware information routing from common representation space in lifelong learning. arXiv 2023, arXiv:2302.11346. [Google Scholar]
  179. Hou, S.; Pan, X.; Loy, C.C.; Wang, Z.; Lin, D. Lifelong learning via progressive distillation and retrospection. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 437–452. [Google Scholar]
  180. Wang, F.Y.; Zhou, D.W.; Ye, H.J.; Zhan, D.C. Foster: Feature boosting and compression for class-incremental learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 398–414. [Google Scholar]
  181. Chaudhry, A.; Dokania, P.K.; Ajanthan, T.; Torr, P.H. Riemannian walk for incremental learning: Understanding forgetting and intransigence. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 532–547. [Google Scholar]
  182. Wang, L.; Zhang, M.; Jia, Z.; Li, Q.; Bao, C.; Ma, K.; Zhu, J.; Zhong, Y. Afec: Active forgetting of negative transfer in continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2021; Volume 34, pp. 22379–22391. [Google Scholar]
  183. Verwimp, E.; De Lange, M.; Tuytelaars, T. Rehearsal revealed: The limits and merits of revisiting samples in continual learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 9385–9394. [Google Scholar]
  184. Bonicelli, L.; Boschini, M.; Porrello, A.; Spampinato, C.; Calderara, S. On the effectiveness of lipschitz-driven rehearsal in continual learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 31886–31901. [Google Scholar]
  185. Yu, L.; Hu, T.; Hong, L.; Liu, Z.; Weller, A.; Liu, W. Continual learning by modeling intra-class variation. arXiv 2022, arXiv:2210.05398. [Google Scholar]
  186. Buzzega, P.; Boschini, M.; Porrello, A.; Abati, D.; Calderara, S. Dark experience for general continual learning: A strong, simple baseline. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 15920–15930. [Google Scholar]
  187. Boschini, M.; Bonicelli, L.; Buzzega, P.; Porrello, A.; Calderara, S. Class-incremental continual learning into the extended der-verse. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 5497–5512. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  188. Prabhu, A.; Torr, P.H.; Dokania, P.K. Gdumb: A simple approach that questions our progress in continual learning. In Proceedings of the Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020; Proceedings, Part II 16; Springer: Berlin/Heidelberg, Germany, 2020; pp. 524–540. [Google Scholar]
  189. Ayub, A.; Wagner, A.R. EEC: Learning to encode and regenerate images for continual learning. arXiv 2021, arXiv:2101.04904. [Google Scholar]
  190. Ostapenko, O.; Puscas, M.; Klein, T.; Jahnichen, P.; Nabi, M. Learning to remember: A synaptic plasticity driven framework for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2019; pp. 11321–11329. [Google Scholar]
  191. Kemker, R.; Kanan, C. Fearnet: Brain-inspired model for incremental learning. arXiv 2017, arXiv:1711.10563. [Google Scholar]
  192. Riemer, M.; Klinger, T.; Bouneffouf, D.; Franceschini, M. Scalable recollections for continual lifelong learning. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2019; Volume 33, pp. 1352–1359. [Google Scholar]
  193. Rostami, M.; Kolouri, S.; Pilly, P.K. Complementary learning for overcoming catastrophic forgetting using experience replay. In Proceedings of the 28th International Joint Conference on Artificial Intelligence; International Joint Conferences on Artificial Intelligence: Marina Del Rey, CA, USA, 2019; pp. 3339–3345. [Google Scholar]
  194. Pfülb, B.; Gepperth, A.; Bagus, B. Continual learning with fully probabilistic models. arXiv 2021, arXiv:2104.09240x. [Google Scholar]
  195. Gopalakrishnan, S.; Singh, P.R.; Fayek, H.; Ramasamy, S.; Ambikapathi, A. Knowledge capture and replay for continual learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2022; pp. 10–18. [Google Scholar]
  196. Ye, F.; Bors, A.G. Learning latent representations across multiple data domains using lifelong VAEGAN. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 777–795. [Google Scholar]
  197. Nguyen, C.V.; Li, Y.; Bui, T.D.; Turner, R.E. Variational continual learning. arXiv 2017, arXiv:1710.10628. [Google Scholar]
  198. Seff, A.; Beatson, A.; Suo, D.; Liu, H. Continual learning in generative adversarial nets. arXiv 2017, arXiv:1705.08395. [Google Scholar]
  199. He, C.; Wang, R.; Shan, S.; Chen, X. Exemplar-supported generative reproduction for class incremental learning. In Proceedings of the BMVC; BMVA (British Machine Vision Association): Durham, UK, 2018; Volume 1, p. 2. [Google Scholar]
  200. Xiang, Y.; Fu, Y.; Ji, P.; Huang, H. Incremental learning using conditional adversarial networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 6619–6628. [Google Scholar]
  201. Cong, Y.; Zhao, M.; Li, J.; Wang, S.; Carin, L. Gan memory with no forgetting. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 16481–16494. [Google Scholar]
  202. Liu, X.; Wu, C.; Menta, M.; Herranz, L.; Raducanu, B.; Bagdanov, A.D.; Jui, S.; de Weijer, J.v. Generative feature replay for class-incremental learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops; IEEE: New York, NY, USA, 2020; pp. 226–227. [Google Scholar]
  203. Ostapenko, O.; Lesort, T.; Rodriguez, P.; Arefin, M.R.; Douillard, A.; Rish, I.; Charlin, L. Continual learning with foundation models: An empirical study of latent replay. In Proceedings of the Conference on Lifelong Learning Agents; PMLR: Cambridge, MA, USA, 2022; pp. 60–91. [Google Scholar]
  204. Wang, Z.; Liu, L.; Kong, Y.; Guo, J.; Tao, D. Online continual learning with contrastive vision transformer. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 631–650. [Google Scholar]
  205. Wang, Y.; Huang, Z.; Hong, X. S-prompts learning with pre-trained transformers: An occam’s razor for domain incremental learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 5682–5695. [Google Scholar]
  206. Wang, Z.; Zhang, Z.; Ebrahimi, S.; Sun, R.; Zhang, H.; Lee, C.Y.; Ren, X.; Su, G.; Perot, V.; Dy, J.; et al. Dualprompt: Complementary prompting for rehearsal-free continual learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 631–648. [Google Scholar]
  207. Wang, Z.; Zhang, Z.; Lee, C.Y.; Zhang, H.; Sun, R.; Ren, X.; Su, G.; Perot, V.; Dy, J.; Pfister, T. Learning to prompt for continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2022; pp. 139–149. [Google Scholar]
  208. Smith, J.S.; Karlinsky, L.; Gutta, V.; Cascante-Bonilla, P.; Kim, D.; Arbelle, A.; Panda, R.; Feris, R.; Kira, Z. Coda-prompt: Continual decomposed attention-based prompting for rehearsal-free continual learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2023; pp. 11909–11919. [Google Scholar]
  209. McDonnell, M.D.; Gong, D.; Parvaneh, A.; Abbasnejad, E.; van den Hengel, A. RanPAC: Random Projections and Pre-trained Models for Continual Learning. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); NeurIPS: San Diego, CA, USA, 2023; Volume 36, pp. 12022–12053. [Google Scholar]
  210. Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. In Proceedings of the The Tenth International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2022. [Google Scholar]
  211. Zenke, F.; Poole, B.; Ganguli, S. Continual Learning Through Synaptic Intelligence. In Proceedings of the 34th International Conference on Machine Learning (ICML); Precup, D., Teh, Y.W., Eds.; PMLR (Proceedings of Machine Learning Research): Cambridge, MA, USA, 2017; Volume 70, pp. 3987–3995. [Google Scholar]
  212. Aljundi, R.; Babiloni, F.; Elhoseiny, M.; Rohrbach, M.; Tuytelaars, T. Memory aware synapses: Learning what (not) to forget. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2018; pp. 139–154. [Google Scholar]
  213. Rolnick, D.; Ahuja, A.; Schwarz, J.; Lillicrap, T.P.; Wayne, G. Experience Replay for Continual Learning. In Proceedings of the Advances in Neural Information Processing Systems 32 (NeurIPS 2019); Wallach, H.M., Larochelle, H., Beygelzimer, A., d’Alché Buc, F., Fox, E., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2019; pp. 348–358. [Google Scholar]
  214. Rusu, A.A.; Rabinowitz, N.C.; Desjardins, G.; Soyer, H.; Kirkpatrick, J.; Kavukcuoglu, K.; Pascanu, R.; Hadsell, R. Progressive Neural Networks. arXiv 2016, arXiv:1606.04671. [Google Scholar]
  215. Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised contrastive learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 18661–18673. [Google Scholar]
  216. Fini, E.; Da Costa, V.G.T.; Alameda-Pineda, X.; Ricci, E.; Alahari, K.; Mairal, J. Self-Supervised Models are Continual Learners. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2022. [Google Scholar]
  217. Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; de Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-Efficient Transfer Learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (ICML); PMLR (Proceedings of Machine Learning Research): Cambridge, MA, USA, 2019; Volume 97, pp. 2790–2799. [Google Scholar]
  218. Wang, X.; Zhang, Y.; Chen, T.; Gao, S.; Jin, S.; Yang, X.; Xi, Z.; Zheng, R.; Zou, Y.; Gui, T.; et al. TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models. arXiv 2023, arXiv:2310.06762. [Google Scholar]
  219. Chen, H.; Wu, Z.; Han, X.; Jia, M.; Jiang, Y.G. Promptfusion: Decoupling stability and plasticity for continual learning. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024; pp. 196–212. [Google Scholar]
  220. Park, C.W.; Seo, S.W.; Kang, N.; Ko, B.; Choi, B.W.; Park, C.M.; Chang, D.K.; Kim, H.; Kim, H.; Lee, H.; et al. Artificial intelligence in health care: Current applications and issues. J. Korean Med. Sci. 2020, 35, e379. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  221. Zhu, D.; Bu, Q.; Zhu, Z.; Zhang, Y.; Wang, Z. Advancing autonomy through lifelong learning: A survey of autonomous intelligent systems. Front. Neurorobotics 2024, 18, 1385778. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  222. Ciupek, D.; Malawski, M.; Pieciak, T. Federated Learning: A new frontier in the exploration of multi-institutional medical imaging data. arXiv 2025, arXiv:2503.20107. [Google Scholar]
  223. Thakur, G.K.; Thakur, A.; Kulkarni, S.; Khan, N.; Khan, S. Deep learning approaches for medical image analysis and diagnosis. Cureus 2024, 16. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  224. Jeon, J.; Kim, J.; Kim, J.; Kim, K.; Mohaisen, A.; Kim, J.K. Privacy-preserving deep learning computation for geo-distributed medical big-data platforms. In Proceedings of the 2019 49th Annual IEEE/IFIP International Conference on Dependable Systems and Networks–Supplemental Volume (DSN-S); IEEE: New York, NY, USA, 2019; pp. 3–4. [Google Scholar]
  225. Pianykh, O.S.; Langs, G.; Dewey, M.; Enzmann, D.R.; Herold, C.J.; Schoenberg, S.O.; Brink, J.A. Continuous learning AI in radiology: Implementation principles and early applications. Radiology 2020, 297, 6–14. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  226. Pinto-Coelho, L. How artificial intelligence is shaping medical imaging technology: A survey of innovations and applications. Bioengineering 2023, 10, 1435. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  227. Zhu, Z.; Sun, Y.; Honarvar Shakibaei Asli, B. Early Breast Cancer Detection Using Artificial Intelligence Techniques Based on Advanced Image Processing Tools. Electronics 2024, 13, 3575. [Google Scholar] [CrossRef] [Scilit]
  228. Lee, C.S.; Lee, A.Y. Applications of Continual Learning Machine Learning in Clinical Practice. Lancet Digit. Health 2020, 2, e279–e281. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  229. Derakhshani, M.M.; Najdenkoska, I.; van Sonsbeek, T.; Zhen, X.; Mahapatra, D.; Worring, M.; Snoek, C.G.M. LifeLonger: A Benchmark for Continual Disease Classification. In Proceedings of the Medical Image Computing and Computer-Assisted Intervention (MICCAI); Springer: Berlin/Heidelberg, Germany, 2022. [Google Scholar]
  230. Perkonigg, M.; Hofmanninger, J.; Herold, C.J.; Brink, J.A.; Pianykh, O.; Prosch, H.; Langs, G. Dynamic Memory to Alleviate Catastrophic Forgetting in Continual Learning with Medical Imaging. Nat. Commun. 2021, 12, 5678. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  231. da Silva Motta, D.; Badaró, R.; Santos, A.; Kirchner, F. Use of Artificial Intelligence on the Control of Vector-Borne Diseases; IntechOpen: London, UK, 2018. [Google Scholar]
  232. Dasari, S.; Ebert, F.; Tian, S.; Nair, S.; Bucher, B.; Schmeckpeper, K.; Singh, S.; Levine, S.; Finn, C. Robonet: Large-scale multi-robot learning. arXiv 2019, arXiv:1910.11215. [Google Scholar]
  233. Haque, N. Catastrophic Forgetting in LLMs: A Comparative Analysis Across Language Tasks. arXiv 2025, arXiv:2504.01241. [Google Scholar]
  234. Yao, Y.; González-Vélez, H. AI-Powered System to Facilitate Personalized Adaptive Learning in Digital Transformation. Appl. Sci. 2025, 15, 4989. [Google Scholar] [CrossRef] [Scilit]
  235. Li, D.; Chen, Z.; Cho, E.; Hao, J.; Liu, X.; Xing, F.; Guo, C.; Liu, Y. Overcoming catastrophic forgetting during domain adaptation of seq2seq language generation. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 5441–5454. [Google Scholar]
  236. Liu, T.; Ungar, L.; Sedoc, J. Continual learning for sentence representations using conceptors. arXiv 2019, arXiv:1904.09187. [Google Scholar]
  237. Monaikul, N.; Castellucci, G.; Filice, S.; Rokhlenko, O. Continual learning for named entity recognition. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 13570–13577. [Google Scholar]
  238. Li, G.; Zhai, Y.; Chen, Q.; Gao, X.; Zhang, J.; Zhang, Y. Continual few-shot intent detection. In Proceedings of the 29th International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 333–343. [Google Scholar]
  239. Liu, Q.; Yu, X.; He, S.; Liu, K.; Zhao, J. Lifelong intent detection via multi-strategy rebalancing. arXiv 2021, arXiv:2108.04445. [Google Scholar]
  240. Varshney, V.; Patidar, M.; Kumar, R.; Shroff, G.; Vig, L. Prompt Augmented Generative Replay via Supervised Contrastive Training for Lifelong Intent Detection. U.S. Patent App. 18/215,972, 11 January 2024. [Google Scholar]
  241. Qin, C.; Joty, S. Lfpt5: A unified framework for lifelong few-shot language learning based on prompt tuning of t5. arXiv 2021, arXiv:2110.07298. [Google Scholar]
  242. Sun, J.; Wang, S.; Zhang, J.; Zong, C. Distill and replay for continual language learning. In Proceedings of the 28th International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2020; pp. 3569–3579. [Google Scholar]
  243. Cao, Y.; Wei, H.R.; Chen, B.; Wan, X. Continual learning for neural machine translation. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 3964–3974. [Google Scholar]
  244. Shao, C.; Feng, Y. Overcoming catastrophic forgetting beyond continual learning: Balanced training for neural machine translation. arXiv 2022, arXiv:2203.03910. [Google Scholar]
  245. Qin, Y.; Zhang, J.; Lin, Y.; Liu, Z.; Li, P.; Sun, M.; Zhou, J. Elle: Efficient lifelong pre-training for emerging data. arXiv 2022, arXiv:2203.06311. [Google Scholar]
  246. Huang, Y.; Zhang, Y.; Chen, J.; Wang, X.; Yang, D. Continual learning for text classification with information disentanglement based regularization. arXiv 2021, arXiv:2104.05489. [Google Scholar]
  247. de Masson D’Autume, C.; Ruder, S.; Kong, L.; Yogatama, D. Episodic memory in lifelong language learning. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar]
  248. Wang, Z.; Mehta, S.V.; Póczos, B.; Carbonell, J. Efficient meta lifelong-learning with limited memory. arXiv 2020, arXiv:2010.02500. [Google Scholar]
  249. Xu, K.; Verma, S.; Finn, C.; Levine, S. Continual learning of control primitives: Skill discovery via reset-games. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 4999–5010. [Google Scholar]
  250. Mi, F.; Chen, L.; Zhao, M.; Huang, M.; Faltings, B. Continual learning for natural language generation in task-oriented dialog systems. arXiv 2020, arXiv:2010.00910. [Google Scholar]
  251. Li, Z.; Qu, L.; Haffari, G. Total recall: A customized continual learning method for neural semantic parsers. arXiv 2021, arXiv:2109.05186. [Google Scholar]
  252. Sun, F.K.; Ho, C.H.; Lee, H.Y. Lamol: Language modeling for lifelong language learning. arXiv 2019, arXiv:1909.03329. [Google Scholar]
  253. Zhang, Y.; Wang, X.; Yang, D. Continual sequence generation with adaptive compositional modules. arXiv 2022, arXiv:2203.10652. [Google Scholar]
  254. Wang, R.; Yu, T.; Zhao, H.; Kim, S.; Mitra, S.; Zhang, R.; Henao, R. Few-shot class-incremental learning for named entity recognition. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 571–582. [Google Scholar]
  255. Geng, B.; Yuan, F.; Xu, Q.; Shen, Y.; Xu, R.; Yang, M. Continual learning for task-oriented dialogue system with iterative network pruning, expanding and masking. arXiv 2021, arXiv:2107.08173. [Google Scholar]
  256. Shen, Y.; Zeng, X.; Jin, H. A progressive model to enable continual learning for semantic slot filling. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP); Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 1279–1284. [Google Scholar]
  257. Wang, C.; Pan, H.; Liu, Y.; Chen, K.; Qiu, M.; Zhou, W.; Huang, J.; Chen, H.; Lin, W.; Cai, D. Mell: Large-scale extensible user intent classification for dialogue systems with meta lifelong learning. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining; ACM: New York, NY, USA, 2021; pp. 3649–3659. [Google Scholar]
  258. Wu, T.; Li, X.; Li, Y.F.; Haffari, G.; Qi, G.; Zhu, Y.; Xu, G. Curriculum-meta learning for order-robust continual relation extraction. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI: Washington, DC, USA, 2021; Volume 35, pp. 10363–10369. [Google Scholar]
  259. Madotto, A.; Lin, Z.; Zhou, Z.; Moon, S.; Crook, P.A.; Liu, B.; Yu, Z.; Cho, E.; Fung, P.; Wang, Z. Continual learning in task-oriented dialogue systems. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 7452–7467. [Google Scholar]
  260. Ermis, B.; Zappella, G.; Wistuba, M.; Rawal, A.; Archambeau, C. Memory efficient continual learning with transformers. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 10629–10642. [Google Scholar]
  261. Zhu, Q.; Li, B.; Mi, F.; Zhu, X.; Huang, M. Continual prompt tuning for dialog state tracking. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 1124–1137. [Google Scholar]
  262. Liu, M.; Chang, S.; Huang, L. Incremental prompting: Episodic memory prompt for lifelong event detection. In Proceedings of the 29th International Conference on Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 2157–2165. [Google Scholar]
  263. Yin, W.; Li, J.; Xiong, C. Contintin: Continual learning from task instructions. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Association for Computational Linguistics: Stroudsburg, PA, USA, 2022; pp. 3062–3072. [Google Scholar]
  264. Xia, C.; Yin, W.; Feng, Y.; Yu, P.S. Incremental few-shot text classification with multi-round new classes: Formulation, dataset and system. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 1351–1360. [Google Scholar]
  265. Garcia, X.; Constant, N.; Parikh, A.; Firat, O. Towards continual learning for multilingual machine translation via vocabulary substitution. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2021; pp. 1184–1192. [Google Scholar]
  266. Yan, S.; Hong, L.; Xu, H.; Han, J.; Tuytelaars, T.; Li, Z.; He, X. Generative negative text replay for continual vision-language pretraining. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2022; pp. 22–38. [Google Scholar]
  267. Greco, C.; Plank, B.; Fernández, R.; Bernardi, R. Psycholinguistics meets continual learning: Measuring catastrophic forgetting in visual question answering. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 3601–3605. [Google Scholar]
  268. Martínez-Plumed, F.; Ferri, C.; Hernández-Orallo, J.; Ramírez-Quintana, M.J. Forgetting and consolidation for incremental and cumulative knowledge acquisition systems. arXiv 2015, arXiv:1502.05615. [Google Scholar]
  269. Christakopoulou, K.; Lalama, A.; Adams, C.; Qu, I.; Amir, Y.; Chucri, S.; Vollucci, P.; Soldo, F.; Bseiso, D.; Scodel, S.; et al. Large language models for user interest journeys. arXiv 2023, arXiv:2305.15498. [Google Scholar]
  270. Wang, X.J.; Lee, C.P.; Mutlu, B. LearnMate: Enhancing Online Education with LLM-Powered Personalized Learning Plans and Support. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems; ACM: New York, NY, USA, 2025; pp. 1–10. [Google Scholar]
  271. Sabeima, M.; Lamolle, M.; Nanne, M.F. Towards personalized adaptive learning in e-learning recommender systems. Int. J. Adv. Comput. Sci. Appl. 2022, 13, 14–20. [Google Scholar] [CrossRef] [Scilit]
  272. Joy, J.; Raj, N.S.; VG, R. Ontology-based E-learning content recommender system for addressing the pure cold-start problem. ACM J. Data Inf. Qual. 2021, 13, 1–27. [Google Scholar] [CrossRef] [Scilit]
  273. Liu, Z.; Wang, Y.; Vaidya, S.; Ruehle, F.; Halverson, J.; Soljacic, M.; Hou, T.; Tegmark, M. KAN: Kolmogorov–arnold networks. In Proceedings of the International Conference on Learning Representations; ICLR: Appleton, WI, USA, 2025; Volume 2025, pp. 70367–70413. [Google Scholar]
  274. Bountouni, N.; Koussouris, S.; Vasileiou, A.; Kazazis, S.A. A Holistic Framework for Safeguarding of SMEs: A Case Study. In Proceedings of the 2023 19th International Conference on the Design of Reliable Communication Networks (DRCN); IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar]
  275. Asmar, M.; Tuqan, A. Integrating machine learning for sustaining cybersecurity in digital banks. Heliyon 2024, 10, e37571. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  276. Ahmed, U.; Nazir, M.; Sarwar, A.; Ali, T.; Aggoune, E.H.M.; Shahzad, T.; Khan, M.A. Signature-based intrusion detection using machine learning and deep learning approaches empowered with fuzzy clustering. Sci. Rep. 2025, 15, 1726. [Google Scholar] [CrossRef] [Scilit]
  277. Dohare, S.; Hernandez-Garcia, J.F.; Lan, Q.; Rahman, P.; Mahmood, A.R.; Sutton, R.S. Loss of plasticity in deep continual learning. Nature 2024, 632, 768–774. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  278. Jones, R.; Omar, M.; Mohammed, D.; Nobles, C.; Dawson, M. Harnessing the speed and accuracy of machine learning to advance cybersecurity. In Proceedings of the 2023 Congress in Computer Science, Computer Engineering, & Applied Computing (CSCE); IEEE: New York, NY, USA, 2023; pp. 418–421. [Google Scholar]
  279. Rahul-Vigneswaran, K.; Poornachandran, P.; Soman, K. A compendium on network and host based intrusion detection systems. In Proceedings of the ICDSMLA 2019: Proceedings of the 1st International Conference on Data Science, Machine Learning and Applications; Springer: Berlin/Heidelberg, Germany, 2020; pp. 23–30. [Google Scholar]
  280. Stokes, J.W.; Wang, D.; Marinescu, M.; Marino, M.; Bussone, B. Attack and defense of dynamic analysis-based, adversarial neural malware detection models. In Proceedings of the MILCOM 2018 IEEE Military Communications Conference (MILCOM); IEEE: New York, NY, USA, 2018; pp. 1–8. [Google Scholar]
  281. Sameen, M.; Han, K.; Hwang, S.O. PhishHaven-An efficient real-time AI phishing URLs detection system. IEEE Access 2020, 8, 83425–83443. [Google Scholar] [CrossRef] [Scilit]
  282. Zhang, Q.; Chen, M.; Bukharin, A.; He, P.; Cheng, Y.; Chen, W.; Zhao, T. AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning. In Proceedings of the International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2023. [Google Scholar]
  283. Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient Finetuning of Quantized LLMs. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); NeurIPS: San Diego, CA, USA, 2023. [Google Scholar]
  284. Liu, S.Y.; Wang, C.Y.; Yin, H.; Molchanov, P.; Wang, Y.C.F.; Cheng, K.T.; Chen, M.H. DoRA: Weight-Decomposed Low-Rank Adaptation. In Proceedings of the 41st International Conference on Machine Learning (ICML); OpenReview.net: Newton Highlands, MA, USA, 2024. [Google Scholar]
  285. Chen, Y.; Qian, S.; Tang, H.; Lai, X.; Liu, Z.; Han, S.; Jia, J. LongLoRA: Efficient Fine-Tuning of Long-Context Large Language Models. In International Conference on Learning Representations (ICLR); ICLR: Appleton, WI, USA, 2024. [Google Scholar]
  286. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
Figure 1. PRISMA 2020 flow diagram summarizing the literature identification, screening, eligibility assessment, and study selection process adopted in this review. Records were retrieved from multiple scientific databases, screened according to predefined inclusion and exclusion criteria, and assessed through full-text review. After removing duplicates and excluding irrelevant or ineligible studies, 285 publications were selected for qualitative analysis and discussion in this review.
Figure 1. PRISMA 2020 flow diagram summarizing the literature identification, screening, eligibility assessment, and study selection process adopted in this review. Records were retrieved from multiple scientific databases, screened according to predefined inclusion and exclusion criteria, and assessed through full-text review. After removing duplicates and excluding irrelevant or ineligible studies, 285 publications were selected for qualitative analysis and discussion in this review.
Mathematics 14 02774 g001
Figure 2. Conceptual overview of CL illustrating the relationship between sequential learning, theoretical foundations, methodological paradigms, and practical applications. (A) CL involves sequentially learning a stream of tasks whose data distributions evolve over time while preserving knowledge acquired from previous tasks. (B) The core objective of CL is to balance the stability–plasticity dilemma by maintaining previously learned knowledge (stability), efficiently acquiring new knowledge (plasticity), and preserving generalization across both within-task and cross-task distribution shifts. (C) To address catastrophic forgetting, representative CL methods operate at different stages of the learning pipeline, including replay-based approaches that revisit previous data, architecture-based methods that expand or isolate model components, representation learning methods that preserve transferable feature embeddings, regularization-based methods that constrain parameter or function updates, and optimization-based methods that reduce gradient interference during sequential training. (D) CL has broad applicability across diverse real-world scenarios, where increasing task complexity and domain-specific requirements necessitate scalable lifelong adaptation in areas such as computer vision, natural language processing (NLP), robotics, reinforcement learning, healthcare, and multimodal foundation models. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Figure 2. Conceptual overview of CL illustrating the relationship between sequential learning, theoretical foundations, methodological paradigms, and practical applications. (A) CL involves sequentially learning a stream of tasks whose data distributions evolve over time while preserving knowledge acquired from previous tasks. (B) The core objective of CL is to balance the stability–plasticity dilemma by maintaining previously learned knowledge (stability), efficiently acquiring new knowledge (plasticity), and preserving generalization across both within-task and cross-task distribution shifts. (C) To address catastrophic forgetting, representative CL methods operate at different stages of the learning pipeline, including replay-based approaches that revisit previous data, architecture-based methods that expand or isolate model components, representation learning methods that preserve transferable feature embeddings, regularization-based methods that constrain parameter or function updates, and optimization-based methods that reduce gradient interference during sequential training. (D) CL has broad applicability across diverse real-world scenarios, where increasing task complexity and domain-specific requirements necessitate scalable lifelong adaptation in areas such as computer vision, natural language processing (NLP), robotics, reinforcement learning, healthcare, and multimodal foundation models. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Mathematics 14 02774 g002
Figure 3. Classification of CL approaches according to the characteristics of their training and inference configurations.
Figure 3. Classification of CL approaches according to the characteristics of their training and inference configurations.
Mathematics 14 02774 g003
Figure 4. Unified illustration of the theoretical foundations of continual learning. (A) The stability–plasticity dilemma illustrates the trade-off between preserving previously acquired knowledge (stability) and adapting to new tasks (plasticity). Excessive plasticity results in catastrophic forgetting, whereas excessive stability limits adaptation to new tasks. (B) Representative continual learning strategies for mitigating catastrophic forgetting, including replay-based methods, regularization-based methods, and architecture-based methods, which preserve previous knowledge through memory replay, parameter regularization, and structural adaptation, respectively.
Figure 4. Unified illustration of the theoretical foundations of continual learning. (A) The stability–plasticity dilemma illustrates the trade-off between preserving previously acquired knowledge (stability) and adapting to new tasks (plasticity). Excessive plasticity results in catastrophic forgetting, whereas excessive stability limits adaptation to new tasks. (B) Representative continual learning strategies for mitigating catastrophic forgetting, including replay-based methods, regularization-based methods, and architecture-based methods, which preserve previous knowledge through memory replay, parameter regularization, and structural adaptation, respectively.
Mathematics 14 02774 g004
Figure 5. Illustration of regularization-based continual learning for mitigating catastrophic forgetting during sequential task learning. After learning the initial task, the previously trained model is frozen and serves as a reference while a new model is optimized on incoming tasks. Regularization-based approaches preserve prior knowledge by constraining either the model parameters or the model outputs during optimization. Weight regularization methods (e.g., Elastic Weight Consolidation, Synaptic Intelligence (SI), and Memory-Aware Synapses) penalize changes to parameters that are important for previously learned tasks, whereas function regularization methods preserve intermediate feature representations or output logits through knowledge distillation and consistency constraints. By limiting excessive parameter updates while allowing sufficient plasticity for new tasks, these methods maintain previously acquired knowledge and alleviate catastrophic forgetting during continual adaptation. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Figure 5. Illustration of regularization-based continual learning for mitigating catastrophic forgetting during sequential task learning. After learning the initial task, the previously trained model is frozen and serves as a reference while a new model is optimized on incoming tasks. Regularization-based approaches preserve prior knowledge by constraining either the model parameters or the model outputs during optimization. Weight regularization methods (e.g., Elastic Weight Consolidation, Synaptic Intelligence (SI), and Memory-Aware Synapses) penalize changes to parameters that are important for previously learned tasks, whereas function regularization methods preserve intermediate feature representations or output logits through knowledge distillation and consistency constraints. By limiting excessive parameter updates while allowing sufficient plasticity for new tasks, these methods maintain previously acquired knowledge and alleviate catastrophic forgetting during continual adaptation. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Mathematics 14 02774 g005
Figure 6. Overview of replay-based continual learning strategies for mitigating catastrophic forgetting during sequential task learning. The left network is first trained on Task A (e.g., puppy and cat) and extracts representative feature embeddings. Knowledge from previous tasks is preserved either by storing representative samples in a memory buffer (experience replay), generating pseudo-samples using a generative model (generative replay), or retaining informative feature representations (feature replay). During Task B, replayed samples or reconstructed features are combined with data from newly introduced classes (e.g., panda and bunny) to jointly optimize the model, thereby maintaining previously acquired knowledge while learning new concepts. By revisiting representative information from earlier tasks, replay-based methods effectively reduce catastrophic forgetting and improve knowledge retention across long task sequences. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Figure 6. Overview of replay-based continual learning strategies for mitigating catastrophic forgetting during sequential task learning. The left network is first trained on Task A (e.g., puppy and cat) and extracts representative feature embeddings. Knowledge from previous tasks is preserved either by storing representative samples in a memory buffer (experience replay), generating pseudo-samples using a generative model (generative replay), or retaining informative feature representations (feature replay). During Task B, replayed samples or reconstructed features are combined with data from newly introduced classes (e.g., panda and bunny) to jointly optimize the model, thereby maintaining previously acquired knowledge while learning new concepts. By revisiting representative information from earlier tasks, replay-based methods effectively reduce catastrophic forgetting and improve knowledge retention across long task sequences. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Mathematics 14 02774 g006
Figure 7. Illustration of representative strategies used in class-incremental learning to mitigate catastrophic forgetting at different representation levels. During Task A, the model is trained to recognize the initial classes (e.g., puppy and cat). When Task B introduces new classes (e.g., panda and bunny), replay-based methods preserve representative samples from previously learned tasks by storing or regenerating past data, thereby reducing forgetting in the input space. Knowledge distillation complements replay by transferring information from the previous model to the updated model through feature-level and output-level consistency constraints. Feature distillation preserves intermediate feature representations learned from previous tasks, whereas label distillation maintains the predictive behavior of earlier classifiers by matching soft output distributions. By jointly exploiting replay and knowledge distillation, CIL methods effectively retain previously acquired knowledge while continuously integrating new classes without requiring complete retraining. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Figure 7. Illustration of representative strategies used in class-incremental learning to mitigate catastrophic forgetting at different representation levels. During Task A, the model is trained to recognize the initial classes (e.g., puppy and cat). When Task B introduces new classes (e.g., panda and bunny), replay-based methods preserve representative samples from previously learned tasks by storing or regenerating past data, thereby reducing forgetting in the input space. Knowledge distillation complements replay by transferring information from the previous model to the updated model through feature-level and output-level consistency constraints. Feature distillation preserves intermediate feature representations learned from previous tasks, whereas label distillation maintains the predictive behavior of earlier classifiers by matching soft output distributions. By jointly exploiting replay and knowledge distillation, CIL methods effectively retain previously acquired knowledge while continuously integrating new classes without requiring complete retraining. Figure created by the authors based on concepts summarized in Wang et al. [11], and redrawn with modifications for this review.
Mathematics 14 02774 g007
Table 1. Representative database-specific search strategy used during the literature search. Search queries were adapted slightly according to the syntax supported by each digital library while preserving the same search intent.
Table 1. Representative database-specific search strategy used during the literature search. Search queries were adapted slightly according to the syntax supported by each digital library while preserving the same search intent.
DatabaseRepresentative Search Query
IEEE Xplore(“continual learning” OR “lifelong learning” OR “incremental learning”) AND (“catastrophic forgetting” OR “experience replay” OR “regularization” OR “knowledge distillation”)
ACM Digital Library(“continual learning” OR “lifelong learning”) AND (“foundation models” OR “prompt learning” OR “Parameter-Efficient Fine-Tuning”)
SpringerLink(“continual learning”) AND (“vision transformer” OR “foundation model” OR “multimodal learning”)
ScienceDirect(“continual learning”) AND (“experience replay” OR “regularization” OR “benchmark”)
Web of ScienceTS = (“continual learning” OR “incremental learning”) AND TS = (“catastrophic forgetting” OR “foundation model”)
ScopusTITLE-ABS-KEY (“continual learning” OR “lifelong learning”) AND (“multimodal” OR “prompt learning” OR “PEFT”)
Google Scholar“continual learning” “foundation model” “prompt learning” “Parameter-Efficient Fine-Tuning”
Table 2. Representative examples of excluded studies and reasons for exclusion.
Table 2. Representative examples of excluded studies and reasons for exclusion.
Representative Study TypeReason for Exclusion
Transfer learning studiesFocused exclusively on transfer learning without continual or lifelong adaptation.
Multi-task learning studiesAddressed simultaneous multi-task optimization rather than sequential continual learning.
Hardware implementation studiesPrimarily focused on hardware acceleration or system implementation without methodological contributions to continual learning.
Workshop papersExcluded because they lacked sufficient methodological detail or comprehensive experimental validation.
Duplicate publicationsEarlier preprint versions were excluded when a corresponding peer-reviewed journal or conference paper was available.
Studies outside the review scopeFocused on unrelated topics that did not directly address continual learning algorithms, evaluation, theory, or applications.
Table 3. Comparison of representative continual learning survey papers.
Table 3. Comparison of representative continual learning survey papers.
SurveyYearMain FocusCL CoverageModern TrendsMain Limitation
Wang et al. [11]2024General CL theory and methodsTIL, DIL, and CILLimited discussion of prompting and PEFTMinimal focus on foundation-model adaptation and modern multimodal CL
Van de Ven et al. [18]2022Taxonomy of CL scenariosTIL, DIL, and CILDoes not cover recent CL trendsPrimarily focused on conceptual categorization of CL settings
Bidaki et al. [19]2025Online CLStreaming and online CLBenchmark-oriented discussionNarrow scope centered on online learning settings
Zhou et al. [20]2024Class-incremental learningMainly CILLimited multimodal and foundation-model discussionRestricted primarily to CIL strategies and benchmarks
Wickramasinghe et al. [21]2023Overview of CL methodsGeneral CL settingsCovers traditional CL methodsLimited synthesis of transformer-, prompt-, and foundation-model-based continual learning methods
This review2026Comprehensive review of modern continual learningTIL, DIL, CIL, data-incremental, online, multimodal, and federated CLFoundation models, prompt learning, PEFT, diffusion models, and multimodal continual learningIntroduces an expanded taxonomy of continual learning scenarios and methodological categories, provides a unified analysis of evaluation protocols, benchmark fragmentation, and reproducibility, critically examines foundation-model-based continual learning, and synthesizes emerging research directions for scalable and deployment-oriented lifelong learning
Table 4. Comparison of continual learning, transfer learning, multi-task learning, and online learning.
Table 4. Comparison of continual learning, transfer learning, multi-task learning, and online learning.
FeatureCLTransfer LearningMulti-Task LearningOnline Learning
Task AvailabilitySequentialOne-time transferSimultaneousSingle task
FocusLearning without
forgetting
Knowledge transferShared
representation
Incremental
updates
Addresses ForgettingYesNoNoNo
Data DistributionNon-stationaryVariesVariesStationary
Table 5. CL scenarios.
Table 5. CL scenarios.
ScenarioTask Label
Know
Output
Space
Data
Distribution
Example Application
Task-IncrementalYesVariesChangesMulti-task NLP,
robotics
Domain-IncrementalNOSameChangesHandwriting recognition,
IoT sensors
Class-IncrementalNoExpandsChangesImage classification,
object detection
Instance-
Incremental
N/ASameSame (new
data)
Spam filtering, online
analytics
Unsupervised/OtherN/AN/AChangesClustering, RL in dynamic
settings
Table 6. Overview of TIL.
Table 6. Overview of TIL.
AspectDetails
DefinitionModels learn a sequence of distinct tasks, with task identity
provided during both training and inference.
Core ChallengeMaintaining task-specific performance without interference
between tasks (catastrophic forgetting).
Inference RequirementTask identity is known, allowing the model to use task-specific
components (e.g., separate output heads).
Key Techniques- Task-specific output heads.
- Parameter isolation (dedicated parameters for each task).
- Regularization to preserve important parameters.
Advantages- Robust retention of task-specific knowledge.
- Simplified learning due to known task boundaries and identities.
Challenges- Scalability issues with a growing number of tasks.
- Limited knowledge transfer between tasks.
Example Applications- Sequential learning of different object categories (e.g., animals, vehicles).
- Robotics: Learning distinct tasks like grasping and navigation.
- Diagnostic systems for different modalities (e.g., X-rays, MRIs).
Evaluation Metrics- Task-specific accuracy.
- Memory and computational efficiency for handling multiple tasks.
Future Directions- Modular architectures with adaptive parameter sharing and efficient knowledge transfer.
Table 7. Overview of DIL.
Table 7. Overview of DIL.
AspectDetails
DefinitionModels learn to adapt to new data distributions (domains)
over time while maintaining the same task objective.
Core ChallengeAdapting to new domains without forgetting knowledge of
previously learned domains (catastrophic forgetting).
Inference RequirementTask identity is unknown; the model must generalize across
domains without explicit domain information.
Key Techniques- Domain adaptation methods (e.g., feature alignment).
- Regularization techniques to retain domain-invariant features.
- Memory replay or dynamic models to balance old and new
knowledge.
Advantages- Allows systems to handle non-stationary data distributions.
- Maintains consistent task performance across multiple domains.
Challenges- Catastrophic forgetting when adapting to new domains.
- Handling domain-specific biases while ensuring generalization.
- Computational and memory constraints as new domains increase.
Example Applications- Object recognition in different environmental conditions
(e.g., sunny, foggy, rainy).
- Medical imaging systems adapting to scans from different
hospitals or devices.
- NLP tasks such as sentiment analysis across different domains
(e.g., movie reviews, product reviews).
Evaluation Metrics- Performance consistency across domains.
-Forgetting rate for previously learned domains.
- Domain generalization ability on unseen domains.
Future Directions- Efficient methods for domain adaptation without overfitting
to new domains.
- Scalable approaches to handle increasing numbers of domains.
- Techniques to balance domain-specific and domain-invariant
learning.
Table 8. Overview of CIL.
Table 8. Overview of CIL.
AspectDetails
DefinitionModels learn new classes sequentially, and the task identity is not
provided during inference.
Core ChallengeCatastrophic forgetting-new learning overwrites knowledge of
previously learned classes.
Inference RequirementModel must classify inputs across all learned classes without
knowledge of task identity.
Key Techniques- Memory replay (storing/replaying previous class examples).
- KD (preserving learned representations).
- Dynamic architecture (expanding capacity for new classes).
Advantages- Enables incremental learning without full retraining.
- Efficient handling of scenarios where new class data is
available over time.
Challenges- Handling class imbalance, as new classes often have fewer
examples.
- Managing memory and computational costs as the number
of classes increases.
Example Applications- Extending image classifiers with new object categories.
- Autonomous vehicles learning new traffic signs and objects.
- Healthcare models adapting to diagnose new diseases.
Evaluation Metrics- Accuracy across all classes (old and new).
- Forgetting rate (performance drop on previously learned
classes).
Future Directions- Scalable memory-efficient replay methods.
- Adaptive architectures that balance stability and plasticity.
- Improved algorithms for mitigating class imbalance and
preserving older class knowledge.
Table 9. Overview of data-incremental learning.
Table 9. Overview of data-incremental learning.
AspectDetails
DefinitionModels learn incrementally from a stream of data instances,
which may belong to existing or new classes, without
explicit task boundaries.
Core ChallengeAdapting to new data while retaining knowledge of
previously learned data, especially without clear
transitions or task identities.
Inference RequirementThe model must classify instances across all learned
classes without explicit knowledge of when new data or
classes were introduced.
Key Techniques- Memory replay (storing or generating past data).
- Regularization techniques to preserve critical parameters.
- Dynamic architectures for flexible capacity adjustment.
Advantages- Handles continuously evolving data streams.
- Allows for learning without task-specific information or
retraining.
Challenges- Managing catastrophic forgetting as new data arrives.
- Handling class imbalance and unstructured data streams.
- Resource efficiency for memory and computational costs.
Example Applications- Object recognition systems that adapt to new categories
dynamically.
- Recommendation systems updating preferences with new
user data and items.
- Continuous monitoring systems in healthcare,
incorporating evolving signals from wearable devices.
Evaluation Metrics- Accuracy across all classes (old and new).
- Forgetting rate (performance drop on previously learned
data).
- Adaptation speed to new data.
Future Directions- Hybrid methods combining memory replay with adaptive
architectures.
- Scalable solutions for handling large and imbalanced data
streams.
- Techniques for efficient data prioritization and
representation learning.
Table 10. Overview of other emerging paradigms in CL.
Table 10. Overview of other emerging paradigms in CL.
ParadigmDescriptionKey ChallengesKey TechniquesExample Applications
Few-Shot CLModels learn new tasks or classes with minimal labeled data while retaining prior knowledge.- Adapting with limited data.
- Avoiding catastrophic forgetting.
- Meta-learning.
- Episodic memory.
- Generative replay.
- Rare disease diagnosis.
- Few-shot object recognition.
Unsupervised CLModels learn from data streams without explicit labels by discovering patterns or structures.- Extracting meaningful features from unlabeled data.
- Balancing old and new pattern representations.
- Self-supervised learning.
- Contrastive learning.
- Clustering methods.
- Video surveillance anomaly detection.
- Social media trend analysis.
Meta-Continual LearningCombines meta-learning with CL to enable rapid adaptation to new tasks.- Balancing fast adaptation with knowledge retention.
- Stability–plasticity trade-off.
- Gradient-based meta-learning.
- Memory-augmented neural networks.
- Personalized AI assistants.
- Adaptive recommendation systems.
Federated CLModels learn incrementally across distributed nodes while preserving privacy.- Handling heterogeneous data distributions across nodes.
- Avoiding forgetting across distributed devices.
- Privacy concerns.
- Decentralized learning algorithms.
- Secure aggregation protocols.
- Adaptive synchronization methods.
- Personalized healthcare monitoring.
- Mobile device personalization.
Multi-Agent CLMultiple agents learn and adapt in a shared environment while interacting and collaborating.- Coordinating knowledge transfer between agents.
- Managing inter-agent dependencies and scalability.
- Communication protocols.
- Shared memory systems.
- Ensemble learning.
- Collaborative robotics.
- Distributed sensor networks.
Table 11. Overview of major theoretical foundations of CL.
Table 11. Overview of major theoretical foundations of CL.
ConceptCore IdeaMain ChallengeRepresentative Strategies
Stability–Plasticity DilemmaBalancing retention of prior knowledge with adaptation to new informationExcessive stability limits adaptation, while excessive plasticity causes forgettingRegularization, replay mechanisms, adaptive architectures
Catastrophic ForgettingLearning new tasks degrades performance on earlier tasksParameter interference and overlapping representationsReplay methods, parameter isolation, knowledge distillation, regularization
Forward and Backward TransferLeveraging previous knowledge to improve future learning and vice versaAvoiding negative transfer across tasksShared representations, multi-task learning, transferable feature learning
Representation LearningLearning reusable and task-invariant feature representationsSeparating task-specific and generalizable featuresSelf-supervised learning, contrastive learning, feature disentanglement
Neuroscientific InspirationDrawing inspiration from biological memory and adaptation mechanismsTranslating biological principles into scalable AI systemsSynaptic consolidation, rehearsal mechanisms, dynamic expansion
Practical and Ethical ConsiderationsEnsuring reliable and responsible continual adaptationResource constraints, fairness, privacy, and safetyLightweight models, federated learning, fairness-aware training
Table 12. Comparison between catastrophic forgetting and negative transfer.
Table 12. Comparison between catastrophic forgetting and negative transfer.
AspectDescription
Catastrophic ForgettingLearning new tasks reduces performance on previously learned tasks.
Negative TransferKnowledge from previous tasks impairs learning or generalization on future tasks.
Primary CauseParameter overwriting and representation drift.
Primary Cause of Negative TransferConflicting task representations and gradient interference.
Typical SolutionsReplay, regularization, parameter isolation, knowledge distillation.
Negative Transfer MitigationGradient projection, modular networks, prompt tuning, task similarity estimation, parameter-efficient adaptation.
Table 13. Unique challenges in foundation-model CL.
Table 13. Unique challenges in foundation-model CL.
ChallengeCauseOpen Research Direction
Prompt interferenceCompetition among task-specific prompts during long continual learning sequencesPrompt routing, prompt composition, dynamic prompt allocation
Multimodal alignment collapseDrift in shared embedding space during continual multimodal adaptationAlignment-preserving continual optimization and modality-aware replay
Adapter accumulationIncreasing number of task-specific adapters and LoRA modulesAdapter compression, parameter sharing, dynamic module selection
Instruction driftInstruction-following ability changes after continual fine-tuningInstruction-aware continual learning and alignment regularization
Long-term reasoning degradationSequential updates alter pretrained reasoning capabilitiesKnowledge consolidation and reasoning-aware continual adaptation
Evaluation inconsistencyLack of benchmarks for foundation-model continual learningUnified benchmarks including multimodal reasoning and long-term adaptation
Table 14. Overview of catastrophic forgetting.
Table 14. Overview of catastrophic forgetting.
AspectDescription
DefinitionThe significant loss of performance on previously learned tasks when a neural network learns new tasks.
CauseOverwriting of neural network parameters due to global updates during training on new tasks.
Key MechanismParameter Drift: Critical parameters for previous tasks are modified to optimize new task learning.
Factors Exacerbating Forgetting- Overlapping representations shared by different tasks.
- Sequential data access without revisiting earlier tasks.
- Lack of task awareness during inference in class-/domain-incremental settings.
Examples- A model trained to classify animals forgetting how to classify vehicles after learning new classes.
- An object detection model in autonomous driving failing to recognize stop signs after adapting to new road signs.
Table 15. Overview of mitigation strategies.
Table 15. Overview of mitigation strategies.
Mitigation StrategyDescriptionExamples
Regularization MethodsIntroduce constraints during training to prevent significant updates to parameters crucial for earlier tasks.- EWC: Penalizes parameter changes.
- Synaptic Intelligence: Tracks parameter importance.
Replay-Based MethodsRetain and replay data from previous tasks during training on new tasks.- Experience Replay: Stores a subset of prior task data.
- Generative Replay: Generates synthetic data from past tasks.
Dynamic ArchitecturesExpand or adapt the network architecture to allocate new resources for each task.- Progressive Neural Networks: Adds new parameters per task.
- Dynamically expandable networks.
Representation LearningLearns generalizable features that can be reused across tasks, reducing task-specific interference.- Self-supervised pretraining.
- Disentangled representations.
Hybrid ApproachesCombine multiple strategies, such as regularization with replay or dynamic architectures.- Replay with EWC to balance plasticity and stability.
Evaluation MetricsDescription
Forgetting RateMeasures the drop in performance on previously learned tasks after learning new ones.
AccuracyAssesses performance across all tasks (old and new).
Knowledge TransferEvaluates how well the model uses previous knowledge to improve learning on new tasks.
Table 16. Comparison between experience replay and generative replay.
Table 16. Comparison between experience replay and generative replay.
AspectExperience ReplayGenerative Replay
Memory usageStores selected raw samples or compressed examples.Stores a generative model that synthesizes previous data.
PrivacyMay be problematic when previous data are sensitive.Avoids direct storage of raw samples but may still leak information if not properly controlled.
Replay qualityHigh fidelity because original samples are replayed.Depends on generator quality, diversity, and label consistency.
Computational costRelatively low compared with training a generator.Higher due to training and maintaining a generative model.
Best suited forClass-incremental learning and reinforcement learning when memory is available.Privacy-sensitive or memory-constrained settings where raw data cannot be stored.
Table 17. Practical comparison of representative continual learning paradigms from a deployment perspective. The ratings summarize representative trends reported in the literature and should not be interpreted as absolute rankings.
Table 17. Practical comparison of representative continual learning paradigms from a deployment perspective. The ratings summarize representative trends reported in the literature and should not be interpreted as absolute rankings.
Method CategoryMemoryComp.ScalabilityBenchmark
Performance
Suitable CL
Scenarios
Foundation
Models
Major Limitation
Regularization-BasedLowLowHighMediumTIL, DILLimitedPerformance degrades under long task sequences and severe domain shifts.
Replay-BasedHighMediumMediumHighCIL, DIL, OnlineModerateRequires large replay memory and raises privacy concerns.
Architecture-BasedHighMediumLowHighTIL, CILModerateModel size continuously increases as new tasks are added.
Optimization-BasedLowHighMediumMediumOnline, CILLimitedHigh optimization cost due to gradient conflict management.
Representation LearningMediumMediumHighHighDIL, SSL, MultimodalHighSensitive to representation drift under significant distribution shifts.
Prompt-BasedLowLowHighHighCIL, MultimodalExcellentPrompt interference may occur during long continual adaptation.
PEFTLowLowHighHighFoundation-Model CLExcellentAccumulation of adapters increases long-term management complexity.
Federated Continual LearningMediumHighMediumMediumDistributed CLModerateCommunication overhead and heterogeneous client distributions.
Table 18. Representative quantitative comparison of major continual learning paradigms reported on widely used continual learning benchmarks. Values indicate representative performance trends reported in the literature rather than direct head-to-head comparisons because different studies employ different backbone architectures, replay budgets, datasets, and evaluation protocols.
Table 18. Representative quantitative comparison of major continual learning paradigms reported on widely used continual learning benchmarks. Values indicate representative performance trends reported in the literature rather than direct head-to-head comparisons because different studies employ different backbone architectures, replay budgets, datasets, and evaluation protocols.
Method CategoryFinal
Accuracy
ForgettingMemory
Cost
Training
Cost
ScalabilityOfficial
Code
Representative Methods
RegularizationMediumHighVery LowLowHighYesEWC [12], SI [211], and MAS [212]
ReplayHighLowHighMediumMediumYesER [213], DER++ [186], and GDumb [188]
Architecture-BasedHighVery LowVery HighMediumLowPartialPNN [214] and Expert Gate [61]
Optimization-BasedMedium–HighMediumLowHighMediumYesGEM [57] and A-GEM [164]
Representation LearningHighMediumMediumMediumHighYesSupCon [215] and self-supervised CL [216]
Prompt LearningHighLowVery LowLowHighYesL2P [207], DualPrompt [206], and CODA-Prompt [208]
PEFTHighLowLowLowHighYesLoRA [210] and adapters [217]
Foundation ModelsVery HighMediumMediumVery HighMediumYesRanPAC [209] and TRACE [218]
Table 19. Representative performance reported in recent continual learning literature.
Table 19. Representative performance reported in recent continual learning literature.
MethodSplit CIFAR-100Split TinyImageNetImageNet-RAverage Forgetting
EWC [12]58–6542–4835–40High
SI [211]60–6643–4936–41High
ER [213]70–7856–6348–54Low
DER++ [186]73–8160–6752–58Very Low
A-GEM [164]63–7048–5540–46Medium
L2P [207]82–8670–7562–66Very Low
DualPrompt [206]84–8872–7764–69Very Low
LoRA-based CL [210]83–8771–7663–68Low
Table 20. Proposed standardized evaluation checklist for continual learning research.
Table 20. Proposed standardized evaluation checklist for continual learning research.
Evaluation ComponentRecommended Reporting Practice
Continual learning scenarioClearly specify whether the study follows task-, domain-, class-, online-, or data-incremental learning.
Benchmark datasetsReport all datasets, task order, number of tasks, number of classes, and train/test splits.
Model architectureDescribe backbone network, parameter initialization, pretrained weights, and continual learning components.
Memory budgetReport replay memory size, storage strategy, memory sampling policy, and whether memory size is fixed or adaptive.
Evaluation metricsReport final average accuracy (FAA), average forgetting (AF), backward transfer (BWT), forward transfer (FWT), and task-wise performance whenever applicable.
Computational efficiencyReport training time, inference time, model parameters, FLOPs, GPU memory usage, and computational overhead.
Statistical analysisReport mean and standard deviation over multiple random seeds together with statistical significance tests whenever appropriate.
Baseline comparisonCompare against representative methods from replay-based, regularization-based, architecture-based, optimization-based, and parameter-efficient continual learning.
Implementation detailsProvide optimizer, learning rate, batch size, number of epochs, hardware configuration, and software framework.
ReproducibilityRelease source code, pretrained models, dataset preprocessing scripts, random seeds, and configuration files whenever possible.
Table 21. Comparison of representative CL benchmarks and their characteristics.
Table 21. Comparison of representative CL benchmarks and their characteristics.
BenchmarkDomainTypical ScenarioCommon MetricsFoundation ModelsMain Limitations
Split CIFAR-100VisionClass-ILAccuracy, ForgettingLimitedSmall-scale dataset with limited semantic diversity.
Split Tiny-ImageNetVisionClass-ILAccuracy, BWT, FWTModerateLower visual diversity than ImageNet and limited realism.
Split ImageNetVisionClass-ILAccuracy, ForgettingExcellentHigh computational cost and long training time.
CORe50VisionDomain-ILAccuracyModerateLimited object diversity and relatively small scale.
DomainNetVisionDomain-ILAccuracy, Domain GeneralizationExcellentDomain imbalance and high computational requirements.
CLVision BenchmarkVisionMultiple CL SettingsFAA, AF, BWTExcellentProtocols vary across benchmark configurations.
CLUE/NLP BenchmarksNLPTask-IL, Domain-ILAccuracy, F1ExcellentLimited long-horizon continual adaptation.
LLM Continual Learning BenchmarksLLMsInstruction Continual LearningTask Accuracy, ForgettingNativeLack of standardized protocols and high computational cost.
Table 22. Representative quantitative performance of influential continual learning methods reported in the literature. The results are reproduced from the respective original papers and are intended to illustrate representative performance trends rather than direct comparisons, since experimental protocols differ across studies.
Table 22. Representative quantitative performance of influential continual learning methods reported in the literature. The results are reproduced from the respective original papers and are intended to illustrate representative performance trends rather than direct comparisons, since experimental protocols differ across studies.
MethodCategoryBenchmarkFinal Avg. Accuracy (%)Avg. Forgetting (%)
EWC [12]RegularizationSplit CIFAR-100≈58≈18
iCaRL [66]ReplaySplit CIFAR-100≈64≈13
DER++ [186]ReplaySplit CIFAR-100≈74≈7
L2P [207]Prompt-basedSplit ImageNet-R≈81Low
DualPrompt [206]Prompt-basedSplit ImageNet-R≈84Very Low
CODA-Prompt [208]Prompt-basedSplit ImageNet-R≈87Very Low
DER++ [219]ReplayCORe50≈87Low
Table 23. Comparison of major continual learning method categories with computational considerations.
Table 23. Comparison of major continual learning method categories with computational considerations.
Method CategoryKey CharacteristicsMain ChallengesReplay MemoryParameter OverheadComputational OverheadTypical Applications
Regularization-Based MethodsConstrain parameter updates to preserve previous knowledge; memory-efficient and easy to integrateLimited performance under severe domain shifts and long task sequencesNoneLowLowTask-Incremental Learning, resource-constrained systems, privacy-sensitive applications
Replay-Based MethodsReplay stored or generated samples to reinforce previous knowledge; strong retention performanceReplay buffer management, privacy concerns, and storage overheadRequired (fixed-size buffer or synthetic replay)LowModerateClass-incremental learning, reinforcement learning, streaming adaptation
Architecture-Based MethodsAllocate task-specific modules or expandable subnetworks to reduce interferencePoor scalability due to parameter growth and increasing model complexityNoneHigh (grows with tasks)Moderate–HighTask-Incremental Learning with explicit task boundaries
Optimization-Based MethodsModify gradient updates to balance stability and plasticity during trainingHigh optimization complexity and gradient computation overheadNoneLowHighGradient-constrained continual adaptation and stability-focused learning
Representation-Learning MethodsLearn transferable and domain-invariant feature representations across tasksRepresentation drift under highly heterogeneous task distributionsOptionalLow–ModerateModerateDomain-incremental learning and self-supervised continual adaptation
Prompt-Based and PEFT MethodsAdapt pretrained foundation models using prompts, adapters, or low-rank updatesPrompt interference, adapter scalability, and long-term stabilityNoneVery Low (small trainable modules)LowFoundation models, multimodal systems, and large-scale deployment
Federated and Privacy-Aware CLEnable continual learning across distributed clients without centralized data sharingClient drift, communication overhead, and heterogeneous data distributionsOptional (local buffers)ModerateHigh (communication + synchronization)Healthcare, finance, edge AI, and mobile systems
Table 24. Comparison of resource requirements across major continual learning paradigms.
Table 24. Comparison of resource requirements across major continual learning paradigms.
Method CategoryTraining MemoryInference MemoryLong-Term StoragePrimary Source of Resource Consumption
Regularization-BasedLow–ModerateLowLowParameter importance statistics (e.g., Fisher information and importance weights).
Replay-BasedHighLowHighReplay buffer or generative model used to preserve previous knowledge.
Architecture-BasedModerateHighModerateAdditional task-specific subnetworks, adapters, or classifier heads.
Optimization-BasedModerate–HighLowLowGradient manipulation and optimization statistics.
Representation-LearningModerateLow–ModerateLow–ModerateFeature representations and auxiliary embedding spaces.
Prompt-Based and PEFTLowLow–ModerateModerateTask-specific prompts, adapters, or low-rank parameter updates.
Federated Continual LearningModerateModerateModerateLocal models, communication buffers, and client synchronization.
Table 25. Summary of key applications of CL, focusing on their description, benefits, and examples.
Table 25. Summary of key applications of CL, focusing on their description, benefits, and examples.
Application AreaDescriptionKey BenefitsExamples
Healthcare
and Medical
Imaging
Enables dynamic adaptation
to evolving medical
knowledge, diseases, and
patient data over time.
Personalized diagnostics,
improved adaptability, and
long-term patient monitoring.
Radiology systems adapting
to new imaging techniques
or emerging diseases like
novel cancer types.
Robotics and
Autonomous
Systems
Allows robots and autonomous
systems to learn new tasks,
adapt to dynamic environments,
and retain prior knowledge.
Efficient task performance,
knowledge transfer, and
adaptability in real-world
scenarios.
Household robots learning
new cleaning techniques
while retaining old
capabilities like object
recognition.
Natural
Language
Processing
(NLP)
Helps models stay updated with
evolving language patterns,
domain-specific knowledge,
and user preferences.
Better understanding of new
language constructs,
improved domain adaptation,
and enhanced usability.
Chatbots adapting to new
slang or technical jargon
while maintaining general
conversational abilities.
Recommender
Systems
Adapts to changing user
preferences and updates
content or product catalogs
dynamically.
Improved user engagement,
personalized recommendations,
and scalability for diverse user
bases.
Streaming platforms
suggesting trending shows
based on current
preferences without
forgetting past ones.
CybersecurityLearns from new attack
patterns and threat vectors
while retaining the ability
to recognize older threats.
Improved security, real-time
threat detection, and reduced
vulnerability to emerging
cyberattacks.
Intrusion detection systems
identifying novel malware
while protecting against
traditional viruses.
Table 26. Representative continual learning benchmarks and evaluation metrics across major application domains.
Table 26. Representative continual learning benchmarks and evaluation metrics across major application domains.
ApplicationRepresentative BenchmarksCommon Evaluation MetricsDomain-Specific Continual Learning Challenge
Computer VisionSplit MNIST, Split CIFAR-100, Tiny ImageNet, CORe50Average Accuracy, Forgetting, BWT, FWTLarge class expansion, domain shifts, long task sequences
HealthcareBraTS, ISIC, CAMUS, EchoNet-Dynamic, CheXpertDice, IoU, HD95, AUC, Sensitivity, Average ForgettingPrivacy constraints, scarce annotations, clinical reliability
RoboticsMeta-World, RoboSuite, ManiSkill, HabitatTask Success Rate, Cumulative Reward, Adaptation SpeedEmbodied interaction, non-stationary environments, safety
Natural Language ProcessingCLINC150, AG News, Amazon Reviews, SuperGLUEAccuracy, Macro-F1, BLEU, ROUGE, PerplexityVocabulary expansion, instruction drift, long-context retention
Autonomous DrivingBDD100K, nuScenes, Waymo Open DatasetmAP, mIoU, Recall, Adaptation AccuracyWeather variation, continual object discovery, safety-critical deployment
Table 27. Comparative view of the challenges in CL, their impacts, examples, and potential solutions.
Table 27. Comparative view of the challenges in CL, their impacts, examples, and potential solutions.
ChallengeDescriptionImpactExamplesPotential Solutions
Catastrophic
Forgetting
Overwriting of previous
knowledge when
learning new tasks.
Loss of performance on
earlier tasks, limiting
multi-task applications.
A model trained on new
object classes forgets
previously learned ones.
Replay methods,
regularization techniques
(e.g., EWC, SI),
parameter isolation
(e.g., PNNs, PackNet).
Scalability to
Real-World
Tasks
Difficulty in handling
diverse, undefined, and
open-ended tasks
found in real-world
environments.
Limits practical
applications, especially
in dynamic or multi-
domain environments.
A robot operating in a
dynamic home environment
fails to generalize across
diverse tasks.
Dynamic architectures
(e.g., expandable networks),
meta-learning, unsupervised
task detection.
Memory
Constraints
Storing data from
previous tasks is often
infeasible for large-scale
or resource-limited
applications.
Limits model ability to
effectively retain and
replay past information.
Replay-based methods
requiring storage of
vast datasets for
continual adaptation.
Efficient memory management
techniques, synthetic replay
using generative models, data
pruning.
Computational
Overhead
Increased computational
demands for training
and inference due to replay,
regularization, or parameter
isolation techniques.
Hinders real-time
applications on edge
devices or systems
with limited resources.
On-device CL
in IoT systems is slowed by
high computational
requirements.
Lightweight models, parameter
optimization, pruning, and
efficient task-specific parameter
allocation.
Bias
Amplification
Sequential learning may
reinforce biases present
in earlier data or tasks.
Skewed model behavior,
disproportionately
affecting certain
demographic groups.
A financial model favoring
certain demographics due
to biased historical data.
Fairness-aware training, regular
bias audits, diversity-focused
data augmentation.
Transparency
and
Explainability
Models evolving continuously
can become opaque, making
their decision-making hard
to interpret.
Erodes trust, particularly
in sensitive applications
like healthcare or finance.
Difficulty auditing a
continually adapting
medical diagnostic
system.
Explainability frameworks,
interpretable architecture
designs, and model debugging
tools.
Privacy
Concerns
Replay-based methods
storing or processing user
data may violate privacy
regulations.
Non-compliance with
privacy laws (e.g., GDPR,
HIPAA), leading to legal
and ethical implications.
Retaining user data for
replay in recommendation
systems could breach user
consent.
Privacy-preserving methods
like federated learning,
data anonymization, and
synthetic data generation.
Unintended
Consequences
Autonomous learning systems
may exhibit behaviors or
decisions not aligned with
human intentions or societal
norms.
Potential safety risks,
ethical conflicts, or
misaligned system
behavior in real-world
scenarios.
A self-learning robot adopts
unsafe behaviors while
optimizing a task
autonomously.
Strict behavioral constraints,
ethical guidelines for
autonomous systems, and
robust oversight mechanisms
during model deployment.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ullah, Z.; Hong, M.; Kim, J. Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions. Mathematics 2026, 14, 2774. https://doi.org/10.3390/math14152774

AMA Style

Ullah Z, Hong M, Kim J. Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions. Mathematics. 2026; 14(15):2774. https://doi.org/10.3390/math14152774

Chicago/Turabian Style

Ullah, Zahid, Minki Hong, and Jihie Kim. 2026. "Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions" Mathematics 14, no. 15: 2774. https://doi.org/10.3390/math14152774

APA Style

Ullah, Z., Hong, M., & Kim, J. (2026). Modern Continual Learning with Foundation Models, Evaluation Challenges, and Future Directions. Mathematics, 14(15), 2774. https://doi.org/10.3390/math14152774

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop