1. Introduction
Human learning naturally unfolds over time through practice, feedback, interaction, reflection, and application across different settings. Understanding and supporting that process depends on meaningful evidence of what learners know, can do, and how they develop over time. Educational assessment provides a systematic means of generating, interpreting, and using that evidence. Historically, educational systems have accomplished this through discrete assessment events—including examinations, assignments, projects, and classroom performances—that can be efficiently administered, evaluated, and translated into grades, transcripts, and credentials. These assessment events provide efficient and potentially valuable summaries of achievement. Because they capture learning at discrete points in time, however, any individual assessment offers only a partial view of the developmental processes through which learners build, refine, and apply knowledge and competencies [
1].
The prevalence of episodic assessment reflects both evolving conceptions of assessment and the historical conditions under which large educational systems developed. Industrial-era institutions required approaches that supported standardization, administrative efficiency, and the evaluation of large numbers of learners [
2,
3,
4,
5,
6,
7]. Discrete assessment events provided a practical means of eliciting, evaluating, recording, and communicating evidence of student performance. Their continued use reflects these important educational and institutional functions. At the same time, because each assessment captures performance at a particular point in time, episodic assessment provides only a partial representation of learning trajectories that unfold across activities, contexts, and time.
Research spanning psychometrics, assessment, and the learning sciences has progressively broadened what counts as meaningful evidence of learning. This work emphasizes that competence develops through repeated performance, feedback, revision, collaboration, and participation in contextualized activity [
8,
9,
10,
11,
12,
13]. Programmatic assessment similarly demonstrates the value of combining multiple observations over time rather than relying on a single high-stakes event [
14,
15,
16]. These traditions do not eliminate the need for examinations, assignments, or professional judgment. Instead, they show the value of interpreting individual assessment events as parts of a larger evidentiary record.
Learning now occurs across an expanding range of physical, digital, educational, professional, and social environments. Learners write and revise within digital platforms, interact with AI tutors, participate in simulations and collaborative projects, construct portfolios, and apply knowledge in workplaces. These activities generate potentially meaningful evidence of learning, but that evidence typically remains dispersed across systems, represented in incompatible forms, and disconnected from later educational decisions.
Advances in artificial intelligence, learning analytics, multimodal analytics, learner modeling, natural language processing, and distributed computation render it increasingly feasible to connect, integrate, and interpret meaningful evidence of learning across contexts and over time [
17,
18,
19,
20,
21]. These capacities, however, do not automatically produce valid, equitable, or educationally useful assessments. Learning evidence remains distributed across heterogeneous platforms, represented in incompatible formats, interpreted using different models, and governed under different institutional policies. Realizing the potential of these advances, therefore, requires an infrastructure that can connect, interpret, and govern evidence across contexts and over time while supporting valid, equitable, and trustworthy educational decisions [
18,
22,
23,
24].
This paper introduces the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), a conceptual sociotechnical framework for connecting, interpreting, and governing evidence of learning across contexts and time. AIM-CAI is not proposed as a centralized learner database, a fully developed technical system, or a mechanism for comprehensive learner monitoring. It describes a federated evidence infrastructure through which locally governed systems may contribute selected evidence to context-sensitive and probabilistic interpretations of learner development.
In this paper, continuous assessment refers to the longitudinal accumulation and interpretation of evidence generated through meaningful learning activities. Evidence may be generated at different intervals, through different methods, and for different purposes as learners engage in diverse educational experiences. The defining characteristic is that interpretations of learner development can be updated as relevant evidence accumulates over time, rather than remaining tied to a single assessment occasion. Episodic assessment therefore becomes one important source of evidence within a broader longitudinal evidence system.
The framework integrates six functions required to move responsibly from learning activity to educational interpretation: generating evidence across distributed environments; structuring evidence in forms that preserve its source and context; translating evidence across systems and competency representations; evaluating its relevance through psychometric and inferential models; representing development longitudinally through dynamic competency profiles; and governing how evidence and inferences are accessed, interpreted, contested, and used. These functional capabilities are represented as six interconnected layers: learning evidence generation, learning evidence representation, learning evidence translation, evidence interpretation, learner representation, and governance and trust.
The individual components of AIM-CAI build on established traditions, including evidence-centered design, formative and programmatic assessment, stealth assessment, learning analytics, learner modeling, semantic interoperability, and privacy-preserving computation [
8,
9,
14,
25,
26,
27,
28]. The contribution of AIM-CAI lies in specifying how these capabilities may operate together as an integrated infrastructure for longitudinal evidence interpretation. Its novelty is therefore architectural and integrative rather than based on the invention of each constituent component.
AIM-CAI is presented as a reference architecture and research agenda for advancing AI-mediated continuous assessment. It organizes the functional capabilities required to connect, interpret, represent, and govern learning evidence across educational contexts while identifying the empirical, psychometric, technical, and governance challenges that remain. Continued research and bounded implementations will refine these capabilities, evaluate their educational utility, and strengthen the evidence needed to support trustworthy educational decision-making. Throughout this work, the central purpose remains unchanged: using trustworthy evidence to provide more timely, meaningful, and informed support for learning.
This manuscript is organized into eight sections.
Section 2 examines the historical and practical limitations of assessment systems organized around bounded, cumulative events.
Section 3 explores how advances in artificial intelligence, learning analytics, multimodal analytics, blockchain technologies, and learner modeling are reshaping the feasibility conditions of continuous assessment.
Section 4 introduces the proposed AIM-CAI framework, while
Section 5 examines the governance, interoperability, and trust architectures required to support continuous inferential ecosystems responsibly.
Section 6 considers the broader implications for educational institutions, credentialing systems, and lifelong learning.
Section 7 outlines a potential research agenda addressing unresolved psychometric, technical, governance, and sociotechnical challenges along with possible empirical validation and pilot portfolios. Finally,
Section 8 concludes the paper.
2. The Limits of Episodic Assessment
Although episodic assessment provides valuable evidence of learner achievement, each assessment event captures only one observation within a broader developmental process. Over the past several decades, however, research across assessment and the learning sciences has progressively expanded conceptions of learning, emphasizing that knowledge and competencies develop through practice, feedback, revision, collaboration, and participation across contexts and over time [
1,
8,
9,
10,
11,
12,
13]. This broader conception of learning also expands the kinds of evidence needed to understand learner development, highlighting several inferential and institutional boundaries that shape the evidence episodic assessment can reasonably provide. Inferential boundaries concern the conclusions that can reasonably be drawn from available evidence, whereas institutional boundaries concern where evidence of learning is generated, recognized, and incorporated into educational decision-making.
One important inferential boundary emerges from the developmental nature of learning. Assessments based on short-term performance in decontextualized tasks provide valuable evidence of some forms of knowledge and skill, but they are less well suited to representing adaptive expertise, collaboration, creativity, or the transfer of learning across contexts [
8,
11,
29]. Formative assessment research similarly demonstrates that learning develops through ongoing interaction, feedback, revision, and participation, with each learning experience contributing additional evidence of learner development [
9,
30]. Episodic assessment provides valuable but necessarily incomplete evidence of competencies that emerge and evolve over time.
The developmental nature of learning also has important implications for how assessment evidence supports learning. Research on formative assessment demonstrates that learning improves when assessment is integrated into instruction and evidence is used to provide timely feedback that informs subsequent learning [
9]. Black and Wiliam [
9] argued that the distinction between formative and summative assessment lies not simply in the timing of assessment events but in whether the resulting evidence informs learners’ next steps. Similarly, Hattie and Timperley [
30] emphasized that effective feedback reduces the gap between current understanding and desired performance. When evidence becomes available only after instructional sequences have concluded, opportunities for revision, adaptation, and continued learning are more limited.
A second inferential boundary concerns the alignment between assessment tasks and the competencies they are intended to represent. Traditional assessments often rely on structured, decontextualized tasks that provide valuable evidence of specific knowledge and skills but may be less well suited to representing complex judgment, problem-solving, collaboration, and adaptability in authentic settings [
11]. Authentic assessment research emphasizes that meaningful competence is often demonstrated through performance in contexts that resemble the situations in which knowledge is ultimately applied [
31]. This perspective aligns with situated learning theory, which views learning as developing through participation in authentic social and professional contexts rather than through isolated demonstrations of knowledge [
13]. Evidence generated through traditional examinations provides an important perspective on learner competence, but it represents only part of the broader capabilities demonstrated across authentic educational, professional, and social settings.
A third inferential boundary concerns the sufficiency of evidence for supporting valid conclusions about learner competence. Assessment is fundamentally a process of evidentiary reasoning in which observations of learner performance are interpreted to support inferences about underlying knowledge and competencies [
8]. Pellegrino et al. [
8] emphasized that valid assessment depends on coherent alignment among models of learning, the tasks used to elicit evidence, and the interpretive frameworks used to draw conclusions from that evidence. Because episodic assessment relies on a limited number of observations collected at particular points in time, the available evidence may provide an insufficient basis for drawing valid inferences about developmental growth or complex competencies.
Several assessment approaches have sought to address these inferential boundaries by drawing on evidence from multiple observations rather than relying on a single assessment event. Among the most influential is programmatic assessment, which proposes that meaningful judgments about learner competence should emerge from the aggregation of multiple low-stakes observations collected across contexts, evaluators, and time [
14,
15,
16]. Similar principles are reflected in competency-based education (CBE) and frameworks for assessing twenty-first century skills, both of which recognize that complex competencies require multiple forms of evidence to support valid interpretation [
29,
32]. These approaches reflect growing recognition that learner competence is best understood through the accumulation of evidence across learning experiences rather than through isolated assessment snapshots.
Traditional assessment systems remain closely tied to formal educational programs, courses, and institutional credentialing processes. Contemporary learners, however, increasingly develop knowledge and competencies through workplace participation, collaborative networks, online communities, self-directed learning, and other experiences that extend beyond formal educational settings. Much of this evidence, however, remains disconnected from formal assessment and credentialing systems, even as portfolios, micro-credentials, and alternative credentials have expanded the range of recognized learning experiences [
33,
34].
These inferential and institutional boundaries emerged under historical conditions that rendered continuous observation and interpretation of learner development impractical. Sustained assessment requires multidimensional evidence, ongoing interpretation, and continuous feedback—activities that traditionally exceeded the operational capacity of educational institutions operating at scale. Episodic assessment, therefore, became an effective and scalable approach for generating evidence to support educational decision-making.
Although episodic assessment continues to provide valuable evidence for educational decision-making, contemporary conceptions of learning increasingly emphasize evidence that is developmental, contextual, longitudinal, and distributed across learning experiences. Until recently, the practical challenges of capturing, integrating, and interpreting such diverse evidence limited the feasibility of extending assessment beyond episodic events. Advances described in the following section increasingly change these feasibility conditions, motivating the development of infrastructures capable of connecting, interpreting, and governing learning evidence across contexts and over time.
3. Technological Advances Changing the Feasibility Conditions for Continuous Assessment
Recent advances in artificial intelligence, learning analytics, educational data mining, natural language processing, multimodal learning analytics, learner modeling, and related technologies are changing the practical feasibility of collecting, integrating, and interpreting evidence of learning across contexts and over time. Many of the limitations discussed in the previous section reflected historical and operational constraints rather than educational ideals. Continuously observing, interpreting, and documenting complex learning processes across large numbers of learners was simply impractical. These practical constraints shaped assessment systems around episodic events that could be administered, evaluated, and documented efficiently. Today, many of these operational constraints are diminishing as advances in AI and related technologies enable the scalable interpretation of rich, heterogeneous evidence generated through authentic learning activities. As a result, assessment can increasingly be viewed as an ongoing process of evidentiary reasoning supported by diverse forms of learning evidence rather than as a sequence of isolated testing events.
Expanding the capacity to collect and interpret evidence does not, by itself, support valid educational inferences. Richer evidence must still be interpreted through valid inferential models, aligned with theories of learning, and governed in ways that preserve transparency, learner agency, and appropriate human oversight. The technologies discussed in this section contribute complementary capabilities for generating, interpreting, and governing evidence within a continuous assessment infrastructure.
3.1. Expanding Sources of Learning Evidence
Contemporary digital learning environments expand the sources of evidence available for understanding learner development. As learners write, revise, collaborate, solve problems, navigate simulations, and interact with digital tools, they generate process-oriented evidence that documents how learning unfolds over time. Research in learning analytics and educational data mining has shown that these digital traces can support richer interpretations of learning processes and learner development [
18,
19]. Process-oriented evidence captures how learners engage, revise, collaborate, and solve problems over time, providing insights that complement the products of learning traditionally represented in assignments, examinations, and other assessment artifacts.
Language provides one of the richest sources of evidence about learner thinking and development. Advances in natural language processing and writing analytics have expanded the capacity to interpret linguistic, semantic, and discourse-level patterns in learner writing, supporting probabilistic inferences about conceptual understanding, reasoning, writing quality, and metacognitive strategy use [
21,
35,
36,
37]. Intelligent writing support systems, such as Writing Pal, further demonstrate how natural language processing can simultaneously provide adaptive feedback and generate evidence about learners’ writing processes and strategies [
21]. Collectively, these developments illustrate how language itself can serve as a continuous source of evidence reflecting evolving cognitive and metacognitive activity throughout learning.
Dialogue and collaboration provide additional sources of evidence about learner understanding and knowledge construction. Dialogue-based intelligent tutoring systems, such as AutoTutor, demonstrate how conversational interactions can provide evidence of conceptual understanding, explanatory reasoning, and misconceptions as learning unfolds [
37]. Collaborative learning analytics extends this perspective to peer interaction and group learning by examining patterns of participation, collaborative reasoning, and knowledge construction across extended interactions using machine learning and sequence analysis techniques [
18,
20]. These approaches illustrate how dialogue and collaboration generate process-oriented evidence that supports continuous interpretation of learning.
Multimodal interactions provide additional sources of evidence about how learning unfolds across cognitive, social, affective, and behavioral dimensions. Multimodal learning analytics (MMLA) integrates evidence from speech, gaze, movement, physiological signals, collaborative interactions, video analysis, and digital activity logs to examine learner engagement, cognitive load, emotional response, collaboration, and self-regulation throughout extended learning activities [
20,
38]. For example, eye-tracking systems provide evidence about attention allocation during complex problem-solving, while process-trace analytics in simulations and virtual laboratories capture how learners navigate tasks, test hypotheses, and regulate their learning over time [
39,
40]. Together, these approaches demonstrate how multiple streams of observable evidence can be integrated to support richer interpretations of learner development.
These developments expand both the breadth and continuity of evidence available to support educational interpretation. Contemporary learning environments can capture sequences of learner actions, revisions, decisions, interactions, and strategic behaviors, providing a richer representation of learning trajectories and competency development over time. AI systems contribute by organizing, integrating, and modeling these heterogeneous evidence streams, supporting longitudinal evidence accumulation and probabilistic inferences about learner development at scale.
3.2. Advances in Interpreting Learning Evidence
Expanding the sources of learning evidence also expands the need for principled methods to interpret that evidence. Educational decisions depend on valid inferences about learner knowledge, skills, and development supported by rich learning evidence. Research in learner modeling, intelligent tutoring systems, evidence-centered design, and embedded assessment has established many of the theoretical and computational foundations for interpreting diverse forms of learning evidence. Although recent advances in generative AI have accelerated interest in these approaches, the underlying principles have developed over several decades.
Evidence-Centered Design (ECD) provides one of the most influential frameworks for interpreting learning evidence [
25,
41]. Consistent with the Knowing What Students Know assessment triangle proposed by Pellegrino and colleagues [
8], ECD conceptualizes assessment as a process of evidentiary reasoning that connects observations of learner behavior to probabilistic claims about underlying competencies. Both frameworks emphasize that educational assessment depends on aligning theories of learning (cognition), observable evidence, and principled interpretation. The framework further emphasizes alignment among the competencies being assessed, the evidence required to support interpretive claims, and the tasks designed to elicit that evidence. Because competencies represent latent constructs, they are inferred from observable behaviors through probabilistic reasoning. ECD therefore provides a principled foundation for interpreting the diverse evidence streams generated in contemporary digital learning environments.
Because learning cannot be observed directly, educational interpretation always involves uncertainty. Observable behaviors provide evidence that increases or decreases confidence in claims about learner competencies, but they do not establish those competencies with certainty. Thus, continuous assessment depends on accumulating evidence across multiple observations, contexts, and learning activities to strengthen the validity of interpretive claims.
Embedded assessment integrates evidence collection directly into learning activities and digital environments. Stealth assessment operationalizes this approach by gathering evidence continuously as learners interact with video games, simulations, intelligent tutoring systems, and collaborative learning tasks [
27,
42,
43]. These environments capture process-oriented evidence that is interpreted through evidence models and learner models to update probabilistic estimates of learner competencies in real time. By integrating assessment within authentic learning activities, embedded assessment supports continuous interpretation of learner development as learning unfolds.
Continuous learner modeling provides a mechanism for updating interpretations of learner knowledge as evidence accumulates over time. Intelligent tutoring systems demonstrated the feasibility of this approach by using cognitive models and Bayesian Knowledge Tracing (BKT) to estimate learners’ mastery of specific knowledge components during problem-solving [
26,
44]. Learner models operationalize evidentiary reasoning by continuously updating probabilistic estimates of learner competencies as new evidence becomes available from solution paths, successful and unsuccessful attempts, hint requests, and other learning behaviors. This continuous updating supports adaptive feedback and personalized instructional support. Although many early systems focused on well-defined domains such as algebra, they established the practical foundations for continuously updating probabilistic models of learner understanding throughout the learning process.
Recent advances in artificial intelligence extend these interpretive capabilities to increasingly open-ended learning environments. Machine learning techniques support the analysis of evidence generated through complex simulations, collaborative dialogue, writing processes, coding activities, and multimodal interactions. These systems identify patterns associated with conceptual understanding, self-regulated learning, collaboration, and strategic problem-solving across extended learning experiences. The resulting inferences remain probabilistic and become progressively more informative through the accumulation of evidence over time. Contemporary computational psychometric approaches, including Bayesian networks, hidden Markov models (HMMs), and related learner modeling techniques, support this process by continuously updating competency profiles as new evidence becomes available [
44,
45,
46,
47]. These developments provide the inferential foundation for continuous assessment by enabling diverse forms of learning evidence to be interpreted, integrated, and accumulated across learning activities, instructional contexts, and educational experiences, supporting longitudinal models of learner development.
3.3. Trusted Evidence Infrastructure
Continuous assessment depends on trusted infrastructures that support the generation, interpretation, and governance of learning evidence across educational systems. As learning increasingly spans institutions, workplaces, digital platforms, and informal environments, assessment infrastructures must support evidence provenance, integrity, interoperability, portability, and governance. These capabilities enable learning evidence to persist, move across educational contexts, and remain interpretable while preserving its provenance, context, integrity, and meaning over time.
Blockchain technologies have emerged as one approach for strengthening trust in distributed learning records. By combining cryptographic verification with distributed ledgers, blockchain infrastructures can provide tamper resistance, provenance tracking, and verifiable educational records across multiple organizations [
48,
49]. These capabilities make blockchain particularly relevant for continuous assessment systems that integrate evidence from diverse learning environments.
Trusted evidence infrastructures enable learning evidence to extend beyond the boundaries of individual courses, institutions, and credentialing systems. Evidence generated across formal education, workplace learning, professional development, and informal learning experiences can be accumulated into longitudinal learner records while preserving information about the origin, context, timing, conditions, and integrity of that evidence. Blockchain-based systems, including the Blockchain of Learning Logs (BOLL) and BEMPAS, illustrate how distributed infrastructures can support portable learner records, automated credentialing, access control, and cross-institutional verification through cryptographic methods and smart contracts [
49,
50,
51]. These capabilities support the recognition of learning as a cumulative and continuously documented process that spans diverse educational contexts.
The effectiveness of trusted evidence infrastructures depends on more than just secure recordkeeping. Continuous assessment requires common data standards, interoperable competency frameworks, governance structures, privacy protections, and audit mechanisms that define how evidence is generated, validated, interpreted, shared, and used across educational systems. These governance mechanisms establish the conditions under which learning evidence can be trusted, appropriately interpreted, and responsibly incorporated into educational decisions. Trusted evidence infrastructures also support learner agency by enabling learners to inspect, manage, and selectively share evidence about their learning while providing transparency into how that evidence is interpreted and used.
Although trusted evidence infrastructures have advanced considerably, several implementation challenges remain. Blockchain-based educational systems often operate as fragmented ecosystems with limited interoperability and cross-platform compatibility [
49,
50]. Scalability also remains an important consideration, particularly in environments that generate large volumes of continuous learning evidence requiring secure verification and storage [
50,
52]. Continued progress will depend on technical advances and governance frameworks that support secure, interoperable, and sustainable evidence ecosystems [
23,
24,
48,
50].
4. AI-Mediated Continuous Assessment Infrastructure (AIM-CAI)
The preceding sections established that continuous assessment depends on three complementary capabilities: generating rich and continuous evidence of learning, interpreting that evidence through principled inferential models, and governing trusted learning evidence across educational systems. These capabilities establish the technical and conceptual foundation for continuous assessment as an evidence infrastructure rather than as a sequence of isolated assessment events.
The AI-Mediated Continuous Assessment Infrastructure (AIM-CAI) framework synthesizes these capabilities into an integrated functional architecture for continuous assessment. The architecture emerged through the synthesis of assessment, learner modeling, learning analytics, and educational infrastructure literature reviewed in the preceding sections. AIM-CAI is organized around the functional capabilities required to transform learning activities into meaningful educational decisions. Artificial intelligence expands the operational feasibility of this infrastructure by integrating heterogeneous evidence, continuously updating probabilistic learner models, and supporting the scalable interpretation of learning processes. Within AIM-CAI, AI functions as an enabling capability that supports evidence-centered educational decision-making while preserving the principles of validity, transparency, and human oversight.
Figure 1 presents the six functional layers of AIM-CAI. This framework serves as a reference architecture to organize the functional capabilities required to support continuous assessment across distributed educational environments. Each layer corresponds to a distinct function within the evidentiary chain linking learning activities to educational decisions. Collectively, the layers describe how learning evidence is generated, prepared for interpretation, connected across systems, interpreted through inferential models, accumulated into longitudinal learner representations, and governed throughout its lifecycle. The layers, therefore, represent functional requirements of a continuous assessment infrastructure. Individual technologies, instructional models, institutional workflows, and implementation strategies contribute to these functions while remaining adaptable to different educational contexts.
The six layers comprise (1) evidence generation, which captures learning activity across educational contexts; (2) evidence representation, which organizes and normalizes heterogeneous evidence into comparable representations; (3) evidence translation, which supports interoperability across systems, competency frameworks, and educational contexts; (4) evidence interpretation, which applies evidentiary reasoning and probabilistic learner modeling to generate competency inferences; (5) learner representation, which accumulates evidence into longitudinal models of learner development; and (6) governance and trust, which provides the provenance, interoperability, transparency, learner agency, privacy, and institutional oversight required to support trustworthy educational decision-making.
AIM-CAI provides a functional reference architecture for organizing continuous assessment rather than prescribing a single technical implementation. The six layers represent the functional capabilities required to transform distributed learning activities into trustworthy educational decisions. Each layer performs a distinct function in the evidentiary chain and provides the foundation for the subsequent layer. The framework is intended to support the coordinated generation, exchange, interpretation, representation, and governance of learning evidence across multiple educational systems while preserving institutional autonomy and distributed governance. Learning evidence can accumulate across courses, institutions, workplaces, and lifelong learning experiences while retaining the provenance, context, and interpretability required for meaningful educational decisions. As assessment practices, learner models, interoperability standards, and AI technologies continue to evolve, these functional capabilities provide a stable architectural foundation at scale for integrating new methods and implementations within a coherent evidence infrastructure.
The architecture exhibits six defining characteristics that emerge consistently from the assessment, learner modeling, and educational infrastructure literature synthesized throughout this paper. Assessment functions continuously through the accumulation of evidence across learning experiences rather than through isolated assessment events. Competency estimation remains inferential because learner knowledge and skills are represented as latent constructs supported by observable evidence. Educational decisions remain probabilistic because evidence strengthens confidence in competency claims without eliminating uncertainty (e.g., measurement error, incomplete evidence, incorrect inference). Learning evidence is distributed across institutions, workplaces, digital platforms, and informal learning environments. Learner representations evolve longitudinally as evidence accumulates over time. Finally, assessment functions as an interoperable infrastructure that coordinates evidence generation, interpretation, representation, and governance across educational ecosystems. The following subsections elaborate on each functional layer of the AIM-CAI framework.
4.3. Learning Evidence Translation Layer
The Learning Evidence Translation layer serves as the conceptual core of AIM-CAI. Existing educational ecosystems are characterized by profound heterogeneity across platforms, schemas, competency frameworks, and institutional boundaries. Traditional interoperability approaches rely on predefined standards such as SCORM, IMS Learning Tools Interoperability (LTI), and the Experience API (xAPI). While these support technical exchange, they depend on prior agreement on common standards and have faced persistent adoption and fragmentation challenges (discussed further in
Section 5.2).
The Learning Evidence Translation layer enables evidence representations generated within one educational context to be interpreted and exchanged across other systems, competency frameworks, and institutional environments. Educational ecosystems are inherently heterogeneous, encompassing diverse learning management systems, assessment platforms, competency frameworks, institutional data models, and instructional practices. Learning evidence, therefore, requires translation across these heterogeneous representations while preserving its educational meaning and contextual integrity. The purpose of this layer is to support semantic interoperability so that evidence generated in one context remains meaningful and usable in another.
Existing interoperability standards, including SCORM, LTI, and xAPI, support the technical exchange of learning information across educational systems. Continuous assessment, however, also requires preserving the educational meaning of evidence as it moves across diverse competency frameworks, instructional models, and institutional contexts. The Learning Evidence Translation layer, therefore, extends beyond technical interoperability by supporting semantic translation among heterogeneous evidence representations while preserving local educational context.
Artificial intelligence expands the operational feasibility of semantic translation through machine learning, natural language processing, and large language models that assist with ontology alignment, semantic mapping, and probabilistic reconciliation across competency frameworks. Rather than requiring institutions to adopt identical competency models, standardized data structures, or centralized learner repositories, AI-assisted translation enables locally meaningful representations of learning evidence to be connected across distributed educational ecosystems while preserving institutional flexibility and contextual specificity. Because semantic translation influences subsequent competency inferences, AI-generated mappings should be treated as provisional rather than authoritative. Within the Learning Evidence Translation layer, implementations should provide mechanisms for documenting the provenance, version, and degree of semantic equivalence associated with competency mappings while allowing those mappings to be reviewed and refined through appropriate institutional governance. Rather than treating competency correspondences as binary matches, translation relationships may vary in scope, specificity, performance expectations, and assessment requirements. Consequently, translation outcomes should preserve information about the confidence and completeness of each mapping so that downstream interpretive processes can account for partial or uncertain correspondences. Where semantic equivalence cannot be established with sufficient confidence, evidence should remain linked to its original competency framework or require additional human review before informing consequential educational decisions such as accreditation, certification, or credentialing.
Semantic translation has become an important architectural capability in other information-intensive domains. Healthcare informatics, for example, coordinates heterogeneous patient records, clinical ontologies, and diagnostic vocabularies across decentralized systems without requiring identical local data structures. Technologies such as the RDF, ontology alignment, and knowledge graph architectures support semantic interoperability while preserving institutional flexibility and local context. Similar approaches have begun to emerge within education. Musa et al. [
58], for example, proposed a graph-based ontology alignment framework that integrates learning analytics data across learning management systems, MOOCs, and other educational platforms through semantic mapping and reconciliation techniques.
The Learning Evidence Translation layer also introduces important interpretive challenges. Semantic translation remains probabilistic because educational concepts often vary across institutional, disciplinary, and cultural contexts. Ontology alignment systems and large language models remain susceptible to semantic drift, ambiguous mappings, and other interpretive errors that may influence downstream competency inferences [
59]. In addition, translation models trained on historically biased educational data may reproduce inequities through probabilistic classification and mapping processes [
60]. Therefore, the Learning Evidence Translation layer supports inferential compatibility under conditions of uncertainty while preserving the contextual information required for the subsequent Evidence Interpretation layer.
4.4. Evidence Interpretation Layer
The Evidence Interpretation layer transforms translated learning evidence into probabilistic inferences about learner knowledge, skills, and competencies. Because these constructs are latent, they cannot be observed directly. Instead, the layer applies evidentiary reasoning and probabilistic learner modeling to relate observable learning evidence to defensible educational interpretations [
8,
41]. Its purpose is to support valid educational interpretation by determining which evidence contributes to competency claims, how evidence is weighted, how uncertainty is represented, and how competency estimates are updated as additional evidence becomes available across learning activities, educational contexts, and time.
The Evidence Interpretation layer is grounded in established theories of educational measurement. The ECD conceptualizes assessment as an evidentiary argument that connects observations of learner behavior to claims about underlying competencies through explicitly defined interpretive models [
41]. Similarly, Kane’s argument-based approach to validity emphasizes that educational claims depend on the interpretive warrants that justify competency inferences rather than on evidence collection alone [
61]. Together, these perspectives establish the theoretical foundation for treating continuous assessment as an ongoing process of evidentiary reasoning supported by accumulating evidence rather than as a sequence of isolated measurements.
Within the Evidence Interpretation layer, implementations must evaluate the dependence among evidence sources and the availability of evidence before updating competency estimates. For example, records generated within the same task, learning episode, or platform may provide redundant or overlapping information. Related observations can be identified through provenance, timing, and contextual metadata, and then grouped such that their combined contribution does not artificially inflate model confidence. Implementations should also account for learners’ opportunities to generate evidence. Limited evidentiary coverage from sparse or uneven records increases uncertainty estimates. Estimates should therefore reflect the quality, diversity, relevance, and independence of observations, along with the range of opportunities from which they were generated. These considerations reduce advantages associated with greater platform use and support fairer interpretations across learners.
Behavioral evidence frequently contains construct-irrelevant variance arising from interface characteristics, temporary contextual influences, social dynamics, fatigue, cultural differences, or other factors unrelated to the competencies being assessed. Multimodal learning analytics and computational psychometrics, therefore, emphasize that behavioral telemetry requires theoretically grounded interpretation to support valid educational inferences [
20,
60,
61]. The quality of competency inferences depends fundamentally on the quality of both the available evidence and the interpretive models applied to that evidence. Poorly calibrated probabilistic models, noisy observations, or insufficient theoretical grounding may reduce inferential accuracy [
20,
62,
63]. The Evidence Interpretation layer, therefore, provides the functional capability for distinguishing meaningful educational signals from contextual artifacts, ambiguity, redundant observations, and noise while supporting psychometric calibration, uncertainty estimation, and validation.
The Evidence Interpretation layer also supports continuous updating of competency inferences as new evidence becomes available. Bayesian networks, dynamic Bayesian networks (DBN), knowledge tracing approaches, and related probabilistic learner models estimate evolving learner states through sequential evidence accumulation [
44,
45,
46]. The resulting competency inferences do not constitute direct measurements of learner capability. Instead, they represent probabilistic interpretations conditioned on incomplete and heterogeneous evidence that remain open to revision as additional evidence becomes available. This continuous updating allows persistent developmental patterns to emerge while distinguishing durable learning from temporary fluctuations in performance. For example, consider a learner completing a programming assignment within an introductory engineering course. Observable evidence may include the correctness of the submitted solution, the sequence of code revisions, the frequency of hint requests, response times, and reflective explanations accompanying the final submission. Individually, none of these observations directly demonstrates competency in recursion or algorithmic reasoning. For instance, if the learner adds a valid base case and revises the recursive call so that the input progresses toward that base case, these actions may support the claim that the learner recognizes key structural requirements for recursive termination within the task. They would not, however, establish that the learner can independently design a recursive algorithm or transfer this understanding to a novel problem. Similarly, a reflective explanation that each recursive call must reduce the problem and that the base case terminates execution may provide evidence of conceptual understanding of decomposition and termination, but it does not establish that the learner can implement these principles correctly without support. If the learner identifies and corrects a recursive-call error after receiving a general hint to inspect termination, the observation may instead support a more limited claim concerning partial understanding, productive debugging, and effective use of scaffolding rather than unaided mastery. The specificity of the support is also consequential: a hint that identifies the exact correction would warrant a substantially narrower inference about the learner’s independent understanding.
Within the architectural capability represented by the Evidence Interpretation layer, implementations should evaluate the evidentiary relevance of these heterogeneous observations and synthesize them through evidence models grounded in theories of learning and competency. These models may then support probabilistic learner models in estimating the extent to which the observed behaviors provide evidence for claims about underlying knowledge and skills. As additional evidence accumulates across assignments, tutoring interactions, and other learning activities, competency estimates may be updated, increasing or decreasing confidence in the learner’s developing understanding while explicitly representing the remaining uncertainty. This inferential process is consistent with the principles of ECD, in which observable behaviors serve as evidence supporting probabilistic claims about latent competencies rather than functioning as direct measures of learner ability.
The probabilistic nature of competency inferences makes the representation of uncertainty an essential functional capability represented by the Evidence Interpretation layer. Confidence estimates should reflect not only the amount of available evidence but also its quality, diversity, independence, and contextual relevance. Where evidence is sparse, unevenly distributed, or unavailable, uncertainty should remain explicit rather than being interpreted as evidence of diminished learner competence. This transparency supports more informed educational decision-making while preserving opportunities for human review, additional evidence collection, and continued learner development.
The validity of continuous assessment, therefore, depends on the appropriateness of the interpretations and educational decisions supported by these inferential processes rather than on evidence collection alone [
64]. Within AIM-CAI, the Evidence Interpretation layer provides the inferential foundation upon which longitudinal learner representations are constructed, enabling subsequent architectural layers to maintain evolving representations of learner development while supporting trustworthy educational decisions.
4.5. Learner Representation Layer
The Learner Representation layer maintains longitudinal representations of learner development by integrating competency inferences generated through the Evidence Interpretation layer. Rather than storing isolated assessment outcomes, this layer maintains an evolving representation of learner knowledge, skills, and competencies that is continuously updated as new evidence becomes available. The resulting learner representation reflects both current competency estimates and the uncertainty associated with those estimates, supporting ongoing learning, instructional decision-making, advising, and credentialing.
Traditional educational records, including transcripts and grade point averages (GPA), summarize achievement at discrete points in time. The Learner Representation layer instead preserves developmental trajectories that reflect how competencies evolve across learning activities, educational contexts, and time. Learner representations therefore remain dynamic, context-sensitive, probabilistic, uncertainty-aware, and continuously revisable rather than functioning as static archival records or singular classifications of achievement.
Research in learner modeling (e.g., Bayesian), knowledge tracing (e.g., deep knowledge tracing), and dynamic Bayesian networks supports this approach by demonstrating how competency representations can be updated sequentially as new evidence accumulates [
44,
46]. More recent sequence modeling approaches extend these capabilities by representing increasingly complex temporal dependencies among learning activities and learner behaviors [
47]. These methods illustrate computational approaches for maintaining longitudinal learner representations within continuous assessment systems while supporting the continuous refinement of competency estimates as additional evidence becomes available.
The purpose of the Learner Representation layer extends beyond maintaining computational learner models. Longitudinal learner representations provide a shared evidence base that supports personalized learning, instructional feedback, advising, credentialing, and learner reflection while preserving uncertainty, evidence provenance, and opportunities for continued development. As additional evidence accumulates throughout a learner’s educational journey, these representations evolve continuously, supporting educational decisions while remaining open to revision and refinement.
Maintaining longitudinal learner representations also introduces important psychometric challenges. Competency estimates remain probabilistic because learner knowledge and skills are inferred from incomplete and heterogeneous evidence [
8]. Learner representations inherit the evidentiary and calibration limitations described in
Section 4.4, including poor-quality evidence, inadequate calibration, or insufficient theoretical grounding, which may degrade the accuracy and stability [
20,
62,
63]. Reliable learner representations require trustworthy evidence, sound evidentiary interpretation, ongoing psychometric calibration, and appropriate uncertainty estimation across the preceding architectural functions.
4.6. Governance and Trust Layer
The Governance and Trust layer establishes the policies, controls, and institutional mechanisms that ensure learning evidence and learner representations remain trustworthy throughout AIM-CAI. Unlike the preceding layers, which transform learning evidence into progressively richer educational representations, the Governance and Trust layer operates across the entire architecture by establishing the conditions under which evidence is generated, represented, translated, interpreted, maintained, shared, and ultimately used to support educational decisions. Governance and trust, therefore, function as cross-cutting architectural capabilities that span every layer of the framework.
The Governance and Trust layer preserves the provenance, integrity, interoperability, privacy, transparency, learner agency, and accountability required for trustworthy continuous assessment. It documents how evidence was generated, how competency inferences were produced, how uncertainty is represented, and how learner representations evolve over time. These capabilities support human oversight, interpretive auditing, institutional accountability, and responsible educational decision-making while enabling learners to understand, manage, and selectively share evidence about their learning. In addition to these overarching principles, the Governance and Trust layer represents the architectural capability through which governance requirements for the stewardship and use of learning evidence can be specified. Within AIM-CAI, these requirements should define who is authorized to access, interpret, share, and reuse learning evidence; the educational purposes for which evidence may be used; the duration for which evidence and associated competency inferences are retained; and the procedures through which learners may inspect, correct, contest, or request review of both learning evidence and resulting competency interpretations. Collectively, these governance requirements are intended to support institutional safeguards that promote transparency, accountability, and the appropriate use of learning evidence throughout its lifecycle.
Learner consent represents one important component of trustworthy governance but should not be regarded as sufficient on its own. Educational participation often involves institutional requirements and power relationships that may limit the practical freedom to decline certain forms of data collection or analysis. Accordingly, trustworthy governance should extend beyond consent by incorporating clear institutional policies governing the necessity, proportionality, and legitimacy of evidence collection and use. Where learners do not authorize optional uses of their learning evidence, those data should be excluded from the corresponding analytical processes or used only in accordance with applicable institutional policies and legal requirements. Consequently, governance depends not only on obtaining consent but also on ensuring that educational decisions remain fair, transparent, and subject to appropriate human oversight. Because continuous assessment operates across distributed educational ecosystems, governance also extends beyond individual institutions. Federated architectures provide one approach for coordinating learning evidence while preserving institutional autonomy and reducing the need to centralize learner data. Federated learning approaches demonstrate that predictive educational models can be developed collaboratively while maintaining local control of learner information through distributed computation and privacy-preserving model exchange [
28]. This architectural approach supports continuous assessment across educational institutions while preserving institutional flexibility, learner agency, and distributed stewardship of learning evidence.
Within AIM-CAI, governance requirements should distinguish among educational uses according to their potential impact on learners. Lower-risk applications, such as formative feedback, adaptive learning support, or instructional recommendations, may rely primarily on probabilistic competency estimates accompanied by appropriate uncertainty information. Within AIM-CAI, higher-risk applications—including accreditation, certification, selection, progression, or other decisions that substantially affect educational opportunity—are expected to incorporate additional governance safeguards. These safeguards include human review of competency inferences, traceability of the evidentiary basis supporting decisions, documentation of the interpretive process, and mechanisms through which learners may challenge or request review of consequential decisions. Accordingly, AIM-CAI treats risk-proportionate governance as a core architectural principle, whereby governance requirements become progressively more rigorous as the educational consequences of AI-supported inferences increase. Trust within continuous assessment depends on more than technical security. Trustworthy infrastructures require common evidence standards, psychometric validation, interoperability, auditability, learner consent, and governance mechanisms that define how evidence is generated, interpreted, shared, and used. These mechanisms ensure that learner representations remain transparent, explainable, accountable, and open to human review as they inform educational decisions.
4.7. AIM-CAI: A Functional Reference Architecture
AIM-CAI is intended to serve as a reference architecture that organizes the functional capabilities essential for AI-mediated continuous assessment. Its purpose is to offer a shared framework to guide research, engineering, institutional implementation, and governance, without prescribing a single technical implementation or fixed system architecture. As AI, educational technologies, assessment methods, and governance practices continue to evolve, this framework supplies a common vocabulary for coordinating future development while remaining intentionally extensible and open to refinement.
The six interconnected layers of AIM-CAI collectively describe the functional capabilities necessary to support continuous assessment. Learning evidence is generated across diverse educational environments, represented in forms that preserve educational meaning, translated between systems and competency frameworks, interpreted through principled evidentiary reasoning, aggregated into evolving learner representations, and overseen by trustworthy institutional mechanisms. Artificial intelligence expands the operational feasibility of these functions by supporting evidence generation, interpretation, and coordination while preserving the central role of educational assessment, psychometric validity, learner agency, and human judgment.
The following scenario illustrates how the functional capabilities represented by the six architectural layers of AIM-CAI may be implemented within an authentic learning environment. Consider a first-year engineering student enrolled in an introductory programming course supported by the AIM-CAI framework. As the student engages in quizzes, programming assignments, virtual laboratory activities, discussion forums, and interactions with an AI tutor, a wide array of learning evidence—including assessment results, coding logs, hint requests, response times, and reflective writing—is continuously collected and standardized. Learning evidence from multiple learning platforms is semantically translated into a shared competency framework, enabling probabilistic learner models to estimate mastery of programming concepts such as variables, loops, functions, debugging, and recursion. These competency estimates are incorporated into a longitudinal learner representation that reflects learning progress, confidence estimates, and evidence history while remaining adaptable as new evidence becomes available. Throughout this process, governance mechanisms ensure the provenance of evidence, maintain transparency, and uphold learner agency, thereby enabling instructors to audit AI-generated competency inferences. As a result, the continuously updated learner model can provide personalized recommendations—such as targeted practice on recursion—and alert instructors to opportunities for focused intervention, thereby supporting ongoing learner development.
5. Governance, Trust, and Federated Infrastructure
The Governance and Trust layer introduced in the AIM-CAI framework establishes the functional mechanisms that support trustworthy continuous assessment. While
Section 4 outlined the layered architecture of AIM-CAI, this section explores the organizational, technical, governance, and policy conditions necessary for its implementation. Because continuous assessment relies on distributed evidence generation, probabilistic inference, semantic interoperability, and longitudinal learner representation, governance extends beyond technical implementation to address learner agency, privacy, accountability, semantic stability, algorithmic bias, institutional stewardship, and public trust. Governance, therefore, functions as an integral component of the continuous assessment infrastructure rather than as a policy framework applied after deployment.
Traditional educational assessment systems have concentrated governance largely within institutional boundaries. Grades, transcripts, and testing records are generated and controlled by individual institutions operating within relatively stable administrative structures. Continuous assessment infrastructures fundamentally alter this model by distributing evidence generation, interpretation, and competency inference across platforms, organizations, and contexts. Governance, therefore, becomes a systems-level challenge involving technical standards, semantic alignment, privacy protection, accountability structures, and epistemic authority.
Governance challenges within continuous assessment ecosystems extend beyond technical accuracy to include questions of interpretation, uncertainty, authority, and educational action. Stewardship-oriented approaches emphasize the need for transparency, provenance, accountable oversight, learner agency, and institutional responsibility when AI systems participate in generating competency inferences and informing educational decisions [
65].
5.2. Why Standards-First Architectures Are Unlikely to Succeed
One of the central governance challenges facing continuous assessment infrastructure involves interoperability across heterogeneous educational ecosystems. Historically, educational technology systems have sought interoperability primarily through standards-based architectures such as SCORM, xAPI, LTI, and Caliper Analytics. These frameworks sought to standardize how educational content, learner interactions, and assessment records are represented and exchanged across systems.
However, the historical evolution of educational interoperability standards reveals persistent fragmentation and instability. SCORM, originally developed through the U.S. Department of Defense Advanced Distributed Learning (ADL) initiative, was designed primarily for browser-based learning objects and coarse-grained tracking of completion, time-on-task, and assessment scores. While SCORM standardized content packaging and sequencing, it remained fundamentally constrained by rigid course-centric architectures and limited telemetry capabilities.
The Experience API (xAPI) attempted to address these limitations by enabling fine-grained event logging across diverse learning environments through flexible “Actor–Verb–Object” statements stored within a Learning Record Store (LRS). Yet the flexibility of xAPI simultaneously produced severe semantic fragmentation. Different developers and platforms frequently implemented inconsistent verbs, object structures, and activity representations, resulting in highly heterogeneous telemetry ecosystems lacking stable semantic equivalence. Similar fragmentation has emerged across competing interoperability initiatives such as cmi5, Caliper Analytics, and LTI.
This fragmentation reflects a broader sociotechnical reality: innovation in educational technology consistently outpaces the development and adoption of interoperability standards. Formal standards require sustained collaboration, institutional coordination, and consensus-building processes that often unfold over many years. During that time, AI systems, learner modeling approaches, multimodal analytics, educational platforms, and competency frameworks continue to evolve. As a result, educational ecosystems remain heterogeneous even as interoperability becomes increasingly important. These conditions increase the need for architectural approaches that support semantic translation across diverse and evolving systems while preserving the educational meaning of learning evidence.
For this reason, AIM-CAI suggests that AI-mediated semantic interoperability is likely to become more viable than universal standardization. Rather than forcing all educational systems to adopt identical schemas or competency ontologies, AI-mediated translation layers increasingly support probabilistic semantic alignment across heterogeneous ecosystems. As discussed in
Section 4, interoperability increasingly emerges through dynamic interpretation and ontology reconciliation rather than rigid schema compliance.
6. Institutional and Lifelong Learning Implications
The emergence of AI-mediated continuous assessment infrastructure carries implications that extend far beyond assessment itself. Historically, educational institutions maintained near-exclusive authority over the production, interpretation, and certification of learning evidence. At present, episodic assessment occurs primarily within institutional boundaries through formally sanctioned courses, examinations, and credentialing systems. As a result, institutions control not only instructional delivery but also the mechanisms for recognizing, documenting, and legitimizing competence. The development of continuous, distributed, and infrastructure-based assessment systems may fundamentally alter this arrangement.
AIM-CAI supports the generation, representation, translation, interpretation, governance, and longitudinal use of learning evidence across heterogeneous learning environments rather than exclusively within formal institutional settings. Learning evidence may emerge through workplace participation, AI tutoring systems, collaborative online communities, simulations, portfolios, mobile learning environments, authentic learning experiences, and self-directed learning activities. Thus, assessment no longer belongs exclusively to educational institutions. Instead, institutions increasingly become participants within broader ecosystems of competency interpretation and validation.
6.2. Assessment Beyond Institutional Boundaries
The second major implication of AIM-CAI involves the decoupling of assessment from exclusively institutional environments. Historically, educational institutions have maintained authority over what counted as legitimate evidence of learning because they controlled both instruction and assessment. However, as learning increasingly occurs across distributed digital ecosystems, workplaces, online communities, video streaming sites, and AI-supported environments, institutional boundaries are less able to contain meaningful competency development.
This transition is particularly visible within workforce learning and professional development contexts. Workplace-based assessment models, for instance, Programmatic Assessment by Cees van der Vleuten and Lambert Schuwirth, already rely heavily on longitudinal evidence gathered across authentic practice environments rather than isolated examinations. Similarly, AI-based tutoring systems, professional simulations, collaborative platforms, game-based learning, and self-directed online learning increasingly generate rich evidence on problem-solving, communication, collaboration, and strategic reasoning outside formal classrooms.
As continuous assessment infrastructures mature, educational institutions may therefore lose their historical monopoly over competency validation. This does not eliminate the role of institutions; rather, institutions increasingly shift from functioning as exclusive issuers of learning legitimacy toward functioning as validators, interpreters, and governors within broader competency ecosystems. Universities may increasingly contribute expertise in evidence interpretation, psychometric validation, governance, and credential integration while recognizing that learning itself occurs across distributed contexts extending beyond institutional control.
This transition also creates new opportunities for lifelong learning systems. Traditional educational structures often fragment learning into isolated phases associated with degrees, semesters, or formal programs. Continuous assessment infrastructures such as AIM-CAI instead support more persistent and longitudinal learner representations capable of extending across career transitions, reskilling efforts, and evolving professional trajectories. In this sense, competency development increasingly becomes lifelong rather than institutionally episodic.
7. Research Agenda and Open Challenges
The AIM-CAI reference architecture proposed in this paper organizes the functional capabilities needed to support AI-mediated continuous assessment. It provides a shared framework that can evolve alongside advances in educational assessment, artificial intelligence, and digital learning technologies. Although these advances have made continuous assessment increasingly feasible, important psychometric, technical, governance, and theoretical challenges remain. Addressing these challenges requires interdisciplinary research spanning psychometrics, learning sciences, artificial intelligence, human–computer interaction, data governance, and critical data studies.
The AIM-CAI framework comprises multiple interacting layers of evidence generation, semantic interpretation, probabilistic inference, learner modeling, and governance. Each functional layer reveals a corresponding set of research and engineering challenges related to validity, reliability, interpretability, fairness, interoperability, and institutional legitimacy. Thus, the future development of AI-mediated continuous assessment systems depends not only on technical advancement but also on rigorous empirical validation, governance innovation, and sustained critical scrutiny.
7.4. Interoperability, Governance, and Cross-Contextual Inference
A central assumption underlying AI-mediated continuous assessment infrastructure, AIM-CAI, is that evidence generated across heterogeneous environments can be meaningfully interpreted, translated, and reconciled across contexts and time. However, substantial open questions remain regarding the feasibility, governance, and legitimacy of cross-contextual competency inference. Educational ecosystems remain deeply fragmented across institutions, platforms, domains, cultures, and competency frameworks. Even within highly structured domains such as healthcare informatics, semantic interoperability systems remain vulnerable to ontology drift, alignment instability, and interpretive inconsistency. Educational competencies are often considerably more ambiguous, context-dependent, and socially constructed than medical diagnoses or technical standards.
Nevertheless, governance challenges associated with continuous assessment infrastructures remain among the least resolved areas of research [
24]. Existing educational privacy frameworks, such as the Family Educational Rights and Privacy Act (FERPA) (
https://studentprivacy.ed.gov/faq/what-ferpa, accessed on 9 June 2026) and the Children’s Online Privacy Protection Act (COPPA) (
https://www.ftc.gov/business-guidance/privacy-security/childrens-privacy, accessed on 9 June 2026), were designed primarily for institutionally centralized records rather than distributed, probabilistic, and continuously updated evidence ecosystems. As continuous assessment infrastructures increasingly depend on workplace participation, collaborative interaction, AI tutoring systems, multimodal evidence, and self-directed learning environments, substantial uncertainty emerges regarding ownership of learner evidence, rights to algorithmic inference, portability of competency representations, and governance of longitudinal learner models.
These tensions become especially significant within federated and privacy-preserving architectures. Although secure enclaves, federated learning, differential privacy, and trusted research environments may reduce many surveillance risks, they also introduce new governance complexities related to auditability, accountability, semantic consistency, and coordination failure. Distributed systems may obscure the origins of inferential errors, who governs semantic translation processes, and which institutional actors remain responsible when harmful competency classifications emerge.
Many of the central research challenges associated with AI-mediated continuous assessment infrastructure are best understood not as isolated technical problems but as ongoing tensions among competing infrastructural priorities. As summarized in
Table 1, future research must address how continuous assessment ecosystems balance interoperability and local autonomy, privacy and utility analytics, learner agency and predictive automation, transparency and model complexity, and distributed governance and institutional accountability. These challenges also raise broader epistemic and sociotechnical questions concerning power, representation, and institutional authority, as continuous assessment systems increasingly shape how competence is defined, represented, and legitimized across educational ecosystems. Without careful governance design, continuous inferential systems risk privileging dominant cultural norms, rhetorical traditions, and behavioral expectations while marginalizing non-traditional or culturally situated forms of learning and expression.
Future research must therefore investigate governance architectures capable of preserving epistemic pluralism, supporting learner agency, and maintaining institutional accountability while still enabling meaningful interoperability across distributed learning ecosystems. This includes research on federated governance structures, trusted research environments, permissioned access systems, participatory governance models, semantic accountability mechanisms, and learner-centered consent architectures capable of operating within continuously evolving educational infrastructures.