Next Article in Journal
Mapping IT Reference Frameworks for Governance, Service Management, and Quality Assurance: A Scoping Review
Previous Article in Journal
PC-PLF: Path-Conditioned Per-Layer LoRA Fusion for Open-Vocabulary ROADWork Segmentation
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

AI-Mediated Continuous Assessment Infrastructure (AIM-CAI): Connecting Learning Evidence Across Contexts and Time

by
Danielle S. McNamara
* and
Mohammad Nehal Hasnine
Learning Engineering Institute, Arizona State University, Tempe, AZ 85281, USA
*
Author to whom correspondence should be addressed.
Information 2026, 17(8), 806; https://doi.org/10.3390/info17080806
Submission received: 18 June 2026 / Revised: 12 August 2026 / Accepted: 18 August 2026 / Published: 21 August 2026
(This article belongs to the Section Information Applications)

Abstract

Educational assessment systems have primarily relied on episodic forms of assessment, including examinations, assignments, grades, and credentials. These approaches provide efficient and scalable summaries of achievement and yet capture only part of the developmental process through which learners build competence. Moreover, learning increasingly unfolds across digital platforms, workplaces, collaborative networks, and AI-mediated environments, generating rich evidence of learner development that remains fragmented across systems and contexts. Advances in artificial intelligence, learning analytics, multimodal analytics, learner modeling, and semantic interoperability make it increasingly feasible to connect, integrate, and interpret this evidence across contexts and over time. This paper introduces the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), a sociotechnical framework supporting longitudinal, probabilistic interpretation of distributed evidence of learning. Within AIM-CAI, continuous assessment refers to the ongoing accumulation and dynamic interpretation of evidence generated through learning activities. The framework integrates distributed evidence systems, evidence serialization mechanisms, AI-mediated semantic translation, probabilistic learner models, dynamic competency profiles, and federated governance architectures to support context-sensitive interpretations of learner development while maintaining human judgment, privacy, accountability, and learner agency. The authors examine implications for assessment, credentialing, lifelong learning, institutional roles, interoperability, and governance and outline a research agenda addressing key psychometric, ethical, and governance challenges, including validity, fairness, surveillance, semantic instability, and ownership of learning evidence.

1. Introduction

Human learning naturally unfolds over time through practice, feedback, interaction, reflection, and application across different settings. Understanding and supporting that process depends on meaningful evidence of what learners know, can do, and how they develop over time. Educational assessment provides a systematic means of generating, interpreting, and using that evidence. Historically, educational systems have accomplished this through discrete assessment events—including examinations, assignments, projects, and classroom performances—that can be efficiently administered, evaluated, and translated into grades, transcripts, and credentials. These assessment events provide efficient and potentially valuable summaries of achievement. Because they capture learning at discrete points in time, however, any individual assessment offers only a partial view of the developmental processes through which learners build, refine, and apply knowledge and competencies [1].
The prevalence of episodic assessment reflects both evolving conceptions of assessment and the historical conditions under which large educational systems developed. Industrial-era institutions required approaches that supported standardization, administrative efficiency, and the evaluation of large numbers of learners [2,3,4,5,6,7]. Discrete assessment events provided a practical means of eliciting, evaluating, recording, and communicating evidence of student performance. Their continued use reflects these important educational and institutional functions. At the same time, because each assessment captures performance at a particular point in time, episodic assessment provides only a partial representation of learning trajectories that unfold across activities, contexts, and time.
Research spanning psychometrics, assessment, and the learning sciences has progressively broadened what counts as meaningful evidence of learning. This work emphasizes that competence develops through repeated performance, feedback, revision, collaboration, and participation in contextualized activity [8,9,10,11,12,13]. Programmatic assessment similarly demonstrates the value of combining multiple observations over time rather than relying on a single high-stakes event [14,15,16]. These traditions do not eliminate the need for examinations, assignments, or professional judgment. Instead, they show the value of interpreting individual assessment events as parts of a larger evidentiary record.
Learning now occurs across an expanding range of physical, digital, educational, professional, and social environments. Learners write and revise within digital platforms, interact with AI tutors, participate in simulations and collaborative projects, construct portfolios, and apply knowledge in workplaces. These activities generate potentially meaningful evidence of learning, but that evidence typically remains dispersed across systems, represented in incompatible forms, and disconnected from later educational decisions.
Advances in artificial intelligence, learning analytics, multimodal analytics, learner modeling, natural language processing, and distributed computation render it increasingly feasible to connect, integrate, and interpret meaningful evidence of learning across contexts and over time [17,18,19,20,21]. These capacities, however, do not automatically produce valid, equitable, or educationally useful assessments. Learning evidence remains distributed across heterogeneous platforms, represented in incompatible formats, interpreted using different models, and governed under different institutional policies. Realizing the potential of these advances, therefore, requires an infrastructure that can connect, interpret, and govern evidence across contexts and over time while supporting valid, equitable, and trustworthy educational decisions [18,22,23,24].
This paper introduces the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), a conceptual sociotechnical framework for connecting, interpreting, and governing evidence of learning across contexts and time. AIM-CAI is not proposed as a centralized learner database, a fully developed technical system, or a mechanism for comprehensive learner monitoring. It describes a federated evidence infrastructure through which locally governed systems may contribute selected evidence to context-sensitive and probabilistic interpretations of learner development.
In this paper, continuous assessment refers to the longitudinal accumulation and interpretation of evidence generated through meaningful learning activities. Evidence may be generated at different intervals, through different methods, and for different purposes as learners engage in diverse educational experiences. The defining characteristic is that interpretations of learner development can be updated as relevant evidence accumulates over time, rather than remaining tied to a single assessment occasion. Episodic assessment therefore becomes one important source of evidence within a broader longitudinal evidence system.
The framework integrates six functions required to move responsibly from learning activity to educational interpretation: generating evidence across distributed environments; structuring evidence in forms that preserve its source and context; translating evidence across systems and competency representations; evaluating its relevance through psychometric and inferential models; representing development longitudinally through dynamic competency profiles; and governing how evidence and inferences are accessed, interpreted, contested, and used. These functional capabilities are represented as six interconnected layers: learning evidence generation, learning evidence representation, learning evidence translation, evidence interpretation, learner representation, and governance and trust.
The individual components of AIM-CAI build on established traditions, including evidence-centered design, formative and programmatic assessment, stealth assessment, learning analytics, learner modeling, semantic interoperability, and privacy-preserving computation [8,9,14,25,26,27,28]. The contribution of AIM-CAI lies in specifying how these capabilities may operate together as an integrated infrastructure for longitudinal evidence interpretation. Its novelty is therefore architectural and integrative rather than based on the invention of each constituent component.
AIM-CAI is presented as a reference architecture and research agenda for advancing AI-mediated continuous assessment. It organizes the functional capabilities required to connect, interpret, represent, and govern learning evidence across educational contexts while identifying the empirical, psychometric, technical, and governance challenges that remain. Continued research and bounded implementations will refine these capabilities, evaluate their educational utility, and strengthen the evidence needed to support trustworthy educational decision-making. Throughout this work, the central purpose remains unchanged: using trustworthy evidence to provide more timely, meaningful, and informed support for learning.
This manuscript is organized into eight sections. Section 2 examines the historical and practical limitations of assessment systems organized around bounded, cumulative events. Section 3 explores how advances in artificial intelligence, learning analytics, multimodal analytics, blockchain technologies, and learner modeling are reshaping the feasibility conditions of continuous assessment. Section 4 introduces the proposed AIM-CAI framework, while Section 5 examines the governance, interoperability, and trust architectures required to support continuous inferential ecosystems responsibly. Section 6 considers the broader implications for educational institutions, credentialing systems, and lifelong learning. Section 7 outlines a potential research agenda addressing unresolved psychometric, technical, governance, and sociotechnical challenges along with possible empirical validation and pilot portfolios. Finally, Section 8 concludes the paper.

2. The Limits of Episodic Assessment

Although episodic assessment provides valuable evidence of learner achievement, each assessment event captures only one observation within a broader developmental process. Over the past several decades, however, research across assessment and the learning sciences has progressively expanded conceptions of learning, emphasizing that knowledge and competencies develop through practice, feedback, revision, collaboration, and participation across contexts and over time [1,8,9,10,11,12,13]. This broader conception of learning also expands the kinds of evidence needed to understand learner development, highlighting several inferential and institutional boundaries that shape the evidence episodic assessment can reasonably provide. Inferential boundaries concern the conclusions that can reasonably be drawn from available evidence, whereas institutional boundaries concern where evidence of learning is generated, recognized, and incorporated into educational decision-making.
One important inferential boundary emerges from the developmental nature of learning. Assessments based on short-term performance in decontextualized tasks provide valuable evidence of some forms of knowledge and skill, but they are less well suited to representing adaptive expertise, collaboration, creativity, or the transfer of learning across contexts [8,11,29]. Formative assessment research similarly demonstrates that learning develops through ongoing interaction, feedback, revision, and participation, with each learning experience contributing additional evidence of learner development [9,30]. Episodic assessment provides valuable but necessarily incomplete evidence of competencies that emerge and evolve over time.
The developmental nature of learning also has important implications for how assessment evidence supports learning. Research on formative assessment demonstrates that learning improves when assessment is integrated into instruction and evidence is used to provide timely feedback that informs subsequent learning [9]. Black and Wiliam [9] argued that the distinction between formative and summative assessment lies not simply in the timing of assessment events but in whether the resulting evidence informs learners’ next steps. Similarly, Hattie and Timperley [30] emphasized that effective feedback reduces the gap between current understanding and desired performance. When evidence becomes available only after instructional sequences have concluded, opportunities for revision, adaptation, and continued learning are more limited.
A second inferential boundary concerns the alignment between assessment tasks and the competencies they are intended to represent. Traditional assessments often rely on structured, decontextualized tasks that provide valuable evidence of specific knowledge and skills but may be less well suited to representing complex judgment, problem-solving, collaboration, and adaptability in authentic settings [11]. Authentic assessment research emphasizes that meaningful competence is often demonstrated through performance in contexts that resemble the situations in which knowledge is ultimately applied [31]. This perspective aligns with situated learning theory, which views learning as developing through participation in authentic social and professional contexts rather than through isolated demonstrations of knowledge [13]. Evidence generated through traditional examinations provides an important perspective on learner competence, but it represents only part of the broader capabilities demonstrated across authentic educational, professional, and social settings.
A third inferential boundary concerns the sufficiency of evidence for supporting valid conclusions about learner competence. Assessment is fundamentally a process of evidentiary reasoning in which observations of learner performance are interpreted to support inferences about underlying knowledge and competencies [8]. Pellegrino et al. [8] emphasized that valid assessment depends on coherent alignment among models of learning, the tasks used to elicit evidence, and the interpretive frameworks used to draw conclusions from that evidence. Because episodic assessment relies on a limited number of observations collected at particular points in time, the available evidence may provide an insufficient basis for drawing valid inferences about developmental growth or complex competencies.
Several assessment approaches have sought to address these inferential boundaries by drawing on evidence from multiple observations rather than relying on a single assessment event. Among the most influential is programmatic assessment, which proposes that meaningful judgments about learner competence should emerge from the aggregation of multiple low-stakes observations collected across contexts, evaluators, and time [14,15,16]. Similar principles are reflected in competency-based education (CBE) and frameworks for assessing twenty-first century skills, both of which recognize that complex competencies require multiple forms of evidence to support valid interpretation [29,32]. These approaches reflect growing recognition that learner competence is best understood through the accumulation of evidence across learning experiences rather than through isolated assessment snapshots.
Traditional assessment systems remain closely tied to formal educational programs, courses, and institutional credentialing processes. Contemporary learners, however, increasingly develop knowledge and competencies through workplace participation, collaborative networks, online communities, self-directed learning, and other experiences that extend beyond formal educational settings. Much of this evidence, however, remains disconnected from formal assessment and credentialing systems, even as portfolios, micro-credentials, and alternative credentials have expanded the range of recognized learning experiences [33,34].
These inferential and institutional boundaries emerged under historical conditions that rendered continuous observation and interpretation of learner development impractical. Sustained assessment requires multidimensional evidence, ongoing interpretation, and continuous feedback—activities that traditionally exceeded the operational capacity of educational institutions operating at scale. Episodic assessment, therefore, became an effective and scalable approach for generating evidence to support educational decision-making.
Although episodic assessment continues to provide valuable evidence for educational decision-making, contemporary conceptions of learning increasingly emphasize evidence that is developmental, contextual, longitudinal, and distributed across learning experiences. Until recently, the practical challenges of capturing, integrating, and interpreting such diverse evidence limited the feasibility of extending assessment beyond episodic events. Advances described in the following section increasingly change these feasibility conditions, motivating the development of infrastructures capable of connecting, interpreting, and governing learning evidence across contexts and over time.

3. Technological Advances Changing the Feasibility Conditions for Continuous Assessment

Recent advances in artificial intelligence, learning analytics, educational data mining, natural language processing, multimodal learning analytics, learner modeling, and related technologies are changing the practical feasibility of collecting, integrating, and interpreting evidence of learning across contexts and over time. Many of the limitations discussed in the previous section reflected historical and operational constraints rather than educational ideals. Continuously observing, interpreting, and documenting complex learning processes across large numbers of learners was simply impractical. These practical constraints shaped assessment systems around episodic events that could be administered, evaluated, and documented efficiently. Today, many of these operational constraints are diminishing as advances in AI and related technologies enable the scalable interpretation of rich, heterogeneous evidence generated through authentic learning activities. As a result, assessment can increasingly be viewed as an ongoing process of evidentiary reasoning supported by diverse forms of learning evidence rather than as a sequence of isolated testing events.
Expanding the capacity to collect and interpret evidence does not, by itself, support valid educational inferences. Richer evidence must still be interpreted through valid inferential models, aligned with theories of learning, and governed in ways that preserve transparency, learner agency, and appropriate human oversight. The technologies discussed in this section contribute complementary capabilities for generating, interpreting, and governing evidence within a continuous assessment infrastructure.

3.1. Expanding Sources of Learning Evidence

Contemporary digital learning environments expand the sources of evidence available for understanding learner development. As learners write, revise, collaborate, solve problems, navigate simulations, and interact with digital tools, they generate process-oriented evidence that documents how learning unfolds over time. Research in learning analytics and educational data mining has shown that these digital traces can support richer interpretations of learning processes and learner development [18,19]. Process-oriented evidence captures how learners engage, revise, collaborate, and solve problems over time, providing insights that complement the products of learning traditionally represented in assignments, examinations, and other assessment artifacts.
Language provides one of the richest sources of evidence about learner thinking and development. Advances in natural language processing and writing analytics have expanded the capacity to interpret linguistic, semantic, and discourse-level patterns in learner writing, supporting probabilistic inferences about conceptual understanding, reasoning, writing quality, and metacognitive strategy use [21,35,36,37]. Intelligent writing support systems, such as Writing Pal, further demonstrate how natural language processing can simultaneously provide adaptive feedback and generate evidence about learners’ writing processes and strategies [21]. Collectively, these developments illustrate how language itself can serve as a continuous source of evidence reflecting evolving cognitive and metacognitive activity throughout learning.
Dialogue and collaboration provide additional sources of evidence about learner understanding and knowledge construction. Dialogue-based intelligent tutoring systems, such as AutoTutor, demonstrate how conversational interactions can provide evidence of conceptual understanding, explanatory reasoning, and misconceptions as learning unfolds [37]. Collaborative learning analytics extends this perspective to peer interaction and group learning by examining patterns of participation, collaborative reasoning, and knowledge construction across extended interactions using machine learning and sequence analysis techniques [18,20]. These approaches illustrate how dialogue and collaboration generate process-oriented evidence that supports continuous interpretation of learning.
Multimodal interactions provide additional sources of evidence about how learning unfolds across cognitive, social, affective, and behavioral dimensions. Multimodal learning analytics (MMLA) integrates evidence from speech, gaze, movement, physiological signals, collaborative interactions, video analysis, and digital activity logs to examine learner engagement, cognitive load, emotional response, collaboration, and self-regulation throughout extended learning activities [20,38]. For example, eye-tracking systems provide evidence about attention allocation during complex problem-solving, while process-trace analytics in simulations and virtual laboratories capture how learners navigate tasks, test hypotheses, and regulate their learning over time [39,40]. Together, these approaches demonstrate how multiple streams of observable evidence can be integrated to support richer interpretations of learner development.
These developments expand both the breadth and continuity of evidence available to support educational interpretation. Contemporary learning environments can capture sequences of learner actions, revisions, decisions, interactions, and strategic behaviors, providing a richer representation of learning trajectories and competency development over time. AI systems contribute by organizing, integrating, and modeling these heterogeneous evidence streams, supporting longitudinal evidence accumulation and probabilistic inferences about learner development at scale.

3.2. Advances in Interpreting Learning Evidence

Expanding the sources of learning evidence also expands the need for principled methods to interpret that evidence. Educational decisions depend on valid inferences about learner knowledge, skills, and development supported by rich learning evidence. Research in learner modeling, intelligent tutoring systems, evidence-centered design, and embedded assessment has established many of the theoretical and computational foundations for interpreting diverse forms of learning evidence. Although recent advances in generative AI have accelerated interest in these approaches, the underlying principles have developed over several decades.
Evidence-Centered Design (ECD) provides one of the most influential frameworks for interpreting learning evidence [25,41]. Consistent with the Knowing What Students Know assessment triangle proposed by Pellegrino and colleagues [8], ECD conceptualizes assessment as a process of evidentiary reasoning that connects observations of learner behavior to probabilistic claims about underlying competencies. Both frameworks emphasize that educational assessment depends on aligning theories of learning (cognition), observable evidence, and principled interpretation. The framework further emphasizes alignment among the competencies being assessed, the evidence required to support interpretive claims, and the tasks designed to elicit that evidence. Because competencies represent latent constructs, they are inferred from observable behaviors through probabilistic reasoning. ECD therefore provides a principled foundation for interpreting the diverse evidence streams generated in contemporary digital learning environments.
Because learning cannot be observed directly, educational interpretation always involves uncertainty. Observable behaviors provide evidence that increases or decreases confidence in claims about learner competencies, but they do not establish those competencies with certainty. Thus, continuous assessment depends on accumulating evidence across multiple observations, contexts, and learning activities to strengthen the validity of interpretive claims.
Embedded assessment integrates evidence collection directly into learning activities and digital environments. Stealth assessment operationalizes this approach by gathering evidence continuously as learners interact with video games, simulations, intelligent tutoring systems, and collaborative learning tasks [27,42,43]. These environments capture process-oriented evidence that is interpreted through evidence models and learner models to update probabilistic estimates of learner competencies in real time. By integrating assessment within authentic learning activities, embedded assessment supports continuous interpretation of learner development as learning unfolds.
Continuous learner modeling provides a mechanism for updating interpretations of learner knowledge as evidence accumulates over time. Intelligent tutoring systems demonstrated the feasibility of this approach by using cognitive models and Bayesian Knowledge Tracing (BKT) to estimate learners’ mastery of specific knowledge components during problem-solving [26,44]. Learner models operationalize evidentiary reasoning by continuously updating probabilistic estimates of learner competencies as new evidence becomes available from solution paths, successful and unsuccessful attempts, hint requests, and other learning behaviors. This continuous updating supports adaptive feedback and personalized instructional support. Although many early systems focused on well-defined domains such as algebra, they established the practical foundations for continuously updating probabilistic models of learner understanding throughout the learning process.
Recent advances in artificial intelligence extend these interpretive capabilities to increasingly open-ended learning environments. Machine learning techniques support the analysis of evidence generated through complex simulations, collaborative dialogue, writing processes, coding activities, and multimodal interactions. These systems identify patterns associated with conceptual understanding, self-regulated learning, collaboration, and strategic problem-solving across extended learning experiences. The resulting inferences remain probabilistic and become progressively more informative through the accumulation of evidence over time. Contemporary computational psychometric approaches, including Bayesian networks, hidden Markov models (HMMs), and related learner modeling techniques, support this process by continuously updating competency profiles as new evidence becomes available [44,45,46,47]. These developments provide the inferential foundation for continuous assessment by enabling diverse forms of learning evidence to be interpreted, integrated, and accumulated across learning activities, instructional contexts, and educational experiences, supporting longitudinal models of learner development.

3.3. Trusted Evidence Infrastructure

Continuous assessment depends on trusted infrastructures that support the generation, interpretation, and governance of learning evidence across educational systems. As learning increasingly spans institutions, workplaces, digital platforms, and informal environments, assessment infrastructures must support evidence provenance, integrity, interoperability, portability, and governance. These capabilities enable learning evidence to persist, move across educational contexts, and remain interpretable while preserving its provenance, context, integrity, and meaning over time.
Blockchain technologies have emerged as one approach for strengthening trust in distributed learning records. By combining cryptographic verification with distributed ledgers, blockchain infrastructures can provide tamper resistance, provenance tracking, and verifiable educational records across multiple organizations [48,49]. These capabilities make blockchain particularly relevant for continuous assessment systems that integrate evidence from diverse learning environments.
Trusted evidence infrastructures enable learning evidence to extend beyond the boundaries of individual courses, institutions, and credentialing systems. Evidence generated across formal education, workplace learning, professional development, and informal learning experiences can be accumulated into longitudinal learner records while preserving information about the origin, context, timing, conditions, and integrity of that evidence. Blockchain-based systems, including the Blockchain of Learning Logs (BOLL) and BEMPAS, illustrate how distributed infrastructures can support portable learner records, automated credentialing, access control, and cross-institutional verification through cryptographic methods and smart contracts [49,50,51]. These capabilities support the recognition of learning as a cumulative and continuously documented process that spans diverse educational contexts.
The effectiveness of trusted evidence infrastructures depends on more than just secure recordkeeping. Continuous assessment requires common data standards, interoperable competency frameworks, governance structures, privacy protections, and audit mechanisms that define how evidence is generated, validated, interpreted, shared, and used across educational systems. These governance mechanisms establish the conditions under which learning evidence can be trusted, appropriately interpreted, and responsibly incorporated into educational decisions. Trusted evidence infrastructures also support learner agency by enabling learners to inspect, manage, and selectively share evidence about their learning while providing transparency into how that evidence is interpreted and used.
Although trusted evidence infrastructures have advanced considerably, several implementation challenges remain. Blockchain-based educational systems often operate as fragmented ecosystems with limited interoperability and cross-platform compatibility [49,50]. Scalability also remains an important consideration, particularly in environments that generate large volumes of continuous learning evidence requiring secure verification and storage [50,52]. Continued progress will depend on technical advances and governance frameworks that support secure, interoperable, and sustainable evidence ecosystems [23,24,48,50].

4. AI-Mediated Continuous Assessment Infrastructure (AIM-CAI)

The preceding sections established that continuous assessment depends on three complementary capabilities: generating rich and continuous evidence of learning, interpreting that evidence through principled inferential models, and governing trusted learning evidence across educational systems. These capabilities establish the technical and conceptual foundation for continuous assessment as an evidence infrastructure rather than as a sequence of isolated assessment events.
The AI-Mediated Continuous Assessment Infrastructure (AIM-CAI) framework synthesizes these capabilities into an integrated functional architecture for continuous assessment. The architecture emerged through the synthesis of assessment, learner modeling, learning analytics, and educational infrastructure literature reviewed in the preceding sections. AIM-CAI is organized around the functional capabilities required to transform learning activities into meaningful educational decisions. Artificial intelligence expands the operational feasibility of this infrastructure by integrating heterogeneous evidence, continuously updating probabilistic learner models, and supporting the scalable interpretation of learning processes. Within AIM-CAI, AI functions as an enabling capability that supports evidence-centered educational decision-making while preserving the principles of validity, transparency, and human oversight.
Figure 1 presents the six functional layers of AIM-CAI. This framework serves as a reference architecture to organize the functional capabilities required to support continuous assessment across distributed educational environments. Each layer corresponds to a distinct function within the evidentiary chain linking learning activities to educational decisions. Collectively, the layers describe how learning evidence is generated, prepared for interpretation, connected across systems, interpreted through inferential models, accumulated into longitudinal learner representations, and governed throughout its lifecycle. The layers, therefore, represent functional requirements of a continuous assessment infrastructure. Individual technologies, instructional models, institutional workflows, and implementation strategies contribute to these functions while remaining adaptable to different educational contexts.
The six layers comprise (1) evidence generation, which captures learning activity across educational contexts; (2) evidence representation, which organizes and normalizes heterogeneous evidence into comparable representations; (3) evidence translation, which supports interoperability across systems, competency frameworks, and educational contexts; (4) evidence interpretation, which applies evidentiary reasoning and probabilistic learner modeling to generate competency inferences; (5) learner representation, which accumulates evidence into longitudinal models of learner development; and (6) governance and trust, which provides the provenance, interoperability, transparency, learner agency, privacy, and institutional oversight required to support trustworthy educational decision-making.
AIM-CAI provides a functional reference architecture for organizing continuous assessment rather than prescribing a single technical implementation. The six layers represent the functional capabilities required to transform distributed learning activities into trustworthy educational decisions. Each layer performs a distinct function in the evidentiary chain and provides the foundation for the subsequent layer. The framework is intended to support the coordinated generation, exchange, interpretation, representation, and governance of learning evidence across multiple educational systems while preserving institutional autonomy and distributed governance. Learning evidence can accumulate across courses, institutions, workplaces, and lifelong learning experiences while retaining the provenance, context, and interpretability required for meaningful educational decisions. As assessment practices, learner models, interoperability standards, and AI technologies continue to evolve, these functional capabilities provide a stable architectural foundation at scale for integrating new methods and implementations within a coherent evidence infrastructure.
The architecture exhibits six defining characteristics that emerge consistently from the assessment, learner modeling, and educational infrastructure literature synthesized throughout this paper. Assessment functions continuously through the accumulation of evidence across learning experiences rather than through isolated assessment events. Competency estimation remains inferential because learner knowledge and skills are represented as latent constructs supported by observable evidence. Educational decisions remain probabilistic because evidence strengthens confidence in competency claims without eliminating uncertainty (e.g., measurement error, incomplete evidence, incorrect inference). Learning evidence is distributed across institutions, workplaces, digital platforms, and informal learning environments. Learner representations evolve longitudinally as evidence accumulates over time. Finally, assessment functions as an interoperable infrastructure that coordinates evidence generation, interpretation, representation, and governance across educational ecosystems. The following subsections elaborate on each functional layer of the AIM-CAI framework.

4.1. Learning Evidence Generation Layer

Within AIM-CAI, evidence is generated continuously across heterogeneous learning environments. The Evidence Generation layer provides the architectural capability for capturing observable learning activities across distributed educational environments and representing them as sources of assessment evidence. Its purpose is to support the continuous accumulation of evidence generated through authentic learning experiences while preserving the contextual information needed for subsequent interpretation. Evidence may originate from learning management systems, intelligent tutoring systems, collaborative learning environments, workplace experiences, portfolios, simulations, laboratories, conversations, or instructor observations [39,40,53,54]. The architectural role of this layer is to capture these heterogeneous observations while preserving their provenance, context, timing, and conditions for later interpretation.
The Evidence Generation layer accommodates and organizes evidence generated across physical, digital, and hybrid learning environments. Representative forms of learning evidence include writing processes, conversational interactions, simulations, educational games, collaborative activities, multimodal learning experiences, and workplace performance [20,27,36,37,53]. Educational AI and multimodal learning analytics expand the ability to capture and organize heterogeneous evidence streams, including behavioral traces, interaction logs, writing revisions, dialogue, gesture, gaze, speech, and other observable learning activities that provide evidence of learner development.
The overarching layer supports the continuous accumulation of learning evidence as learner development unfolds across multiple contexts and extended periods of time [53]. Distributed learning evidence expands the range of authentic learning activities that contribute to competency inference across educational institutions, workplaces, digital platforms, and informal learning environments [8,10,55]. These diverse evidence streams vary substantially in reliability, granularity, interpretability, and theoretical relevance. A learning management system clickstream, a collaborative dialogue transcript, a workplace observation, an authentic performance, and a multimodal physiological signal each provide fundamentally different forms of evidence and therefore require different interpretive models. Multimodal learning analytics research similarly cautions that behavioral and sensor data may introduce construct-irrelevant variance when interpreted without sufficient theoretical grounding [20]. Capturing learning evidence, therefore, extends beyond recording observations to preserving the provenance, context, timing, and conditions under which evidence is generated. These contextual attributes provide the foundation for subsequent architectural layers to represent, translate, interpret, and govern learning evidence while supporting principled competency inference.

4.2. Learning Evidence Representation Layer

The Learning Evidence Representation layer preserves the educational meaning and context of heterogeneous learning evidence while establishing representations that support subsequent interpretation and exchange across educational systems. Learning evidence generated through educational activities varies substantially in format, granularity, modality, and context. Supporting interpretation and interoperability, therefore, requires representations that preserve contextual information, provenance, timing, and relationships while organizing evidence into forms suitable for subsequent processing. This capability ensures that learning evidence remains interpretable, comparable, and interoperable without diminishing the richness of the original learning activity.
Heterogeneous learning evidence cannot support continuous assessment unless it can be organized into representations that preserve both meaning and context. Writing revisions, conversational dialogue, simulation traces, multimodal observations, portfolio artifacts, workplace evaluations, and instructor observations each produce evidence in fundamentally different forms. Learning analytics research has, therefore, emphasized the importance of event representation, metadata, temporal ordering, and contextualization as prerequisites for meaningful analysis and longitudinal interpretation [18,20,53].
Representing learning evidence involves organizing observations into machine-readable representations that preserve relationships among learners, activities, time, instructional context, and evidence source. Temporal ordering is particularly important because many forms of learning evidence, including writing revisions, collaborative dialogue, and interaction sequences, derive much of their educational meaning from the relationships among events rather than from isolated observations [18,21,56,57]. These representations may include standardized event models, metadata, timestamps, provenance records, semantic annotations, and contextual descriptors that enable evidence generated across diverse environments to be interpreted consistently within subsequent architectural layers. Representing learning evidence requires balancing standardization with contextual fidelity. Excessive normalization may obscure meaningful differences among learning activities, whereas insufficient structure limits interoperability and computational interpretation.

4.3. Learning Evidence Translation Layer

The Learning Evidence Translation layer serves as the conceptual core of AIM-CAI. Existing educational ecosystems are characterized by profound heterogeneity across platforms, schemas, competency frameworks, and institutional boundaries. Traditional interoperability approaches rely on predefined standards such as SCORM, IMS Learning Tools Interoperability (LTI), and the Experience API (xAPI). While these support technical exchange, they depend on prior agreement on common standards and have faced persistent adoption and fragmentation challenges (discussed further in Section 5.2).
The Learning Evidence Translation layer enables evidence representations generated within one educational context to be interpreted and exchanged across other systems, competency frameworks, and institutional environments. Educational ecosystems are inherently heterogeneous, encompassing diverse learning management systems, assessment platforms, competency frameworks, institutional data models, and instructional practices. Learning evidence, therefore, requires translation across these heterogeneous representations while preserving its educational meaning and contextual integrity. The purpose of this layer is to support semantic interoperability so that evidence generated in one context remains meaningful and usable in another.
Existing interoperability standards, including SCORM, LTI, and xAPI, support the technical exchange of learning information across educational systems. Continuous assessment, however, also requires preserving the educational meaning of evidence as it moves across diverse competency frameworks, instructional models, and institutional contexts. The Learning Evidence Translation layer, therefore, extends beyond technical interoperability by supporting semantic translation among heterogeneous evidence representations while preserving local educational context.
Artificial intelligence expands the operational feasibility of semantic translation through machine learning, natural language processing, and large language models that assist with ontology alignment, semantic mapping, and probabilistic reconciliation across competency frameworks. Rather than requiring institutions to adopt identical competency models, standardized data structures, or centralized learner repositories, AI-assisted translation enables locally meaningful representations of learning evidence to be connected across distributed educational ecosystems while preserving institutional flexibility and contextual specificity. Because semantic translation influences subsequent competency inferences, AI-generated mappings should be treated as provisional rather than authoritative. Within the Learning Evidence Translation layer, implementations should provide mechanisms for documenting the provenance, version, and degree of semantic equivalence associated with competency mappings while allowing those mappings to be reviewed and refined through appropriate institutional governance. Rather than treating competency correspondences as binary matches, translation relationships may vary in scope, specificity, performance expectations, and assessment requirements. Consequently, translation outcomes should preserve information about the confidence and completeness of each mapping so that downstream interpretive processes can account for partial or uncertain correspondences. Where semantic equivalence cannot be established with sufficient confidence, evidence should remain linked to its original competency framework or require additional human review before informing consequential educational decisions such as accreditation, certification, or credentialing.
Semantic translation has become an important architectural capability in other information-intensive domains. Healthcare informatics, for example, coordinates heterogeneous patient records, clinical ontologies, and diagnostic vocabularies across decentralized systems without requiring identical local data structures. Technologies such as the RDF, ontology alignment, and knowledge graph architectures support semantic interoperability while preserving institutional flexibility and local context. Similar approaches have begun to emerge within education. Musa et al. [58], for example, proposed a graph-based ontology alignment framework that integrates learning analytics data across learning management systems, MOOCs, and other educational platforms through semantic mapping and reconciliation techniques.
The Learning Evidence Translation layer also introduces important interpretive challenges. Semantic translation remains probabilistic because educational concepts often vary across institutional, disciplinary, and cultural contexts. Ontology alignment systems and large language models remain susceptible to semantic drift, ambiguous mappings, and other interpretive errors that may influence downstream competency inferences [59]. In addition, translation models trained on historically biased educational data may reproduce inequities through probabilistic classification and mapping processes [60]. Therefore, the Learning Evidence Translation layer supports inferential compatibility under conditions of uncertainty while preserving the contextual information required for the subsequent Evidence Interpretation layer.

4.4. Evidence Interpretation Layer

The Evidence Interpretation layer transforms translated learning evidence into probabilistic inferences about learner knowledge, skills, and competencies. Because these constructs are latent, they cannot be observed directly. Instead, the layer applies evidentiary reasoning and probabilistic learner modeling to relate observable learning evidence to defensible educational interpretations [8,41]. Its purpose is to support valid educational interpretation by determining which evidence contributes to competency claims, how evidence is weighted, how uncertainty is represented, and how competency estimates are updated as additional evidence becomes available across learning activities, educational contexts, and time.
The Evidence Interpretation layer is grounded in established theories of educational measurement. The ECD conceptualizes assessment as an evidentiary argument that connects observations of learner behavior to claims about underlying competencies through explicitly defined interpretive models [41]. Similarly, Kane’s argument-based approach to validity emphasizes that educational claims depend on the interpretive warrants that justify competency inferences rather than on evidence collection alone [61]. Together, these perspectives establish the theoretical foundation for treating continuous assessment as an ongoing process of evidentiary reasoning supported by accumulating evidence rather than as a sequence of isolated measurements.
Within the Evidence Interpretation layer, implementations must evaluate the dependence among evidence sources and the availability of evidence before updating competency estimates. For example, records generated within the same task, learning episode, or platform may provide redundant or overlapping information. Related observations can be identified through provenance, timing, and contextual metadata, and then grouped such that their combined contribution does not artificially inflate model confidence. Implementations should also account for learners’ opportunities to generate evidence. Limited evidentiary coverage from sparse or uneven records increases uncertainty estimates. Estimates should therefore reflect the quality, diversity, relevance, and independence of observations, along with the range of opportunities from which they were generated. These considerations reduce advantages associated with greater platform use and support fairer interpretations across learners.
Behavioral evidence frequently contains construct-irrelevant variance arising from interface characteristics, temporary contextual influences, social dynamics, fatigue, cultural differences, or other factors unrelated to the competencies being assessed. Multimodal learning analytics and computational psychometrics, therefore, emphasize that behavioral telemetry requires theoretically grounded interpretation to support valid educational inferences [20,60,61]. The quality of competency inferences depends fundamentally on the quality of both the available evidence and the interpretive models applied to that evidence. Poorly calibrated probabilistic models, noisy observations, or insufficient theoretical grounding may reduce inferential accuracy [20,62,63]. The Evidence Interpretation layer, therefore, provides the functional capability for distinguishing meaningful educational signals from contextual artifacts, ambiguity, redundant observations, and noise while supporting psychometric calibration, uncertainty estimation, and validation.
The Evidence Interpretation layer also supports continuous updating of competency inferences as new evidence becomes available. Bayesian networks, dynamic Bayesian networks (DBN), knowledge tracing approaches, and related probabilistic learner models estimate evolving learner states through sequential evidence accumulation [44,45,46]. The resulting competency inferences do not constitute direct measurements of learner capability. Instead, they represent probabilistic interpretations conditioned on incomplete and heterogeneous evidence that remain open to revision as additional evidence becomes available. This continuous updating allows persistent developmental patterns to emerge while distinguishing durable learning from temporary fluctuations in performance. For example, consider a learner completing a programming assignment within an introductory engineering course. Observable evidence may include the correctness of the submitted solution, the sequence of code revisions, the frequency of hint requests, response times, and reflective explanations accompanying the final submission. Individually, none of these observations directly demonstrates competency in recursion or algorithmic reasoning. For instance, if the learner adds a valid base case and revises the recursive call so that the input progresses toward that base case, these actions may support the claim that the learner recognizes key structural requirements for recursive termination within the task. They would not, however, establish that the learner can independently design a recursive algorithm or transfer this understanding to a novel problem. Similarly, a reflective explanation that each recursive call must reduce the problem and that the base case terminates execution may provide evidence of conceptual understanding of decomposition and termination, but it does not establish that the learner can implement these principles correctly without support. If the learner identifies and corrects a recursive-call error after receiving a general hint to inspect termination, the observation may instead support a more limited claim concerning partial understanding, productive debugging, and effective use of scaffolding rather than unaided mastery. The specificity of the support is also consequential: a hint that identifies the exact correction would warrant a substantially narrower inference about the learner’s independent understanding.
Within the architectural capability represented by the Evidence Interpretation layer, implementations should evaluate the evidentiary relevance of these heterogeneous observations and synthesize them through evidence models grounded in theories of learning and competency. These models may then support probabilistic learner models in estimating the extent to which the observed behaviors provide evidence for claims about underlying knowledge and skills. As additional evidence accumulates across assignments, tutoring interactions, and other learning activities, competency estimates may be updated, increasing or decreasing confidence in the learner’s developing understanding while explicitly representing the remaining uncertainty. This inferential process is consistent with the principles of ECD, in which observable behaviors serve as evidence supporting probabilistic claims about latent competencies rather than functioning as direct measures of learner ability.
The probabilistic nature of competency inferences makes the representation of uncertainty an essential functional capability represented by the Evidence Interpretation layer. Confidence estimates should reflect not only the amount of available evidence but also its quality, diversity, independence, and contextual relevance. Where evidence is sparse, unevenly distributed, or unavailable, uncertainty should remain explicit rather than being interpreted as evidence of diminished learner competence. This transparency supports more informed educational decision-making while preserving opportunities for human review, additional evidence collection, and continued learner development.
The validity of continuous assessment, therefore, depends on the appropriateness of the interpretations and educational decisions supported by these inferential processes rather than on evidence collection alone [64]. Within AIM-CAI, the Evidence Interpretation layer provides the inferential foundation upon which longitudinal learner representations are constructed, enabling subsequent architectural layers to maintain evolving representations of learner development while supporting trustworthy educational decisions.

4.5. Learner Representation Layer

The Learner Representation layer maintains longitudinal representations of learner development by integrating competency inferences generated through the Evidence Interpretation layer. Rather than storing isolated assessment outcomes, this layer maintains an evolving representation of learner knowledge, skills, and competencies that is continuously updated as new evidence becomes available. The resulting learner representation reflects both current competency estimates and the uncertainty associated with those estimates, supporting ongoing learning, instructional decision-making, advising, and credentialing.
Traditional educational records, including transcripts and grade point averages (GPA), summarize achievement at discrete points in time. The Learner Representation layer instead preserves developmental trajectories that reflect how competencies evolve across learning activities, educational contexts, and time. Learner representations therefore remain dynamic, context-sensitive, probabilistic, uncertainty-aware, and continuously revisable rather than functioning as static archival records or singular classifications of achievement.
Research in learner modeling (e.g., Bayesian), knowledge tracing (e.g., deep knowledge tracing), and dynamic Bayesian networks supports this approach by demonstrating how competency representations can be updated sequentially as new evidence accumulates [44,46]. More recent sequence modeling approaches extend these capabilities by representing increasingly complex temporal dependencies among learning activities and learner behaviors [47]. These methods illustrate computational approaches for maintaining longitudinal learner representations within continuous assessment systems while supporting the continuous refinement of competency estimates as additional evidence becomes available.
The purpose of the Learner Representation layer extends beyond maintaining computational learner models. Longitudinal learner representations provide a shared evidence base that supports personalized learning, instructional feedback, advising, credentialing, and learner reflection while preserving uncertainty, evidence provenance, and opportunities for continued development. As additional evidence accumulates throughout a learner’s educational journey, these representations evolve continuously, supporting educational decisions while remaining open to revision and refinement.
Maintaining longitudinal learner representations also introduces important psychometric challenges. Competency estimates remain probabilistic because learner knowledge and skills are inferred from incomplete and heterogeneous evidence [8]. Learner representations inherit the evidentiary and calibration limitations described in Section 4.4, including poor-quality evidence, inadequate calibration, or insufficient theoretical grounding, which may degrade the accuracy and stability [20,62,63]. Reliable learner representations require trustworthy evidence, sound evidentiary interpretation, ongoing psychometric calibration, and appropriate uncertainty estimation across the preceding architectural functions.

4.6. Governance and Trust Layer

The Governance and Trust layer establishes the policies, controls, and institutional mechanisms that ensure learning evidence and learner representations remain trustworthy throughout AIM-CAI. Unlike the preceding layers, which transform learning evidence into progressively richer educational representations, the Governance and Trust layer operates across the entire architecture by establishing the conditions under which evidence is generated, represented, translated, interpreted, maintained, shared, and ultimately used to support educational decisions. Governance and trust, therefore, function as cross-cutting architectural capabilities that span every layer of the framework.
The Governance and Trust layer preserves the provenance, integrity, interoperability, privacy, transparency, learner agency, and accountability required for trustworthy continuous assessment. It documents how evidence was generated, how competency inferences were produced, how uncertainty is represented, and how learner representations evolve over time. These capabilities support human oversight, interpretive auditing, institutional accountability, and responsible educational decision-making while enabling learners to understand, manage, and selectively share evidence about their learning. In addition to these overarching principles, the Governance and Trust layer represents the architectural capability through which governance requirements for the stewardship and use of learning evidence can be specified. Within AIM-CAI, these requirements should define who is authorized to access, interpret, share, and reuse learning evidence; the educational purposes for which evidence may be used; the duration for which evidence and associated competency inferences are retained; and the procedures through which learners may inspect, correct, contest, or request review of both learning evidence and resulting competency interpretations. Collectively, these governance requirements are intended to support institutional safeguards that promote transparency, accountability, and the appropriate use of learning evidence throughout its lifecycle.
Learner consent represents one important component of trustworthy governance but should not be regarded as sufficient on its own. Educational participation often involves institutional requirements and power relationships that may limit the practical freedom to decline certain forms of data collection or analysis. Accordingly, trustworthy governance should extend beyond consent by incorporating clear institutional policies governing the necessity, proportionality, and legitimacy of evidence collection and use. Where learners do not authorize optional uses of their learning evidence, those data should be excluded from the corresponding analytical processes or used only in accordance with applicable institutional policies and legal requirements. Consequently, governance depends not only on obtaining consent but also on ensuring that educational decisions remain fair, transparent, and subject to appropriate human oversight. Because continuous assessment operates across distributed educational ecosystems, governance also extends beyond individual institutions. Federated architectures provide one approach for coordinating learning evidence while preserving institutional autonomy and reducing the need to centralize learner data. Federated learning approaches demonstrate that predictive educational models can be developed collaboratively while maintaining local control of learner information through distributed computation and privacy-preserving model exchange [28]. This architectural approach supports continuous assessment across educational institutions while preserving institutional flexibility, learner agency, and distributed stewardship of learning evidence.
Within AIM-CAI, governance requirements should distinguish among educational uses according to their potential impact on learners. Lower-risk applications, such as formative feedback, adaptive learning support, or instructional recommendations, may rely primarily on probabilistic competency estimates accompanied by appropriate uncertainty information. Within AIM-CAI, higher-risk applications—including accreditation, certification, selection, progression, or other decisions that substantially affect educational opportunity—are expected to incorporate additional governance safeguards. These safeguards include human review of competency inferences, traceability of the evidentiary basis supporting decisions, documentation of the interpretive process, and mechanisms through which learners may challenge or request review of consequential decisions. Accordingly, AIM-CAI treats risk-proportionate governance as a core architectural principle, whereby governance requirements become progressively more rigorous as the educational consequences of AI-supported inferences increase. Trust within continuous assessment depends on more than technical security. Trustworthy infrastructures require common evidence standards, psychometric validation, interoperability, auditability, learner consent, and governance mechanisms that define how evidence is generated, interpreted, shared, and used. These mechanisms ensure that learner representations remain transparent, explainable, accountable, and open to human review as they inform educational decisions.

4.7. AIM-CAI: A Functional Reference Architecture

AIM-CAI is intended to serve as a reference architecture that organizes the functional capabilities essential for AI-mediated continuous assessment. Its purpose is to offer a shared framework to guide research, engineering, institutional implementation, and governance, without prescribing a single technical implementation or fixed system architecture. As AI, educational technologies, assessment methods, and governance practices continue to evolve, this framework supplies a common vocabulary for coordinating future development while remaining intentionally extensible and open to refinement.
The six interconnected layers of AIM-CAI collectively describe the functional capabilities necessary to support continuous assessment. Learning evidence is generated across diverse educational environments, represented in forms that preserve educational meaning, translated between systems and competency frameworks, interpreted through principled evidentiary reasoning, aggregated into evolving learner representations, and overseen by trustworthy institutional mechanisms. Artificial intelligence expands the operational feasibility of these functions by supporting evidence generation, interpretation, and coordination while preserving the central role of educational assessment, psychometric validity, learner agency, and human judgment.
The following scenario illustrates how the functional capabilities represented by the six architectural layers of AIM-CAI may be implemented within an authentic learning environment. Consider a first-year engineering student enrolled in an introductory programming course supported by the AIM-CAI framework. As the student engages in quizzes, programming assignments, virtual laboratory activities, discussion forums, and interactions with an AI tutor, a wide array of learning evidence—including assessment results, coding logs, hint requests, response times, and reflective writing—is continuously collected and standardized. Learning evidence from multiple learning platforms is semantically translated into a shared competency framework, enabling probabilistic learner models to estimate mastery of programming concepts such as variables, loops, functions, debugging, and recursion. These competency estimates are incorporated into a longitudinal learner representation that reflects learning progress, confidence estimates, and evidence history while remaining adaptable as new evidence becomes available. Throughout this process, governance mechanisms ensure the provenance of evidence, maintain transparency, and uphold learner agency, thereby enabling instructors to audit AI-generated competency inferences. As a result, the continuously updated learner model can provide personalized recommendations—such as targeted practice on recursion—and alert instructors to opportunities for focused intervention, thereby supporting ongoing learner development.

5. Governance, Trust, and Federated Infrastructure

The Governance and Trust layer introduced in the AIM-CAI framework establishes the functional mechanisms that support trustworthy continuous assessment. While Section 4 outlined the layered architecture of AIM-CAI, this section explores the organizational, technical, governance, and policy conditions necessary for its implementation. Because continuous assessment relies on distributed evidence generation, probabilistic inference, semantic interoperability, and longitudinal learner representation, governance extends beyond technical implementation to address learner agency, privacy, accountability, semantic stability, algorithmic bias, institutional stewardship, and public trust. Governance, therefore, functions as an integral component of the continuous assessment infrastructure rather than as a policy framework applied after deployment.
Traditional educational assessment systems have concentrated governance largely within institutional boundaries. Grades, transcripts, and testing records are generated and controlled by individual institutions operating within relatively stable administrative structures. Continuous assessment infrastructures fundamentally alter this model by distributing evidence generation, interpretation, and competency inference across platforms, organizations, and contexts. Governance, therefore, becomes a systems-level challenge involving technical standards, semantic alignment, privacy protection, accountability structures, and epistemic authority.
Governance challenges within continuous assessment ecosystems extend beyond technical accuracy to include questions of interpretation, uncertainty, authority, and educational action. Stewardship-oriented approaches emphasize the need for transparency, provenance, accountable oversight, learner agency, and institutional responsibility when AI systems participate in generating competency inferences and informing educational decisions [65].

5.1. The Governance Challenge of Continuous Assessment

Continuous assessment infrastructures create forms of educational visibility and inferential power that extend well beyond traditional testing systems. Rather than evaluating learners through isolated examinations, these systems potentially capture interaction traces, behavioral telemetry, conversational exchanges, revision histories, multimodal signals, and longitudinal learning trajectories over extended periods of time. As a result, continuous assessment infrastructures risk evolving into pervasive systems of algorithmic monitoring if governance mechanisms are not embedded directly within their architecture.
Critical learning analytics scholarship has repeatedly warned about these risks. Selwyn [66] argued that learning analytics systems often transform complex educational processes into simplified behavioral proxies optimized for computational visibility rather than pedagogical meaning. Similarly, Knox et al. [67] described how educational AI systems may contribute to ‘machine behaviourism’, in which learners increasingly adapt themselves to algorithmically preferred forms of participation and performance. Within continuous assessment environments, learners may begin regulating their own behavior in response to invisible predictive systems designed to infer competence, engagement, or risk status from continuous streams of interaction data.
These concerns align closely with broader critiques of algorithmic governmentality and surveillance capitalism. Rouvroy [68] articulated that algorithmic systems increasingly govern human behavior through predictive profiling and anticipatory classification rather than explicit institutional control. Within education, continuous assessment infrastructures risk producing persistent “data doubles” that shape how learners are categorized, supported, or restricted across educational pathways [69]. Such systems may unintentionally normalize behavioral conformity by rewarding learners whose interaction patterns align with dominant algorithmic expectations while penalizing exploratory, culturally situated, or non-linear approaches to learning.
These governance concerns are not merely ethical abstractions. They emerge directly from the infrastructural properties, such as distributed evidence collection, probabilistic inference, and longitudinal learner modeling, that make continuous assessment possible. Consequently, trustworthy continuous assessment systems require governance architectures designed explicitly around transparency, accountability, learner agency, and privacy-preserving computation [65].

5.2. Why Standards-First Architectures Are Unlikely to Succeed

One of the central governance challenges facing continuous assessment infrastructure involves interoperability across heterogeneous educational ecosystems. Historically, educational technology systems have sought interoperability primarily through standards-based architectures such as SCORM, xAPI, LTI, and Caliper Analytics. These frameworks sought to standardize how educational content, learner interactions, and assessment records are represented and exchanged across systems.
However, the historical evolution of educational interoperability standards reveals persistent fragmentation and instability. SCORM, originally developed through the U.S. Department of Defense Advanced Distributed Learning (ADL) initiative, was designed primarily for browser-based learning objects and coarse-grained tracking of completion, time-on-task, and assessment scores. While SCORM standardized content packaging and sequencing, it remained fundamentally constrained by rigid course-centric architectures and limited telemetry capabilities.
The Experience API (xAPI) attempted to address these limitations by enabling fine-grained event logging across diverse learning environments through flexible “Actor–Verb–Object” statements stored within a Learning Record Store (LRS). Yet the flexibility of xAPI simultaneously produced severe semantic fragmentation. Different developers and platforms frequently implemented inconsistent verbs, object structures, and activity representations, resulting in highly heterogeneous telemetry ecosystems lacking stable semantic equivalence. Similar fragmentation has emerged across competing interoperability initiatives such as cmi5, Caliper Analytics, and LTI.
This fragmentation reflects a broader sociotechnical reality: innovation in educational technology consistently outpaces the development and adoption of interoperability standards. Formal standards require sustained collaboration, institutional coordination, and consensus-building processes that often unfold over many years. During that time, AI systems, learner modeling approaches, multimodal analytics, educational platforms, and competency frameworks continue to evolve. As a result, educational ecosystems remain heterogeneous even as interoperability becomes increasingly important. These conditions increase the need for architectural approaches that support semantic translation across diverse and evolving systems while preserving the educational meaning of learning evidence.
For this reason, AIM-CAI suggests that AI-mediated semantic interoperability is likely to become more viable than universal standardization. Rather than forcing all educational systems to adopt identical schemas or competency ontologies, AI-mediated translation layers increasingly support probabilistic semantic alignment across heterogeneous ecosystems. As discussed in Section 4, interoperability increasingly emerges through dynamic interpretation and ontology reconciliation rather than rigid schema compliance.

5.3. Semantic Interoperability and Translation Under Uncertainty

The shift toward AI-mediated interoperability introduces both opportunities and substantial governance risks. Semantic interoperability systems increasingly rely on semantic web technologies, such as the Resource Description Framework (RDF), the Web Ontology Language (OWL), knowledge graphs, and ontology alignment mechanisms, to reconcile heterogeneous educational representations. Recent developments in large language models and graph embeddings further enable probabilistic mapping across localized competency systems and distributed learning environments.
Some limitations of semantic interoperability, however, become visible particularly when examining analogous efforts in healthcare informatics. Clinical systems such as the Unified Medical Language System (UMLS) and Systematized Nomenclature of Medicine Clinical Terms (SNOMED CT) have spent decades attempting to reconcile heterogeneous clinical vocabularies and semantic structures [70,71]. Yet even within medicine—where concepts are often more stable and formally defined than educational competencies—ontology alignment systems remain vulnerable to semantic ambiguity, inconsistent categorization, and recurring alignment failures [72,73,74].
These limitations are highly relevant to educational assessment because educational competencies are often culturally situated, context-dependent, and epistemically contested. Competencies such as creativity, collaboration, critical thinking, and argumentation lack clear, consistent definitions, unlike medical diagnoses. The framework proposed here assumes that broad interoperability across educational systems is more likely to succeed when infrastructures support probabilistic alignment among heterogeneous representations of learning and competence rather than relying upon universally standardized semantic definitions.
This distinction is critical because it reframes interoperability as a dynamic and continuously negotiated process rather than a fixed technical state. AI-mediated translation layers, therefore, function not as universal semantic authorities but as probabilistic mediation systems capable of reconciling localized representations while preserving contextual variation and institutional autonomy.

5.4. Federated and Privacy-Preserving Architectures

Because continuous assessment infrastructures rely on highly sensitive and longitudinal learner data, centralized architectures introduce substantial political, legal, and ethical risks. Centralized learner databases create vulnerable targets for data breaches, intensify surveillance concerns, and concentrate institutional power over the representation of learners’ identities and competencies. As a result, federated and privacy-preserving architectures increasingly offer a more viable alternative.
Federated learning approaches enable predictive models and learner modeling systems to operate across distributed institutional environments without centralizing raw learner data. Rather than moving data into centralized repositories, federated architectures move analytic computation to local environments and share only aggregated model updates or privacy-preserving outputs. Emerging educational infrastructures such as SafeInsights (https://www.safeinsights.org/, accessed on 17 June 2026) increasingly adopt similar principles by utilizing secure enclaves, synthetic data sandboxes, and “zero-copy” analytics models that enable researchers to run analyses within protected institutional environments without extracting identifiable learner records.
These architectures parallel developments in healthcare Trusted Research Environments (TREs) and Personal Health Trains (PHTs), where sensitive patient data remain locally governed while supporting distributed computation and collaborative analysis [75]. Such systems increasingly rely on privacy-enhancing technologies, including differential privacy, federated learning, secure enclaves, and secure multi-party computation to reduce disclosure risks while preserving analytic functionality.
Federated systems are not without limitations. Differential privacy techniques often introduce tradeoffs between privacy and analytic precision, particularly for small or marginalized learner populations. Federated systems also present challenges related to error localization, debugging, and fault reproducibility [76]. Moreover, decentralized governance structures may diffuse accountability when erroneous or discriminatory learner inferences emerge across distributed systems [77,78].
These limitations reinforce the importance of governance within distributed continuous assessment ecosystems. As learning evidence, learner representations, and AI-supported inferences increasingly span institutions and platforms, governance becomes essential for preserving accountability, semantic consistency, auditability, and human oversight. Within the AIM-CAI reference architecture, governance is treated as an integral component of trustworthy continuous assessment because the interpretation and use of learning evidence require transparency, accountability, and human oversight across distributed educational ecosystems.

5.5. Epistemic and Sociotechnical Risks

Continuous assessment infrastructures also raise broader epistemic and sociotechnical concerns related to educational authority, cultural representation, and algorithmic normalization. Data colonialism scholars indicate that contemporary digital systems increasingly treat human activity itself as a resource for extraction, monetization, and behavioral prediction [79]. Within education, continuous assessment infrastructures may extend this logic by transforming learning activity into persistent streams of computationally extractable behavioral data.
These concerns are especially significant because educational AI systems are frequently trained on historically biased datasets and culturally dominant linguistic patterns. As a result, automated assessment systems may undervalue non-Western rhetorical traditions, culturally situated epistemologies, or alternative forms of expression and reasoning [80,81,82]. Continuous assessment systems, therefore, risk reinforcing epistemic colonialism by embedding dominant assumptions about competence, communication, and legitimate knowledge directly into algorithmic infrastructures.
Alternative governance perspectives grounded in relational and decolonial approaches to data governance emphasize learner sovereignty, epistemic pluralism, and participatory design. Such approaches indicate that educational data should not be treated merely as extractable institutional property but as relational and community-embedded representations requiring collective stewardship and contextual interpretation. Within AIM-CAI continuous assessment infrastructures, this implies that local educational communities must retain meaningful authority over how competencies are defined, represented, and interpreted.

5.6. Infrastructural Governance and Trust

To address these risks, governance within AI-mediated continuous assessment infrastructures must become infrastructural rather than merely regulatory. Governance mechanisms must be embedded directly within technical architectures, semantic translation systems, inferential models, and data-sharing protocols, rather than appended solely through external policy statements.
The National Institute of Standards and Technology (NIST) AI Risk Management Framework (AI RMF) (https://www.nist.gov/itl/ai-risk-management-framework, accessed on 8 June 2026) provides one important model for operationalizing this approach. The framework conceptualizes trustworthy AI through iterative governance functions involving Govern, Map, Measure, and Manage. Within continuous assessment ecosystems, these functions require tracing data lineage, documenting inferential assumptions, auditing algorithmic behavior, monitoring distributional drift, and establishing mechanisms for human oversight and intervention. Governance, therefore, operates across all infrastructural layers rather than functioning as an external administrative process. A trustworthy continuous assessment infrastructure requires:
  • Auditability.
  • Transparency.
  • Probabilistic uncertainty estimation.
  • Human override mechanisms.
  • Semantic accountability.
  • Fairness monitoring.
  • Learner agency.
It is important that governance architectures remain adaptive as the educational ecosystems they govern continue to evolve rapidly. AI-mediated continuous assessment infrastructures, therefore, require governance systems capable not only of regulating current technologies but also of responding continuously to emerging forms of evidence generation, learner modeling, semantic translation, and probabilistic inference.
Governance within AI-mediated continuous assessment infrastructures cannot be reduced solely to technical standards, regulatory compliance, or algorithmic auditing. Educational infrastructures are deeply sociotechnical systems shaped by institutional workflows, faculty practices, administrative structures, and learner participation. Prior research in educational technology and learning analytics repeatedly demonstrates that systems designed without sustained stakeholder involvement often fail to align with authentic educational contexts, increase faculty burden, disrupt institutional practice, or generate low levels of trust and adoption [83,84]. Hence, any continuous assessment infrastructure, including AIM-CAI, requires participatory, human-centered governance models that involve instructors, administrators, learners, and institutional stakeholders throughout the design, implementation, interpretation, and oversight processes, rather than treating educational communities as passive recipients of externally developed AI systems. Consistent with the principle of risk-proportionate governance discussed in Section 4.6, auditing, monitoring, documentation, and institutional oversight should be calibrated to the potential consequences of AI-supported educational decisions.

6. Institutional and Lifelong Learning Implications

The emergence of AI-mediated continuous assessment infrastructure carries implications that extend far beyond assessment itself. Historically, educational institutions maintained near-exclusive authority over the production, interpretation, and certification of learning evidence. At present, episodic assessment occurs primarily within institutional boundaries through formally sanctioned courses, examinations, and credentialing systems. As a result, institutions control not only instructional delivery but also the mechanisms for recognizing, documenting, and legitimizing competence. The development of continuous, distributed, and infrastructure-based assessment systems may fundamentally alter this arrangement.
AIM-CAI supports the generation, representation, translation, interpretation, governance, and longitudinal use of learning evidence across heterogeneous learning environments rather than exclusively within formal institutional settings. Learning evidence may emerge through workplace participation, AI tutoring systems, collaborative online communities, simulations, portfolios, mobile learning environments, authentic learning experiences, and self-directed learning activities. Thus, assessment no longer belongs exclusively to educational institutions. Instead, institutions increasingly become participants within broader ecosystems of competency interpretation and validation.

6.1. From Static Credentials to Continuous Credentialing

One of the most immediate implications of a continuous assessment infrastructure involves the transformation of educational credentialing systems. Traditional credentials such as transcripts, grade point averages, and course completions function primarily as static archival records summarizing isolated performances within bounded institutional contexts [55]. These records provide limited insight into longitudinal competency development, contextual performance variation, or evolving learner capabilities [8].
Continuous assessment infrastructures, in contrast, support more dynamic forms of competency representation. As discussed in Section 4, probabilistic learner models increasingly generate continuously updated competency profiles reflecting longitudinal evidence accumulation across contexts and time. Within AIM-CAI, credentialing shifts from static certification toward continuous competency interpretation. Rather than representing learning as a completed event frozen within a transcript, competency representations become developmental, revisable, and continuously informed by new evidence streams. Because credentialing can substantially affect educational and professional opportunities, its use within AIM-CAI should be subject to the risk-proportionate governance established in Section 4.6.
This shift may significantly reshape how educational and professional systems interpret learner capability. Employers, institutions, and learners themselves may increasingly rely on dynamic evidence trajectories rather than singular degree completions or isolated grades. This does not imply the disappearance of degrees or institutional credentials. Instead, degrees increasingly function as one layer within broader ecosystems of competency evidence and longitudinal learner representation.
Continuous credentialing raises substantial governance challenges. Dynamic competency systems may create pressure toward perpetual evaluation and algorithmic ranking if governance architectures are insufficiently constrained. Therefore, institutions must carefully distinguish between continuous evidence interpretation and continuous surveillance. Competency representations should remain probabilistic, context-sensitive, and learner-centered rather than functioning as permanent or deterministic reputational scores.

6.2. Assessment Beyond Institutional Boundaries

The second major implication of AIM-CAI involves the decoupling of assessment from exclusively institutional environments. Historically, educational institutions have maintained authority over what counted as legitimate evidence of learning because they controlled both instruction and assessment. However, as learning increasingly occurs across distributed digital ecosystems, workplaces, online communities, video streaming sites, and AI-supported environments, institutional boundaries are less able to contain meaningful competency development.
This transition is particularly visible within workforce learning and professional development contexts. Workplace-based assessment models, for instance, Programmatic Assessment by Cees van der Vleuten and Lambert Schuwirth, already rely heavily on longitudinal evidence gathered across authentic practice environments rather than isolated examinations. Similarly, AI-based tutoring systems, professional simulations, collaborative platforms, game-based learning, and self-directed online learning increasingly generate rich evidence on problem-solving, communication, collaboration, and strategic reasoning outside formal classrooms.
As continuous assessment infrastructures mature, educational institutions may therefore lose their historical monopoly over competency validation. This does not eliminate the role of institutions; rather, institutions increasingly shift from functioning as exclusive issuers of learning legitimacy toward functioning as validators, interpreters, and governors within broader competency ecosystems. Universities may increasingly contribute expertise in evidence interpretation, psychometric validation, governance, and credential integration while recognizing that learning itself occurs across distributed contexts extending beyond institutional control.
This transition also creates new opportunities for lifelong learning systems. Traditional educational structures often fragment learning into isolated phases associated with degrees, semesters, or formal programs. Continuous assessment infrastructures such as AIM-CAI instead support more persistent and longitudinal learner representations capable of extending across career transitions, reskilling efforts, and evolving professional trajectories. In this sense, competency development increasingly becomes lifelong rather than institutionally episodic.

6.3. Personalized Pathways and AI-Mediated Advising

Continuous competency inference also has important implications for personalization and learner support. Traditional educational pathways are typically organized around standardized sequences of courses and time-based progression structures designed for administrative scalability rather than individual developmental variation. Although adaptive learning systems have attempted to personalize instruction for decades, many personalization models, such as Item Response Theory (IRT) [85] and Rule-based models [86], have remained limited to narrow domains or simplistic behavioral adaptation.
AI-mediated continuous assessment infrastructures potentially support more sophisticated forms of personalization grounded in longitudinal competency inference. Within AIM-CAI (as illustrated in Figure 1), evidence generated across multiple environments feeds into dynamic learner models that continuously update estimates of learner strengths, weaknesses, developmental trajectories, and contextual performance patterns. This creates opportunities for AI-mediated advising systems capable of supporting individualized learning pathways, targeted interventions, and adaptive credentialing structures.
The AIM-CAI framework proposed here conceptualizes personalization differently from many earlier adaptive learning systems. Personalization is not driven primarily by learner preferences, demographic categories, or simplistic engagement metrics. Instead, personalization increasingly emerges through continuous competency interpretation informed by probabilistic evidence accumulation across contexts and time. In AIM-CAI, AI-mediated advising systems may help learners identify emerging strengths, developmental gaps, transferable competencies, and alternative learning opportunities aligned with evolving goals and trajectories.
Nevertheless, personalization infrastructures introduce substantial risks related to predictive determinism and algorithmic path dependency. Continuous competency systems may unintentionally constrain learners’ opportunities by reinforcing early probabilistic classifications or narrowing educational pathways based on predictive models. Hence, personalized learning infrastructures require governance mechanisms that preserve learner agency, support exploratory learning, and prevent premature foreclosure of future opportunities based on incomplete or historically biased evidence.

6.4. Faculty Roles and Institutional Transformation

The effective implementation, legitimacy, and long-term sustainability of continuous assessment infrastructures depend on their alignment with authentic instructional practice and institutional workflows. Introducing new assessment infrastructures inevitably redistributes faculty labor, and technologies that ignore how teachers actually work have historically been oversold and underused [87]. Therefore, instructors and institutional stakeholders must function not merely as users of AI-mediated assessment systems but as active participants in shaping how evidence, competency, and learner development are represented and governed within continuous assessment ecosystems.
The emergence of continuous assessment infrastructures also reshapes faculty roles and institutional responsibilities. Traditional educational systems devote substantial faculty labor to grading, administrative evaluation, and episodic performance assessment [1,88]. As AI-mediated inferential systems increasingly support evidence collection, pattern detection, and probabilistic learner modeling, faculty roles may shift away from routine evaluative tasks toward more interpretive, relational, and design-oriented functions. Within AIM-CAI ecosystems, faculty increasingly function as mentors, evidence interpreters, learning architects, and governance participants rather than solely as graders or content deliverers. Their expertise is critical for validating competency interpretations, contextualizing learner trajectories, designing meaningful evidence-generating experiences, and identifying situations in which algorithmic systems fail to adequately represent learner development.
This transformation also reshapes the role of educational institutions. Universities may increasingly function as trusted participants within broader learning and competency ecosystems, contributing to psychometric validation, competency governance, interdisciplinary learning design, evidence interpretation, and the stewardship of trusted credentials. Continuous assessment, therefore, elevates institutional responsibility for ensuring that competency inferences remain valid, equitable, interpretable, and socially accountable.

6.5. Lifelong Competency Ecosystems

The convergence of these transformations suggests the emergence of lifelong competency ecosystems that extend beyond traditional course- and degree-based educational structures. Continuous assessment infrastructures increasingly support learner representations that persist across educational institutions, workplaces, professional communities, and self-directed learning environments. In this model, learning trajectories become continuous rather than segmented into isolated institutional episodes.
Within AIM-CAI, lifelong competency should not be interpreted as a purely technical process for optimizing workforce efficiency. As emphasized throughout Section 4 and Section 5, AIM-CAI also raises profound governance, epistemic, and sociotechnical questions regarding who defines competence, who controls learner representations, and how algorithmic systems shape educational opportunity. Thus, the future of lifelong learning ecosystems depends not only on advances in AI and interoperability but also on governance architectures capable of preserving learner agency, epistemic pluralism, privacy, and institutional accountability.
The emergence of AI-mediated continuous assessment infrastructure suggests that educational systems are entering a transition in which assessment increasingly operates as a distributed interpretive infrastructure rather than as an isolated institutional procedure. As shown in Figure 1, evidence generation, evidence representation, semantic interpretation, probabilistic inference, and competency development increasingly occur across interconnected ecosystems extending beyond the traditional boundaries of schools and universities. This transition fundamentally reshapes the role of institutions, credentials, faculty, and learners within the broader landscape of education and lifelong learning.

6.6. Practical Considerations for Institutional Adoption

Although AIM-CAI is presented as a reference architecture rather than a prescribed implementation, institutions considering AI-mediated continuous assessment may benefit from an incremental approach to adoption. Initial efforts should focus on strengthening governance, identifying meaningful sources of learning evidence, and establishing clear competency frameworks before introducing AI-supported inference. Institutions should also establish policies governing evidence stewardship, learner access, data retention, human oversight, and the appropriate use of competency inferences for different educational purposes. Early implementations are likely to be most appropriate in lower-risk contexts, such as formative feedback, learner advising, or instructional improvement, where educational benefits can be evaluated while governance processes mature. Experience gained through these bounded implementations can inform subsequent expansion to more complex applications, provided that appropriate validation, transparency, and institutional oversight are maintained.

7. Research Agenda and Open Challenges

The AIM-CAI reference architecture proposed in this paper organizes the functional capabilities needed to support AI-mediated continuous assessment. It provides a shared framework that can evolve alongside advances in educational assessment, artificial intelligence, and digital learning technologies. Although these advances have made continuous assessment increasingly feasible, important psychometric, technical, governance, and theoretical challenges remain. Addressing these challenges requires interdisciplinary research spanning psychometrics, learning sciences, artificial intelligence, human–computer interaction, data governance, and critical data studies.
The AIM-CAI framework comprises multiple interacting layers of evidence generation, semantic interpretation, probabilistic inference, learner modeling, and governance. Each functional layer reveals a corresponding set of research and engineering challenges related to validity, reliability, interpretability, fairness, interoperability, and institutional legitimacy. Thus, the future development of AI-mediated continuous assessment systems depends not only on technical advancement but also on rigorous empirical validation, governance innovation, and sustained critical scrutiny.

7.1. Validity and Reliability in Continuous Inference

One of the most fundamental open questions concerns validity within continuous and distributed assessment ecosystems. Traditional educational measurement relied heavily on controlled testing conditions, standardized tasks, and relatively stable psychometric assumptions designed to support validity and reliability claims. However, the proposed continuous assessment infrastructure operates across highly heterogeneous learning environments characterized by dynamic interactions, multimodal evidence streams, and substantial contextual variability. Competency interpretations within such systems are therefore inherently probabilistic, relying on learner models that continuously update inferences from accumulating evidence across contexts and time.
This shift raises important questions regarding how validity should be conceptualized when competency inferences emerge from temporally distributed evidence rather than isolated assessment events. Although frameworks such as Evidence-Centered Design and Expanded Evidence-Centered Design provide important foundations for evidentiary reasoning [41,89], substantial research is still needed to determine how validity arguments should operate within continuously evolving inferential systems.
Similarly, reliability becomes more complex under conditions of continuous evidence accumulation. Dynamic Bayesian networks, knowledge tracing systems, and multimodal analytics models remain highly sensitive to noisy evidence, unstable telemetry, and contextual variability [20,39,62,63,90]. Learner behavior may fluctuate due to fatigue, scaffolding, interface confusion, or social conditions rather than genuine changes in competence. Therefore, future research must examine the following:
  • How longitudinal evidence streams affect inferential stability;
  • How probabilistic confidence estimates should be represented;
  • How continuous systems distinguish developmental growth from transient behavioral variation.

7.2. Interpretability, Transparency, and Human Oversight

A second major challenge involves interpretability and transparency within increasingly complex AI-mediated inferential systems. Many contemporary learner modeling approaches rely on machine learning architectures that achieve high predictive performance while remaining difficult for educators, learners, and institutions to interpret. Deep learning systems, transformer architectures, and multimodal fusion models may generate competency estimates without providing meaningful explanations regarding how evidence was weighted, interpreted, or translated into inferential conclusions.
This challenge is significant because educational assessment carries substantial social and institutional consequences. Competency inferences may influence learners’ opportunities, placement decisions, intervention systems, and credentialing pathways. As such, continuous assessment infrastructures require research into Explainable AI (XAI) and interpretable AI systems that support meaningful human oversight. A deeper understanding of the explainable and interpretable aspects of AI systems will increase transparency, enabling stakeholders to understand precisely how inputs are transformed into meaningful outputs. Future research should therefore investigate:
  • Interpretable probabilistic learner models;
  • Transparent semantic translation mechanisms;
  • Explainable and uncertainty-aware dashboards;
  • Human-in-the-loop governance systems capable of intervening when inferential confidence deteriorates or when algorithmic systems generate questionable classifications.
Interpretability is not merely a technical feature; it is foundational to institutional trust and the legitimacy of continuous assessment systems. Learners, educators, and institutions must be able to understand how competency claims are generated, evaluated, revised, and governed. Therefore, future research must extend beyond questions of privacy, interoperability, and technical governance to examine how continuous assessment systems communicate uncertainty, support learner agency, and ensure accountability in the interpretation and use of competency inferences. As AI systems assume a greater role in generating and synthesizing assessment evidence, governance frameworks must address not only how evidence is collected and shared, but also how competency claims are interpreted, contested, and used in consequential educational decisions [65].

7.3. Bias, Fairness, and Algorithmic Harm

Continuous assessment infrastructures also raise profound concerns regarding fairness and algorithmic bias. Because AI-mediated learner models are trained on historical educational data, they may reproduce or amplify existing inequities affecting marginalized populations. Furthermore, predictive systems may disproportionately classify certain learners as at risk [60], introduce cultural biases into the assessments [80,81], or reinforce dominant behavioral norms embedded within training data and competency frameworks [67].
These risks become especially significant under conditions of continuous and longitudinal assessment. Unlike isolated examinations, continuous and longitudinal assessment systems accumulate persistent inferential profiles over time. Early probabilistic classifications may therefore shape future educational opportunities through self-reinforcing feedback loops. For example, learners identified as low-performing or high-risk may receive narrower pathways, reduced opportunities, or intensified monitoring, which ultimately reproduce the very outcomes the systems attempt to predict.
Research on algorithmic fairness in education increasingly demonstrates that traditional bias mitigation approaches remain insufficient for complex, distributed inferential systems [28,60,65,77,78]. Therefore, future research must examine:
  • How bias propagates through semantic translation layers;
  • How federated systems affect fairness auditing;
  • How uncertainty should be communicated across demographic groups;
  • How governance architectures can preserve learner agency and epistemic pluralism.
In addition, continuous assessment systems may unintentionally normalize narrow forms of participation and interaction. Learners may increasingly adapt their behavior toward algorithmically preferred interaction patterns, leading to behavioral conformity and self-surveillance. Research is therefore needed not only on predictive accuracy but also on the sociotechnical effects of continuous inferential environments on learner autonomy, creativity, exploration, and identity formation.

7.4. Interoperability, Governance, and Cross-Contextual Inference

A central assumption underlying AI-mediated continuous assessment infrastructure, AIM-CAI, is that evidence generated across heterogeneous environments can be meaningfully interpreted, translated, and reconciled across contexts and time. However, substantial open questions remain regarding the feasibility, governance, and legitimacy of cross-contextual competency inference. Educational ecosystems remain deeply fragmented across institutions, platforms, domains, cultures, and competency frameworks. Even within highly structured domains such as healthcare informatics, semantic interoperability systems remain vulnerable to ontology drift, alignment instability, and interpretive inconsistency. Educational competencies are often considerably more ambiguous, context-dependent, and socially constructed than medical diagnoses or technical standards.
Nevertheless, governance challenges associated with continuous assessment infrastructures remain among the least resolved areas of research [24]. Existing educational privacy frameworks, such as the Family Educational Rights and Privacy Act (FERPA) (https://studentprivacy.ed.gov/faq/what-ferpa, accessed on 9 June 2026) and the Children’s Online Privacy Protection Act (COPPA) (https://www.ftc.gov/business-guidance/privacy-security/childrens-privacy, accessed on 9 June 2026), were designed primarily for institutionally centralized records rather than distributed, probabilistic, and continuously updated evidence ecosystems. As continuous assessment infrastructures increasingly depend on workplace participation, collaborative interaction, AI tutoring systems, multimodal evidence, and self-directed learning environments, substantial uncertainty emerges regarding ownership of learner evidence, rights to algorithmic inference, portability of competency representations, and governance of longitudinal learner models.
These tensions become especially significant within federated and privacy-preserving architectures. Although secure enclaves, federated learning, differential privacy, and trusted research environments may reduce many surveillance risks, they also introduce new governance complexities related to auditability, accountability, semantic consistency, and coordination failure. Distributed systems may obscure the origins of inferential errors, who governs semantic translation processes, and which institutional actors remain responsible when harmful competency classifications emerge.
Many of the central research challenges associated with AI-mediated continuous assessment infrastructure are best understood not as isolated technical problems but as ongoing tensions among competing infrastructural priorities. As summarized in Table 1, future research must address how continuous assessment ecosystems balance interoperability and local autonomy, privacy and utility analytics, learner agency and predictive automation, transparency and model complexity, and distributed governance and institutional accountability. These challenges also raise broader epistemic and sociotechnical questions concerning power, representation, and institutional authority, as continuous assessment systems increasingly shape how competence is defined, represented, and legitimized across educational ecosystems. Without careful governance design, continuous inferential systems risk privileging dominant cultural norms, rhetorical traditions, and behavioral expectations while marginalizing non-traditional or culturally situated forms of learning and expression.
Future research must therefore investigate governance architectures capable of preserving epistemic pluralism, supporting learner agency, and maintaining institutional accountability while still enabling meaningful interoperability across distributed learning ecosystems. This includes research on federated governance structures, trusted research environments, permissioned access systems, participatory governance models, semantic accountability mechanisms, and learner-centered consent architectures capable of operating within continuously evolving educational infrastructures.

7.5. Over-Surveillance and the Limits of Continuous Assessment

Future research must confront a fundamental tension at the center of any continuous assessment infrastructure: the same systems capable of supporting more authentic, longitudinal, and context-sensitive interpretations of learning also create unprecedented possibilities for educational surveillance and behavioral monitoring.
Continuous assessment infrastructures may improve formative support, reduce reliance on isolated examinations, and better represent developmental learning trajectories. However, these same infrastructures may also normalize constant monitoring, predictive profiling, and algorithmic governance within education. This tension cannot be resolved solely through technical optimization because it reflects competing visions of what education itself should become. Thus, future research must examine not only how continuous assessment systems can be built but also the following:
  • When should they be used;
  • Where limits should be imposed;
  • Which forms of learner monitoring remain pedagogically and ethically unacceptable?
The framework proposed in this paper, therefore, should not be interpreted as a deterministic blueprint for the future of education. Rather, it is intended as a conceptual infrastructure model and research agenda designed to support ongoing scholarly, technical, and institutional exploration of how AI-mediated continuous assessment systems might be developed responsibly, governed democratically, and constrained appropriately within evolving educational ecosystems.

7.6. Research and Implementation Opportunities

The research challenges described throughout this section illustrate that advancing AI-mediated continuous assessment will require coordinated progress across multiple functional capabilities rather than isolated technical innovations. Improvements in evidence generation, learner representations, semantic interoperability, inferential validity, governance, and human oversight are interdependent, and advances in one area often depend on progress in others. The AIM-CAI reference architecture provides a common framework for organizing these complementary research and implementation efforts.
Because each architectural layer represents a distinct functional capability, researchers and institutions can investigate individual components of the architecture through bounded implementation studies while contributing to the refinement of the broader framework. Such studies may focus on individual layers, examine interactions across multiple layers, or evaluate integrated implementations within authentic educational settings. Beyond evaluating individual architectural capabilities, future research should also examine implementation within authentic educational settings, including faculty workflows, learner trust, interpretability, advising practices, organizational adoption, and long-term sustainability. Collectively, these efforts provide an empirical pathway for evaluating, refining, and extending the AIM-CAI reference architecture.
As new assessment methods, AI capabilities, governance approaches, and educational technologies emerge, additional functional capabilities and architectural refinements may also be identified. The framework is therefore intended to evolve alongside the field, providing a shared reference architecture that supports collaboration across disciplines and institutions while guiding the continued development of trustworthy AI-mediated continuous assessment.

8. Conclusions

Educational assessment has traditionally relied on examinations, assignments, transcripts, and credentials that summarize learning at discrete points in time. Converging advances in artificial intelligence, learning analytics, learner modeling, multimodal evidence, and semantic interoperability are making it increasingly feasible to represent learning as a continuous process across contexts and time.
This paper introduced the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), a conceptual reference architecture that organizes the functional capabilities required to support AI-mediated continuous assessment. It organizes the functional capabilities that enable learning evidence to be generated, represented, interpreted, and governed across distributed educational environments. As a shared architectural framework, AIM-CAI provides a common vocabulary for coordinating research, engineering, implementation, and governance across the evolving continuous assessment ecosystem.
The conceptual shifts introduced by AIM-CAI are summarized in Table 2. Collectively, these shifts extend traditional assessment by treating examinations, assignments, projects, workplace experiences, AI-supported learning, and other authentic learning activities as complementary forms of learning evidence that contribute to an evolving understanding of learner development. Assessment becomes a longitudinal process of accumulating, interpreting, and governing learning evidence across contexts and time while preserving the educational value of individual assessment events within a broader evidentiary record.
The primary contribution of AIM-CAI is architectural. It brings together evidence generation, semantic translation, evidentiary reasoning, learner representation, and governance within a single reference architecture while identifying the functional capabilities required to support AI-mediated continuous assessment. By organizing these complementary capabilities, the framework provides a foundation for designing, comparing, evaluating, and refining AI-mediated continuous assessment systems while supporting collaboration across disciplines and institutions.
Continued empirical research will refine the architecture, identify additional functional capabilities as they emerge, and strengthen the evidence needed to support valid, equitable, interpretable, and trustworthy continuous assessment. As educational ecosystems continue to evolve, assessment is becoming an increasingly integrated capability of learning rather than an isolated institutional event. AIM-CAI provides a shared architectural foundation for coordinating that transition while supporting more continuous, transparent, learner-centered, and trustworthy educational decision-making. AIM-CAI is intended to serve as a foundation for future research and bounded implementations rather than as a blueprint for immediate deployment.

Author Contributions

Conceptualization, D.S.M.; investigation, D.S.M. and M.N.H.; resources, D.S.M.; writing—original draft preparation, D.S.M.; writing—review and editing, D.S.M. and M.N.H.; supervision, D.S.M.; project administration, D.S.M.; funding acquisition, D.S.M. All authors have read and agreed to the published version of the manuscript.

Funding

The research reported here was supported by the Institute of Education Sciences, U.S. Department of Education, through Grants R305N210041 and R305T240035 to Arizona State University and Grant NSF IIS 2153481 to Rice University and Arizona State University. The opinions expressed are those of the authors and do not represent views of the Institute of Education Sciences, the U.S. Department of Education, or the National Science Foundation.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

During the preparation of this manuscript, the authors used various generative AI tools, including ChatGPT (5.5), Claude (4.8 and 5), and Gemini (3.5), for brainstorming, literature searches, copyediting, and revisions. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
AIEDArtificial Intelligence in Education
AIM-CAIAI-Mediated Continuous Assessment Infrastructure
GPAGrade Point Average
CBECompetency-Based Education
MMLAMultimodal Learning Analytics
ECDEvidence-Centered Design
HMMHidden Markov Model
BKTBayesian Knowledge Tracing
BOLLBlockchain of Learning Logs
BEMPASBlockchain-based Employee Performance Assessment System
e-ECDExpanded Evidence-Centered Design
SCORMSharable Content Object Reference Model
IMSIMS Global Learning Consortium (now 1EdTech Consortium)
LTILearning Tools Interoperability
xAPIExperience API
LRSLearning Record Store
RDFResource Description Framework
MOOCMassive Open Online Course
DBNDynamic Bayesian Networks
ADLAdvanced Distributed Learning
OWLWeb Ontology Language
SNOMED CTSystematized Nomenclature of Medicine Clinical Terms
UMLSUnified Medical Language System
TRETrusted Research Environment
PHTPersonal Health Train
NISTNational Institute of Standards and Technology
AI RMFAI Risk Management Framework
IRTItem Response Theory
XAIExplainable AI
FERPAFamily Educational Rights and Privacy Act
COPPAChildren’s Online Privacy Protection Act

References

  1. French, S.; Dickerson, A.; Mulder, R.A. A Review of the Benefits and Drawbacks of High-Stakes Final Examinations in Higher Education. High. Educ. 2024, 88, 893–918. [Google Scholar] [CrossRef] [Scilit]
  2. Tyack, D.B. The One Best System: A History of American Urban Education; Harvard University Press: Cambridge, MA, USA, 1974; Volume 95, ISBN 9780674637825. [Google Scholar]
  3. Weber, M. The Theory of Social and Economic Organization; Free Press: New York, NY, USA, 1947; ISBN 978-0-684-83640-9. [Google Scholar]
  4. Beniger, J. The Control Revolution: Technological and Economic Origins of the Information Society; Harvard University Press: Cambridge, MA, USA, 1986; ISBN 978-0-674-16985-2. [Google Scholar]
  5. Madaus, G.F.; O’Dwyer, L.M. A Short History of Performance Assessment: Lessons Learned. Phi Delta Kappan 1999, 80, 688–695. [Google Scholar]
  6. Broadfoot, P.M. Education, Assessment and Society; Open University Press: London, UK, 1996; ISBN 978-0-335-19601-6. [Google Scholar]
  7. Stobart, G. Testing Times: The Uses and Abuses of Assessment; Routledge: London, UK, 2008. [Google Scholar] [CrossRef] [Scilit]
  8. Pellegrino, J.W.; Chudowsky, N.; Glaser, R. (Eds.) Knowing What Students Know. In The Science and Design of Educational Assessment; National Research Council, National Academy Press: Washington, DC, USA, 2001. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Black, P.; Wiliam, D. Assessment and Classroom Learning. Assess. Educ. Princ. Policy Pract. 1998, 5, 7–74. [Google Scholar] [CrossRef] [Scilit]
  10. Shepard, L.A. The Role of Assessment in a Learning Culture. Educ. Res. 2000, 29, 4–14. [Google Scholar] [CrossRef]
  11. Wiggins, G.P. Assessing Student Performance: Exploring the Purpose and Limits of Testing; Jossey-Bass/Wiley: San Francisco, CA, USA, 1993; ISBN 10: 1555425925. [Google Scholar]
  12. Brown, J.S.; Collins, A.; Duguid, P. Situated Cognition and the Culture of Learning. Educ. Res. 1989, 18, 32–42. [Google Scholar] [CrossRef]
  13. Lave, J.; Wenger, E. Situated Learning: Legitimate Peripheral Participation; Cambridge University Press: Cambridge, UK, 1991. [Google Scholar] [CrossRef] [Scilit]
  14. Schuwirth, L.W.T.; van der Vleuten, C.P.M. Programmatic Assessment: From Assessment of Learning to Assessment for Learning. Med. Teach. 2011, 33, 478–485. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. van der Vleuten, C.P.M.; Schuwirth, L.W.T. Assessing Professional Competence: From Methods to Programmes. Med. Educ. 2005, 39, 309–317. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. van der Vleuten, C.P.M.; Schuwirth, L.W.T.; Driessen, E.W.; Dijkstra, J.; Tigelaar, D.; Baartman, L.K.J.; van Tartwijk, J. A Model for Programmatic Assessment Fit for Purpose. Med. Teach. 2012, 34, 205–214. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. How, M.-L.; Hung, W.L.D. Educational Stakeholders’ Independent Evaluation of an Artificial Intelligence-Enabled Adaptive Learning System Using Bayesian Network Predictive Simulations. Educ. Sci. 2019, 9, 110. [Google Scholar] [CrossRef] [Scilit]
  18. Gašević, D.; Greiff, S.; Shaffer, D.W. Towards Strengthening Links between Learning Analytics and Assessment: Challenges and Potentials of a Promising New Bond. Comput. Hum. Behav. 2022, 134, 107304. [Google Scholar] [CrossRef] [Scilit]
  19. Siemens, G.; Baker, R.S.J.D. Learning Analytics and Educational Data Mining: Towards Communication and Collaboration. In Proceedings of the 2nd International Conference on Learning Analytics and Knowledge (LAK), Vancouver, BC, Canada, 29 April–2 May 2012; pp. 252–254. [Google Scholar] [CrossRef] [Scilit]
  20. Blikstein, P.; Worsley, M. Multimodal Learning Analytics and Education Data Mining: Using Computational Technologies to Measure Complex Learning Tasks. J. Learn. Anal. 2016, 3, 220–238. [Google Scholar] [CrossRef] [Scilit]
  21. McNamara, D.S.; Crossley, S.A.; Roscoe, R. Natural Language Processing in an Intelligent Writing Strategy Tutoring System. Behav. Res. Methods 2013, 45, 499–515. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Timmerman, K.; Doom, T. Infrastructure for Continuous Assessment of Retained Relevant Knowledge. ACM Inroads 2017, 8, 73–77. [Google Scholar] [CrossRef] [Scilit]
  23. Miao, F.; Holmes, W. Guidance for Generative AI in Education and Research; UNESCO: Paris, France, 2023; Available online: https://unesdoc.unesco.org/ark:/48223/pf0000386693?locale=en (accessed on 24 June 2026).
  24. Alfaleh, M. Sustainable AI-Driven Assessment in Higher Education: A Systematic Review of Fairness, Transparency, Pedagogical Innovation, and Governance. Sustainability 2026, 18, 785. [Google Scholar] [CrossRef] [Scilit]
  25. Mislevy, R.J.; Haertel, G.D. Implications of Evidence-Centered Design for Educational Testing. Educ. Meas. Issues Pract. 2006, 25, 6–20. [Google Scholar] [CrossRef] [Scilit]
  26. Koedinger, K.R.; Corbett, A.T.; Perfetti, C. The Knowledge-Learning-Instruction Framework: Bridging the Science-Practice Chasm to Enhance Robust Student Learning. Cogn. Sci. 2012, 36, 757–798. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Shute, V.J. Stealth Assessment in Computer-Based Games to Support Learning. In Computer Games and Instruction; Tobias, S., Fletcher, J.D., Eds.; Information Age Publishing: Charlotte, NC, USA, 2011; pp. 503–524. [Google Scholar] [CrossRef] [Scilit]
  28. Khalil, M.; Shakya, R.; Liu, Q. Towards Privacy-Preserving Data-Driven Education: The Potential of Federated Learning. In Proceedings of the 2025 International Conference on New Trends in Computing Sciences (ICTCS), Amman, Jordan, 16–18 April 2025; pp. 113–118. [Google Scholar] [CrossRef] [Scilit]
  29. Griffin, P.; Care, E. Assessment and Teaching of 21st Century Skills: Methods and Approach; Springer: Dordrecht, The Netherlands, 2015. [Google Scholar] [CrossRef] [Scilit]
  30. Hattie, J.; Timperley, H. The Power of Feedback. Rev. Educ. Res. 2007, 77, 81–112. [Google Scholar] [CrossRef] [Scilit]
  31. Gulikers, J.T.M.; Bastiaens, T.J.; Kirschner, P.A. A Five-Dimensional Framework for Authentic Assessment. Educ. Technol. Res. Dev. 2004, 52, 67–86. [Google Scholar] [CrossRef] [Scilit]
  32. Burnette, D.M. The Renewal of Competency-Based Education: A Review of the Literature. J. Contin. High. Educ. 2016, 64, 84–93. [Google Scholar] [CrossRef] [Scilit]
  33. Kato, S.; Galán-Muros, V.; Weko, T. The Emergence of Alternative Credentials. In OECD Education Working Papers No. 216; OECD Publishing: Paris, France, 2020. [Google Scholar] [CrossRef]
  34. Wheelahan, L.; Moodie, G. Gig Qualifications for the Gig Economy: Micro-Credentials and the ‘Hungry Mile’. High. Educ. 2022, 83, 1279–1295. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Crossley, S.A.; McNamara, D.S. Predicting Second Language Writing Proficiency: The Roles of Cohesion and Linguistic Sophistication. J. Res. Read. 2012, 35, 115–135. [Google Scholar] [CrossRef] [Scilit]
  36. McNamara, D.S.; Crossley, S.A.; McCarthy, P.M. Linguistic Features of Writing Quality. Writ. Commun. 2010, 27, 57–86. [Google Scholar] [CrossRef] [Scilit]
  37. Graesser, A.C.; Chipman, P.; Haynes, B.C.; Olney, A. AutoTutor: An Intelligent Tutoring System with Mixed-Initiative Dialogue. IEEE Trans. Educ. 2005, 48, 612–618. [Google Scholar] [CrossRef]
  38. Blikstein, P. Multimodal Learning Analytics. In Proceedings of the Third International Conference on Learning Analytics and Knowledge (LAK), Leuven, Belgium, 8–12 April 2013; pp. 102–106. [Google Scholar] [CrossRef] [Scilit]
  39. Azevedo, R.; Taub, M.; Mudrick, N.V. Using Multi-Channel Trace Data to Infer and Foster Self-Regulated Learning between Humans and Advanced Learning Technologies. In Handbook of Self-Regulation of Learning and Performance, 2nd ed.; Schunk, D.H., Greene, J.A., Eds.; Routledge: New York, NY, USA, 2018; Volume 2, pp. 254–270. ISBN 9781138903197. [Google Scholar]
  40. Emerson, A.; Cloude, E.B.; Azevedo, R.; Lester, J. Multimodal Learning Analytics for Game-based Learning. Br. J. Educ. Technol. 2020, 51, 1505–1526. [Google Scholar] [CrossRef] [Scilit]
  41. Mislevy, R.J.; Steinberg, L.S.; Almond, R.G. On the Structure of Educational Assessments. Meas. Interdiscip. Res. Perspect. 2003, 1, 3–62. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Shute, V.J.; Ventura, M. Stealth Assessment: Measuring and Supporting Learning in Video Games; MIT Press: Cambridge, MA, USA, 2013. [Google Scholar] [CrossRef] [Scilit]
  43. Shute, V.; Lu, X.; Rahimi, S. Stealth Assessment. In The Routledge Encyclopedia of Education; Spector, J.M., Ed.; Taylor & Francis Group: London, UK, 2021; pp. 1–9. [Google Scholar] [CrossRef] [Scilit]
  44. Corbett, A.T.; Anderson, J.R. Knowledge Tracing: Modeling the Acquisition of Procedural Knowledge. User Model. User-Adapt. Interact. 1995, 4, 253–278. [Google Scholar] [CrossRef] [Scilit]
  45. Almond, R.G.; Mislevy, R.J.; Steinberg, L.S.; Yan, D.; Williamson, D.M. Bayesian Networks in Educational Assessment; Springer: Berlin/Heidelberg, Germany, 2015. [Google Scholar] [CrossRef] [Scilit]
  46. Käser, T.; Klingler, S.; Schwing, A.G.; Gross, M. Dynamic Bayesian Networks for Student Modeling. IEEE Trans. Learn. Technol. 2017, 10, 450–462. [Google Scholar] [CrossRef] [Scilit]
  47. Mai, N.T.; Cao, W.; Liu, W. Interpretable Knowledge Tracing via Transformer-Bayesian Hybrid Networks: Learning Temporal Dependencies and Causal Structures in Educational Data. Appl. Sci. 2025, 15, 9605. [Google Scholar] [CrossRef] [Scilit]
  48. Pfeiffer, A.; Bezzina, S.; Wernbacher, T.; Kriglstein, S. The Role of Blockchain Technologies in Digital Assessment. In Proceedings of the EDULEARN20 Proceedings, Online, 6–7 July 2020; pp. 395–403. [Google Scholar] [CrossRef] [Scilit]
  49. Ocheja, P.; Flanagan, B.; Ueda, H.; Ogata, H. Managing Lifelong Learning Records through Blockchain. Res. Pract. Technol. Enhanc. Learn. 2019, 14, 4. [Google Scholar] [CrossRef] [Scilit]
  50. Ocheja, P.; Agbo, F.J.; Oyelere, S.S.; Flanagan, B.; Ogata, H. Blockchain in Education: A Systematic Review and Practical Case Studies. IEEE Access 2022, 10, 99525–99540. [Google Scholar] [CrossRef] [Scilit]
  51. Sifah, E.B.; Xia, H.; Cobblah, C.N.A.; Xia, Q.; Gao, J.; Du, X. BEMPAS: A Decentralized Employee Performance Assessment System Based on Blockchain for Smart City Governance. IEEE Access 2020, 8, 99528–99539. [Google Scholar] [CrossRef] [Scilit]
  52. Widayanti, R.; Rahardja, U.; Oganda, F.P.; Hardini, M.; Devana, V.T. Students Formative Assessment Framework (Faus) Using the Blockchain. In Proceedings of the 2021 3rd International Conference on Cybernetics and Intelligent System (ICORIS), Makassar, Indonesia, 25–26 October 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  53. Hicks, P.J.; Margolis, M.J.; Carraccio, C.L.; Clauser, B.E.; Donnelly, K.; Fromme, H.B.; Gifford, K.A.; Poynter, S.E.; Schumacher, D.J.; Schwartz, A.; et al. A Novel Workplace-Based Assessment for Competency-Based Decisions and Learner Feedback. Med. Teach. 2018, 40, 1143–1150. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Pan, Z.; Biegley, L.; Taylor, A.; Zheng, H. A Systematic Review of Learning Analytics: Incorporated Instructional Interventions on Learning Management Systems. J. Learn. Anal. 2024, 11, 52–72. [Google Scholar] [CrossRef] [Scilit]
  55. Boud, D.; Falchikov, N. Aligning Assessment with Long-term Learning. Assess. Eval. High. Educ. 2006, 31, 399–413. [Google Scholar] [CrossRef] [Scilit]
  56. Jovanović, J.; Gašević, D.; Dawson, S.; Pardo, A.; Mirriahi, N. Learning Analytics to Unveil Learning Strategies in a Flipped Classroom. Internet High. Educ. 2017, 33, 74–85. [Google Scholar] [CrossRef] [Scilit]
  57. Fincham, E.; Gašević, D.; Jovanović, J.; Pardo, A. From Study Tactics to Learning Strategies: An Analytical Method for Extracting Interpretable Representations. IEEE Trans. Learn. Technol. 2018, 12, 59–72. [Google Scholar] [CrossRef]
  58. Musa, M.H.; Salam, S.; Fesol, S.F.A.; Shabarudin, M.S.; Rusdi, J.F.; Norasikin, M.A.; Ahmad, I. Integrating and Retrieving Learning Analytics Data from Heterogeneous Platforms Using Ontology Alignment: Graph-Based Approach. MethodsX 2025, 14, 103092. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  59. Qiang, Z.; Wang, W.; Taylor, K. Agent-OM: Leveraging LLM Agents for Ontology Matching. arXiv 2023, arXiv:2312.00326. [Google Scholar]
  60. Baker, R.S.; Hawn, A. Algorithmic Bias in Education. Int. J. Artif. Intell. Educ. 2022, 32, 1052–1092. [Google Scholar] [CrossRef] [Scilit]
  61. Kane, M. The Argument-Based Approach to Validation. Sch. Psychol. Rev. 2013, 42, 448–457. [Google Scholar] [CrossRef] [Scilit]
  62. Onisko, A.; Druzdzel, M.J. Sensitivity of Bayesian Networks to Noise in Their Parameters. Entropy 2024, 26, 963. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  63. Uglanova, I. Model Criticism of Bayesian Networks in Educational Assessment: A Systematic Review. Pract. Assess. Res. Eval. 2021, 26, 22. [Google Scholar] [CrossRef]
  64. Messick, S. Validity. In Educational Measurement, 3rd ed.; Linn, R.L., Ed.; American Council on Education: New York, NY, USA, 1989; pp. 13–103. ISBN 9780029224007. [Google Scholar]
  65. McNamara, D.S.; Huynh, L. From Prediction to Stewardship: Framing Educational Data Science in the Age of Generative AI. Information 2026, 17, 610. [Google Scholar] [CrossRef] [Scilit]
  66. Selwyn, N. What’s the Problem with Learning Analytics? J. Learn. Anal. 2019, 6, 11–19. [Google Scholar] [CrossRef] [Scilit]
  67. Knox, J.; Williamson, B.; Bayne, S. Machine Behaviourism: Future Visions of ‘Learnification’ and ‘Datafication’ across Humans and Digital Technologies. Learn. Media Technol. 2020, 45, 31–45. [Google Scholar] [CrossRef] [Scilit]
  68. Rouvroy, A. The End(s) of Critique: Data Behaviourism versus Due Process. In Privacy, Due Process and the Computational Turn; Routledge: London, UK, 2013; pp. 143–165. [Google Scholar] [CrossRef] [Scilit]
  69. Haggerty, K.D.; Ericson, R.V. The Surveillant Assemblage. Br. J. Sociol. 2000, 51, 605–622. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  70. Cornet, R.; de Keizer, N. Forty Years of SNOMED: A Literature Review. BMC Med. Inf. Decis. Mak. 2008, 8, S2. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  71. Bodenreider, O. The Unified Medical Language System (UMLS): Integrating Biomedical Terminology. Nucleic Acids Res. 2004, 32, D267–D270. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  72. Wang, L.L.; Bhagavatula, C.; Neumann, M.; Lo, K.; Wilhelm, C.; Ammar, W. Ontology Alignment in the Biomedical Domain Using Entity Definitions and Context. In Proceedings of the BioNLP 2018 Workshop, Melbourne, Australia, 19 July 2018; pp. 47–55. [Google Scholar] [CrossRef] [Scilit]
  73. Xue, X.; Hang, Z.; Tang, Z. Interactive Biomedical Ontology Matching. PLoS ONE 2019, 14, 4. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  74. van Damme, P.; Fernández-Breis, J.T.; Benis, N.; Miñarro-Gimenez, J.A.; De Keizer, N.F.; Cornet, R. Performance Assessment of Ontology Matching Systems for FAIR Data. J. Biomed. Semant. 2022, 13, 19. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  75. Zhang, P.; Kamel Boulos, M.N. Privacy-by-Design Environments for Large-Scale Health Research and Federated Learning from Data. Int. J. Environ. Res. Public Health 2022, 19, 11876. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Gill, W.; Anwar, A.; Gulzar, M.A. FedDebug: Systematic Debugging for Federated Learning Applications. In Proceedings of the 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), Melbourne, Australia, 14–20 May 2023; pp. 512–523. [Google Scholar] [CrossRef] [Scilit]
  77. Mittelstadt, B.D.; Allo, P.; Taddeo, M.; Wachter, S.; Floridi, L. The Ethics of Algorithms: Mapping the Debate. Big Data Soc. 2016, 3, 2. [Google Scholar] [CrossRef] [Scilit]
  78. Diakopoulos, N. Accountability in Algorithmic Decision Making. Commun. ACM 2016, 59, 56–62. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Couldry, N.; Mejias, U.A. The Costs of Connection: How Data Are Colonizing Human Life and Appropriating It for Capitalism; Stanford University Press: Stanford, CA, USA, 2019; ISBN 978-1-5036-0974-7. [Google Scholar]
  80. Yang, K.; Raković, M.; Li, Y.; Guan, Q.; Gašević, D.; Chen, G. Unveiling the Tapestry of Automated Essay Scoring: A Comprehensive Investigation of Accuracy, Fairness, and Generalizability. Proc. AAAI Conf. Artif. Intell. 2024, 38, 22466–22474. [Google Scholar] [CrossRef] [Scilit]
  81. Jadhav, R.; Danve, J.; Shaw, S. Implicit Grading Bias in Large Language Models: How Writing Style Affects Automated Assessment across Math, Programming, and Essay Tasks. arXiv 2026, arXiv:2603.18765. [Google Scholar]
  82. Inoue, A.B. Antiracist Writing Assessment Ecologies: Teaching and Assessing Writing for a Socially Just Future; Parlor Press: Anderson, SC, USA, 2015. [Google Scholar] [CrossRef] [Scilit]
  83. Buckingham Shum, S.; Ferguson, R.; Martinez-Maldonado, R. Human-Centred Learning Analytics. J. Learn. Anal. 2019, 6, 1–9. [Google Scholar] [CrossRef] [Scilit]
  84. Gašević, D.; Tsai, Y.-S.; Dawson, S.; Pardo, A. How Do We Start? An Approach to Learning Analytics Adoption in Higher Education. Int. J. Inf. Learn. Technol. 2019, 36, 342–353. [Google Scholar] [CrossRef] [Scilit]
  85. Mair, P. Item Response Theory. In Modern Psychometrics with R; Mair, P., Ed.; Springer International Publishing: Cham, Switzerland, 2018; pp. 95–159. [Google Scholar] [CrossRef] [Scilit]
  86. Raj, N.S.; Renumol, V.G. A Rule-Based Approach for Adaptive Content Recommendation in a Personalized Learning Environment: An Experimental Analysis. In Proceedings of the 2019 IEEE Tenth International Conference on Technology for Education (T4E), Goa, India, 9–11 December 2019; pp. 138–141. [Google Scholar] [CrossRef] [Scilit]
  87. Cuban, L. Oversold and Underused: Computers in the Classroom; Harvard University Press: Cambridge, MA, USA, 2001. [Google Scholar] [CrossRef] [Scilit]
  88. Bilgin, A.A.; Rowe, A.D.; Clark, L. Academic Workload Implications of Assessing Student Learning in Work-Integrated Learning. Asia-Pac. J. Coop. Educ. 2017, 18, 167–183. [Google Scholar]
  89. Arieli-Attali, M.; Ward, S.; Thomas, J.; Deonovic, B.; von Davier, A.A. The Expanded Evidence-Centered Design (e-ECD) for Learning and Assessment Systems: A Framework for Incorporating Learning Goals and Processes within Assessment Design. Front. Psychol. 2019, 10, 853. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  90. Reichenberg, R. Dynamic Bayesian Networks in Educational Measurement: Reviewing and Advancing the State of the Field. Appl. Meas. Educ. 2018, 31, 335–350. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Six functional layers of the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), illustrating the functional capabilities required to generate, represent, translate, interpret, maintain, and govern learning evidence within continuous assessment ecosystems.
Figure 1. Six functional layers of the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI), illustrating the functional capabilities required to generate, represent, translate, interpret, maintain, and govern learning evidence within continuous assessment ecosystems.
Information 17 00806 g001
Table 1. Governance and interoperability tensions in the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI).
Table 1. Governance and interoperability tensions in the AI-Mediated Continuous Assessment Infrastructure (AIM-CAI).
Governance
Tension
DescriptionExample RisksPotential Governance Responses
Interoperability vs. Local AutonomySystems require cross-platform evidence compatibility while preserving institutional and contextual flexibilitySemantic fragmentation, incompatible learner modelsAI-mediated translation layers, federated standards
Privacy vs. Utility AnalyticsPrivacy-preserving systems reduce data visibility but may weaken inferential precisionLoss of signal for small populations, fairness auditing challengesDifferential privacy, secure enclaves, federated learning
Learner Agency vs. Predictive AutomationContinuous inference may constrain learner opportunity through persistent profilingPath dependency, predictive determinismHuman oversight, learner control, contestability mechanisms
Innovation vs. Governance StabilityEducational technologies evolve faster than standards frameworksFragmented ecosystems, governance lagAdaptive governance architectures, semantic interoperability
Transparency vs. Model ComplexityHighly predictive AI systems may be difficult to interpretBlack-box inference, institutional distrustExplainable AI, audit trails, uncertainty visualization
Distributed Governance vs. AccountabilityFederated systems diffuse responsibility across actorsCoordination failure, unclear liabilityFederated accountability structures, auditability systems
Continuous Support vs. SurveillanceContinuous evidence may support learning, but also normalize monitoringBehavioral conformity, algorithmic governmentalityPermissioned access, bounded monitoring, participatory governance
Table 2. Conceptual Shifts from Episodic Assessment to AI-Mediated Continuous Assessment.
Table 2. Conceptual Shifts from Episodic Assessment to AI-Mediated Continuous Assessment.
Traditional Episodic AssessmentAI-Mediated Continuous Assessment (AIM-CAI)
Learning is documented primarily through examinations, assignments, and final grades.Learning generates evidence across laboratories, writing, collaboration, simulations, workplace experiences, AI-supported learning, portfolios, and other authentic activities.
Individual assessments provide isolated observations of learner performance.Multiple sources of evidence accumulate over time to support evolving interpretations of learner development.
Evidence is interpreted within a single course or institutional context.Evidence may be represented, translated, and interpreted across courses, institutions, workplaces, and lifelong learning experiences while preserving educational meaning and context.
Grades and transcripts summarize achievement at discrete points in time.Learner representations evolve continuously as additional evidence strengthens, refines, or challenges previous competency interpretations.
Feedback is typically associated with individual assignments or courses.Longitudinal evidence supports ongoing feedback, advising, instructional decision-making, learner reflection, and future learning opportunities.
Credentials primarily certify completed learning.Credentials may be supported by transparent longitudinal evidence documenting how competencies developed over time.
Governance focuses primarily on institutional records and assessment events.Governance spans the generation, interpretation, sharing, and use of learning evidence, supporting transparency, learner agency, human oversight, and accountability across educational ecosystems.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

McNamara, D.S.; Hasnine, M.N. AI-Mediated Continuous Assessment Infrastructure (AIM-CAI): Connecting Learning Evidence Across Contexts and Time. Information 2026, 17, 806. https://doi.org/10.3390/info17080806

AMA Style

McNamara DS, Hasnine MN. AI-Mediated Continuous Assessment Infrastructure (AIM-CAI): Connecting Learning Evidence Across Contexts and Time. Information. 2026; 17(8):806. https://doi.org/10.3390/info17080806

Chicago/Turabian Style

McNamara, Danielle S., and Mohammad Nehal Hasnine. 2026. "AI-Mediated Continuous Assessment Infrastructure (AIM-CAI): Connecting Learning Evidence Across Contexts and Time" Information 17, no. 8: 806. https://doi.org/10.3390/info17080806

APA Style

McNamara, D. S., & Hasnine, M. N. (2026). AI-Mediated Continuous Assessment Infrastructure (AIM-CAI): Connecting Learning Evidence Across Contexts and Time. Information, 17(8), 806. https://doi.org/10.3390/info17080806

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop