Previous Article in Journal
Decentering Musical Creativity: Towards Situated Musical Practices in Asian Classrooms
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Preventive and Restorative Academic Integrity in AI-Assisted Higher Education: A Critical Conceptual Synthesis and Integrated Governance Framework

1
Faculty of Agronomy, University of Craiova, 200585 Craiova, Romania
2
Faculty of Mining, University of Petroșani, 332006 Petroșani, Romania
3
Faculty of Geodesy, Technical University of Civil Engineering Bucharest, 020396 Bucharest, Romania
*
Authors to whom correspondence should be addressed.
Educ. Sci. 2026, 16(9), 1365; https://doi.org/10.3390/educsci16091365
Submission received: 30 June 2026 / Revised: 2 August 2026 / Accepted: 21 August 2026 / Published: 24 August 2026

Abstract

Generative artificial intelligence (AI) complicates academic integrity by blurring the boundary between assistance and authorship, enabling cognitive delegation, and introducing algorithmic mediation into assessment and institutional decisions. This article presents a critical conceptual synthesis, not a systematic review. A purposive corpus of 64 academic and policy sources was examined through comparative coding, negative-case analysis, source-to-concept tracing, and normative interpretation. The synthesis distinguishes integrity of academic work, integrity of assessment, and institutional procedural integrity. It proposes an integrated governance framework connecting preventive governance, ethical AI literacy, governed verification, restorative accountability, and proportionate discipline. The framework argues that policy clarity must reach the assessed task; AI literacy supports judgement but cannot neutralize strategic misconduct; automated indicators require corroboration, competent human review, reasons, and appeal; and restorative processes are appropriate only when harm, affected parties, voluntary participation, responsibility, repair, and reintegration are substantively addressed. Its principal contribution is a case-to-governance feedback mechanism through which integrity cases generate institutional learning and trigger revision of policy, assessment design, literacy provision, and technology oversight. Seven testable conceptual propositions, a response-selection pathway, role-specific duties, staged implementation responsibilities, and evaluation indicators are provided. The framework remains a normative and empirically testable proposal rather than a validated intervention or universally effective solution.

1. Introduction

1.1. Continuity and New Challenges

Academic integrity has long been understood as more than detection and punishment. Earlier scholarship already emphasized institutional culture, teaching, assessment design, shared values, prevention, and student learning (Bertram Gallant, 2008; Bretag, 2013, 2016; Bretag et al., 2019; Eaton, 2021; International Center for Academic Integrity, 2021; Macfarlane et al., 2014; Sutherland-Smith, 2008). Generative AI does not invalidate these foundations. It changes the scale and form of mediation by allowing systems to participate in idea generation, drafting, coding, problem solving, feedback, translation, and evaluation.
These capabilities make the provenance of text and reasoning more difficult to reconstruct and permit cognitive delegation that cannot be evaluated through conventional plagiarism categories alone (Cotton et al., 2024; Kasneci et al., 2023; Lim et al., 2024; Lo, 2023; Perkins, 2023; Rudolph et al., 2023; Tlili et al., 2023). AI-enabled grading, proctoring, risk prediction, similarity analysis, and content detection can also affect academic standing through processes that are opaque, biased, or weakly contestable (Eubanks, 2018; European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Floridi & Cowls, 2019; Holmes et al., 2022; Jobin et al., 2019; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024; O’Neil, 2016). The revised analysis therefore rejects both technological determinism and the assumption that every AI-ethics problem is an academic-integrity problem.
The continuity between established academic-integrity scholarship and the present AI-related debate is analytically important because it prevents novelty claims from being built on an inaccurate account of the field. Long before generative systems became widely available, research had already shown that integrity is produced through the interaction of institutional culture, assessment design, teaching practice, student support, and formal regulation. Values-based approaches also emphasized that honesty, trust, fairness, respect, responsibility, and courage must be enacted through educational arrangements rather than stated only in codes of conduct. The current problem is therefore not the disappearance of these principles. It is the need to translate them into environments in which tools can participate in activities that were previously treated as visible evidence of student thinking (Bertram Gallant, 2008; Bretag, 2013, 2016; Bretag et al., 2019; Eaton, 2021; International Center for Academic Integrity, 2021; Macfarlane et al., 2014; Sutherland-Smith, 2008).
Generative AI intensifies ambiguity because it can intervene at several points in the production process. A student may use a system to generate ideas, reorganize an argument, translate a passage, produce code, summarize sources, simulate feedback, or draft a final answer. These functions are not ethically equivalent, and their acceptability depends on the learning outcome, the instructions for the task, the degree of human control, and the transparency of the contribution. A uniform prohibition or permission statement therefore fails to capture the difference between assistance that supports learning and delegation that substitutes for the competence being assessed. The integrity question moves from a simple inquiry about whether a tool was used to a more discriminating inquiry about what the tool did, what the learner retained responsibility for, and what evidence can demonstrate the learner’s own understanding (Cotton et al., 2024; Kasneci et al., 2023; Lim et al., 2024; Lo, 2023; Perkins, 2023; Rudolph et al., 2023; Tlili et al., 2023).
The challenge is also epistemic. Generative outputs can be fluent while containing fabricated references, unsupported claims, hidden bias, or reasoning that cannot be reconstructed from the final text. This means that apparent originality is not necessarily evidence of independent intellectual work, just as surface similarity is not necessarily evidence of misconduct. Institutions must therefore distinguish textual provenance from epistemic responsibility. The relevant obligations include checking claims, validating sources, explaining substantive choices, and accepting responsibility for the final submission. These obligations remain human even when a system performs a large share of the drafting or analytical work. The more extensively a tool participates in a task, the more important it becomes to define what the learner must be able to explain, verify, and defend.
A second change concerns the institutional use of AI. Academic integrity is no longer affected only by systems used by students. Universities may use automated or semi-automated tools for grading, proctoring, similarity analysis, AI-content detection, risk prediction, and case prioritization. These systems can shape the evidence presented in an integrity procedure and can influence decisions with significant academic consequences. Their use introduces questions of validity, bias, privacy, explanation, reviewer competence, and appeal. An institution that demands transparency from students while relying on opaque or weakly validated tools creates an asymmetry that undermines procedural legitimacy. Responsible governance must therefore address both student use of AI and institutional use of algorithmic evidence (Eubanks, 2018; European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Floridi & Cowls, 2019; Holmes et al., 2022; Jobin et al., 2019; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024; O’Neil, 2016).
The continuity-and-change perspective also protects against technological determinism. Generative AI does not inevitably destroy academic integrity, nor does it automatically improve learning. Its effects are mediated by incentives, access, disciplinary norms, staff capability, assessment design, and institutional responses. The same function may support learning in one task and invalidate evidence in another. Translation assistance, for example, may improve equitable participation when language is not the target outcome, yet it may be inappropriate when linguistic competence is being assessed. Similarly, code generation may be used as an object of critique in an advanced course while constituting excessive delegation in an introductory programming task. Ethical judgement therefore requires contextual rules rather than abstract claims about the technology as a whole.
This article treats the present moment as a governance transition rather than a complete break with the past. Established principles remain indispensable, but they must be operationalized through task-level clarity, differentiated responsibility, AI-resilient assessment, governed verification, and proportionate responses. The resulting agenda is demanding because it requires institutions to coordinate policy, pedagogy, evidence, and due process. It is also educationally constructive. By focusing on the conditions under which responsible human–AI collaboration can be demonstrated, institutions can protect the validity of assessment without reducing integrity to surveillance or treating every uncertainty as intentional deception.

1.2. Scope, Gap, and Article Positioning

The article distinguishes three analytically connected levels: integrity of academic work, concerning contribution, authorship, disclosure, attribution, verification, and responsibility for claims; integrity of assessment, concerning valid, fair, accessible, and interpretable evidence of learning; and institutional procedural integrity, concerning transparent, reviewable, and proportionate decisions affecting students. Broader algorithmic governance is included only where it shapes these levels.
A precise scope is necessary because debates about AI in education often combine distinct ethical objects under a single label. Questions about admissions algorithms, student profiling, personalized learning, automated grading, academic authorship, privacy, and institutional procurement all involve responsible AI, but they do not all constitute academic-integrity problems in the same sense. This article therefore uses academic integrity as the central object and includes broader algorithmic governance only when it directly affects academic work, the validity of assessment, or procedures used to make consequential decisions about students. The boundary is not intended to deny the importance of other AI-ethics issues. It is intended to preserve conceptual depth and prevent the framework from becoming an undifferentiated catalogue of educational risks.
The first level, integrity of academic work, concerns contribution, authorship, attribution, disclosure, verification, and responsibility for claims. It asks what the learner or researcher actually did, what was delegated, whether assistance was permitted and disclosed, and whether the final work can be defended. The second level, integrity of assessment, concerns whether a task and its associated judgement provide valid, fair, accessible, and interpretable evidence of the intended learning outcomes. A student may comply with a rule and yet complete a poorly designed task that does not produce meaningful evidence of learning. Conversely, an institution may introduce intrusive verification in an attempt to protect integrity while creating new validity or equity problems. The third level, institutional procedural integrity, concerns the legitimacy of rules and decisions: whether evidence is reliable, reasons are documented, human oversight is substantive, and appeal is genuinely available.
These levels are related through responsibility and evidence, but they are not interchangeable. A problem in assessment design should not automatically be reclassified as student misconduct. An ambiguous task instruction may mitigate individual responsibility without making all conduct acceptable. A statistically generated detector score may prompt inquiry but cannot, by itself, establish authorship or intent. Separating the levels enables a more proportionate analysis because it identifies where a failure occurred and which actor had the capacity and duty to prevent or correct it. It also makes institutional responsibility visible without dissolving student agency.
The literature gap addressed by the article is not an absence of work on any single component. Academic-integrity scholarship offers values, cultural models, and educational approaches; AI-governance frameworks offer principles of fairness, transparency, accountability, privacy, and human agency; AI-literacy research defines technical, critical, and ethical competencies; assessment scholarship addresses validity, feedback, and evaluative judgement; and the restorative-justice literature addresses harm, participation, responsibility, repair, and reintegration (Chan, 2023; Chiu et al., 2024; Dawson et al., 2024; European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Holmes et al., 2022; International Center for Academic Integrity, 2021; Jobin et al., 2019; Karp, 2019b; Laupichler et al., 2022; Long & Magerko, 2020; Macfarlane et al., 2014; National Institute of Standards and Technology, 2023, 2024; Ng et al., 2021; UNESCO, 2021, 2023, 2024a, 2024b; Wachtel, 2016). The unresolved issue is how these components should interact when a concrete AI-related integrity concern arises and how lessons from cases should alter institutional practice.
Existing frameworks frequently remain at one of three levels of abstraction. Some provide broad ethical principles but do not specify how those principles should shape course rules, task instructions, evidence standards, or appeal procedures. Others offer practical guidance for assessment or AI literacy but do not address serious, repeated, or strategic misconduct. Still others invoke restorative language without distinguishing reflection, coaching, or resubmission from a genuine restorative process. The present framework addresses this fragmentation by specifying connecting mechanisms, boundary conditions, actor duties, and response-selection criteria. Its central innovation is relational rather than component-based: each element is familiar, but the model explains how clarity, capability, verification, accountability, and institutional learning should be connected.
Article positioning also requires restraint regarding evidence. The framework is normative and conceptually testable. It does not claim that the proposed relationships have already been validated across institutions, disciplines, or cultural settings. Statements about expected effects are therefore formulated as propositions with counter-cases and illustrative indicators. This distinction is crucial for a Q1 review article because it separates what the literature supports, what the authors infer through synthesis, and what future empirical work must test. It also prevents practical recommendations from being presented as universal solutions.
The resulting position is neither detection-centred nor permissive. Legitimate verification and discipline remain available, but they are placed within a broader system of preventive governance, educational support, due process, and feedback. Restoration is treated as one possible response rather than a substitute for all formal procedures. AI literacy is treated as a capacity rather than a guarantee of conduct. Institutional responsibility is coordinated and differentiated rather than vague or collective. These distinctions define the analytical space in which the article’s contribution can be evaluated.
In Figure 1, the analytical boundary of the review is represented through three connected levels: integrity of academic work, integrity of assessment, and institutional procedural integrity. This structure follows directly from the preceding scope definition and shows that broader AI governance enters the analysis only where it affects assessment, integrity procedures, or student rights.
The literature provides substantial but partly separate guidance on academic-integrity culture, responsible AI, AI literacy, assessment reform, restorative justice, and institutional governance. The unresolved problem is relational: how should prevention, literacy, verification, restoration, and discipline interact, and how can case experience become institutional learning?
In Table 1, six literature strands are compared in terms of their established contribution, persistent limitation, and the specific relationship developed by the present framework. The comparison links the identified gap to the article’s positioning and clarifies that the contribution lies in integration, mechanisms, and institutional feedback rather than in presenting familiar components as individually new.
In Figure 2, the conceptual movement from detection-centred responses to prevention-centred practice and integrated ethical governance is shown. The progression connects the historical discussion to the proposed model while making clear that legitimate verification and discipline are retained within prevention, due process, differentiated responsibility, and institutional learning.

1.3. Aim, Research Questions, and Contribution

The study aims to construct a delimited, normatively explicit, and empirically testable framework for AI-related academic integrity in higher education. Four research questions guide source selection, analysis, framework construction, and discussion:
  • RQ1. How does generative AI modify academic integrity at the levels of academic work, assessment, and institutional procedure?
  • RQ2. Which duties and mechanisms connect preventive governance, ethical AI literacy, governed verification, and restorative accountability?
  • RQ3. Under what conditions should institutions use educational, restorative, disciplinary, or combined responses to AI-related integrity breaches?
  • RQ4. Which principles should remain stable across contexts, and which elements require disciplinary, legal, cultural, or institutional adaptation?
The article contributes a conceptual framework rather than an effectiveness claim. Its originality lies in connecting differentiated responsibility to proportional response selection, embedding restorative justice within due process safeguards, and making case-to-governance feedback a central institutional mechanism.
The aim of the study is to construct a framework that is sufficiently delimited to guide institutional reasoning and sufficiently explicit to be examined empirically. This requires more than assembling recommendations. A governance framework must identify the object of regulation, the actors who hold duties, the mechanisms through which actions are expected to produce outcomes, the conditions under which those mechanisms may fail, and the evidence needed to evaluate implementation. The four research questions therefore function as an analytical sequence. They move from conceptual change, to connecting duties and mechanisms, to response selection, and finally to contextual adaptation.
RQ1 addresses how generative AI modifies academic integrity at the levels of work, assessment, and procedure. The wording deliberately avoids asking whether AI is simply beneficial or harmful. Instead, it directs attention to changes in authorship, cognitive delegation, traceability, evidence of learning, and algorithmic mediation. This allows continuity with established integrity scholarship to be retained while identifying the features that require revision. It also prevents the analysis from attributing every institutional difficulty to student misconduct or every student use of AI to unacceptable outsourcing.
RQ2 concerns the duties and mechanisms that connect preventive governance, ethical AI literacy, governed verification, and restorative accountability. The question is designed to overcome the common practice of presenting policy, literacy, assessment, and response as parallel recommendations. A mechanistic account asks how task-level clarity reduces ambiguity, how literacy supports interpretation and disclosure, how verification can be conducted without converting probabilistic suspicion into proof, and how case analysis can return evidence to policy and assessment design. It also requires role-specific attribution because a mechanism cannot operate if no actor has the authority, resources, or responsibility to implement it.
RQ3 addresses proportional response selection. AI-related concerns vary in intention, impact, rule clarity, degree of delegation, evidence quality, repetition, and potential for repair. A single response is therefore neither fair nor educationally defensible. The research question creates space for educational correction where ambiguity or weak judgement predominates, restorative processes where harm and voluntary participation can be meaningfully addressed, disciplinary procedures where deliberate or serious conduct requires formal adjudication, and combined pathways where several purposes must be served. The purpose is not to weaken standards, but to connect consequences to established facts, learning outcomes, due process, and the nature of the harm.
RQ4 distinguishes stable principles from adaptable implementation. Human responsibility, validity, fairness, reason-giving, contestability, privacy protection, and non-coercive restorative participation are treated as minimum commitments. However, the form of disclosure, the authority of committees, acceptable evidence, record-retention rules, facilitation practices, sanctions, and available resources differ across jurisdictions and disciplines. The question therefore protects the framework from both universalism and relativism. It supports contextual adaptation without allowing local variation to remove basic procedural and ethical safeguards.
The contribution is organized at conceptual, procedural, and institutional levels. Conceptually, the article distinguishes three levels of integrity and defines epistemic responsibility in human–AI co-production. Procedurally, it links differentiated responsibility to proportional response selection and places automated indicators within corroborated, reviewable evidence chains. Institutionally, it introduces a case-to-governance feedback mechanism through which recurring ambiguity, inequitable access, assessment weaknesses, and procedural failures can trigger documented change. These contributions are expressed through seven provisional propositions rather than claims of proven effectiveness.
A further contribution lies in the framework’s evidential discipline. Literature-derived principles, policy-derived obligations, and author-generated proposals are not treated as equivalent. The methodological design documents how sources were selected and how concepts were converted into mechanisms. Counter-positions are used to test the boundaries of the preferred interpretation. Practical recommendations are linked to indicators and cautions. This structure supports scholarly scrutiny because readers can identify which parts of the framework are inherited, adapted, or newly proposed.
The intended outcome is a model that can support institutional deliberation while also generating a research agenda. Universities can use the framework to review policies, assessment practices, literacy provision, case procedures, and technology procurement. Researchers can use the propositions to design studies of clarity, conduct, validity, appeal, restoration, equity, and institutional learning. The article therefore contributes neither a simple checklist nor a claim of final resolution. It offers a coherent set of relationships that can be challenged, refined, and tested across settings.
The research questions also provide the organizing logic for interpretation. Each analytical section identifies the question it informs, while the discussion returns to all four questions and separates the literature-derived findings from author-generated inference. This structure reduces the risk that the review becomes a sequence of themes without an explicit answer. It also creates a transparent basis for judging whether the conclusions remain within the evidence.
The framework’s practical usefulness should be assessed against the same logic. A university should be able to identify which part of the model addresses a particular concern, which actors hold duties, what evidence is needed, and what remains uncertain. If the model cannot support these decisions or generates unmanageable complexity, its propositions should be revised. The aim is therefore both explanatory and evaluative.

2. Design and Method

2.1. Rationale for a Critical Conceptual Synthesis

Critical conceptual synthesis was selected because the research questions concern conceptual boundaries, relationships among theoretical traditions, allocation of duties, and criteria for institutional action. The objective is not effect estimation, exhaustive mapping, or study-quality aggregation. A scoping review could map breadth, and a policy analysis could compare formal rules, but neither would by itself justify the normative relationships required for an integrated governance model.
The choice of review design follows from the nature of the research problem. The article does not seek to estimate the effect of a single intervention, calculate prevalence, or aggregate comparable outcomes. It asks how concepts drawn from several traditions can be delimited, related, and converted into criteria for institutional action. This is a theory-building and normative task. A critical conceptual synthesis is appropriate because it permits close comparison of definitions, assumptions, duties, mechanisms, and tensions while retaining the ability to formulate propositions that can later be tested.
A systematic review would require a protocol, exhaustive or highly reproducible search procedures, documented duplicate removal, study-level quality appraisal, and a transparent screening record. Those features are valuable when a review question concerns a defined evidence base and comparable empirical studies. They were not fully available for the present project, and it would be methodologically misleading to reconstruct them retrospectively. The revised article therefore does not use systematic-review language to imply precision that the underlying audit trail cannot support. The structured search improves transparency, but the corpus remains purposive and conceptually selected.
A scoping review would be useful for mapping the volume, distribution, and characteristics of scholarship on AI and academic integrity. However, mapping alone would not answer the normative questions at the centre of this study. Knowing that the literature contains themes of policy, literacy, assessment, or restoration does not establish how those themes should interact, when one should constrain another, or who is responsible for implementation. The present design moves beyond description by examining conceptual tensions and constructing conditional relationships. This is especially important where values such as privacy, transparency, prevention, trust, consistency, and restoration can conflict.
An integrative review could combine theoretical and empirical sources and might appear close to the present design. The distinction lies in the primary analytical purpose. The current study does not aim to synthesize empirical findings into a comprehensive account of what is known. It uses heterogeneous sources according to their different functions: empirical studies illuminate patterns and risks; conceptual scholarship defines constructs; policy documents state normative expectations; and practice frameworks describe procedures. The synthesis then asks what institutional architecture can be justified when these sources are considered together. Heterogeneity is therefore not treated as if all sources provided the same level or type of evidence.
Policy analysis would also be insufficient on its own. Formal guidance from UNESCO, OECD, the European Union, NIST, and universities provides important principles and operational expectations, but policy documents often remain at a system level. They may not address the validity of specific assessments, the educational meaning of authorship, the limits of AI literacy, or the substantive requirements of restorative justice. Conversely, academic scholarship may identify ethical concerns without specifying governance responsibilities. A critical conceptual synthesis permits these bodies of work to be compared without collapsing policy authority into empirical evidence (European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Holmes et al., 2022; National Institute of Standards and Technology, 2023, 2024; OECD, 2023, 2026; UNESCO, 2021, 2023).
The word critical refers to more than disagreement with detection-centred approaches. It denotes attention to power, surveillance, information asymmetry, unequal access, disability and language effects, institutional capacity, and vendor influence. The analysis examines who can define acceptable use, who must disclose, who controls the evidence, who has access to appeal, and whose burdens increase when new safeguards are introduced. These questions are necessary because a formally neutral rule may produce unequal consequences, and a nominally human-reviewed decision may remain dominated by opaque technology.
The conceptual component also requires reflexivity. Prevention, AI literacy, and restoration were initial sensitising domains, but they were not accepted as sufficient merely because they framed the original problem. The analysis searched for counter-cases in which literacy did not prevent misconduct, restoration was inappropriate, verification was legitimate, or assessment redesign created new inequities. This process led to additional mechanisms, including epistemic responsibility, proportional attribution, governed verification, and institutional feedback. The final framework therefore differs from a simple restatement of the initial pillars.
The design supports a specific kind of contribution: a transparent, arguable, and empirically testable normative proposal. It cannot establish causal effectiveness, and it does not claim universal transferability. Its quality depends instead on conceptual clarity, traceability to relevant sources, engagement with counter-positions, internal coherence, explicit boundary conditions, and usefulness for future inquiry. These criteria are made visible throughout the methodology, appendices, propositions, and implementation matrices.
Conceptual synthesis also permits the analysis of category boundaries. Terms such as authorship, originality, transparency, literacy, harm, and restoration are used differently across studies. A review that merely aggregates occurrences would obscure these differences. Critical comparison allows the study to determine when concepts are compatible, when they require translation, and when a tension should remain visible rather than be resolved artificially.
The design remains accountable to evidence because conceptual freedom is not unrestricted. Interpretations must be traceable to sources, counter-positions must be considered, and propositions must identify empirical tests. This combination of normative reasoning and evidential discipline is particularly suited to a rapidly changing field in which definitive intervention studies remain limited but institutional decisions cannot be postponed.

2.2. Normative Foundations, Critical Lens, and Scope Conditions

The framework draws on four normative traditions: academic-integrity values of honesty, trust, fairness, respect, responsibility, and courage (International Center for Academic Integrity, 2021); assessment principles of validity, transparency, evaluative judgement, accessibility, and social justice (Boud & Falchikov, 2007; Carless & Boud, 2018; Dawson et al., 2024; McArthur, 2016; Nicol & Macfarlane-Dick, 2006; Tai et al., 2018); responsible-AI principles of human agency, accountability, explainability, privacy, risk management, and contestability (European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Floridi & Cowls, 2019; Holmes et al., 2022; Jobin et al., 2019; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024; UNESCO, 2021); and restorative-justice requirements concerning harm, participation, responsibility, repair, and reintegration (Bussu & Karp, 2026; Evans & Vaandering, 2022; Karp, 2019a, 2019b; Wachtel, 2016; Zehr, 2015).
The critical lens examines surveillance, information and appeal asymmetries, unequal access to AI, linguistic and disability-related disadvantage, institutional capacity, and vendor influence. Stable minimum principles are distinguished from implementation details that require legal, cultural, disciplinary, and institutional adaptation.
The framework uses a plural normative foundation because no single ethical tradition adequately captures the problem. Academic-integrity values provide the moral vocabulary of honesty, trust, fairness, respect, responsibility, and courage. Assessment scholarship contributes validity, transparency, evaluative judgement, accessibility, and social justice. Responsible-AI frameworks add human agency, accountability, privacy, explainability, risk management, and contestability. Restorative justice contributes attention to harm, affected parties, participation, responsibility, repair, and reintegration (Dawson et al., 2024; European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Floridi & Cowls, 2019; Holmes et al., 2022; International Center for Academic Integrity, 2021; Jobin et al., 2019; Karp, 2019a, 2019b; McArthur, 2016; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024; UNESCO, 2021; Wachtel, 2016; Zehr, 2015). These traditions overlap, but their priorities are not identical.
Pluralism is necessary because AI-related integrity decisions often involve competing legitimate claims. Transparency can support accountability while exposing private data or encouraging disproportionate record collection. Verification can protect the validity of assessment while undermining trust or creating accessibility burdens. Consistent sanctions can support fairness while ignoring ambiguity or institutional contribution. Restorative participation can support repair while becoming coercive if refusal carries an adverse inference. The framework therefore does not present ethical principles as a list that can always be maximized simultaneously. It treats them as decision considerations whose relationship must be explained.
Proportionality provides one organizing principle. A safeguard or response should be connected to the educational purpose, the seriousness and reliability of the concern, the stakes of the decision, and the availability of less intrusive alternatives. Requiring a detailed prompt log may be justified in a high-stakes assessment where extensive AI use is permitted, and contribution must be demonstrated, but it may be excessive for routine spelling assistance. An oral clarification may support fact-finding, but it should not become an unstructured interrogation that disadvantages students with anxiety, disability, or limited proficiency in the language of instruction. Proportionality requires institutions to justify both action and burden.
Procedural justice provides a second organizing principle. Students should know the rules that apply, understand the evidence used, have a meaningful opportunity to respond, receive reasons for consequential decisions, and have access to independent review. Human oversight is substantive only when reviewers have competence, time, authority, and relevant documentation. A person who merely confirms an automated output does not provide genuine human judgement. These requirements apply whether the institution uses AI-content detection, proctoring, similarity analysis, predictive analytics, or other automated indicators.
The critical lens focuses on how formal arrangements distribute power and risk. Students may be required to disclose detailed use while providers withhold system documentation. Staff may be held accountable for decisions without sufficient time or training. Institutions may procure tools whose model changes, error rates, or data practices cannot be independently audited. Students with fewer resources may rely on free tools with greater privacy risks, while those with better access can use more capable systems. Language models may interact differently with non-native writing, and oral verification may create unequal burdens. These asymmetries must be treated as integrity concerns when they affect evidence or procedure.
The framework also distinguishes stable principles from scope conditions. Stable principles include human answerability for consequential decisions, intelligible rules, valid evidence, privacy protection, non-discrimination, reason-giving, contestability, and voluntary restorative participation. Scope conditions determine how those principles are implemented. Professional accreditation may restrict permissible AI functions in one programme more than another. Data-protection law may limit record collection. Cultural expectations may shape dialogue, authority, and repair. Institutional capacity may influence timelines and facilitation models. Adaptation is therefore required, but it should not eliminate the stable commitments.
A further scope condition concerns educational stage. Expectations for novice learners should reflect their prior guidance, opportunities to practise, and capacity to interpret complex rules. The same undisclosed AI function may warrant different responses in an introductory course and a final professional assessment because the learning outcomes, prior instruction, and consequences differ. This does not create arbitrary leniency. It connects responsibility to the conditions under which informed and intentional action was possible.
The normative foundation is intentionally non-consequentialist in one respect: procedural legitimacy is not justified only by whether misconduct declines. An institution could reduce reported cases through intrusive surveillance, inaccessible appeals, or discouraged reporting, yet such an outcome would not demonstrate integrity. Evaluation must therefore include process, fairness, access, reasons, and institutional learning, not only case counts. This position aligns the framework with education as a relational and rights-sensitive practice rather than a narrow system of behavioural control.
The plural foundation also clarifies the status of educational purpose. Education is not used to excuse misconduct or replace rights with informal discretion. It provides a criterion for evaluating whether policy, assessment, verification, and response support valid learning and responsible participation. A measure that increases control but weakens learning, fairness, or contestability requires stronger justification.
Scope conditions should be recorded in implementation documents. Institutions should state the legal, disciplinary, cultural, accessibility, and resource assumptions that shape a policy or procedure. This transparency enables comparison and prevents a locally adapted practice from being presented as universally appropriate. It also supports later review when conditions change.

2.3. Source Identification, Eligibility, and Corpus

The purposive corpus contains 64 academic and policy publications. The primary contemporary window was 2016–July 2026; foundational pre-2016 sources were retained where they define academic integrity, assessment, plagiarism, algorithmic ethics, or restorative justice. Sources were identified iteratively through Scopus, Web of Science, Google Scholar, publisher platforms, and official UNESCO, OECD, European Union, and NIST repositories.
The source-identification strategy was designed to support conceptual coverage rather than statistical representation. The corpus needed to include the principal traditions required by the research questions: academic integrity, assessment and assessment security, generative AI in education, AI literacy, responsible-AI governance, algorithmic accountability, and restorative justice. Searches were therefore iterative. Findings in one domain generated additional terms and citation paths in another. For example, the literature on AI-content detection led to work on validity and due process, while the literature on restorative practice led to questions about consent, facilitation, records, and power imbalance.
Scopus and Web of Science were used to identify peer-reviewed scholarship across education, ethics, technology, and higher-education research. Google Scholar supported broader citation tracing and the identification of books, chapters, reports, and emerging work that may not be consistently indexed. Publisher platforms were used to verify bibliographic details and access complete texts where available. Official repositories were used for UNESCO, OECD, European Union, and NIST documents because policy status and version control are important when the source informs governance obligations rather than only conceptual discussion (European Union, 2024; National Institute of Standards and Technology, 2023, 2024; OECD, 2023, 2026; UNESCO, 2021, 2023, 2024a, 2024b; UNESCO International Institute for Higher Education in Latin America and the Caribbean, 2023).
Eligibility was based on conceptual function. A source was retained when it directly contributed to at least one research question by defining a construct, documenting a relevant practice or risk, offering an ethical or policy principle, identifying a counter-position, or describing implementation. Sources that addressed general AI adoption without a meaningful link to academic work, assessment, integrity procedure, literacy, or restoration were excluded. This decision protected the analytical boundary of the article. It also prevented the corpus from being dominated by broad discussions of AI in education that would not help explain the proposed framework.
The primary temporal window, 2016 to July 2026, captures the rapid development of AI-in-education scholarship and contemporary governance. Foundational pre-2016 sources were retained when they defined concepts that remain necessary, including academic-integrity values, plagiarism, assessment, evaluative judgement, algorithmic ethics, and restorative justice. The temporal design therefore distinguishes foundational relevance from technological recency. A recent publication was not automatically preferred if an earlier source offered the clearer theoretical foundation, and an older source was not retained merely because it was highly cited.
The English-language restriction reflects the interpretive demands of conceptual synthesis rather than a claim that English-language scholarship is superior. The study required close analysis of definitions, normative distinctions, and procedural conditions. The team could not guarantee equivalent precision across additional languages within the available process. This choice introduces geographic and epistemic bias, particularly because concepts of authorship, authority, trust, disclosure, and repair vary across cultures. The limitation is therefore addressed in the conclusions and research agenda rather than framed as a quality filter.
Citation impact was not used as a threshold. High-impact sources can identify established debates, but such a criterion would systematically disadvantage new scholarship and work from less visible regions or journals. Quality was instead considered through transparent provenance, relevance, clarity of contribution, and the role assigned to the source. Empirical studies were not treated as if they established universal causal effects. Policy documents were not treated as empirical evidence. Conceptual works were used to define or relate ideas. This role-sensitive interpretation is documented in the retained-source matrix.
The corpus of 64 sources is purposive and finite. The stopping rule was conceptual sufficiency: searching continued until the major categories, counter-positions, and relationships required by the research questions were represented and additional sources primarily repeated already captured functions. This is not the same as thematic saturation in an empirical qualitative study, nor does it demonstrate exhaustive retrieval. It is a pragmatic theory-building criterion. The audit trail allows readers to see how each retained source contributes, but it does not support calculation of inclusion probabilities or publication bias.
Source selection was also linked to negative-case analysis. Literature was sought that challenged preferred assumptions, including evidence that literacy does not guarantee ethical conduct, that process-rich assessment can increase burden or inequity, that verification and sanctions may be legitimate, that human review can reproduce automation bias, and that restorative approaches can be coercive or unsuitable. Including such sources reduces the risk that the corpus simply confirms the initial three-pillar structure. It also strengthens the empirical testability of the final propositions because each is accompanied by conditions under which the expected relationship may not hold.
In Table 2, the methodological architecture is presented by pairing each operational decision with a transparency safeguard. This organization connects the rationale for a critical conceptual synthesis to the practical measures used to distinguish purposive theory building from systematic-review claims.
In Table 3, the principal concept domains, illustrative search combinations, and eligibility functions are set out. The table therefore links the source-identification strategy to the research questions and shows how each search domain contributed a distinct analytical function to the corpus.

2.4. Analytical Procedure and Audit Trail

Analysis proceeded in six stages. Research questions and scope were defined; sensitising concepts were extracted from the established literature; sources were compared using open and focused coding; negative cases and counter-positions were examined; concepts were reorganised into mechanisms; and seven provisional propositions were formulated and checked against boundary cases. The initial pillars were therefore inputs to inquiry, not predetermined findings.
The analytical procedure was organized to make the movement from source material to framework propositions visible. The first stage established the research questions and three-level scope. This step prevented later coding from expanding into every ethical issue associated with AI in education. It also provided an initial decision rule: a concept was relevant when it clarified academic work, assessment validity, or institutional procedure. The scope could be refined during analysis, but any change had to be justified in relation to the research questions.
The second stage identified sensitising concepts from the established literature. These included authorship, contribution, attribution, academic-integrity values, assessment validity, evaluative judgement, AI literacy, human agency, transparency, privacy, contestability, harm, participation, repair, and institutional responsibility. Sensitising concepts are starting points rather than findings. They direct attention without determining the final categories. Treating them in this way was important because prevention, literacy, and restoration were present in the initial problem framing and could otherwise have been reproduced circularly as the inevitable outcome.
The third stage used open and focused coding. Open coding identified recurrent problems and relationships such as ambiguity, cognitive delegation, concealment, advantage, verification, process evidence, accessibility, surveillance, opacity, appeal, workload, provider influence, facilitation, coercion, reintegration, and policy feedback. Focused coding then compared how these ideas were defined and used across source traditions. For example, transparency in AI governance often concerns system explanation and data practices, whereas transparency in academic integrity may concern task rules and disclosure of assistance. The synthesis retained these distinctions rather than treating identical terminology as having identical meaning.
The fourth stage examined counter-positions and boundary cases. This stage was not an optional balance exercise; it tested the emerging framework. Claims that prevention should replace detection were challenged by cases requiring legitimate verification. Claims that literacy supports integrity were challenged by deliberate strategic misconduct. Claims that oral or process-based assessment increases validity were challenged by accessibility, language, workload, and reliability concerns. Claims that restoration is preferable were challenged by disputed facts, coercion, serious harm, professional risk, and repeated conduct. These cases shaped the response matrix and the caution attached to each proposition.
The fifth stage reorganized categories into mechanisms. A thematic list would show that policy, literacy, assessment, verification, and restoration are important, but it would not explain their relationship. Mechanism mapping asked what change is expected, through which process, for whom, and under what conditions. Task-level clarity was connected to reduced ambiguity when students can understand the rule and when guidance is accessible. AI literacy was connected to better judgement and verification but not to guaranteed compliance. Case feedback was connected to institutional learning only when patterns led to documented and implemented change.
The sixth stage formulated seven provisional propositions. Each proposition includes a mechanism, a boundary condition or counter-case, and an illustrative indicator. This structure is deliberately more demanding than a general recommendation. It makes the framework open to empirical qualification or rejection. If task-level guidance increases but ambiguity-related concerns do not decline, researchers must examine comprehension, enforcement, incentives, or measurement. If restorative participation appears positive but repair, fairness, or recurrence do not improve, the assumed mechanism may be incomplete. The propositions therefore create a bridge between conceptual synthesis and future research.
The audit trail consists of the source matrix, coding memos, framework comparison, negative-case log, proposition matrix, and reviewer-response synchronization record (Appendix G). The purpose is not to claim exact reproducibility in the manner of a registered systematic review. Conceptual interpretation necessarily involves judgement. The purpose is to allow readers to reconstruct the principal decisions, see which sources support which domains, and identify where the authors have moved from literature-derived claims to integrative inference or normative proposal.
No stage is described as validation. The framework has not undergone expert consensus, stakeholder deliberation, institutional case testing, or intervention evaluation. Internal coherence and boundary testing improve conceptual credibility, but they do not establish practical effectiveness. The methodology therefore supports transparent theory building. It also defines the next steps required for stronger claims: empirical studies, cross-context comparison, equity audits, implementation research, and deliberative assessment with students, educators, integrity practitioners, and policy specialists.
Analytical memos were used to distinguish descriptive observations, normative principles, and integrative decisions. This distinction prevents a policy statement from being treated as empirical evidence and prevents an author proposal from appearing to be a consensus finding. The same distinction informs the language used in the results and discussion, where verbs such as identifies, interprets, argues, and proposes signal different evidential status.
The audit trail is also relevant to reviewer and reader trust. Because conceptual synthesis cannot be reproduced mechanically, transparency about judgement is especially important. Readers should be able to see why a concept was included, how a counter-case modified the model, and where uncertainty remains. This form of traceability supports critical engagement rather than asking readers to accept the framework as a completed solution. Publication-oriented accessibility and visual-quality checks are documented in Appendix C, while Appendix E inventories the analytical roles of all 12 figures and 18 main tables.
In Figure 3, the six stages of the critical conceptual synthesis and the associated audit trail are shown as a connected analytical process. The sequence explains how the study moved from scope definition and source construction to counter-case testing, mechanism development, provisional propositions, and explicit empirical limits.
In Figure 4, the 64-source corpus is compared across its principal analytical domains. The graph supports methodological transparency by showing the relative contribution of AI literacy and education, AI governance and ethics, academic integrity, assessment, restorative justice, and other cross-cutting sources without implying a quality ranking.

2.5. Counter-Positions, Boundary Cases, and Methodological Limitations

The synthesis deliberately includes strategic misconduct, legitimate uses of verification and sanctions, limitations of AI literacy, accessibility and workload costs of process-rich assessment, automation bias in human review, and critiques of restorative practice where voluntariness, disputed facts, power imbalance, or serious harm make restoration inappropriate.
The design remains limited by purposive selection, English-language restriction, heterogeneous source types, interpretive coding, possible confirmation bias, and rapid technological and regulatory change. The framework has not been empirically tested and cannot establish causal effectiveness. These constraints convert recommendations into propositions and evaluation questions rather than validated solutions.
Counter-positions are integral to the analysis because AI-related integrity debates are vulnerable to polarized claims. One position treats detection and sanction as obsolete, while another treats technological control as the primary solution. The framework rejects both extremes. Verification remains legitimate when it is targeted, valid, transparent, corroborated, proportionate, and reviewable. Sanctions remain legitimate where conduct is deliberate, repeated, serious, unsafe, or incompatible with voluntary restoration. The problem is not the existence of verification or discipline; it is their use without evidential and procedural safeguards (Bretag et al., 2019; Cotton et al., 2024; Dawson, 2021; Dawson et al., 2024; European Union, 2024; National Institute of Standards and Technology, 2023, 2024).
A second counter-position concerns AI literacy. The educational literature often presents literacy as a response to misuse, but knowledge cannot be assumed to produce conduct. A student may understand disclosure rules and still conceal use because of competition, workload, pressure, or anticipated advantage. An educator may understand system limitations yet rely on automated indicators because of time constraints or institutional expectations. Literacy is therefore necessary for informed judgement but insufficient without clear rules, valid assessment, incentives, authority, resources, and accountability. This limitation prevents the framework from shifting institutional responsibility onto individuals.
A third boundary case concerns assessment redesign. Process portfolios, oral defences, contextual projects, reflective commentary, and in-class components can improve the interpretability of learning evidence, but none is inherently AI-proof. Each can introduce new risks. Portfolios may be fabricated and burdensome. Oral assessment may amplify anxiety, disability, language effects, or inter-rater variation. Context-rich projects may depend on unequal access to resources. Supervised components may increase surveillance and logistics. The framework therefore evaluates combinations of methods and safeguards rather than ranking formats as universally secure (Boud & Falchikov, 2007; Corbin et al., 2025; Dawson, 2021; Dawson et al., 2024; McArthur, 2016; Xia et al., 2024).
Restorative justice presents another important boundary. Reflection, retraining, or resubmission may be educational, but they are not restorative without attention to harm, affected parties, responsibility, voluntary participation, repair, and reintegration. Even a well-designed restorative process may be unsuitable when facts are seriously disputed, participation is unsafe, power imbalance cannot be managed, professional or legal obligations require formal adjudication, or the conduct is repeated and strategic. Refusal to participate must not be treated as evidence of guilt or an aggravating factor. These limitations protect restoration from becoming either symbolic or coercive (Bussu & Karp, 2026; Evans & Vaandering, 2022; Karp, 2019a, 2019b; Wachtel, 2016; Zehr, 2015).
Human oversight is also subject to a counter-case. The presence of a person in a decision chain is often treated as sufficient protection against automated harm. In practice, reviewers may lack training, time, authority, or access to documentation. They may defer to a system because of automation bias or institutional pressure. Substantive human oversight therefore requires more than an override button. It requires competence, evidential independence, written reasons, accountability, and an appeal structure capable of examining both the human and technological components of the decision (Eubanks, 2018; European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Holmes et al., 2022; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024; O’Neil, 2016).
Methodologically, the purposive corpus limits claims about comprehensiveness. English-language restriction may underrepresent regional traditions and non-Western conceptions of authorship, trust, authority, and repair. The combination of books, empirical studies, conceptual articles, policy frameworks, and practice guidance creates evidential heterogeneity. Interpretive coding introduces the possibility of confirmation bias. Rapid technological and regulatory change may make task-level recommendations obsolete. These limitations are not resolved by a larger number of references; they require transparent interpretation and future cross-context research.
The conceptual design also limits causal inference. The article cannot demonstrate that policy clarity reduces breaches, that literacy improves conduct, that AI-resilient assessment increases validity, that governed verification reduces procedural error, or that restorative processes improve reintegration. These are propositions. Their evaluation requires appropriate designs, including longitudinal studies, comparative assessment research, case audits, appeal analysis, vignette experiments, implementation studies, and participant research. Indicators must distinguish implementation from outcome and reported cases from actual conduct.
A final limitation concerns institutional capacity. The framework assumes that universities can provide staff development, accessibility support, trained facilitators, data-protection review, procurement expertise, and independent appeals. Many institutions face resource constraints. A conceptually desirable procedure may fail if workloads, authority, or infrastructure are absent. Implementation must therefore begin with minimum safeguards and realistic sequencing. Resource limitation may explain adaptation, but it should not justify opaque evidence, automated proof, inaccessible appeal, or coerced participation.
Another limitation concerns the absence of direct stakeholder participation in framework construction. Students, educators, integrity practitioners, accessibility specialists, and policymakers may identify practical or cultural issues not visible in the published literature. Future deliberative studies should therefore examine the framework’s legitimacy, clarity, and feasibility. Such work may revise terminology, duties, or response criteria.
The limitations also influence presentation. Claims are calibrated, figures are labelled as conceptual or heuristic, and recommendations include cautions. This discipline is not a weakness of the article. It is part of the contribution because it shows how institutions can act under uncertainty without presenting normative proposals as validated interventions.

3. Conceptual Boundaries and Epistemic Responsibility

3.1. Three Levels of Integrity

Integrity of academic work concerns what the learner or researcher contributed and can defend. Integrity of assessment concerns whether the task and judgement produce valid and fair evidence of the intended learning. Institutional procedural integrity concerns the rules, evidence, reasons, human oversight, and appeal used in consequential decisions. The levels interact but are not interchangeable.
The three-level model provides a diagnostic structure for locating a concern before assigning responsibility or selecting a response. Integrity of academic work focuses on the submission and the process through which it was produced. The key questions are what the learner contributed, which functions were delegated, whether assistance was permitted and disclosed, whether sources and claims were verified, and whether the learner can explain the reasoning. This level preserves the importance of individual agency while recognizing that contemporary academic work may involve multiple human and technological contributors.
Integrity of assessment shifts attention from conduct to the quality of the task and judgement. An assessment is not valid merely because it is difficult to complete with AI. It must generate evidence of the intended learning outcome, support reliable interpretation, and avoid unjustified disadvantage. A task may be highly controlled yet educationally weak, or it may permit AI while still providing strong evidence through process documentation, critique, demonstration, or complementary assessment. The central issue is therefore not resistance to technology but the credibility and fairness of the inference made from student performance (Boud & Falchikov, 2007; Carless & Boud, 2018; Dawson, 2021; Dawson et al., 2024; McArthur, 2016; Nicol & Macfarlane-Dick, 2006; Tai et al., 2018).
Institutional procedural integrity concerns how rules and decisions are constructed and applied. It includes policy clarity, data governance, evidence standards, substantive human review, documentation, reason-giving, timeliness, consistency, privacy, and appeal. A decision can be substantively correct yet procedurally deficient if the student could not understand the allegation, access the evidence, respond meaningfully, or seek independent review. Procedural integrity also requires institutions to examine the tools and vendors that shape evidence, rather than treating technology as a neutral background.
The levels interact through causal and normative pathways. Ambiguous task instructions can contribute to inappropriate use at the level of academic work. A poorly aligned assessment can create incentives for delegation or make contribution difficult to interpret. An unreliable detector can distort procedural evidence. Conversely, a well-designed task can support transparent use, and a fair procedure can distinguish misunderstanding from deliberate concealment. The model therefore encourages decision-makers to ask whether the concern originates primarily in conduct, assessment, institutional procedure, or a combination.
Separating the levels avoids two common errors. The first is over-individualization, in which every concern is attributed to the student even when policy, teaching, assessment, or technology created avoidable ambiguity. The second is over-distribution, in which institutional contribution is used to erase individual agency. The framework instead supports coordinated and differentiated responsibility. A vague rule may mitigate responsibility, but deliberate advantage-seeking can remain relevant. An institutional failure may require policy revision even when a student also receives a proportionate consequence.
The three-level model is also useful for evidence selection. At the work level, evidence may include drafts, an AI-use statement, prompt excerpts, source checks, or an explanation of reasoning. At the assessment level, evidence may include learning-outcome mapping, moderation records, accessibility review, and convergent performance across tasks. At the procedural level, evidence may include tool validation, reviewer notes, reasons, appeal outcomes, and differential-impact audits. Matching evidence to the level of inquiry prevents a single indicator from being asked to prove more than it can support.
The model has implications for policy writing. Institutional regulations should state common rights and responsibilities, but programmes and courses must translate them into disciplinary and pedagogical terms. Task instructions should identify permitted functions, disclosure expectations, and required evidence. Case procedures should state how concerns will be established and reviewed. When these layers are coherent, students can connect general values to concrete action. When they conflict, institutional ambiguity becomes part of the proportionality analysis.
Finally, the model clarifies the limits of the article. Broader questions of admissions, resource allocation, or general learning analytics are included only when they affect assessment evidence, integrity procedures, or academic standing. This boundary protects analytical coherence while acknowledging that procedural integrity may connect academic-integrity governance to wider responsible-AI obligations (European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Holmes et al., 2022; National Institute of Standards and Technology, 2023, 2024; UNESCO, 2021).
The distinction also supports organisational learning because each level points to a different corrective response. A failure at the academic-work level may require clarification, education, repair, or discipline. A failure at the assessment level may require redesign, moderation, or accessibility adjustment. A failure at the procedural level may require evidence review, reviewer training, policy amendment, procurement action, or appeal. When several levels contribute, responses should be coordinated rather than forcing the case into a single category. This makes the model useful not only for classification but also for assigning action after a decision.
For empirical research, the three levels can guide data collection and case comparison. Studies can examine whether institutions distinguish conduct evidence from assessment-validity evidence and procedural evidence, whether responsibility attribution changes when institutional contribution is made visible, and whether level-specific corrective actions reduce recurrence. Such research would test whether the model improves consistency and explanatory quality rather than assuming that conceptual separation automatically produces better outcomes.
In Table 4, the three levels of academic integrity are compared through their primary object, core questions, and typical evidence. This comparison operationalizes the conceptual boundary developed in the text and provides a diagnostic basis for locating concerns before responsibility or response is determined.

3.2. Authorship, Contribution, Attribution, and Responsibility

Authorship, contribution, attribution, responsibility, and intellectual property are related but distinct. Authorship denotes accountable intellectual contribution under disciplinary conventions. Contribution identifies functions performed by a person or tool. Attribution acknowledges sources or assistance. Responsibility concerns answerability for claims, reasoning, evidence, and consequences. Intellectual-property status concerns legal rights and does not resolve academic authorship.
Epistemic responsibility requires verification of factual and computational claims, justification of substantive choices, disclosure of material AI assistance where required, interpretation of output in relation to the task, preservation of proportionate process evidence, and acceptance of responsibility for error, bias, fabricated references, or inappropriate data disclosure. AI systems perform cognitive-support functions but are not moral agents.
Authorship, contribution, attribution, responsibility, and intellectual property must be separated because they answer different questions. Authorship is a disciplinary and academic status associated with accountable intellectual contribution. Contribution describes the functions performed by a person or tool. Attribution acknowledges sources, assistance, or influence. Responsibility concerns who must answer for the claims, reasoning, evidence, and consequences of the final work. Intellectual property concerns legal rights and ownership, which may not determine whether an academic contribution satisfies authorship conventions. Conflating these categories produces rules that are difficult to teach and inconsistent to enforce.
Generative AI complicates authorship because it can produce text, code, images, calculations, and recommendations without possessing moral agency or academic responsibility. A system may perform a cognitively significant function, but it cannot verify the final claim, consent to authorship, respond to criticism, or accept consequences. The learner or researcher therefore remains responsible for the submitted work. This does not mean that AI contribution should be ignored. It means that contribution should be disclosed and evaluated without treating the tool as a responsible author (Floridi & Cowls, 2019; Holmes et al., 2022; Kasneci et al., 2023; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024).
Epistemic responsibility provides a practical standard. It requires the human submitter to verify factual, computational, and bibliographic claims; justify substantive choices; interpret outputs in relation to disciplinary evidence; disclose material assistance when required; preserve proportionate process evidence; and correct errors or harm. These obligations focus on answerability rather than a fictional requirement of completely unaided creation. They also align integrity with learning because the student must demonstrate understanding of the material retained in the final work.
Disclosure must be meaningful and proportionate. A binary statement that AI was used may be too vague to explain contribution, while a complete prompt history may be intrusive, burdensome, and misleading. Task-specific rules should identify which functions require disclosure and what level of detail is necessary. A short statement may be sufficient for brainstorming or language feedback. Extensive drafting, synthesis, code generation, or problem solving may require a contribution statement, examples of influence, verification steps, and an explanation of final decisions. Requirements should protect privacy and avoid collecting unrelated personal information.
Attribution also differs from citation. AI outputs may not have a stable author or recoverable source in the conventional sense, and a model’s answer may synthesize patterns from training data without traceable references. Citing the tool does not verify the content. Students must still locate and evaluate authoritative sources for claims that require evidence. Where institutional rules request citation or acknowledgment of a system, that practice should be understood as transparency about assistance, not as a substitute for scholarly sourcing. The distinction is especially important when models fabricate references or reproduce bias.
Intention remains relevant but cannot be inferred solely from the amount of AI-generated text. A student may openly use extensive assistance in a task that permits it and still demonstrate the intended learning. Another student may conceal a limited but decisive use that replaces the target competence. Responsibility therefore depends on function, rule clarity, relationship to the learning outcome, degree of control, disclosure, advantage, and capacity to explain. The framework avoids equating output percentage with misconduct severity.
Authorship conventions also vary by discipline. In writing-intensive subjects, argument construction and language may be central outcomes. In other contexts, language support may be peripheral to the competence being assessed. Coding, mathematical derivation, design, translation, and creative practice each require different distinctions between tool use and assessed contribution. Institutions should establish programme-level conventions and then translate them into task rules rather than relying on a single university-wide statement.
A defensible authorship policy should therefore answer five questions: which AI functions are permitted; which learning processes must remain demonstrably human; what disclosure is required; what evidence may be requested; and who remains accountable for error, bias, data use, and final claims. This structure makes academic expectations teachable and reviewable. It also supports proportionality because decisions can be connected to the actual function of the tool rather than to fear, novelty, or unverifiable assumptions about textual style (Cotton et al., 2024; Fishman, 2009; Kasneci et al., 2023; Perkins, 2023; Tlili et al., 2023).
The distinction between contribution and responsibility can also guide collaborative academic work. Human co-authors, editors, translators, software, and AI systems may each influence a product, but responsible authors must understand and approve the final claims. Contribution statements can document functions without implying equivalent accountability. The same principle can be adapted to student group work, where individual and collective contributions require clear expectations.
Policy should also address correction after submission. When an AI-assisted error, fabricated reference, or privacy breach is discovered, responsibility includes timely acknowledgment and remediation. Institutions should distinguish good-faith correction from concealment. This approach reinforces integrity as an ongoing practice of answerability rather than a one-time declaration of authorship.
In Figure 5, actor-specific duties are connected to coordinated institutional conditions and then to proportional attribution. The diagram extends the discussion of epistemic responsibility by showing that shared conditions do not erase individual duties or decision accountability.

3.3. Functional Classification of Human–AI Co-Production

A functional classification is preferable to a binary human-versus-AI distinction because academic work is often produced through iterative interaction. The relevant unit of analysis is the function performed and its relationship to the learning outcome. Spelling correction, retrieval support, idea generation, translation, feedback, drafting, synthesis, coding, and decision generation involve different forms of delegation. A classification system should therefore describe the work process rather than attempt to infer a single percentage of AI authorship from the final product.
The degree of involvement matters, but it must be interpreted contextually. Lower involvement may include local editing or formatting that does not replace the competence being assessed. Intermediate involvement may include brainstorming, feedback, translation, or code suggestions that shape the work but leave substantive decisions with the learner. Higher involvement may include extensive drafting, problem solving, synthesis, or adoption of output with limited transformation. These categories are not automatic risk scores. A higher-involvement use can be legitimate in a task designed to assess critique of AI output, while a seemingly minor use can be decisive if it supplies the central answer in a closed assessment.
Control is a second dimension. The learner’s role may range from selective acceptance of local suggestions to iterative prompting and substantial revision, or to largely uncritical adoption. Control should be evaluated through the decisions the learner can identify and justify. Prompting activity alone is not proof of intellectual ownership. A student may produce a long prompt history without understanding the output. Conversely, a concise interaction may be followed by extensive independent evaluation. Evidence should therefore focus on reasoning, source use, revision, and capacity to defend the final work.
Disclosure is a third dimension and should be matched to materiality. The purpose is to make contribution interpretable, not to create an administrative archive of every digital interaction. Institutions should distinguish routine assistance from functions that materially shape content, reasoning, or evidence. Disclosure formats may include a short use statement, a contribution table, selected prompts, a description of verification, or an explanation of limitations. The format should be accessible and should not require students to surrender unnecessary personal data or proprietary information.
Evidence of learning is a fourth dimension. Where AI use affects a central learning outcome, the institution may require complementary evidence such as a commentary, selected drafts, a structured oral explanation, reconstruction of a method, or an additional task. Such evidence should be planned in advance whenever possible. Retrospective verification should not become an improvised test with unclear criteria. Assessors need training, moderation, and reasonable adjustments to ensure that the verification process evaluates the intended competence rather than confidence, language fluency, or familiarity with AI terminology.
The final dimension concerns the integrity question created by the use. The issue may be rule compliance, validity of evidence, unfair advantage, concealment, data protection, or responsibility for false claims. These concerns should not be collapsed. A disclosed use may still invalidate an assessment if it replaces the target competence, while an undisclosed use may be minor where rules were ambiguous and the function peripheral. Proportional analysis must consider the combination of function, control, outcome alignment, disclosure, evidence, and intention.
Functional classification also supports curriculum design. Programmes can define how acceptable use progresses across levels of study. Novices may need tightly specified functions and guided verification. Advanced students may be expected to select tools critically, document contribution, evaluate bias, and defend complex human–AI workflows. Progression allows AI literacy and epistemic responsibility to become assessed outcomes rather than generic orientation content. It also reduces inconsistency across courses.
The classification should be reviewed as tools evolve. Brand-based rules become obsolete quickly, whereas function-based rules remain more stable. A system that begins as a writing assistant may later perform research, coding, or multimodal generation. Policies should therefore describe what a tool does in relation to the task, not simply whether a named product is allowed. This approach supports fairer decisions, clearer teaching, and more durable governance (Chan, 2023; Cotton et al., 2024; Dawson, 2021; Kasneci et al., 2023; Lim et al., 2024; Perkins, 2023; Rudolph et al., 2023; Tai et al., 2018; UNESCO, 2023; Xia et al., 2024).
The classification can also improve feedback. Instead of telling a student only that AI use was excessive, an educator can identify the problematic function, explain how it replaced or obscured the learning outcome, and specify what evidence or revision is needed. This form of feedback is more educationally useful than a tool-based prohibition because it teaches the relationship between assistance, contribution, and responsibility. It also creates a common vocabulary for moderation and appeal.
Institutional adoption should involve scenario testing across disciplines. Programme teams can apply the dimensions to representative tasks, compare judgements, identify disagreement, and refine examples. Student participation can reveal where terminology is unclear, or disclosure burdens are unrealistic. The resulting guidance should remain revisable because new systems may combine functions that were previously separate. A functional model is therefore a governance instrument as well as an analytical classification.
In Table 5, human-AI co-production is classified across function, control, disclosure, evidence of learning, and integrity concern. The comparison translates the preceding conceptual distinctions into task-level questions without treating degrees of AI involvement as measured risk scores.
In Table 6, epistemic responsibility is expressed through the obligations to verify, justify, disclose, interpret, and accept responsibility. Each obligation is connected to possible evidence and a boundary caution so that evidential requirements remain proportionate to the task and decision stakes.

4. Procedural Integrity and Attributable Responsibility

4.1. Actor Duties and Proportional Attribution

Distributed responsibility means that integrity outcomes are shaped by multiple actors; it does not dissolve individual attribution. Institutional ambiguity can mitigate responsibility but does not erase deliberate advantage-seeking. Proportional analysis considers rule clarity, educational stage, prior guidance, intention, degree of delegation, relation to the learning outcome, advantage obtained, impact, repetition, evidence quality, and willingness to explain and repair.
Academic integrity in AI-assisted education is shaped by multiple actors, but responsibility must remain attributable. Students control many decisions about use, disclosure, verification, and submission. Educators control task design, instructions, examples, feedback, and first-line judgement. Programmes coordinate disciplinary expectations and progression. Institutions establish policy, support, due process, and technology governance. Integrity and quality units support consistency, case review, and institutional learning. Providers influence documentation, privacy, auditability, and system limitations. The framework therefore replaces vague shared responsibility with coordinated and differentiated duties.
Student duties include understanding task-specific rules, seeking clarification where feasible, disclosing material assistance, verifying claims, protecting sensitive data, and demonstrating the intended learning. These duties are not absolute in isolation. A student cannot reasonably be expected to comply with a rule that is inaccessible, contradictory, or introduced after submission. Nor should students bear sole responsibility for provider opacity or institutional failure to offer guidance. However, ambiguity does not automatically excuse deliberate concealment or planned advantage-seeking. Proportional attribution must consider both the institutional conditions and the student’s agency.
Educator duties extend beyond detecting inappropriate use. Educators should align permitted AI functions with learning outcomes, communicate expectations at task level, provide examples, teach necessary literacy, design valid evidence, and evaluate cases through corroborated information. They should not rely on automated indicators as proof or create inaccessible verification. These responsibilities require workload, training, and institutional support. Holding educators accountable without providing time, authority, or resources would reproduce the same diffusion of responsibility that the framework seeks to avoid.
Programme and institutional duties concern coherence. A university-wide policy should establish values, common definitions, data protections, limits on automated evidence, human review, and appeal. Programme teams should interpret these principles in relation to disciplinary norms, professional obligations, and progression. Courses and tasks should operationalize permissions and disclosure. Institutions should provide staff development, student support, accessibility review, trained case personnel, and periodic policy revision. Inconsistency across these levels creates ambiguity that can undermine both prevention and fair attribution (Chan, 2023; European Union, 2024; Holmes et al., 2022; International Center for Academic Integrity, 2021; National Institute of Standards and Technology, 2023, 2024; OECD, 2023, 2026).
Technology-provider duties are more limited because institutions cannot directly control every design decision, but providers are not ethically neutral. They should document intended use, known limitations, data practices, model changes, accessibility, security, and mechanisms for review. Institutions should translate these expectations into procurement criteria and contracts. Provider failure does not remove institutional judgement, just as institutional procurement failure does not automatically remove student responsibility. The framework identifies where duties intersect without allowing one actor to displace another.
Proportional attribution considers rule clarity, educational stage, prior guidance, intention, degree of delegation, relationship to the learning outcome, advantage obtained, impact on others, repetition, evidence quality, and willingness to explain and repair. These factors should be documented rather than applied intuitively. A decision-maker should be able to explain why a factor mitigates or aggravates responsibility and how it influenced the selected response. This requirement supports consistency while preserving contextual judgement.
Institutional contribution must also produce institutional consequences. If a case reveals ambiguous instructions, inaccessible support, inconsistent course rules, unreliable technology, or inadequate staff training, the response should include corrective action beyond the individual case. Without such feedback, distributed responsibility becomes rhetorical. The institution acknowledges its role but does not change the conditions that produced the concern. The case-to-governance mechanism therefore links attribution to policy and assessment revision.
The model does not seek equal distribution of blame. Responsibility follows control, knowledge, authority, and contribution. A student who knowingly conceals extensive delegation may hold substantial responsibility even when policy could be improved. An institution may hold primary responsibility for an adverse automated decision based on an unvalidated tool. An educator may hold responsibility for an assessment that creates unjustified barriers. Differentiated attribution enables these distinctions and supports more credible, educationally defensible decisions (Bretag et al., 2019; Cotton et al., 2024; European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Holmes et al., 2022; National Institute of Standards and Technology, 2023, 2024; Perkins, 2023).
The allocation of duties should be reflected in institutional documentation. Assessment briefs, policy templates, reviewer forms, procurement criteria, and appeal records should each identify the actor responsible for the relevant decision. This documentation reduces the risk that responsibility is asserted only after a problem occurs. It also allows quality assurance to examine whether the institution provided the guidance, training, authority, and support that its policy assumes.
Empirical evaluation can test whether differentiated duties improve consistency and trust. Vignette studies may compare attribution decisions under clear and ambiguous rules, while case audits can examine whether institutional contribution is acknowledged and corrected. Student and staff interviews can explore whether the model feels fair and understandable. The objective is not to distribute responsibility evenly but to make the reasoning visible, contestable, and connected to preventive action.
In Table 7, the primary duties of students, educators, programmes, integrity units, and technology providers are compared with evidence of fulfilment and limits of responsibility. This allocation supports proportional attribution by identifying where responsibility is coordinated and where it remains actor specific.

4.2. Algorithmic Mediation and Substantive Human Oversight

Algorithmic decision-making is relevant when automated grading, proctoring, similarity analysis, AI-content detection, or risk prediction affects assessment evidence or an integrity decision. Automated indicators are leads requiring corroboration, not proof. Substantive oversight requires competence, time, authority to depart from system outputs, access to documentation, reason-giving, and independent review or appeal.
Algorithmic mediation becomes an academic-integrity issue when a system affects evidence of learning, the detection or interpretation of possible misconduct, or a consequential decision about academic standing. Relevant systems include automated grading, remote proctoring, similarity analysis, AI-content detection, authorship verification, and risk prediction. Their outputs vary in validity and meaning. A similarity percentage identifies overlap, not intent. An AI-content score expresses a probabilistic classification, not authorship. A risk model identifies patterns, not guilt. Governance must therefore begin by defining what an indicator can and cannot support.
Data and tool provenance form the first stage of accountability. Institutions should know which system and version were used, the purpose of deployment, the data processed, the settings applied, and any relevant vendor changes. Procurement should include legal basis, privacy, security, accessibility, bias, documentation, and auditability. A tool should not enter a high-stakes integrity process merely because it is convenient or widely marketed. The institution remains responsible for determining whether its use is educationally and procedurally justified (European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Holmes et al., 2022; National Institute of Standards and Technology, 2023, 2024; UNESCO, 2021).
Evidence quality is the second stage. Validity claims should match the actual use. A system tested on one language, genre, or population may not support inference in another. Error rates should be examined across relevant groups, and limitations should be communicated to reviewers and students. Probabilistic outputs require corroboration through independent evidence such as drafts, source records, task-specific explanation, or other contextual information. The absence of corroboration should reduce the weight of the indicator rather than increase pressure on the student to disprove it.
Substantive human review is the third stage. Reviewers need competence to interpret the tool, time to examine the evidence, authority to depart from the output, and access to documentation. They should record what evidence was considered, how conflicting information was resolved, and why the final inference is proportionate. Training should address automation bias, confirmation bias, differential language effects, disability, privacy, and the distinction between textual features and academic responsibility. A nominal reviewer who simply accepts a score does not provide human oversight (Eubanks, 2018; Floridi & Cowls, 2019; Jobin et al., 2019; Mittelstadt et al., 2016; O’Neil, 2016; Williamson, 2021; Williamson et al., 2020).
Reason-giving is the fourth stage. Consequential decisions should identify the applicable rule, the established facts, the evidential basis, the interpretation of intention and advantage, any institutional contribution, and the rationale for the response. Reasons allow students to understand the decision and enable independent review. They also create institutional data for consistency analysis and policy improvement. Vague statements that a system detected AI use are insufficient because they conceal both technical uncertainty and human judgement.
Contestability is the fifth stage. Students should receive intelligible notice, a meaningful opportunity to respond, access to relevant evidence, reasonable time, and an appeal route independent of the original decision. Contestability is not satisfied by providing a technical report that cannot be interpreted or challenged. Institutions should also protect students from coerced admissions. A request to explain work may be legitimate, but participation should not require waiving procedural rights or accepting an unproven allegation.
Institutional audit completes the chain. Universities should examine overrides, appeals, reversals, differential outcomes, complaints, accessibility effects, data retention, and vendor changes. A high rate of reviewer agreement with a system may indicate accuracy, automation bias, or lack of authority. A low appeal rate may indicate trust or barriers. Audit therefore requires interpretation, not only counting. Findings should inform procurement, reviewer training, task design, and policy.
The legitimacy of algorithmic mediation depends on the whole chain rather than a single technical property. Explainability without appeal is insufficient. Human review without authority is insufficient. Privacy review without validity evidence is insufficient. The framework therefore treats consequential AI-mediated decisions as socio-organizational processes. The institution must be able to explain how technology, evidence, human judgement, and rights operated together in the case.
Institutions should also define thresholds for non-use. A system may be inappropriate when its validity is unknown, its data practices are unacceptable, its effects cannot be audited, or the educational stakes exceed the reliability of the indicator. The decision not to deploy a tool is a legitimate governance outcome. Responsible innovation does not require every available technology to be incorporated into integrity procedures.
Research should examine the interaction between reviewer competence and organisational conditions. Training may not reduce automation bias when workloads are high, or authority is weak. Appeal data can reveal error, but only for cases that reach appeal. Experimental and observational studies should therefore combine technical performance, reviewer behaviour, student experience, and institutional structure. This broader approach reflects the framework’s claim that algorithmic accountability is socio-organisational rather than purely technical.
In Figure 6, the complete accountability chain for an algorithmically mediated integrity decision is presented, from data and tool provenance to institutional audit. The sequence connects the discussion of substantive human oversight to the evidential, procedural, and contestability safeguards required at each stage.
In Table 8, each stage of an algorithmically mediated decision is linked to a minimum safeguard, the evidence required to demonstrate it, and a failure mode that must be avoided. The table thereby converts the accountability chain into an operational review structure.

5. Preventive Governance and AI-Resilient Assessment

5.1. Multi-Level Policy Architecture

Effective prevention requires consistency across institutional, programme, course, and task levels. A generic institutional statement cannot determine whether translation, feedback, coding assistance, or generative drafting is permitted in a particular assessment. Policy must reach the assessed task while preserving common rights, data protections, and due process guarantees.
A multi-level policy architecture is necessary because institutional principles do not automatically resolve task-level questions. University policy can define values, common rights, data protections, limits on automated evidence, and due process guarantees. It cannot determine, for every assessment, whether translation, brainstorming, coding assistance, feedback, or generative drafting is compatible with the learning outcome. Programme, course, and task levels are therefore needed to translate common principles into disciplinary and pedagogical expectations.
At the institutional level, policy should define key terms, distinguish authorship from assistance, establish minimum disclosure principles, protect privacy, prohibit automated proof, require substantive human review, and provide appeal. It should also assign ownership for review, staff development, student support, procurement, accessibility, and case data. A policy that only lists prohibited tools will become obsolete and may encourage covert use. Function-based guidance is more durable because it focuses on what the system does in relation to academic work (Chan, 2023; European Union, 2024; International Center for Academic Integrity, 2021; National Institute of Standards and Technology, 2023, 2024; OECD, 2023, 2026; UNESCO, 2021, 2023; UNESCO International Institute for Higher Education in Latin America and the Caribbean, 2023).
At programme level, disciplinary conventions and progression should be specified. Programmes differ in how they understand original contribution, acceptable collaboration, evidence, professional responsibility, and risk. A health, engineering, law, language, design, or computing programme may require different boundaries because the assessed competencies and external obligations differ. Programme mapping can identify where AI literacy is introduced, practised, and assessed, and where particular forms of independent performance remain essential.
At course level, guidance should explain the pedagogical purpose of AI use and establish recurring expectations. Students should understand why certain functions are encouraged, limited, or prohibited. Examples should include acceptable and unacceptable uses, disclosure models, verification expectations, and available support. Course-level coherence reduces the cognitive burden created when each assignment uses completely different terminology. However, it should not eliminate the flexibility needed for different learning outcomes.
At task level, instructions should be operational. They should state the intended learning outcome, permitted and prohibited functions, required disclosure, process evidence that may be requested, privacy boundaries, and the consequences of non-compliance. Instructions should be accessible before the work begins and should be included in the assessment brief rather than distributed across unrelated policy pages. Where rules change during a course, students need clear notice and reasonable transition.
The feedback level closes the architecture. Cases, appeals, accessibility concerns, student questions, staff experience, and technology changes should be analyzed for recurring patterns. A cluster of cases may reveal unclear wording, an invalid task, unequal access, weak literacy support, or inconsistent judgement. Policy review should document the issue, the action selected, the responsible owner, and whether implementation occurred. Without this loop, policy remains static, and institutions repeatedly treat systemic ambiguity as individual failure.
Communication and comprehension are as important as publication. Institutions should test whether students and staff can identify permitted functions and disclosure requirements in realistic scenarios. Orientation modules, examples, decision trees, and task templates can support understanding, but they should not substitute for dialogue within courses. Comprehension data should be interpreted alongside actual practice because quiz performance does not guarantee conduct.
The architecture also supports proportional attribution. When rules are coherent across levels, and the task instruction is clear, deliberate concealment can be assessed with greater confidence. When levels conflict, institutional ambiguity becomes a mitigating factor and a trigger for corrective action. This relationship demonstrates why policy quality is not merely administrative. It shapes the fairness, educational value, and evidential reliability of integrity decisions.
Policy coherence should be evaluated horizontally as well as vertically. Courses within the same programme may use different assessment purposes and therefore different permissions, but they should use consistent concepts and explain justified variation. Contradictory language can create accidental breaches and reduce trust. Programme moderation can identify these inconsistencies before tasks are released and can ensure that professional or accreditation constraints are communicated clearly.
A mature architecture should also include change management. When a provider introduces new functions or when institutional policy changes, affected tasks, examples, training, and disclosure requirements should be reviewed. Students should not be expected to infer the implications of technical change. Versioned guidance, clear effective dates, and archived prior rules support fairness and auditability. These practices turn policy from a static document into a maintained educational system.
Institutional ownership should be visible to users. Students and staff need to know where to obtain clarification, report inconsistency, request accessibility support, and challenge a rule. A centralized portal may improve access, but responsibility for pedagogical interpretation remains with programmes and educators. Clear escalation routes help prevent uncertainty from being resolved informally and inconsistently. They also provide evidence about where policy language requires revision.
In Figure 7, institutional principles, programme rules, course guidance, and task instructions are connected through a preventive governance process and a review loop. The figure shows how policy becomes effective only when it reaches the assessed task and when case evidence returns to governance for revision.
In Table 9, the policy architecture is specified across institutional, programme, course, task, and feedback levels. For each level, the primary function, minimum content, and responsible owner or review cycle are identified, linking policy clarity to implementation accountability.

5.2. AI-Resilient Assessment

AI-resilient assessment remains capable of producing credible, fair, accessible, and interpretable evidence of intended learning under realistic conditions of AI availability. It does not seek to eliminate AI. It aligns permitted use with learning outcomes, makes material contribution visible, and uses complementary evidence only where justified by validity and decision stakes (Boud & Falchikov, 2007; Carless & Boud, 2018; Corbin et al., 2025; Dawson, 2021; Dawson et al., 2024; McArthur, 2016; Nicol & Macfarlane-Dick, 2006; Tai et al., 2018; Xia et al., 2024).
AI-resilient assessment is defined by the credibility of the evidence it produces, not by the complete exclusion of AI. The concept begins with learning-outcome alignment. Educators must identify which cognitive, practical, communicative, or ethical processes the task is intended to demonstrate and then decide which AI functions support, alter, or replace those processes. This approach avoids the unstable objective of designing AI-proof tasks and instead focuses on whether the learner’s competence remains interpretable under realistic conditions of tool availability (Boud & Falchikov, 2007; Carless & Boud, 2018; Corbin et al., 2025; Dawson, 2021; Dawson et al., 2024; McArthur, 2016; Nicol & Macfarlane-Dick, 2006; Tai et al., 2018; Xia et al., 2024).
Process portfolios can make development and decision points visible. Selected drafts, source notes, code versions, design iterations, or reflective commentary may help assessors understand how the work evolved. However, extensive portfolios can create workload, privacy, and accessibility burdens, and process artefacts can themselves be fabricated. Selective checkpoints, sampled verification, and clear retention rules are therefore preferable to unrestricted surveillance of the entire production process.
Reflective AI-use commentary can support epistemic responsibility by asking students to identify material functions, evaluate output, describe verification, and explain revisions. The commentary should be linked to the learning outcome rather than treated as a generic declaration. Formulaic statements add little evidence, and excessive disclosure can encourage students to report trivial interactions or reveal sensitive information. Prompts should therefore be specific, proportionate, and assessed through clear criteria.
Oral defence or demonstration can provide complementary evidence of understanding and ownership. It is especially useful when the final product alone cannot show how decisions were made. Yet oral assessment can introduce anxiety, language bias, disability barriers, scheduling problems, and inter-rater variation. Structured questions, reasonable adjustments, assessor training, moderation, and limited scope are necessary. An oral component should verify the target competence, not reward confidence or spontaneous fluency.
Context-rich projects can reduce generic answer substitution by requiring application to local data, field conditions, personal design choices, or iterative stakeholder needs. They may also support authentic professional learning. However, they can reproduce inequity when students have unequal access to contexts, equipment, data, or networks. Projects can still be outsourced or heavily generated. Resource provision, process evidence, transparent collaboration rules, and contextual moderation are therefore required.
Supervised components may provide complementary evidence in high-stakes decisions, but they should be used sparingly. Controlled conditions can reduce uncertainty while also increasing surveillance, logistics, cost, and accessibility burdens. Supervision may measure performance under constrained conditions rather than authentic competence. Institutions should explain why less intrusive evidence is insufficient and provide accommodations. The objective is evidential triangulation, not a return to universal closed-book testing.
AI-output critique is a particularly relevant strategy because it makes evaluation of generated material an explicit learning activity. Students can identify errors, bias, missing evidence, weak reasoning, or disciplinary limitations and then improve the output. This approach can develop evaluative judgement and AI literacy. It may, however, advantage students with greater prior access to AI or assess tool familiarity more than disciplinary knowledge. Explicit teaching and alternative pathways are needed.
No strategy should be assigned a universal risk rank. The appropriate combination depends on the learning outcome, stakes, cohort, discipline, accessibility, workload, and available moderation. Validity can decrease when too many safeguards are added, because the burden changes what is being assessed. AI-resilient assessment is therefore an exercise in balanced design. It combines permitted use, transparent contribution, proportionate evidence, and fair judgement while retaining the possibility that certain functions must remain restricted in particular tasks.
Assessment combinations should be designed around complementary evidence rather than duplication. A portfolio and oral defence, for example, should answer different validity questions. The portfolio may show development and decision points, while the oral component may clarify reasoning or ownership. Repeating the same demand in several formats increases burden without necessarily improving interpretation. Design should specify the evidential role of each component and how assessors will integrate the information.
Evaluation of AI-resilient assessment should include students and staff. Students can identify hidden accessibility and workload costs, while staff can assess feasibility, moderation, and reliability. Pilot studies and staged implementation are preferable to institution-wide adoption of untested formats. The goal is a defensible balance among validity, authenticity, transparency, access, and workload. This balance must be reviewed as AI capabilities and disciplinary practices evolve.
In Table 10, six assessment strategies are compared through their potential integrity benefits, principal trade-offs, and required safeguards. The comparison supports the preceding definition of AI-resilient assessment by showing why no format should be treated as inherently AI-proof.
In Figure 8, the six assessment strategies are compared across validity support, contribution visibility, accessibility, staff workload, and reliability. The heatmap is an author-generated decision aid that makes trade-offs visible and must be locally validated rather than interpreted as empirical performance evidence.
In Table 11, the assessment strategies are compared according to best use, primary validity risk, accessibility concern, staff burden, and recommended combination. This matrix complements Figure 8 by translating the heuristic profile into contextual design choices.

5.3. Governed Verification and Institutional Learning

Prevention and verification are not opposites. Similarity reports, draft comparisons, oral clarification, and technical logs may support fact-finding when targeted, transparent, proportionate, corroborated, and reviewable. Problems arise when probabilistic tools are treated as conclusive, surveillance becomes excessive, or verification substitutes for assessment redesign.
Institutions should analyse recurring patterns rather than merely publish rules. Repeated cases in one task may reveal ambiguous instructions, invalid assessment design, unequal access, weak literacy support, or inappropriate technology. Raw case counts are ambiguous and should never be treated as direct measures of integrity culture.
Prevention and verification should be understood as complementary rather than opposing functions. Clear policy and assessment design reduce ambiguity, but they cannot eliminate all uncertainty or deliberate misconduct. Verification may therefore be necessary to establish facts, protect the validity of assessment, and support fair decision-making. Its legitimacy depends on purpose, evidence, proportionality, transparency, human judgement, and appeal. Verification should answer a defined question, not serve as general suspicion.
Similarity reports can identify textual overlap but require interpretation of quotation, common language, source use, and disciplinary practice. Draft comparisons may show development but can be incomplete or manipulated. Oral clarification may test understanding but can introduce accessibility and reliability concerns. Technical logs may document actions but not necessarily intention or intellectual control. AI-content detectors are probabilistic and vulnerable to false classification. Each method provides a limited form of evidence and should be corroborated where consequences are significant (Dawson, 2021; Dawson et al., 2024; European Union, 2024; National Institute of Standards and Technology, 2023, 2024).
Governed verification begins with declared rules. Students should know which evidence may be requested and under what circumstances. Retrospective requirements should be exceptional because students cannot preserve records they were never told to retain. Data minimization should guide evidence collection. Institutions should request only what is necessary for the specific concern and protect unrelated private, personal, or commercially sensitive information. The burden should reflect the stakes of the decision.
Human review must connect evidence to the learning outcome and the alleged breach. A reviewer should distinguish inability to explain a detail from proof of misconduct, especially when stress, disability, language, or time has affected performance. Structured procedures and moderation can improve consistency. Where evidence remains ambiguous, the decision should reflect uncertainty rather than convert it into a presumption against the student. The standard of proof and consequences should be defined in institutional policy.
Verification also has an educational function. A clarification process can reveal misunderstandings about attribution, verification, or permissible use and can lead to coaching, revised work, or a learning plan. However, educational purpose must not be used to bypass due process. Students should understand whether a meeting is formative, investigative, or both. Records, rights, and possible outcomes should be clear. The distinction protects trust and prevents coerced participation.
Institutional learning requires analysis beyond the individual case. Repeated concerns in one assessment may indicate unclear instructions, poor outcome alignment, unrealistic workload, unequal access, or inconsistent teaching. A pattern of detector reversals may indicate tool limitations or reviewer bias. Appeals may reveal procedural weakness. Institutions should code cases in a privacy-preserving manner and review patterns with educators, students, accessibility services, integrity staff, and governance functions.
Case counts must be interpreted cautiously. A decline may reflect effective prevention, reduced reporting, inconsistent enforcement, or loss of trust. An increase may reflect greater misconduct, improved reporting, clearer definitions, or more aggressive detection. Evaluation should therefore combine counts with comprehension, appeal outcomes, reversal patterns, workload, accessibility, recurrence, and documented changes. The case-to-governance loop is successful only when evidence leads to implemented action.
Governed verification thus protects both standards and rights. It recognizes that some concerns require inquiry while rejecting automated suspicion and surveillance-first governance. It also transforms cases into feedback about institutional design. This relationship is central to the integrated framework because it connects response to prevention rather than treating enforcement as the endpoint.
Verification protocols should include stopping rules. Once sufficient reliable evidence has been obtained, additional collection may become disproportionate. Conversely, where evidence remains weak, institutions should not escalate intrusion merely to produce certainty. Decision standards should allow findings such as insufficient evidence, educational concern without misconduct, or institutional ambiguity. These outcomes are more credible than forcing every inquiry into a binary conclusion.
Institutional review should also examine the emotional and relational effects of verification. Even a procedurally correct inquiry can damage trust if communication is accusatory or opaque. Clear notice, respectful questioning, timely decisions, and support can reduce unnecessary harm. Qualitative feedback from students and reviewers should therefore complement formal case metrics. This information can improve scripts, training, and assessment guidance.
Training should include the communication of uncertainty. Reviewers may feel pressure to reach definitive conclusions, yet evidential humility is essential when tools are probabilistic, and process records are incomplete. Decision templates can include the strength and limitations of each source of evidence, alternative explanations, and the standard applied. This practice improves reasons and appeal review.
The institutional-learning function also depends on closing actions. Committees should not only identify lessons but assign owners, deadlines, and verification of implementation. Students and educators should receive appropriate feedback about resulting changes. Visible action can demonstrate that case review is not solely punitive and that institutional responsibility has practical consequences.

6. Ethical AI Literacy: Functions, Roles, and Limits

6.1. AI Literacy as Individual and Institutional Capacity

AI literacy includes understanding system capabilities and limitations, critical evaluation of outputs, data and bias awareness, responsible use, and ethical judgement (Chiu et al., 2024; Laupichler et al., 2022; Long & Magerko, 2020; Ng et al., 2021; Southworth et al., 2023; UNESCO, 2024a, 2024b). In academic integrity it performs four distinct functions: individual capacity; curricular learning outcome; organisational capacity for procurement, oversight, and case review; and preventive support that reduces misunderstanding. It becomes an ethical infrastructure only when these functions are institutionalised.
AI literacy is often defined as a set of knowledge, skills, and dispositions for understanding, evaluating, and using AI. In academic integrity, the concept must be broader because responsible practice depends on individual judgement, curriculum design, institutional oversight, and continuous adaptation. A student who understands model limitations but cannot interpret task rules remains vulnerable to misuse. An educator who understands prompting but cannot design valid assessment lacks an essential professional capability. An institution that offers student training while procuring opaque tools has not established an ethical literacy infrastructure (Chiu et al., 2024; Laupichler et al., 2022; Long & Magerko, 2020; Ng et al., 2021; Southworth et al., 2023; UNESCO, 2024a, 2024b).
Foundational understanding concerns how systems generate and mediate outputs. Users should know that generative models are probabilistic, may fabricate information, can reproduce bias, and do not possess understanding or responsibility in the human sense. This knowledge supports appropriate scepticism. It should be taught through examples relevant to the discipline rather than abstract technical description. Students do not need to become machine-learning specialists, but they need enough understanding to recognize limitations and ask appropriate questions.
Critical evaluation concerns evidence, uncertainty, sources, and context. Students and staff should be able to compare outputs with authoritative material, detect unsupported claims, evaluate reasoning, and identify potential bias. Critical evaluation is central to epistemic responsibility because fluent output can create unwarranted confidence. It also connects AI literacy to established information, feedback, and evaluative-judgement studies rather than treating it as a wholly new competence.
Ethical judgement concerns responsibility, fairness, privacy, disclosure, and the relationship between tool use and learning outcomes. Users must decide not only whether a system can perform a function but whether it should perform that function in the task. They should understand data risks, unequal access, potential harm, and the importance of transparent contribution. Ethical judgement is developed through cases, dialogue, guided practice, and assessment rather than a one-time policy quiz.
Organisational capacity concerns the competence of educators, programme leaders, integrity staff, quality teams, procurement functions, and decision-makers. Institutions need staff who can evaluate vendor claims, interpret automated indicators, design rules, provide accessibility, facilitate restorative processes, and audit differential impacts. Authority and resources are part of literacy in this organisational sense. Knowledge without the ability to act does not create effective oversight.
Continuous renewal is necessary because systems, practices, and regulation change rapidly. Guidance should be reviewed when tool capabilities, privacy terms, model versions, or disciplinary uses change. Staff development should be iterative and linked to real cases. Student literacy should progress across a programme rather than rely on generic induction. Institutions should also learn from appeals, accessibility concerns, and assessment review.
Role specificity prevents literacy from becoming an undifferentiated list. Students require capacities for use, verification, disclosure, and defence of learning. Educators require assessment, communication, evidence, and moderation skills. Programme leaders require curriculum mapping and consistency. Integrity staff require due process and case analysis competence. Procurement and IT teams require security, data, accessibility, documentation, and contestability expertise. Each role contributes to the same integrity system through different duties.
AI literacy becomes an ethical infrastructure only when these capacities are embedded in policy, curriculum, staffing, support, assessment, procurement, and governance. The metaphor should not be used to imply that literacy alone ensures conduct. It describes the institutional conditions that make informed judgement possible. Incentives, pressure, opportunity, norms, and accountability remain necessary parts of the framework.
Curricular integration should include progression and assessment. Introductory learning may focus on system limitations, privacy, and task rules. Intermediate learning can develop verification, disclosure, and critique. Advanced learning can address discipline-specific workflows, professional responsibility, procurement, and governance. Programmes should identify where these outcomes are taught and how students demonstrate them, avoiding both repetition and gaps.
Institutional capacity also depends on communities of practice. Educators, integrity staff, librarians, learning designers, accessibility specialists, and technical teams can share cases and resources. Such collaboration reduces isolated decision-making and allows guidance to reflect several forms of expertise. Communities of practice should be connected to formal governance so that insights lead to policy and resource decisions rather than remaining informal discussion.
In Figure 9, AI literacy is represented as a layered institutional capacity system that progresses from foundational understanding to critical evaluation, ethical judgement, organisational capability, and continuous renewal. The structure connects individual competence to the wider institutional responsibilities discussed in the text.
In Table 12, minimum and advanced AI-literacy competencies are compared across students, educators, programme leaders, integrity staff, and procurement or IT teams. The role-specific structure shows how literacy contributes differently to responsible use, valid assessment, procedural integrity, and institutional deployment.
In Table 13, the individual, curricular, organisational, and preventive functions of AI literacy are distinguished. For each function, the supported outcome is presented alongside what literacy cannot guarantee, reinforcing the argument that competence must be combined with policy, assessment, resources, incentives, and accountability.

6.2. Literacy Is Necessary but Not Sufficient

A knowledgeable student may still conceal AI use because of competition, workload, pressure, perceived unfairness, or anticipated advantage. Conversely, a student may breach a rule without understanding it. Literacy can reduce ambiguity and strengthen judgement, but conduct also depends on incentives, opportunity, assessment design, norms, access, and enforcement (Bretag et al., 2019; Cotton et al., 2024; Eaton, 2021; Perkins, 2023). Institutional responsibility must not be transferred to individuals under the label of literacy.
The distinction between competence and conduct is essential. A person can understand a rule and still decide to breach it. Academic pressure, competition, workload, fear of failure, perceived unfairness, and anticipated advantage may override knowledge. Strategic contract cheating existed before generative AI and demonstrates that misconduct cannot be explained only by misunderstanding. Generative systems change the opportunity structure by making assistance more accessible, private, and difficult to trace, but they do not remove motivation or agency (Bretag et al., 2019; Cotton et al., 2024; Eaton, 2021; Perkins, 2023).
Conversely, a person can breach a rule without informed intention. Guidance may be unclear, inconsistent across courses, inaccessible, or introduced too late. Students may not understand the difference between language support, idea generation, drafting, and substitution of a learning outcome. International students and students using accessibility tools may encounter additional ambiguity when assistance functions overlap. Literacy and policy clarity can reduce these problems, but responsibility must be attributed in relation to prior guidance and realistic opportunities to learn.
A useful behavioural distinction separates capacity, opportunity, motivation, and accountability. Literacy primarily strengthens capacity: the ability to understand systems, evaluate outputs, and interpret rules. Policy and assessment design shape opportunity by making responsible use feasible and inappropriate delegation less attractive. Institutional culture, workload, and incentives influence motivation. Verification and response procedures provide accountability. A framework that addresses only one element will produce incomplete prevention.
Educator conduct is subject to the same limitation. Staff may understand that an AI detector is probabilistic yet rely on it because of workload, managerial pressure, or lack of alternative evidence. They may know that oral verification can disadvantage some students but use it without accommodation because no institutional support exists. Professional development must therefore be accompanied by time, authority, moderation, and procedural requirements. Otherwise, literacy becomes a symbolic expectation rather than operational capacity.
Institutional literacy can also fail. A university may publish sophisticated guidance while procurement decisions remain disconnected from academic governance. Data-protection review may occur without assessment-validity review. Integrity staff may not be included when systems are purchased. Quality units may collect case counts without examining appeals or differential impacts. Organisational structures must connect expertise to decision authority and create routes for evidence to influence policy.
Evaluation should therefore avoid treating course completion or quiz scores as proof of ethical behaviour. Literacy measures can assess knowledge, judgement, and confidence, but conduct requires process evidence, disclosure practice, case characteristics, and contextual data. Self-report is vulnerable to social desirability, while reported misconduct is influenced by detection and enforcement. Mixed methods and longitudinal designs are needed to understand the relationship between literacy and action.
The limitation does not make literacy unimportant. Without literacy, students cannot evaluate output, educators cannot design or judge responsibly, and institutions cannot oversee technology. The framework treats literacy as an enabling condition that interacts with clarity, assessment, support, verification, and proportionate consequences. This calibrated claim is stronger than either extreme: literacy is neither a universal solution nor a peripheral addition.
The practical implication is that literacy initiatives should include ethical dilemmas, disciplinary examples, disclosure practice, source verification, privacy, and reflection on incentives. They should also be evaluated alongside policy comprehension, assessment design, workload, access, and case outcomes. Institutions should ask not only whether participants know the rules, but whether the environment enables and expects responsible action.
This distinction has implications for communication. Institutions should avoid messages that imply misconduct results from ignorance alone or that trained students have no legitimate uncertainty. Guidance should acknowledge pressure, incentives, and evolving norms while maintaining clear expectations. A credible message combines education with support and proportionate accountability. It also makes institutional duties visible, which can strengthen trust and willingness to seek clarification.
Research should examine mediating conditions. The effect of literacy may vary with assessment design, rule clarity, access, workload, belonging, and perceived legitimacy. Studies that measure only pre- and post-training knowledge will miss these interactions. Longitudinal, mixed-method research can identify when literacy supports responsible action and when structural conditions override it. Such findings would allow institutions to target interventions more precisely.
The model also cautions against unequal expectations. Students may be required to master rapidly changing tools while staff guidance remains uncertain. Institutions should sequence expectations with teaching and access, provide alternatives where tools are unavailable, and avoid assuming that frequent use equals literacy. Competence must be demonstrated rather than inferred from familiarity.
A prevention strategy should therefore combine literacy with workload review, support, fair assessment, accessible rules, and credible procedures. Interventions can be targeted to the mechanism identified in case data. If ambiguity is primary, guidance may be appropriate; if strategic advantage predominates, accountability and assessment design may require greater attention.
The framework therefore supports a diagnostic approach to prevention. Before prescribing more training, institutions should examine whether the problem reflects knowledge, ambiguous rules, invalid assessment, unequal access, workload, incentives, or weak accountability. Literacy is the appropriate response when capacity is the limiting mechanism. Other conditions require different interventions. This diagnostic discipline can reduce repetitive training that places responsibility on students without changing the environment.

7. Restorative Accountability and Proportionate Response

7.1. Distinguishing Educational Correction from Restorative Justice

Reflection, retraining, or task revision may be educational, but they are not restorative by themselves. A restorative process requires attention to harm or unfairness, affected people or communities, acceptance of responsibility, meaningful and voluntary participation, feasible repair, and reintegration (Bussu & Karp, 2026; Evans & Vaandering, 2022; Karp, 2019a, 2019b; Wachtel, 2016; Zehr, 2015).
Educational correction and restorative justice can overlap, but they should not be treated as synonyms. Educational correction may include clarification, coaching, resubmission, skills development, or a learning plan. Its primary purpose is to address misunderstanding, weak judgement, or gaps in competence. These responses can be appropriate and proportionate, especially in low-impact first cases. They do not become restorative merely because they are less punitive than a sanction.
Restorative justice begins with harm and relationships. It asks who or what was affected, what responsibility can be acknowledged, what participation is appropriate, what repair is feasible, and how reintegration can occur. In academic integrity, harm may include unfair advantage, invalid assessment evidence, additional burdens on peers or staff, damage to trust, or professional risk. A restorative process should make these effects visible without exaggerating them or using harm language to coerce admission (Bussu & Karp, 2026; Evans & Vaandering, 2022; Karp, 2019a, 2019b; Wachtel, 2016; Zehr, 2015).
Participation must be meaningful and voluntary. A student should understand the purpose, possible outcomes, confidentiality, records, and relationship to formal procedures. Refusal should not be treated as evidence of guilt or an aggravating factor. Affected parties should not be required to participate, and their safety, workload, and preferences should be considered. Facilitation is necessary where power differences, emotion, or conflict may shape dialogue.
Responsibility in a restorative process differs from a forced confession. Facts should be sufficiently established or accepted for constructive participation. The process can distinguish intention from impact and individual conduct from institutional contribution. A student may acknowledge that undisclosed AI use created unfairness while also identifying ambiguous instructions or inadequate support. The institution can accept responsibility for those conditions without erasing the student’s agency. This differentiated account supports both repair and policy learning.
Repair should be specific, feasible, proportionate, and educationally relevant. It may include corrected work, an explanation, contribution to peer learning, completion of a skills activity, or another agreed action. Repair should not be humiliating, excessive, or disguised punishment. It should address the identified harm rather than serve as a generic consequence. Timelines, support, privacy, and completion should be documented.
Reintegration is also essential. A restorative process should avoid permanent stigma and should clarify how the student returns to ordinary academic participation after fulfilling the agreement. Record practices must balance accountability, recurrence, professional obligations, privacy, and the purpose of restoration. An institution that retains indefinite labels or shares case details unnecessarily undermines reintegration.
Educational correction may support restoration by building understanding or enabling repair, but it may also stand alone. Formal discipline may coexist with restorative elements when serious conduct requires adjudication and affected parties still seek repair. The relationship among pathways should be transparent so that students understand which elements are voluntary, which are required, and which rights remain available.
Distinguishing the approaches strengthens rather than weakens educational practice. It allows institutions to use coaching and rework without overstating them as restorative justice, and it protects restorative processes from being reduced to an administrative alternative to sanction. The distinction also supports evaluation because institutions can assess educational learning, restorative quality, procedural fairness, and disciplinary consistency as different outcomes.
The distinction should appear in policy and reporting. Institutions should label a response according to its actual components rather than using restorative terminology for any non-punitive outcome. This protects conceptual integrity and allows meaningful evaluation. It also helps students understand whether participation is voluntary, whether a formal finding exists, and what records or rights apply.
Training for case staff should include examples of educational, restorative, disciplinary, and combined processes. Staff need to recognize when a case lacks an identifiable harm or when dialogue would be unsafe. They also need skills in explaining options without pressure. Clear distinctions support appropriate referral and reduce the risk that restoration becomes a convenient substitute for investigation or adequate student support.
A policy taxonomy should state the purpose, entry conditions, rights, and expected outcomes of each response. This helps prevent informal educational meetings from becoming undocumented investigations and prevents restorative processes from being used where formal adjudication is required. Clear taxonomy also supports accurate institutional reporting.
The distinction has implications for staff roles. Educators may provide correction and feedback, while trained facilitators manage restorative dialogue and authorized bodies determine formal findings. Roles can overlap, but conflicts of interest and power should be considered. Separation or additional safeguards may be necessary when the same person teaches, investigates, and facilitates.
The distinction also protects proportionality. A low-impact misunderstanding may require education without the complexity of a facilitated restorative process. A serious case may require formal adjudication even when educational work is included. Using the correct category prevents over-processing minor concerns and under-processing serious ones. It also allows resources such as trained facilitation to be reserved for cases in which restorative mechanisms can genuinely operate.
For this reason, institutional forms and communications should use precise terminology consistently. Precision protects students, supports staff training, improves evaluation, and ensures that the ethical commitments associated with restorative justice are not diluted through routine administrative usage.

7.2. Selecting an Educational, Restorative, Disciplinary, or Combined Pathway

Response selection should begin with fact establishment rather than a presumption about the appropriate outcome. Decision-makers should identify the tool and function, the applicable rule, the learning outcome, the disclosure provided, and the reliability of the evidence. Automated indicators should not determine the pathway. Where facts remain uncertain, the process should preserve that uncertainty and avoid using a student’s refusal to accept an allegation as proof of misconduct.
The second stage examines conditions. Rule clarity, prior guidance, educational stage, assessment design, literacy support, access, workload, and institutional practice affect how responsibility should be interpreted. These conditions do not automatically excuse conduct, but they can mitigate responsibility and reveal institutional duties. A pathway is proportionate only when it reflects both individual action and the environment in which informed action was possible.
The third stage attributes responsibility by considering intention, degree of delegation, relation to the learning outcome, advantage, harm, repetition, evidence quality, and capacity to explain. Deliberate concealment of a central AI-generated solution differs from poorly disclosed language support under an ambiguous rule. Repeated conduct after clear guidance differs from a first misunderstanding. Professional or safety implications may increase the need for formal procedure even where educational needs remain.
An educational pathway is most appropriate where ambiguity, low impact, first occurrence, or weak judgement predominates. Possible outcomes include clarification, coaching, rework, a learning plan, or targeted literacy development. Educational responses should still be documented sufficiently for consistency and recurrence analysis. They should not be used to avoid addressing institutional failures or to create informal consequences without rights.
A restorative pathway is appropriate when identifiable harm or unfairness can be addressed, responsibility can be meaningfully discussed, participation is voluntary, facilitation is available, and repair is feasible. Intention may be mixed, but the facts must be sufficiently established for dialogue. Restoration is not suitable merely because it appears compassionate or administratively efficient. Eligibility and safeguards must be assessed case by case.
A disciplinary pathway remains necessary where conduct is deliberate, repeated, high-impact, professionally unsafe, or incompatible with voluntary restoration. Formal procedure protects both standards and student rights because evidence, reasons, representation, and appeal are defined. Discipline should be proportionate and may include educational conditions. The framework does not treat sanction as inherently punitive or illegitimate; it treats unreviewable, automated, or excessive sanction as problematic.
Combined pathways may serve several purposes. A formal finding and proportionate consequence may be accompanied by educational work or a voluntary restorative agreement. The relationship should be explicit. Participation in restoration should not be required to obtain an appeal or to reduce a sanction unless the policy clearly and fairly defines the role of acknowledged responsibility. Combined approaches should avoid double punishment and should coordinate records, timelines, and support.
Every pathway should end with reasons, appeal protection, and feedback. Decision-makers should explain why the selected response fits the facts and factors. Institutions should identify whether policy, assessment, literacy, access, or technology requires change. This final step connects individual accountability to prevention and ensures that similar cases do not recur under unchanged conditions (Bretag et al., 2019; Bussu & Karp, 2026; Cotton et al., 2024; Evans & Vaandering, 2022; Karp, 2019a, 2019b; Wachtel, 2016; Zehr, 2015).
Consistency should be supported through structured documentation and moderation. A matrix cannot replace judgement, but it can ensure that relevant factors are considered and reasons are recorded. Periodic review of anonymized decisions can identify unexplained variation across programmes or groups. Where variation reflects legitimate disciplinary differences, those differences should be stated in advance. Where it reflects bias or inconsistent practice, corrective action is required.
Pathway selection should also consider timing. Delayed decisions can increase anxiety, disrupt progression, and make repair less meaningful. Institutions need realistic timelines, communication during delay, and priority rules for high-stakes cases. Efficiency, however, should not be achieved by relying on automated proof or pressuring students into informal resolution. Timeliness is one component of procedural fairness, not a reason to reduce safeguards.
Decision-makers should also consider the interests of affected communities. A case may have implications for group assessment, professional standards, peer workload, or public trust. These interests should be represented without exaggerating harm or turning the process into symbolic punishment. The response should remain connected to established facts and educational purpose.
Review data can improve the pathway matrix over time. Institutions may identify factors associated with recurrence, successful repair, appeal, or disproportionate impact. Adjustments should be transparent and should preserve minimum safeguards. The matrix is therefore a living decision aid, not a fixed tariff of consequences.
Students should receive a clear explanation of available pathways before a decision is finalized where policy permits choice or participation. They should understand which route determines responsibility, which elements are voluntary, what records are created, and how appeal operates. Transparent explanation reduces strategic uncertainty and protects against the perception that less formal pathways require surrender of rights. It also supports informed restorative consent.
In Figure 10, the response-selection pathway begins with fact establishment, contextual analysis, and proportional attribution before branching into educational, restorative, and disciplinary or combined responses. The final stage reconnects each case to reason-giving, appeal, and institutional learning.
In Table 14, educational, restorative, and disciplinary or combined responses are compared across rule clarity, intention, impact, repetition, participation, and outcome. The table links the decision pathway to explicit proportionality criteria and procedural safeguards.

7.3. Restorative Safeguards and Institutional Learning

Eligibility screening protects participants and the integrity of the process. Trained staff should determine whether facts are sufficiently established, whether the case is safe for participation, whether affected parties can be identified appropriately, and whether legal or professional duties require another procedure. Automated evidence should never create a presumption of eligibility or responsibility. Seriously disputed facts, acute safety risks, or unmanaged power imbalances may make restoration unsuitable.
Informed consent should explain the purpose, stages, possible outcomes, confidentiality, records, support, and relationship to formal rights. Participants should be able to decline without adverse inference. Consent is not meaningful when a grade, sanction, or appeal depends on participation in an undisclosed way. Institutions should provide time for advice and should accommodate disability, language, and cultural needs. A support person may be appropriate.
Facilitated dialogue should focus on conduct, effects, needs, responsibilities, and future practice. The facilitator should manage power differences and prevent humiliation, coercion, or token participation. Dialogue does not require direct confrontation in every case. Shuttle processes, written statements, or representative participation may be safer or more appropriate. The form should serve the restorative purpose rather than a fixed script.
Responsibility and harm should be distinguished from intention. A person may not have intended a particular effect but can still acknowledge it and participate in repair. Institutional contribution should also be discussable. Ambiguous rules, inaccessible support, invalid assessment, or inappropriate technology can be part of the account. This does not remove individual responsibility; it creates a more accurate basis for agreement and institutional learning.
The repair agreement should specify actions, timelines, support, monitoring, privacy, and completion. Actions must be proportionate to the established harm and connected to education or restoration. Public apology, unpaid work, or disclosure of private details may be inappropriate or coercive. The agreement should identify what happens if circumstances change or completion becomes impossible. Formal review should be available where disputes arise.
Follow-up and reintegration require deliberate planning. Participants should know how completion is recorded, how recurrence is handled, and when information is removed or restricted. The institution should avoid permanent stigma and uncontrolled disclosure. Support may be needed to restore academic confidence, relationships, or access to learning. Reintegration is not an automatic consequence of completing a task; it is a governance responsibility.
Institutional learning should use anonymized patterns rather than expose individual participants. Cases may reveal recurring ambiguity, assessment weaknesses, unequal access, gaps in literacy, reviewer inconsistency, or technology problems. A quality or integrity committee should assign actions, owners, timelines, and closure criteria. Restorative processes contribute to prevention only when these lessons reach policy, curriculum, assessment, and support.
Evaluation should distinguish participation, completion, perceived fairness, repair, recurrence, equity, and institutional change. Satisfaction alone does not prove repair or deterrence. Low recurrence may reflect selection of low-risk cases. Voluntary participation creates self-selection. Comparative and qualitative research is needed to understand who benefits, who declines, and under what conditions. These safeguards preserve the substantive meaning of restorative justice and prevent it from becoming a symbolic label (Bussu & Karp, 2026; Evans & Vaandering, 2022; Karp, 2019a, 2019b; Wachtel, 2016; Zehr, 2015).
Facilitator competence should be treated as an institutional resource rather than assumed personal skill. Training should address restorative principles, academic policy, power, culture, disability, confidentiality, records, and referral to formal procedures. Supervision and peer review can support quality. Institutions should also monitor facilitator workload and avoid creating incentives to maximize restorative case numbers.
The feedback generated by restorative cases should be separated from confidential personal information. Institutions can record themes such as ambiguous rules, workload, access, assessment design, or technology without exposing identities. Governance committees should receive aggregated findings and document action. This balance allows restorative knowledge to inform prevention while protecting trust and reintegration.
Records require particular care. Institutions should define what is confidential, what is included in formal files, who can access it, and when it is reviewed or removed. Record practices should support accountability and professional obligations without undermining reintegration. Participants should receive this information before consent.
Restorative quality assurance can include case review, facilitator supervision, participant feedback, and equity analysis. However, review should not expose confidential dialogue unnecessarily. Aggregated themes and procedural indicators can support learning while protecting participants. The institution must balance transparency about the programme with privacy within individual cases.
Institutional learning should not be used to extract additional disclosure from participants beyond what is necessary for the restorative process. Aggregation and anonymization must be designed from the start. Where a lesson cannot be reported safely because the group is small or the circumstances are distinctive, privacy should prevail. The feedback mechanism depends on trust, and trust would be weakened if restorative participation became a source of identifiable institutional surveillance.
Where restorative processes are unavailable, institutions should not simulate them through untrained informal meetings. Educational support or formal procedure may be more appropriate until facilitation, consent, record, and safeguarding capacity can be provided responsibly.
In Figure 11, a substantive restorative process is presented through eligibility screening, informed consent, facilitated dialogue, repair agreement, follow-up and reintegration, and anonymized institutional learning. The sequence distinguishes restorative justice from generic reflection, retraining, or task revision.
In Table 15, the requirements of a restorative process are connected to operational questions, minimum safeguards, and unsuitable or high-risk conditions. This structure supports case screening and protects voluntariness, due process, privacy, proportionality, and participant safety.

8. Integrated Preventive–Literacy–Restorative Framework

8.1. Mechanisms of Integration

The revised framework integrates its components through specified mechanisms rather than visual proximity. Preventive governance clarifies expectations and aligns assessment. Ethical AI literacy enables students and staff to interpret those expectations and evaluate outputs. Governed verification supports fact-finding without converting probabilistic indicators into proof. Proportional case assessment selects an educational, restorative, disciplinary, or combined response. Case review closes the loop by converting ambiguity, incentives, assessment vulnerabilities, access inequities, and procedural failures into institutional change.
The framework integrates components through mechanisms rather than by placing them next to one another. Preventive governance creates clear expectations, aligns AI functions with learning outcomes, and establishes due process protections. Ethical AI literacy enables students and staff to interpret expectations, evaluate outputs, protect data, and exercise judgement. Governed verification supports proportionate fact-finding. Response selection connects evidence and responsibility to educational, restorative, disciplinary, or combined pathways. Case review returns lessons to governance.
The first mechanism is clarity. Institutional principles must reach programmes, courses, and assessed tasks. Students need to know which functions are permitted, what must be disclosed, what evidence may be requested, and why the rule relates to the learning outcome. Clarity reduces ambiguity only when guidance is accessible, coherent, and taught. Publication alone is insufficient. The corresponding boundary condition is strategic misconduct: a clear rule may not deter deliberate advantage-seeking.
The second mechanism is capable judgement. AI literacy enables users to recognize limitations, verify claims, disclose contribution, and interpret rules. Organisational literacy enables educators and institutions to design assessment, evaluate tools, review evidence, and provide due process. Literacy does not guarantee conduct because motivation, pressure, opportunity, authority, and resources also matter. It functions as an enabling condition within a broader system.
The third mechanism is evidential validity. AI-resilient assessment combines outcome alignment, transparent contribution, and complementary evidence. Governed verification uses targeted, corroborated, and reviewable methods rather than automated proof. These arrangements increase the interpretability of learning evidence only when burdens remain proportionate and accessible. Excessive process collection or oral testing can reduce fairness and reliability.
The fourth mechanism is differentiated responsibility. Actor-specific duties make institutional and provider contributions visible without erasing student agency. Proportional attribution considers rule clarity, intention, delegation, advantage, harm, repetition, and evidence. This mechanism supports response selection because consequences are connected to what each actor controlled and knew. Vague shared responsibility would instead diffuse accountability.
The fifth mechanism is restorative and disciplinary complementarity. Restoration addresses harm, participation, responsibility, repair, and reintegration where conditions permit. Discipline protects standards and rights where formal adjudication is required. Educational correction addresses misunderstanding and competence. The framework does not assume that one pathway is ethically superior in every case. Proportionality and due process determine the relationship.
The sixth mechanism is case-to-governance feedback. Cases reveal how policy, assessment, literacy, access, technology, and incentives operate in practice. Anonymized patterns should lead to documented revisions and implementation. Feedback fails when case data are incomplete, biased, or interpreted only as evidence of student behaviour. Institutional learning requires ownership, resources, timelines, and audit.
Contextual constraints surround every mechanism. Law, culture, discipline, access, resources, and decision stakes affect implementation. Stable principles remain, but procedures must adapt. The integrated model therefore presents a conditional architecture. Its arrows represent hypotheses and institutional relationships to be tested, not proven causal effects (Chan, 2023; Chiu et al., 2024; European Commission High-Level Expert Group on Artificial Intelligence, 2019; Holmes et al., 2022; Karp, 2019b; National Institute of Standards and Technology, 2023, 2024; Ng et al., 2021; UNESCO, 2023).
The sequence of mechanisms is not strictly linear. Policy influences literacy and assessment, but literacy can also reveal policy ambiguity. Verification can identify assessment weaknesses, while restorative dialogue can reveal access or workload problems. Feedback therefore operates across the model. The visual simplification should be read as an architecture of relationships rather than a fixed chronology.
Implementation can begin at different points depending on institutional readiness, but the components should eventually connect. A university may start with task guidance, assessment review, or appeal reform. The risk is fragmentation if these initiatives develop separately. Governance should therefore identify interfaces, shared terminology, responsible owners, and data flows. Integration is an organisational achievement, not simply a conceptual statement.
Governance should identify failure points between mechanisms. Clear rules may not reach tasks; literacy may not be assessed; verification may not inform policy; restorative lessons may remain confidential without aggregation; procurement may be disconnected from academic quality. Mapping these interfaces can reveal why an apparently comprehensive system fails in practice.
The framework also implies reciprocal accountability. Students are accountable for use and claims, while institutions are accountable for the conditions and procedures through which responsibility is judged. This reciprocity supports trust because expectations apply to all actors in forms appropriate to their authority and control.
The model can also support governance diagnostics. When an outcome is unsatisfactory, institutions can examine which mechanism failed: clarity, capacity, evidence, attribution, response selection, or feedback. This is more useful than attributing failure to a general lack of integrity. Mechanism-based diagnosis identifies an actionable point of intervention and creates a hypothesis for evaluation. Over time, this process can refine both the framework and local practice.
In Figure 12, the article’s central contribution is consolidated by linking preventive governance, ethical AI literacy, and restorative accountability through an integrated mechanism and a case-to-governance feedback loop. The diagram brings together the preceding sections and shows how clarity, capable judgement, differentiated responsibility, proportionate response, and anonymized case learning operate under contextual constraints.

8.2. Conceptual Propositions

The seven propositions translate the integrated framework into claims that can be examined rather than accepted as general principles. Each proposition identifies an expected relationship, the mechanism through which it may occur, a boundary condition, and an illustrative indicator. This structure is necessary because a conceptual model becomes scientifically useful when it can be qualified or contradicted by evidence. The propositions do not convert normative commitments into empirical facts; they specify what future research would need to observe.
P1 proposes that task-level policy clarity reduces ambiguity-related breaches. The mechanism is interpretability: students can connect general principles to permitted functions and disclosure requirements. The proposition may fail when guidance is inaccessible, contradictory, poorly taught, or strategically ignored. Evaluation should therefore measure both the presence of guidance and comprehension, while distinguishing ambiguity-related cases from deliberate conduct. Changes in reporting and enforcement must also be considered.
P2 proposes that AI literacy supports responsible judgement but does not guarantee compliant conduct. The mechanism is improved capacity for verification, disclosure, error recognition, and ethical interpretation. Pressure, competition, perceived unfairness, and advantage may override knowledge. Research should therefore combine literacy measures with process evidence, disclosure behaviour, and contextual data. A finding that knowledge and conduct diverge would refine rather than invalidate the broader framework.
P3 proposes that AI-resilient assessment improves evidential validity when verification is proportionate and accessible. Complementary evidence can make contribution and understanding more interpretable. The boundary condition is burden: additional documentation, oral assessment, or supervision may reduce fairness, reliability, or authenticity. Studies should examine validity, accessibility, staff and student workload, moderation, and differential effects rather than use misconduct counts as the only outcome.
P4 proposes that governed verification is more legitimate than automated suspicion. Corroboration, competent human review, reasons, and appeal are expected to reduce error and procedural harm. The proposition may fail when reviewers exhibit automation bias, lack authority, or operate under time pressure. Relevant indicators include overrides, appeal outcomes, reversal reasons, differential error, and participant understanding. Legitimacy also requires qualitative evidence about whether the procedure was intelligible and contestable.
P5 proposes that differentiated responsibility improves prevention when duties remain attributable. The mechanism is visibility: institutional, educator, student, and provider contributions can be addressed without dissolving agency. The boundary condition is vague collectivism, in which everyone is responsible and therefore no one acts. Evaluation should examine role clarity, compliance, reason-giving, action closure, and consistency of attribution across cases.
P6 proposes that restorative processes support learning and repair when substantive safeguards are present. Dialogue and agreed action may address relational consequences and reintegration. Coercion, serious harm, disputed facts, unsafe participation, or inappropriate case selection may make restoration ineffective or unjust. Research should distinguish generic education from genuine restoration and assess consent, facilitation, repair, participant experience, recurrence, equity, and records.
P7 proposes that case-to-governance feedback produces institutional learning when patterns lead to documented and implemented change. Cases can reveal ambiguity, assessment weakness, access inequity, or procedural failure. Data may be incomplete or biased, and documentation may be symbolic. Evaluation should trace the chain from case analysis to decision, ownership, implementation, and outcome. This proposition represents the framework’s central institutional contribution because it connects response to prevention.
The propositions are interdependent. Policy clarity without literacy may not be understood. Literacy without valid assessment may not preserve evidence of learning. Verification without procedural safeguards may create harm. Restoration without eligibility may become coercive. Feedback without resources may remain symbolic. Empirical work should therefore test both individual relationships and the coherence of the wider system.
The propositions should be interpreted as a research programme rather than a prediction that every institution will observe the same effects. Context may alter the strength, direction, or meaning of a relationship. A clarity intervention may improve comprehension without changing cases because incentives remain strong. A restorative process may support learning but not reduce recurrence. Such findings would help refine mechanisms and scope conditions.
Testing should also consider unintended effects. Task-level rules may become excessively complex. Process evidence may increase surveillance. Appeals may be formally available but burdensome. Feedback systems may encourage managerial counting rather than learning. Empirical studies should therefore include adverse outcomes and participant experience. A proposition is strengthened not by confirming only its preferred effect but by accurately identifying when and why it fails.
Proposition testing should use outcome definitions that reflect the mechanism. P1 concerns ambiguity-related concerns, not all misconduct. P3 concerns evidential validity and fairness, not simply lower detection. P7 concerns implemented institutional change, not the existence of a review meeting. Precise outcomes reduce the risk of apparently confirming a proposition through an unrelated metric.
Researchers should also examine interactions among propositions. Literacy may strengthen the effect of clarity, while accessibility burden may weaken the effect of assessment redesign. Multilevel and realist approaches may be useful because institutional context shapes individual behaviour and procedure. The proposition matrix can therefore support both focused studies and system-level evaluation. Appendix F provides an empirical test matrix for P1–P7, including suggested designs, primary data, outcomes or mechanisms, and key threats to inference.
In Table 16, the integrated framework is translated into seven provisional propositions that connect mechanisms, boundary conditions or counter-cases, and illustrative empirical indicators. The propositions provide the bridge from conceptual synthesis to future testing without presenting the framework as an already validated intervention.

8.3. Conflicts, Stable Principles, and Contextual Adaptation

The framework does not assume harmony. Prevention can become surveillance; transparency can conflict with privacy; process evidence can create unequal workload; restoration can conflict with consistency or public accountability; and distributed responsibility can obscure individual attribution. Institutions should explain how conflicts were balanced, why less intrusive alternatives were insufficient, and how affected students can challenge decisions.
The framework contains unavoidable tensions. Prevention can become surveillance when institutions collect extensive process data or monitor students continuously. Transparency can conflict with privacy when disclosure requirements expose sensitive interactions or personal data. Assessment security can conflict with accessibility when oral, supervised, or process-heavy methods create unequal burdens. Restoration can conflict with consistency or public accountability when cases are handled privately. Distributed responsibility can obscure individual attribution. A credible framework must make these conflicts visible.
Conflict management begins with reason-giving. Institutions should identify the principles in tension, the stakes, the affected groups, the evidence, and the alternatives considered. They should explain why a less intrusive or burdensome option was insufficient. The decision should be reviewable. This process does not guarantee agreement, but it makes ethical judgement accountable and creates evidence for future policy review.
Stable principles include human answerability for consequential academic decisions, intelligible rules, valid evidence, fairness, privacy protection, non-discrimination, substantive human review, written reasons, contestability, and non-coercive restorative participation. These commitments should not disappear because local resources are limited or cultural practices differ. They define the minimum ethical identity of the framework (European Commission High-Level Expert Group on Artificial Intelligence, 2019; European Union, 2024; Floridi & Cowls, 2019; International Center for Academic Integrity, 2021; Jobin et al., 2019; Karp, 2019a, 2019b; McArthur, 2016; Mittelstadt et al., 2016; National Institute of Standards and Technology, 2023, 2024; UNESCO, 2021; Zehr, 2015).
Adaptable elements include disclosure formats, committee structures, timelines, standards of proof, representation, record retention, facilitation methods, sanctions, and resource models. Disciplines may define authorship and independent performance differently. Professional accreditation may require formal reporting. Legal systems may impose specific data or appeal obligations. Cultural expectations may influence dialogue, authority, and repair. Adaptation should be documented rather than assumed.
Accessibility is both a stable commitment and an adaptable practice. Institutions must avoid unjustified disadvantage, but appropriate accommodations depend on the task and student. Oral verification may require alternative formats, additional time, structured questions, or support. Process evidence may require accessible technologies. Privacy protections may affect what records can be collected. Accessibility services should participate in policy and assessment design rather than only respond after a concern arises.
Resource constraints require staged implementation. Institutions may not immediately establish every element, but minimum safeguards should come first: no automated proof, clear rules, human reasons, privacy, and appeal. Curriculum mapping, staff development, restorative facilitation, procurement audit, and advanced indicators can follow. Resource limitation should influence sequencing, not justify procedurally illegitimate decisions.
Cultural adaptation should avoid both universalism and relativism. Concepts of individual authorship, collaboration, teacher authority, apology, trust, and repair vary. Cross-cultural consultation can improve implementation. However, local custom should not justify coercion, opaque evidence, discrimination, or removal of appeal. The framework treats context as a condition for thoughtful application, not as a reason to abandon stable rights.
The framework should also be periodically revised. Technology, regulation, educational practice, and social expectations change. Stable principles provide continuity, while adaptable elements respond to evidence. Case feedback, appeals, student and staff experience, provider changes, and research findings should inform review. Continuous adaptation is therefore part of integrity rather than evidence that policy has failed.
Institutions can use structured ethical deliberation to manage conflicts. A decision record may identify the purpose, affected rights, evidence, alternatives, burden, accommodations, review route, and planned evaluation. This approach creates consistency without pretending that every tension has a formulaic solution. It also enables later audit when outcomes reveal unexpected harm.
Adaptation should be participatory. Students, educators, professional bodies, accessibility specialists, legal and data-protection experts, and cultural stakeholders may identify different risks. Consultation should occur before implementation and during review. Participation does not remove institutional responsibility for final decisions, but it improves legitimacy and reveals assumptions that a narrow governance group may overlook.
Institutions should publish the rationale for major adaptations where possible. Transparency about disciplinary or legal differences helps students understand variation and enables external scrutiny. Confidential case details need not be disclosed, but the principles and procedures should be intelligible.
Periodic review should include a sunset or reconsideration mechanism for intrusive measures. Tools, monitoring practices, and temporary restrictions should not continue indefinitely without evidence. This safeguard is especially important in rapidly changing technological environments where an emergency response can become normalized after its original justification has disappeared.
Local adaptation should be recorded in a framework-specific implementation statement. Such a statement can identify stable principles, contextual decisions, responsible actors, evidence requirements, accommodations, and review dates. This documentation makes adaptation transparent to students and reviewers and supports comparison across programmes. It also prevents informal variation from becoming an unexamined source of unequal treatment.
Adaptation decisions should also be reversible. Institutions should specify review dates and conditions for modification, especially where a local procedure creates new burdens or relies on emerging technology. Reversibility supports responsible experimentation while protecting stable commitments.
In Table 17, stable ethical commitments are distinguished from elements requiring local legal, cultural, disciplinary, procedural, and resource adaptation. This comparison operationalizes contextual sensitivity while preserving minimum commitments to human responsibility, fairness, reason-giving, contestability, and non-coercive participation.

9. Institutional Application and Evaluation

9.1. Responsibilities, Resources, Timelines, and Indicators

Implementation should begin with minimum procedural protections rather than technology procurement. Institutions first need a policy hierarchy, task-guidance templates, disclosure and privacy rules, limits on automated indicators, competent human review, and an accessible appeal route. Programme mapping, staff development, student learning, restorative facilitation, procurement audit, and quality indicators can then be phased according to capacity.
Implementation should begin with governance rather than procurement. The first three months should establish minimum safeguards: a defined scope, a prohibition on treating automated indicators as proof, task-level guidance templates, privacy rules, competent human review, written reasons, and an accessible appeal route. Senior leadership, integrity staff, legal and data-protection functions, educators, accessibility services, and student representatives should share design responsibility while retaining named ownership for each action.
Minimum safeguards require resources. Policy drafting needs legal, educational, and technical expertise. Human review requires trained staff and time. Appeals require independence and administrative support. Accessible guidance requires design and testing. Institutions should therefore identify resource implications explicitly rather than presenting new duties as cost-free. Unfunded expectations are likely to produce nominal compliance, inconsistent judgement, or reliance on automated shortcuts.
The three-to-nine-month stage should focus on curriculum and assessment alignment. Programme teams can map learning outcomes, AI-literacy progression, assessment vulnerability, and professional obligations. Educators and learning designers can revise briefs and rubrics, develop examples, and identify where complementary evidence is needed. Accessibility services should review proposed assessment changes. Staff and student development should use disciplinary cases rather than generic demonstrations.
The nine-to-eighteen-month stage can establish restorative and technology infrastructure. Institutions need eligibility criteria, trained facilitators, consent and record protocols, a tool register, procurement standards, vendor review, and anonymized case coding. Quality units should establish baseline indicators for appeal, differential impact, accessibility, workload, policy comprehension, and case outcomes. This stage requires coordination because restorative, technical, legal, and educational expertise often sits in separate units.
Ongoing assurance should review appeals, reversals, differential outcomes, recurrence, workload, comprehension, model changes, and action closure. Annual public reporting can support accountability, but privacy and interpretive caution are necessary. Institutions should explain what indicators mean and what they cannot establish. A decline in cases is not proof of success, and high participation in training is not proof of competence or conduct.
Named actors improve accountability. Governing bodies approve policy and allocate resources. Programme leaders align curriculum and assessment. Educators communicate task rules and judge evidence. Integrity units support consistency and case handling. Quality teams analyze patterns and implementation. Procurement and IT functions evaluate providers and systems. Students contribute to design, feedback, and review. Providers supply documentation and support. Each output should have an owner, timeline, and evidence of completion.
Indicators should be linked to mechanisms. Task-guidance coverage and comprehension relate to clarity. Literacy participation and assessment relate to capacity. Accessibility review, moderation, and workload relate to AI-resilient assessment. Overrides, reasons, appeals, and differential error relate to governed verification. Consent, agreement completion, repair, and recurrence relate to restoration. Documented policy and assessment changes relate to institutional learning. This alignment prevents data collection from becoming detached from the framework.
Implementation should be iterative. Early experience may reveal that guidance is too complex, staff capacity is insufficient, evidence requirements are burdensome, or appeal routes are unclear. Institutions should treat these findings as information rather than failure. The roadmap supports staged improvement while preserving minimum rights. Its success depends on implementation and review, not the production of documents alone (Chan, 2023; European Union, 2024; Holmes et al., 2022; National Institute of Standards and Technology, 2023, 2024; OECD, 2023, 2026; UNESCO, 2021, 2023, 2024a, 2024b; UNESCO International Institute for Higher Education in Latin America and the Caribbean, 2023).
Risk-based sequencing can help institutions allocate resources. High-stakes decisions involving automated evidence, professional progression, or serious sanctions require early procedural protection. Lower-stakes curriculum and guidance improvements can be piloted and refined. The framework does not support postponing rights until a complete system is affordable. Minimum safeguards should apply from the beginning, while more advanced infrastructure develops over time.
Implementation governance should include escalation and dependency management. A programme cannot revise assessment effectively if institutional disclosure rules remain unclear, and an appeal body cannot review automated evidence without technical documentation. The roadmap should therefore identify dependencies, decision points, and barriers. Regular implementation meetings and transparent action logs can prevent initiatives from becoming isolated or delayed without explanation.
Leadership should communicate that implementation is a quality and educational project, not only an integrity-office initiative. Assessment, curriculum, student support, accessibility, data governance, and procurement all contribute. This framing can improve cooperation and reduce the tendency to assign the problem to a single unit.
Resource planning should include maintenance. Policies, training, tool reviews, and accessibility arrangements require recurring effort. A one-time project budget is insufficient. Institutions should identify sustainable staffing, review cycles, and escalation routes. Where resources are limited, priorities and limitations should be communicated honestly.
Student participation should be resourced rather than assumed. Consultation requires accessible materials, sufficient time, representative recruitment, feedback on how input was used, and protection from tokenism. Student perspectives can identify confusing guidance, hidden workload, or barriers to appeal that staff may not see. Meaningful participation improves implementation quality while institutional leaders retain responsibility for final governance decisions.
In Table 18, the framework is translated into a staged implementation and evaluation roadmap. The table identifies named actors, priority actions and resources, outputs and indicators, and interpretive cautions, thereby connecting the conceptual model to realistic institutional sequencing and accountability. Adaptable task-level AI guidance and case-triage instruments are provided in Appendix D.

9.2. Evaluation Principles

Evaluation should combine implementation, procedural, educational, and equity indicators. Relevant measures include task-specific guidance coverage, comprehension, access to literacy support, documented human review, appeal availability and outcomes, accessibility adjustments, differential impacts, restorative consent and completion, recurrence patterns, workload, and the proportion of case analyses that produce implemented institutional change. Lower case counts can reflect prevention, under-reporting, inconsistent enforcement, or reduced trust and must be interpreted cautiously.
Evaluation should distinguish implementation, process, educational outcome, procedural legitimacy, equity, and institutional learning. A policy may be published but not understood. Training may be delivered but not change judgement. A restorative process may be completed but not repair harm. An appeal route may exist but remain inaccessible. Each domain therefore requires indicators that match the mechanism being evaluated.
Implementation indicators include the proportion of assessments with task-specific guidance, the availability of disclosure templates, staff development, student support, accessibility review, tool registers, trained facilitators, and named appeal routes. These measures show whether infrastructure exists. They do not establish quality or effectiveness. Audits should examine whether outputs are current, accessible, used, and supported by resources.
Procedural indicators include documented human review, written reasons, decision time, appeal availability, appeal outcomes, reversals, and differential error. Low appeal rates are ambiguous because they may reflect trust, lack of awareness, fear, cost, or barriers. Reversal rates may reflect correction or inconsistent first decisions. Qualitative analysis of reasons and participant experience is needed to interpret quantitative patterns (European Union, 2024; National Institute of Standards and Technology, 2023, 2024).
Educational indicators include policy comprehension, AI-literacy competence, quality of verification, disclosure practice, evaluative judgement, and ability to defend learning. Self-report should be combined with tasks, scenarios, process evidence, or observed practice. Measures should distinguish knowledge from conduct and should not assume that confidence equals competence. Longitudinal data can show progression across a programme.
Assessment indicators include validity evidence, moderation, accessibility, workload, reliability, and differential outcomes. A reduction in suspected cases may result from improved assessment or lower detection. More process evidence may improve interpretability while increasing burden. Evaluation should therefore examine the balance among evidence quality, student experience, staff workload, and equity rather than optimize a single metric.
Restorative indicators should include voluntary and informed participation, facilitation quality, agreement completion, perceived fairness, repair, reintegration, recurrence, and institutional learning. Satisfaction alone is insufficient, and selection effects must be considered because unsuitable or unwilling cases are excluded. Qualitative accounts from students, affected parties, and facilitators can identify coercion, power imbalance, or symbolic repair.
Equity indicators should examine differential effects by language, disability, programme, access to tools, and other relevant factors while protecting privacy and avoiding stigmatization. Institutions should be cautious about small groups and sensitive data. Equity analysis should inform redesign, accommodation, training, and procurement rather than merely report disparities.
Institutional-learning indicators should trace whether case patterns lead to decisions, assigned actions, implementation, and review. A documented policy change may be symbolic if teaching and assessment remain unchanged. Action-closure rates, follow-up audits, and stakeholder feedback can provide stronger evidence. Evaluation should therefore treat the case-to-governance loop as a process that must be observed over time.
No single indicator demonstrates integrity. Lower case counts, higher disclosure, frequent appeals, or strong literacy scores each have multiple interpretations. A balanced evaluation framework combines data, reasons, context, and comparison. Its purpose is improvement and accountability, not the production of a simplistic institutional ranking (Bretag et al., 2019; Corbin et al., 2025; Cotton et al., 2024; Dawson et al., 2024; International Center for Academic Integrity, 2021).
Evaluation governance should specify who collects, interprets, and acts on data. The same unit that implements a tool or policy should not be the only source of assurance. Independent quality review, student participation, and periodic external input can reduce confirmation bias. Data access and privacy rules should be defined, especially where case information is sensitive.
Benchmarks should be used cautiously. Cross-institutional comparison may be distorted by different definitions, reporting cultures, assessment practices, and student populations. Internal trend analysis can also mislead when policy or detection changes. Evaluation should prioritize explanatory evidence and local improvement. Where comparative reporting is required, methods and limitations should be transparent.
Evaluation findings should feed directly into governance. Reports should identify decisions, owners, timelines, and follow-up rather than end with descriptive data. Stakeholders should be informed of material changes, and unresolved risks should remain visible. This closes the assurance loop and makes evaluation consequential.
The framework also supports developmental evaluation during implementation. Early feedback can identify unintended burden, confusion, or inequity before practices are institutionalized. Developmental methods are appropriate where technology and policy change rapidly, provided that minimum rights are protected throughout experimentation.
Evaluation should include a schedule for reconsidering indicators themselves. Measures that were useful at launch may become distorted by changing practice or incentives. Staff and students may adapt behaviour to what is counted, and data collection can create burden. Periodic indicator review should examine validity, equity, privacy, cost, and whether each measure still informs a decision. Unused or misleading indicators should be removed.
Findings should be communicated in forms appropriate to different audiences. Governing bodies need assurance and risk information, educators need actionable feedback, students need intelligible explanations, and researchers need methodological detail. Tailored communication can improve use without altering the underlying evidence.

10. Discussion

10.1. Answers to the Research Questions

For RQ1, the synthesis identifies continuity in integrity values but change in the scale of content generation, cognitive delegation, traceability, and algorithmic mediation. The three-level model prevents academic-work questions from being conflated with assessment validity or institutional procedure. For RQ2, the analysis identifies task-level clarity, role-specific AI literacy, governed verification, differentiated responsibility, and case feedback as connecting mechanisms rather than independent recommendations.
For RQ3, the proportional pathway rejects a universal response. Educational correction fits low-impact ambiguity; restorative justice requires identifiable harm, voluntary participation, responsibility, repair, and safeguards; formal discipline remains necessary for deliberate, repeated, serious, unsafe, or otherwise unsuitable cases; and combined pathways may be appropriate when their relationship is transparent and reviewable. For RQ4, stable principles include human responsibility, validity, fairness, reason-giving, contestability, privacy, and non-coercive restoration, while disclosure formats, sanctions, authority structures, records, and resource models require local adaptation.
RQ1 asked how generative AI modifies academic integrity at the levels of academic work, assessment, and institutional procedure. The synthesis identifies continuity in the values and educational foundations of integrity, but change in the scale of content generation, cognitive delegation, weak provenance, and algorithmic mediation. Academic work now requires functional analysis of contribution and epistemic responsibility. Assessment must remain capable of producing valid evidence under realistic AI availability. Institutional procedure must govern automated indicators, human review, reasons, privacy, and appeal.
The answer to RQ1 also clarifies what has not changed. Students remain responsible for submitted claims and must demonstrate the intended learning. Institutions remain responsible for fair assessment and procedure. Authorship still involves accountable intellectual contribution. Technology does not possess moral agency. The innovation lies in how these principles are operationalized when tools can perform functions that were previously treated as evidence of human cognition (Bertram Gallant, 2008; Cotton et al., 2024; Holmes et al., 2022; International Center for Academic Integrity, 2021; Kasneci et al., 2023; Macfarlane et al., 2014; Perkins, 2023).
RQ2 asked which duties and mechanisms connect preventive governance, ethical AI literacy, governed verification, and restorative accountability. The synthesis identifies task-level clarity, capable judgement, evidential validity, differentiated responsibility, proportional response selection, and case-to-governance feedback. These mechanisms require role-specific duties. Students disclose and verify; educators design and communicate; programmes coordinate; institutions protect rights and provide resources; integrity units review patterns; providers document systems.
The answer to RQ2 rejects a simple additive model. Policy does not work merely because it exists. Literacy does not guarantee conduct. Assessment redesign does not eliminate outsourcing. Human review does not protect rights unless it is competent and independent. Restoration does not occur through reflection alone. Each component operates under boundary conditions and depends on the others. The integrated framework therefore represents a system of conditional relationships rather than a checklist.
RQ3 asked when educational, restorative, disciplinary, or combined responses should be used. The synthesis supports educational correction where ambiguity, low impact, first occurrence, or weak judgement predominate. Restoration requires identifiable harm, sufficiently established facts, voluntary participation, facilitation, responsibility, repair, and reintegration. Discipline remains necessary for deliberate, repeated, serious, unsafe, or otherwise unsuitable cases. Combined pathways may be justified when formal accountability and repair serve distinct purposes.
The answer to RQ3 is procedural as well as substantive. Response selection should follow fact establishment, contextual analysis, proportional attribution, reason-giving, and appeal. Automated indicators cannot select the pathway. Institutional contribution should be acknowledged and corrected. Refusal of restoration should not aggravate the case. These safeguards connect fairness to educational purpose.
RQ4 asked which principles should remain stable and which elements require adaptation. Stable principles include human answerability, clarity, validity, fairness, privacy, non-discrimination, substantive review, reasons, contestability, and non-coercive restoration. Adaptable elements include disclosure formats, assessment methods, committee structures, timelines, record rules, sanctions, facilitation, and resources. Adaptation should reflect discipline, law, culture, accessibility, and capacity without removing minimum rights.
Together, the answers define the contribution and its limits. The framework clarifies concepts and proposes relationships, but it does not demonstrate effectiveness. The seven propositions require empirical testing. Cross-cultural and multi-institutional research is necessary. The research questions have therefore produced both a governance model and a structured agenda for evaluating it.
The answers are mutually reinforcing. The three-level scope from RQ1 defines where duties and mechanisms from RQ2 operate. Those mechanisms provide the criteria used in RQ3 to select a response. The stable and adaptable elements from RQ4 determine how the framework can be implemented across contexts. Treating the questions as separate findings would miss this architecture.
The synthesis also identifies unresolved questions. The relative influence of clarity, literacy, assessment design, incentives, and accountability is unknown. The fairness and effectiveness of pathway selection require study. Cultural and disciplinary variation may alter the framework. These uncertainties do not negate the contribution; they define the limits of current knowledge and the empirical work needed to evaluate the proposed relationships.
The responses also demonstrate the value of a critical conceptual synthesis. No single source tradition could answer all four questions. Integrity scholarship defines values, assessment research addresses evidence, AI governance addresses systems and rights, literacy research addresses capacity, and restorative justice addresses harm and repair. Their integration produces a more complete but still provisional account.
The answers should guide revision of the abstract, implications, and conclusions. Each claim should correspond to a research question and remain calibrated to the conceptual design. This alignment reduces repetition and prevents practical recommendations from extending beyond the framework’s evidential basis.

10.2. Theoretical Contribution and Comparison with Existing Frameworks

The originality of the framework does not lie in claiming that prevention, AI literacy, human oversight, or restorative justice are individually new. Its contribution lies in three relationships: differentiated responsibility is linked to proportional response selection; restorative justice is placed within due process rather than outside it; and integrity cases become a source of governance evidence rather than an endpoint. Appendix B compares inherited, adapted, and newly proposed elements across framework families.
The theoretical contribution is deliberately narrow. Prevention, academic-integrity values, AI literacy, human oversight, assessment validity, restorative justice, and institutional responsibility are not new concepts. The article does not claim originality through their simple combination. Its contribution lies in specifying relationships that are often absent or underdeveloped when these framework families are applied separately (Bertram Gallant, 2008; Chan, 2023; European Commission High-Level Expert Group on Artificial Intelligence, 2019; Holmes et al., 2022; Karp, 2019b; Long & Magerko, 2020; Macfarlane et al., 2014; Ng et al., 2021; UNESCO, 2021).
Academic-integrity values and culture frameworks provide the normative foundation of honesty, trust, fairness, respect, responsibility, courage, prevention, and education. Their limitation for the present problem is that cognitive delegation, function-based AI use, and automated integrity decisions require additional operational detail. The proposed framework adapts this tradition through the three integrity levels, epistemic obligations, and differentiated actor duties.
Assessment and assessment-security frameworks contribute validity, authentic evidence, process visibility, evaluative judgement, and legitimate verification. Some approaches emphasize security and control, while others emphasize redesign and learning. The present framework connects both through AI-resilient assessment and governed verification. It treats verification as legitimate but limited and requires accessibility, corroboration, human reasoning, and appeal.
AI-literacy frameworks identify technical, critical, ethical, and social competencies. Their limitation is not conceptual weakness but the risk that competence is treated as sufficient for responsible conduct. The proposed model separates capacity from motivation, opportunity, incentives, authority, and accountability. It also expands literacy from an individual student attribute to an organisational capability involving educators, leaders, integrity staff, procurement, and quality assurance.
Responsible-AI governance contributes human agency, fairness, privacy, transparency, explainability, accountability, and risk management. These frameworks often operate at system or policy level. The article translates them into task and case procedures through the consequential-decision chain, evidence limits, substantive human review, reason-giving, contestability, and institutional audit. This translation connects general AI ethics to academic standing and due process.
Restorative-justice frameworks contribute harm, affected parties, participation, responsibility, repair, and reintegration. In academic-integrity practice, restorative language is sometimes reduced to reflection, retraining, or resubmission. The proposed model preserves the substantive criteria and adds eligibility, consent, facilitation, record safeguards, and a proportional pathway. Restoration is placed within due process rather than outside formal accountability.
The first newly proposed relationship links differentiated responsibility to proportional response selection. Institutional and provider failures become visible without erasing student agency. The second places restorative justice as one pathway among educational, disciplinary, and combined responses, selected through explicit criteria. The third makes case-to-governance feedback central. Cases are treated not only as incidents to resolve but as evidence about policy clarity, assessment design, literacy, access, technology, and procedure.
The framework also contributes an empirical orientation. Seven propositions, boundary conditions, indicators, and suggested designs make the conceptual relationships testable. This does not validate the model, but it creates a clearer standard for future research. The model can be revised if evidence shows that clarity does not reduce ambiguity, literacy does not support judgement, verification remains biased, restoration creates coercion, or feedback fails to produce change.
The comparison also prevents conceptual appropriation. The article does not relabel established integrity, assessment, AI-governance, literacy, or restorative ideas as new. It identifies their origin and function, then explains the adaptation required for AI-assisted academic work. This transparency supports scholarly credibility and allows readers from each tradition to assess whether the integration preserves core principles.
The framework can be challenged at several points. Researchers may question the three-level boundary, the allocation of duties, the mechanism linking cases to governance, or the selection criteria for restoration. Such criticism is productive because the model is explicit enough to revise. A theoretical contribution is valuable not only when accepted but when it makes assumptions and relationships available for disciplined debate.
The framework’s added value may also be judged by parsimony. Integration should clarify decision-making rather than create unnecessary terminology. Core concepts are therefore defined once and operationalized through duties, mechanisms, safeguards, and indicators. Future work should examine whether users can apply the model without excessive complexity.
Comparison with existing frameworks should remain open to revision as new models emerge. The present table represents the literature available within the search window. Later work may provide stronger integration or empirical evidence. The article’s contribution is a documented position in an evolving field, not a claim to final theoretical priority.
The comparative analysis also clarifies transferability. Institutions can adopt individual inherited elements without adopting the complete integrated model, but the expected mechanisms may then differ. For example, literacy without case feedback does not create institutional learning, and human oversight without appeal does not establish contestability. The framework’s contribution is therefore most visible in the connections among components, while modular adoption should be evaluated against the functions that remain absent.
This comparative positioning also creates a clearer basis for peer review. Reviewers can assess whether the inherited concepts are represented accurately, whether the adaptations are justified, and whether the proposed relationships are sufficiently distinct and testable.

10.3. Practical Implications and Research Agenda

Universities require more than a single AI policy. Governance must reach programmes, courses, and tasks and be supported by valid assessment, staff capability, student literacy, procurement controls, due process, and institutional learning. Educators should aim for valid and interpretable evidence of learning rather than AI-proof assessment. Policymakers and quality agencies should evaluate procedural integrity, contestability, accessibility, and documented learning from cases, not only misconduct prevalence.
Empirical research should test task-level clarity, AI-literacy interventions, assessment validity and burden, responsibility attribution under ambiguous rules, automation bias, restorative eligibility and quality, cross-cultural variation, and whether case feedback results in implemented change. Multi-institutional and longitudinal designs are needed before claims of effectiveness or transferability can be made.
The practical implications begin with institutional policy. Universities should replace generic statements with a coherent hierarchy that connects common principles to programmes, courses, and tasks. Policy should define function-based permissions, disclosure, data protection, automated-evidence limits, human review, reasons, and appeal. Student and staff representatives, accessibility services, legal and data-protection functions, educators, and integrity practitioners should participate in design. Guidance should be tested for comprehension and reviewed through cases and appeals.
Curriculum design should treat AI literacy as progressive and disciplinary. Students need opportunities to understand systems, evaluate outputs, verify evidence, protect data, disclose material use, and defend learning. These capacities should be introduced, practised, and assessed at appropriate stages. Educator development should address task design, communication, evidence, moderation, automation bias, accessibility, and response selection. Programme leaders should coordinate expectations to reduce contradiction.
Assessment reform should focus on valid and interpretable evidence rather than AI-proof formats. Educators can combine process evidence, commentary, contextual projects, demonstrations, supervised components, and AI-output critique according to the learning outcome and stakes. Each method requires accessibility and workload analysis. Verification should be planned where possible and should not rely on probabilistic indicators. Moderation and reasonable adjustments are essential.
Case procedures should distinguish educational correction, restorative justice, discipline, and combined responses. Institutions need clear standards of evidence, proportionality factors, trained reviewers, reasons, and independent appeal. Restorative infrastructure requires eligibility screening, voluntary consent, facilitation, repair agreements, records, follow-up, and reintegration. Formal discipline remains necessary in serious or unsuitable cases. The pathways should be coordinated and transparent.
Technology procurement should be integrated with academic governance. Institutions should evaluate purpose, validity, privacy, accessibility, bias, documentation, model change, and contestability. Tool registers and review cycles can support accountability. Provider claims should not substitute for independent judgement. Students should be informed when consequential systems are used and should have access to intelligible evidence and human review (European Union, 2024; National Institute of Standards and Technology, 2023, 2024).
Quality assurance should use balanced indicators. Institutions can monitor guidance coverage, comprehension, literacy, assessment validity, workload, accessibility, human review, appeals, reversals, differential outcomes, restorative quality, recurrence, and action closure. Raw case counts should not be treated as direct evidence of integrity. Public reporting should explain interpretation and protect privacy. The central question is whether evidence leads to implemented improvement.
The research agenda should test P1 through multi-course studies of task-level guidance, comprehension, and case characteristics. Designs should account for changes in reporting and enforcement. P2 requires longitudinal studies of literacy, judgement, disclosure, and conduct under pressure. P3 requires comparative assessment studies of validity, accessibility, workload, moderation, and differential effects. These studies should include students with diverse linguistic, disability, and resource contexts.
P4 requires audits of automated flags, corroboration, reviewer reasons, overrides, appeals, reversals, and differential error. Qualitative work should examine how students and staff understand contestability. P5 can be studied through vignettes and institutional cases that vary rule clarity, actor duties, intention, and institutional contribution. Cross-disciplinary and cross-cultural comparison is especially important because attribution norms differ.
P6 requires research that distinguishes educational interventions from substantive restorative processes. Studies should examine eligibility, consent, facilitation, repair, participant experience, recurrence, equity, and reintegration. Selection effects must be acknowledged. P7 requires longitudinal implementation research tracing whether anonymized case patterns lead to decisions, assigned actions, implementation, and outcomes. Documentation alone should not be treated as learning.
Future research should also examine emerging issues that may alter the framework: multimodal generation, autonomous agents, embedded AI in common software, personalized systems, changing regulation, and new forms of assessment. The function-based approach should remain adaptable, but specific safeguards may require revision. Deliberative research with students, educators, professional bodies, integrity staff, policymakers, and providers can test legitimacy and feasibility.
The practical and research agendas are connected. Implementation creates data about mechanisms, burdens, rights, and outcomes. Research can refine policy and assessment. Institutions should therefore treat governance as an iterative learning process rather than a one-time response to a new technology. The framework provides a structured starting point, while its propositions and limitations define the evidence still required before stronger claims of effectiveness or transferability can be made (Chan, 2023; Chiu et al., 2024; Corbin et al., 2025; Dawson et al., 2024; Holmes et al., 2022; Karp, 2019b; Laupichler et al., 2022; OECD, 2023, 2026; Southworth et al., 2023; UNESCO, 2021, 2023, 2024a, 2024b; UNESCO International Institute for Higher Education in Latin America and the Caribbean, 2023; Xia et al., 2024).
Institutions should also plan for communication during uncertainty. Rules cannot anticipate every new function, and staff may disagree about emerging practices. Interim guidance should identify stable principles, consultation routes, and how ambiguous cases will be handled. Students should not bear the cost of unresolved institutional disagreement. Transparent uncertainty can be more legitimate than false certainty.
Research infrastructure can be embedded in implementation through ethical and privacy-preserving data design. Common definitions, case coding, accessibility measures, and implementation records can support multi-institutional studies. Stakeholder involvement is necessary to ensure that research questions reflect educational and procedural concerns rather than only misconduct detection. Collaboration between institutions can accelerate learning while preserving contextual interpretation.

11. Conclusions

Generative AI intensifies longstanding academic-integrity questions while adding problems of cognitive delegation, weak traceability, and algorithmic mediation. A defensible response requires clear conceptual boundaries among integrity of academic work, integrity of assessment, and institutional procedural integrity.
This critical conceptual synthesis proposes an integrated framework in which preventive governance, ethical AI literacy, governed verification, restorative accountability, and proportionate discipline have complementary but limited roles. Policy clarity must reach the task; literacy must be institutionalised without being mistaken for guaranteed compliance; assessment must remain valid and accessible; automated indicators must not be treated as proof; and restorative practice must satisfy substantive requirements of harm, participation, responsibility, repair, and reintegration.
The central contribution is the feedback mechanism through which cases inform future policy, assessment, literacy provision, procurement, and technology oversight. The seven propositions, comparative matrices, and implementation indicators provide a basis for empirical testing, not evidence of effectiveness. Universal claims would be premature, and contextual adaptation must preserve human responsibility, fairness, reason-giving, contestability, privacy, and continuous review.

Author Contributions

Conceptualization, G.B., M.S., N.D., A.S. and C.D.; methodology, G.B., C.V., P.S., N.D. and L.-O.F.; formal analysis, G.B., E.C. and C.D.; investigation, M.S., P.S. and G.T.; resources, L.-O.F., A.S. and C.D.; data curation, G.B. and N.D.; visualization, G.B., M.S., C.V., P.S., E.C., G.T. and N.D.; supervision, G.B., C.V. and A.S.; project administration, G.B.; writing—original draft preparation, G.B., E.C., G.T. and C.D.; writing—review and editing, G.B. and C.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new empirical data were created or analysed. The conceptual source corpus and framework-comparison audit are documented in Appendix A and Appendix B.

Acknowledgments

During revision, AI-assisted language and formatting tools were used for English refinement, structural organisation, and preparation of draft visual materials. The authors critically reviewed and edited all outputs and take full responsibility for the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Retained-Source Matrix and Analytical Role

The matrix documents the purposive corpus used in the critical conceptual synthesis. Scope indicates the principal geographical or institutional emphasis rather than the nationality of every author. The matrix is not a study-quality ranking and should not be interpreted as a systematic-review evidence table.
Table A1. Characteristics and analytical role of the 64 retained sources.
Table A1. Characteristics and analytical role of the 64 retained sources.
No.SourceClassScopePrimary DomainRole in Synthesis
1Bertram Gallant (2008)BookUSA/InternationalAcademic integrityTeaching-and-learning foundation
2Boud and Falchikov (2007)Edited bookInternationalAssessmentLearning-oriented assessment
3Bretag (2013)ArticleInternationalPlagiarismIntegrity problem framing
4Bretag (2016)HandbookInternationalAcademic integrityField-wide conceptual foundation
5Bretag et al. (2019)Empirical articleAustraliaContract cheatingStrategic misconduct and incentives
6Bussu and Karp (2026)Empirical/conceptual articleHigher education; multi-siteRestorative justiceImplementation conditions and limits
7Carless and Boud (2018)ArticleInternationalFeedback literacyLearner judgement and agency
8Chan (2023)Framework articleHigher educationAI policyPolicy levels and educational use
9Chiu et al. (2024)Framework articleInternationalAI literacyCompetency dimensions
10Corbin et al. (2025)Conceptual articleHigher educationAI and assessmentWicked-problem framing and trade-offs
11Cotton et al. (2024)ArticleUK higher educationAcademic integrity/GenAIPolicy and practice tensions
12Dawson (2021)BookInternationalAssessment securityLegitimate verification and safeguards
13Dawson et al. (2024)ArticleInternationalAssessment validityChallenges cheating-centric framing
14Eaton (2021)BookHigher educationAcademic integrityInstitutional culture and prevention
15Eubanks (2018)BookUSAAlgorithmic inequalityPower and structural harm
16European Commission High-Level Expert Group on Artificial Intelligence (2019)Policy frameworkEuropean UnionTrustworthy AIHuman agency and oversight principles
17European Union (2024)RegulationEuropean UnionAI governanceLegal risk and fundamental-rights context
18Evans and Vaandering (2022)BookEducationRestorative justiceEducational restorative principles
19Farrokhnia et al. (2024)ArticleInternationalChatGPTBenefits, risks, and research agenda
20Fawns (2022)Conceptual articleInternationalSociotechnical pedagogyTechnology–pedagogy entanglement
21Fishman (2009)Conference paperAsia-PacificPlagiarism definitionConceptual delimitation
22Floridi and Cowls (2019)Conceptual articleInternationalAI ethicsNormative principles
23Holmes et al. (2019)Book/reportInternationalAI in educationOpportunities and educational implications
24Holmes et al. (2022)Framework articleInternationalEthics of AI in educationCommunity-wide ethical duties
25Holmes and Tuomi (2022)Review articleEuropean/internationalAI in educationState of practice
26International Center for Academic Integrity (2021)Institutional frameworkInternationalAcademic integrity valuesNormative integrity values
27Jobin et al. (2019)Review articleInternationalAI ethics guidelinesComparative governance principles
28Karp (2019b)BookHigher educationRestorative justiceCollege/university process model
29Karp (2019a)Book chapterHigher educationResponsive regulationProportional response design
30Kasneci et al. (2023)Review/conceptual articleInternationalLarge language modelsEducational opportunities and risks
31Laupichler et al. (2022)Scoping reviewHigher/adult educationAI literacyDefinitions and evidence gaps
32Lim et al. (2024)Review articleInternationalGenerative AIResearch agenda and synthesis
33Lo (2023)Rapid reviewInternationalChatGPT in educationEarly evidence synthesis
34Long and Magerko (2020)Conference paperInternationalAI literacyCore competency definition
35Luckin et al. (2016)ReportInternationalAI in educationHuman-centred educational case
36Macfarlane et al. (2014)Review articleInternationalAcademic integrityPre-AI relational/institutional literature
37McArthur (2016)ArticleInternationalAssessment/social justiceFairness and inclusion
38Mittelstadt et al. (2016)Conceptual articleInternationalAlgorithm ethicsOpacity, responsibility, bias
39National Institute of Standards and Technology (2023)Risk frameworkUSA/internationalAI governanceRisk management functions
40National Institute of Standards and Technology (2024)Risk profileUSA/internationalGenerative AI riskGenAI-specific controls
41Ng et al. (2021)Review articleInternationalAI literacyFour-dimensional literacy model
42Nicol and Macfarlane-Dick (2006)ArticleInternationalFormative assessmentSelf-regulation and feedback
43OECD (2023)Policy reportOECD membersDigital education governanceSystem-level ecosystem context
44OECD (2026)Policy reportOECD membersGenerative AI in educationUpdated use and governance evidence
45O’Neil (2016)BookUSA/internationalAlgorithmic harmRisk of automated inequality
46Perkins (2023)Conceptual articleHigher educationAcademic integrity/LLMsIntegrity considerations
47Rudolph et al. (2023)Conceptual articleHigher educationAssessment/ChatGPTCritique of traditional assessment
48Selwyn (2019)BookInternationalCritical EdTechPower and institutional context
49Southworth et al. (2023)Conceptual articleUSA higher educationAI across curriculumInstitutional literacy model
50Sutherland-Smith (2008)BookHigher educationPlagiarism and learningEducational integrity practices
51Tai et al. (2018)ArticleInternationalEvaluative judgementLearner capacity and assessment
52Tlili et al. (2023)Review/case articleInternationalChatbots in educationOpportunities and ethical risks
53UNESCO (2021)Global recommendationInternationalAI ethicsHuman rights and ethical baseline
54UNESCO (2023)Global guidanceInternationalGenerative AI in educationPolicy and pedagogical recommendations
55UNESCO (2024a)Competency frameworkInternationalStudent AI competencyRole-specific literacy
56UNESCO (2024b)Competency frameworkInternationalTeacher AI competencyRole-specific literacy
57UNESCO International Institute for Higher Education in Latin America and the Caribbean (2023)Policy guideHigher educationGenerative AIInstitutional quick-start guidance
58Wachtel (2016)Practice frameworkInternationalRestorative practiceProcess and participation principles
59Williamson (2021)Critical articleInternationalEducation platformsCommercial and governance interests
60Williamson et al. (2020)Critical articleInternationalDataficationInstitutional power and accountability
61Xia et al. (2024)Review articleInternationalGenerative AI assessmentMaps assessment transformation
62Zawacki-Richter et al. (2019)Systematic reviewInternationalAI in higher educationResearch landscape and gaps
63Zehr (2015)BookInternationalRestorative justiceHarm, obligation, and repair
64Zhai et al. (2021)Review articleInternationalAI in educationHistorical field mapping
Source: Authors’ documentation of the conceptual corpus. Policy documents and academic publications were interpreted according to their distinct evidential and normative roles.

Appendix B. Comparison with Existing Framework Families

The comparison identifies inherited elements, adaptations for generative AI, and the relationships proposed by the present framework. It does not rank the quality of the framework families.
Table A2. Framework comparison and claimed contribution.
Table A2. Framework comparison and claimed contribution.
Framework FamilyElements InheritedLimitation for the Present ProblemAdaptation or Relationship Proposed Here
Academic-integrity values/cultureHonesty, trust, fairness, responsibility, prevention, educationLimited operational detail for cognitive delegation and automated integrity decisionsThree integrity levels; epistemic obligations; differentiated duties
Assessment and assessment-security frameworksValidity, authentic evidence, process visibility, verificationMay emphasise either cheating control or redesign without integrated due processAI-resilient assessment plus accessible, governed verification
AI-literacy frameworksTechnical, critical, ethical, and social competenciesCompetence may be treated as sufficient for responsible conductLiteracy separated from incentives, opportunity, accountability, and institutional capacity
Responsible-AI governanceHuman agency, fairness, privacy, explainability, risk managementOften remains system- or policy-level rather than task- and case-levelConsequential-decision chain, reason-giving, appeal, and procurement duties
Restorative-justice frameworksHarm, participation, responsibility, repair, reintegrationSometimes loosely translated into reflection or retraining in integrity contextsEligibility, consent, facilitation, record safeguards, and proportional pathway
Present integrated frameworkPrevention, literacy, verification, accountability, restoration, disciplineNormative and not yet empirically validatedDifferentiated responsibility + proportional response + case-to-governance feedback

Appendix C. Accessibility and Visual-Quality Verification Checklist

The following checklist records the publication-oriented checks applied to the revised manuscript. It is provided as an author-facing quality-control aid and does not substitute for the journal production process.
  • All 12 figures use high-resolution raster output, concise internal labels, consistent terminology, and captions that state whether content is literature-derived, descriptive, or author-generated.
  • Comparative graphics that use heuristic ratings are explicitly labelled as conceptual decision aids rather than empirical measurements.
  • All 18 main tables have repeated header rows, consistent borders, legible type, and captions that identify their analytical purpose.
  • Heading levels are hierarchical and consistent, enabling navigation and screen-reader interpretation.
  • Colour is not the sole carrier of meaning; figures combine colour with labels, ordering, arrows, and textual descriptions.
  • Figures and tables are introduced and interpreted in the surrounding text rather than inserted as stand-alone decoration.
  • Long tables are permitted to split across pages with repeating headers, and no table row is intentionally used to conceal methodological limitations.
  • The final DOCX was rendered page by page and visually inspected for overlap, clipping, broken glyphs, illegible labels, and inconsistent headers or footers.

Appendix D. Institutional Operational Instruments

The following instruments translate the conceptual framework into auditable institutional practice. They are adaptable templates rather than validated instruments and should be reviewed with students, educators, accessibility services, legal/data-protection functions, and disciplinary specialists.
Table A3. Task-level AI guidance template.
Table A3. Task-level AI guidance template.
Guidance FieldQuestion to AnswerIllustrative EntryQuality Check
Learning purposeWhat competence is the task intended to evidence?Critical synthesis of competing evidenceAI permissions must preserve evidence of this competence
Permitted functionsWhich AI functions may be used and for what purpose?Brainstorming and language feedback; no generative draftingFunctions are described rather than named by brand
Prohibited functionsWhich functions would substitute for the learning outcome?Automated synthesis or generation of the final argumentProhibition is justified by the outcome, not fear of AI
DisclosureWhat must be reported and at what level of detail?Short use statement with material functions and verification stepsRequirements are proportionate and privacy-preserving
Process evidenceWhat evidence may be requested?Selected draft, source audit, and explanation of revisionsEvidence burden is accessible and aligned with decision stakes
Verification and appealHow may understanding be clarified and decisions challenged?Structured clarification meeting; written reasons; appeal routeAutomated indicators are never treated as proof
Table A4. Case triage and proportionality worksheet.
Table A4. Case triage and proportionality worksheet.
FactorQuestions for the Decision-MakerPossible Mitigating ConditionPossible Aggravating Condition
Rule clarityWas the task-level rule clear, accessible, and taught?Ambiguous wording or inaccessible guidanceClear examples and acknowledged understanding
Learning outcomeDid the AI function replace a target competence?Use concerned a peripheral functionUse substituted for the central assessed process
IntentionWhat did the student understand and seek to achieve?Good-faith misunderstanding or poor judgementDeliberate concealment or planned unfair advantage
Degree of delegationHow much of the decisive work was delegated?Limited assistance with substantial human transformationSubstantial output adopted without understanding
Advantage and harmWhat unfair benefit or impact resulted?No material advantage; readily correctedSerious inequity, professional risk, or harm to others
Evidence qualityHow reliable and corroborated is the evidence?Probabilistic indicator without supporting factsMultiple independent and contestable evidence sources
Prior guidanceWhat support and opportunities to learn were provided?No realistic training or examplesRepeated guidance and prior corrective intervention
RepetitionIs this a first or repeated concern?First occurrenceRepeated or escalating pattern
Explanation and repairCan the student explain the work and acknowledge consequences?Open engagement and feasible repairContinued concealment or refusal to address established harm
Institutional contributionDid policy, assessment, access, or technology create avoidable risk?Documented institutional ambiguity or inequityNo relevant institutional failure identified

Appendix E. Analytical Inventory of Figures and Main Tables

The inventory confirms that every visual has a distinct analytical function. It also helps authors, reviewers, and production staff verify cross-references and identify whether a visual is conceptual, descriptive, comparative, procedural, or implementation oriented.
Table A5. Analytical role of the 12 figures.
Table A5. Analytical role of the 12 figures.
FigureTypeAnalytical PurposeRelationship to Reviewer Requests
Figure 1Boundary diagramSeparates academic work, assessment, and procedural integrityPrevents conceptual over-expansion
Figure 2Comparative diagramContrasts detection-centred, prevention-centred, and integrated governanceClarifies article positioning and novelty
Figure 3Method workflowShows synthesis stages, counter-cases, audit trail, and empirical limitsAddresses methodological transparency and circularity
Figure 4Comparative bar graphDescribes the 64-source corpus by analytical domainMakes corpus composition transparent
Figure 5Responsibility diagramSeparates coordinated conditions from attributable actor dutiesAddresses distributed responsibility and proportional attribution
Figure 6Accountability chainDefines safeguards for consequential algorithmic decisionsOperationalises substantive human oversight and appeal
Figure 7Governance cycleConnects institutional, programme, course, task, and review levelsClarifies preventive governance and feedback
Figure 8Comparative heatmapCompares assessment strategies across five decision dimensionsMakes trade-offs visible without claiming measured effects
Figure 9Capacity pyramidDistinguishes technical, critical, ethical, organisational, and renewal layersClarifies AI-literacy functions and institutional responsibility
Figure 10Decision pathwaySelects educational, restorative, disciplinary, or combined responsesOperationalises proportional response selection
Figure 11Restorative processDefines eligibility, consent, facilitation, repair, follow-up, and feedbackPrevents reduction in restoration to reflection or retraining
Figure 12Integrated frameworkConnects the pillars through mechanisms and case-to-governance feedbackStates the central theoretical contribution
Table A6. Analytical role of the 18 main tables.
Table A6. Analytical role of the 18 main tables.
TablePrimary FunctionComparative or Analytical Value
Table 1Gap analysisCompares six literature strands and identifies the contribution developed here
Table 2Review architectureDistinguishes purpose, coverage, search, eligibility, analysis, and quality controls
Table 3Search conceptsLinks search combinations to eligibility and research-question functions
Table 4Three integrity levelsCompares objects, questions, and evidence across conceptual levels
Table 5AI-use classificationCompares involvement, control, disclosure, learning evidence, and integrity concerns
Table 6Epistemic obligationsLinks verification, justification, disclosure, interpretation, and responsibility to evidence
Table 7Actor dutiesCompares student, educator, institutional, quality, and provider responsibilities
Table 8Algorithmic safeguardsMaps safeguards, evidence, and failure modes across the decision chain
Table 9Policy architectureCompares institutional, programme, course, task, and feedback levels
Table 10Assessment strategiesCompares potential benefits, trade-offs, and safeguards
Table 11Assessment decision matrixCompares six strategies across validity, accessibility, workload, and combination
Table 12Role-specific literacyCompares minimum and advanced competencies by educational role
Table 13Literacy functions and limitsSeparates individual, curricular, organisational, and preventive functions
Table 14Response criteriaCompares educational, restorative, and disciplinary/combined emphases
Table 15Restorative safeguardsMaps requirements, operational questions, safeguards, and unsuitable conditions
Table 16Seven propositionsLinks mechanisms, boundary conditions, and empirical indicators
Table 17Stable/adaptable elementsSeparates stable ethical principles from local implementation
Table 18Implementation roadmapIntegrates actors, resources, timelines, outputs, indicators, and cautions

Appendix F. Empirical Testing Agenda

The seven propositions require empirical testing before the framework can support claims of effectiveness. The matrix below identifies feasible designs and cautions that preserve the distinction between implementation evidence, procedural legitimacy, educational outcomes, and misconduct prevalence.
Table A7. Empirical test matrix for the seven conceptual propositions.
Table A7. Empirical test matrix for the seven conceptual propositions.
PropositionSuggested DesignPrimary DataOutcome or MechanismKey Threat to Inference
P1 Policy clarityMulti-course quasi-experiment or interrupted time seriesTask guidance, comprehension, case characteristicsAmbiguity-related concerns and rule comprehensionReporting or enforcement may change with the intervention
P2 AI literacyLongitudinal curriculum studyCompetency measures, process evidence, disclosure behaviourJudgement, verification, and transparent useKnowledge and conduct may diverge under pressure
P3 AI-resilient assessmentComparative assessment-design studyValidity evidence, accessibility, workload, moderationInterpretability of learning evidence and equityAdditional evidence may create burden without improving validity
P4 Governed verificationCase audit and appeal studyFlags, corroboration, reviewer reasons, overrides, appealsError, procedural fairness, and differential impactSelection bias in which cases reach appeal
P5 Differentiated responsibilityVignette experiment plus institutional case studyAttribution judgements, rule clarity, actor dutiesConsistency and proportionality of responsibility attributionNorms may differ by discipline and culture
P6 Restorative processComparative implementation and participant studyConsent, facilitation quality, agreements, recurrence, experienceRepair, learning, fairness, and reintegrationSelf-selection and unsuitable-case exclusion
P7 Case feedbackLongitudinal governance implementation studyCase coding, review decisions, policy/assessment changesDocumented institutional learning and action closureDocumentation may be symbolic rather than implemented

Appendix G. Reviewer-Response-to-Manuscript Synchronization Matrix

This matrix records the principal commitments made in the response letter and the corresponding locations in the revised manuscript. It is intended to prevent discrepancies between the rebuttal and the submitted clean version.
Table A8. Synchronization of reviewer commitments and manuscript evidence.
Table A8. Synchronization of reviewer commitments and manuscript evidence.
Revision CommitmentManuscript EvidenceStatus
Reclassify the article as a critical conceptual synthesisArticle type, Abstract, Section 2.1, Section 2.2, Section 2.3, Section 2.4 and Section 2.5, limitations and conclusionsImplemented
State four explicit research questionsSection 1.3 and Section 10.1Implemented
Delimit three levels of integrityFigure 1, Section 3.1 and Table 4Implemented
Provide transparent source selection and audit trailSection 2.3 and Section 2.4, Table 2 and Table 3, Figure 3 and Figure 4 and Appendix AImplemented
Include counter-positions and negative casesSection 2.5 and boundary conditions in Table 16 and Table A7Implemented
Define epistemic responsibility and human–AI contributionSection 3.2 and Section 3.3, Figure 5 and Table 5 and Table 6Implemented
Preserve attributable duties within distributed responsibilitySection 4.1 and Table 7Implemented
Operationalise substantive human oversight and appealSection 4.2, Figure 6 and Table 8Implemented
Define AI-resilient assessment with trade-offsSection 5.2, Figure 8 and Table 10 and Table 11Implemented
Clarify AI-literacy functions and limitsSection 6, Figure 9 and Table 12 and Table 13Implemented
Ground restorative practice in restorative justiceSection 7, Figure 10 and Figure 11 and Table 14 and Table 15Implemented
Specify mechanisms, propositions, implementation and empirical testingSection 8, Figure 12, Table 16, Table 17 and Table 18 and Appendix B and Appendix FImplemented

References

  1. Bertram Gallant, T. (2008). Academic integrity in the twenty-first century: A teaching and learning imperative. Jossey-Bass. [Google Scholar]
  2. Boud, D., & Falchikov, N. (Eds.). (2007). Rethinking assessment in higher education: Learning for the longer term. Routledge. [Google Scholar]
  3. Bretag, T. (2013). Challenges in addressing plagiarism in education. PLoS Medicine, 10, e1001574. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Bretag, T. (Ed.). (2016). Handbook of academic integrity. Springer. [Google Scholar]
  5. Bretag, T., Harper, R., Burton, M., Ellis, C., Newton, P., van Haeringen, K., Saddiqui, S., & Rozenberg, P. (2019). Contract cheating: A survey of Australian university students. Studies in Higher Education, 44, 1837–1856. [Google Scholar] [CrossRef] [Scilit]
  6. Bussu, A., & Karp, D. R. (2026). Restorative justice in higher education: An analysis of challenges, strategies, and achievements using normalization process theory. Innovative Higher Education, 51, 293–325. [Google Scholar] [CrossRef] [Scilit]
  7. Carless, D., & Boud, D. (2018). The development of student feedback literacy: Enabling uptake of feedback. Assessment & Evaluation in Higher Education, 43, 1315–1325. [Google Scholar] [CrossRef] [Scilit]
  8. Chan, C. K. Y. (2023). A comprehensive AI policy education framework for university teaching and learning. International Journal of Educational Technology in Higher Education, 20, 38. [Google Scholar] [CrossRef] [Scilit]
  9. Chiu, T. K. F., Ahmad, Z., Ismailov, M., & Sanusi, I. T. (2024). What are artificial intelligence literacy and competency? A comprehensive framework to support them. Computers and Education Open, 6, 100171. [Google Scholar] [CrossRef] [Scilit]
  10. Corbin, T., Bearman, M., Boud, D., & Dawson, P. (2025). The wicked problem of AI and assessment. Assessment & Evaluation in Higher Education, 50, 736–752. [Google Scholar] [CrossRef] [Scilit]
  11. Cotton, D. R. E., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61, 228–239. [Google Scholar] [CrossRef] [Scilit]
  12. Dawson, P. (2021). Defending assessment security in a digital world: Preventing e-cheating and supporting academic integrity in higher education. Routledge. [Google Scholar]
  13. Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. Assessment & Evaluation in Higher Education, 49, 1005–1016. [Google Scholar] [CrossRef] [Scilit]
  14. Eaton, S. E. (2021). Plagiarism in higher education: Tackling tough topics in academic integrity. Libraries Unlimited. [Google Scholar]
  15. Eubanks, V. (2018). Automating inequality: How high-tech tools profile, police, and punish the poor. St. Martin’s Press. [Google Scholar]
  16. European Commission High-Level Expert Group on Artificial Intelligence. (2019). Ethics guidelines for trustworthy AI. European Commission. [Google Scholar]
  17. European Union. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. [Google Scholar]
  18. Evans, K., & Vaandering, D. (2022). The little book of restorative justice in education. Good Books. [Google Scholar]
  19. Farrokhnia, M., Banihashem, S. K., Noroozi, O., & Wals, A. (2024). A SWOT analysis of ChatGPT: Implications for educational practice and research. Innovations in Education and Teaching International, 61, 460–474. [Google Scholar] [CrossRef] [Scilit]
  20. Fawns, T. (2022). An entangled pedagogy: Looking beyond the pedagogy–technology dichotomy. Postdigital Science and Education, 4, 711–728. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Fishman, T. (2009, September 28–30). We know it when we see it is not good enough: Toward a standard definition of plagiarism that transcends theft, fraud, and copyright. Proceedings of the 4th Asia Pacific Conference on Educational Integrity, Wollongong, Australia. [Google Scholar]
  22. Floridi, L., & Cowls, J. (2019). A unified framework of five principles for AI in society. Harvard Data Science Review, 1. [Google Scholar] [CrossRef] [Scilit]
  23. Holmes, W., Bialik, M., & Fadel, C. (2019). Artificial intelligence in education: Promises and implications for teaching and learning. Center for Curriculum Redesign. [Google Scholar]
  24. Holmes, W., Porayska-Pomsta, K., Holstein, K., Sutherland, E., Baker, T., Shum, S. B., Santos, O. C., Rodrigo, M. M. T., Cukurova, M., Bittencourt, I. I., & Koedinger, K. R. (2022). Ethics of AI in education: Towards a community-wide framework. International Journal of Artificial Intelligence in Education, 32, 504–526. [Google Scholar] [CrossRef] [Scilit]
  25. Holmes, W., & Tuomi, I. (2022). State of the art and practice in AI in education. European Journal of Education, 57, 542–570. [Google Scholar] [CrossRef] [Scilit]
  26. International Center for Academic Integrity. (2021). The fundamental values of academic integrity (3rd ed.). International Center for Academic Integrity. [Google Scholar]
  27. Jobin, A., Ienca, M., & Vayena, E. (2019). The global landscape of AI ethics guidelines. Nature Machine Intelligence, 1, 389–399. [Google Scholar] [CrossRef] [Scilit]
  28. Karp, D. R. (2019a). Restorative justice and responsive regulation in higher education. In T. Gavrielides (Ed.), Routledge international handbook of restorative justice (pp. 443–457). Routledge. [Google Scholar]
  29. Karp, D. R. (2019b). The little book of restorative justice for colleges and universities (2nd ed.). Good Books. [Google Scholar]
  30. Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
  31. Laupichler, M. C., Aster, A., Schirch, J., & Raupach, T. (2022). Artificial intelligence literacy in higher and adult education: A scoping literature review. Computers and Education: Artificial Intelligence, 3, 100101. [Google Scholar] [CrossRef] [Scilit]
  32. Lim, W. M., Gunasekara, A. N., Pallant, J. L., Pallant, J. I., & Pechenkina, E. (2024). Generative AI and the future of education: Literature synthesis and research agenda. International Journal of Information Management Data Insights, 4, 100232. [Google Scholar]
  33. Lo, C. K. (2023). What is the impact of ChatGPT on education? A rapid review of the literature. Education Sciences, 13, 410. [Google Scholar] [CrossRef] [Scilit]
  34. Long, D., & Magerko, B. (2020, April 25–30). What is AI literacy? Competencies and design considerations. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA. [Google Scholar]
  35. Luckin, R., Holmes, W., Griffiths, M., & Forcier, L. B. (2016). Intelligence unleashed: An argument for AI in education. Pearson. [Google Scholar]
  36. Macfarlane, B., Zhang, J., & Pun, A. (2014). Academic integrity: A review of the literature. Studies in Higher Education, 39, 339–358. [Google Scholar] [CrossRef] [Scilit]
  37. McArthur, J. (2016). Assessment for social justice: The role of assessment in achieving social justice. Assessment & Evaluation in Higher Education, 41, 967–981. [Google Scholar] [CrossRef] [Scilit]
  38. Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3, 2053951716679679. [Google Scholar] [CrossRef] [Scilit]
  39. National Institute of Standards and Technology. (2023). Artificial intelligence risk management framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology.
  40. National Institute of Standards and Technology. (2024). Artificial intelligence risk management framework: Generative artificial intelligence profile (NIST AI 600-1). National Institute of Standards and Technology.
  41. Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. [Google Scholar] [CrossRef] [Scilit]
  42. Nicol, D. J., & Macfarlane-Dick, D. (2006). Formative assessment and self-regulated learning: A model and seven principles of good feedback practice. Studies in Higher Education, 31, 199–218. [Google Scholar] [CrossRef] [Scilit]
  43. OECD. (2023). OECD digital education outlook 2023: Towards an effective digital education ecosystem. OECD Publishing. [Google Scholar]
  44. OECD. (2026). OECD digital education outlook 2026: Exploring effective uses of generative AI in education. OECD Publishing. [Google Scholar] [CrossRef] [Scilit]
  45. O’Neil, C. (2016). Weapons of math destruction: How big data increases inequality and threatens democracy. Crown. [Google Scholar]
  46. Perkins, M. (2023). Academic integrity considerations of AI large language models in the post-pandemic era: ChatGPT and beyond. Journal of University Teaching and Learning Practice, 20, 7. [Google Scholar] [CrossRef] [Scilit]
  47. Rudolph, J., Tan, S., & Tan, S. (2023). ChatGPT: Bullshit spewer or the end of traditional assessments in higher education? Journal of Applied Learning and Teaching, 6, 342–363. [Google Scholar] [CrossRef] [Scilit]
  48. Selwyn, N. (2019). Should robots replace teachers? AI and the future of education. Polity Press. [Google Scholar]
  49. Southworth, J., Migliaccio, K., Glover, J., Reed, D., McCarty, C., Brendemuhl, J., & Thomas, A. (2023). Developing a model for AI across the curriculum: Transforming the higher education landscape via innovation in AI literacy. Computers and Education: Artificial Intelligence, 4, 100127. [Google Scholar] [CrossRef] [Scilit]
  50. Sutherland-Smith, W. (2008). Plagiarism, the internet, and student learning: Improving academic integrity. Routledge. [Google Scholar]
  51. Tai, J., Ajjawi, R., Boud, D., Dawson, P., & Panadero, E. (2018). Developing evaluative judgement: Enabling students to make decisions about the quality of work. Higher Education, 76, 467–481. [Google Scholar] [CrossRef] [Scilit]
  52. Tlili, A., Shehata, B., Adarkwah, M. A., Bozkurt, A., Hickey, D. T., Huang, R., & Agyemang, B. (2023). What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in education. Smart Learning Environments, 10, 15. [Google Scholar] [CrossRef] [Scilit]
  53. UNESCO. (2021). Recommendation on the ethics of artificial intelligence. UNESCO. [Google Scholar]
  54. UNESCO. (2023). Guidance for generative AI in education and research. UNESCO. [Google Scholar]
  55. UNESCO. (2024a). AI competency framework for students. UNESCO. [Google Scholar]
  56. UNESCO. (2024b). AI competency framework for teachers. UNESCO. [Google Scholar]
  57. UNESCO International Institute for Higher Education in Latin America and the Caribbean. (2023). ChatGPT and artificial intelligence in higher education: Quick start guide. UNESCO IESALC. [Google Scholar]
  58. Wachtel, T. (2016). Defining restorative. International Institute for Restorative Practices. [Google Scholar]
  59. Williamson, B. (2021). Making markets through digital platforms: Pearson, edu-business, and the (e)valuation of higher education. Critical Studies in Education, 62, 50–66. [Google Scholar] [CrossRef] [Scilit]
  60. Williamson, B., Bayne, S., & Shay, S. (2020). The datafication of teaching in higher education: Critical issues and perspectives. Teaching in Higher Education, 25, 351–365. [Google Scholar] [CrossRef] [Scilit]
  61. Xia, Q., Weng, X., Ouyang, F., Lin, T.-J., & Chiu, T. K. F. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education, 21, 40. [Google Scholar] [CrossRef] [Scilit]
  62. Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education. International Journal of Educational Technology in Higher Education, 16, 39. [Google Scholar] [CrossRef] [Scilit]
  63. Zehr, H. (2015). The little book of restorative justice (Rev. & updated ed.). Good Books. [Google Scholar]
  64. Zhai, X., Chu, X., Chai, C. S., Jong, M. S.-Y., Istenic, A., Spector, M., Liu, J., Yuan, J., & Li, Y. (2021). A review of artificial intelligence in education from 2010 to 2020. Complexity, 2021, 8812542. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Conceptual scope and three levels of academic integrity in AI-assisted higher education. Source: Authors’ conceptual synthesis.
Figure 1. Conceptual scope and three levels of academic integrity in AI-assisted higher education. Source: Authors’ conceptual synthesis.
Education 16 01365 g001
Figure 2. Comparative positioning from detection-centred responses to integrated ethical governance. Source: Authors’ synthesis; the comparison is conceptual, not an empirical performance ranking.
Figure 2. Comparative positioning from detection-centred responses to integrated ethical governance. Source: Authors’ synthesis; the comparison is conceptual, not an empirical performance ranking.
Education 16 01365 g002
Figure 3. Critical conceptual synthesis and audit-trail workflow. Source: Authors’ conceptual synthesis.
Figure 3. Critical conceptual synthesis and audit-trail workflow. Source: Authors’ conceptual synthesis.
Education 16 01365 g003
Figure 4. Comparative composition of the purposive 64-source corpus. Source: Counts calculated from Appendix A; domain grouping was performed by the authors for descriptive transparency.
Figure 4. Comparative composition of the purposive 64-source corpus. Source: Counts calculated from Appendix A; domain grouping was performed by the authors for descriptive transparency.
Education 16 01365 g004
Figure 5. Differentiated responsibility in human–AI academic co-production. Source: Authors’ conceptual synthesis.
Figure 5. Differentiated responsibility in human–AI academic co-production. Source: Authors’ conceptual synthesis.
Education 16 01365 g005
Figure 6. Algorithmic accountability chain for assessment and integrity decisions. Source: Authors’ conceptual synthesis.
Figure 6. Algorithmic accountability chain for assessment and integrity decisions. Source: Authors’ conceptual synthesis.
Education 16 01365 g006
Figure 7. Multi-level preventive governance process with institutional feedback. Source: Authors’ conceptual synthesis.
Figure 7. Multi-level preventive governance process with institutional feedback. Source: Authors’ conceptual synthesis.
Education 16 01365 g007
Figure 8. Comparative profile of six assessment strategies. Source: Author-generated heuristic for decision support, not empirical measurement. Scores require local validation and adaptation.
Figure 8. Comparative profile of six assessment strategies. Source: Author-generated heuristic for decision support, not empirical measurement. Scores require local validation and adaptation.
Education 16 01365 g008
Figure 9. AI literacy as a layered institutional capacity system. Source: Authors’ conceptual synthesis.
Figure 9. AI literacy as a layered institutional capacity system. Source: Authors’ conceptual synthesis.
Education 16 01365 g009
Figure 10. Proportionate pathway for responding to AI-related academic-integrity concerns. Source: Authors’ conceptual synthesis.
Figure 10. Proportionate pathway for responding to AI-related academic-integrity concerns. Source: Authors’ conceptual synthesis.
Education 16 01365 g010
Figure 11. Restorative-justice process and safeguards for higher-education integrity cases. Source: Authors’ conceptual synthesis.
Figure 11. Restorative-justice process and safeguards for higher-education integrity cases. Source: Authors’ conceptual synthesis.
Education 16 01365 g011
Figure 12. Integrated preventive–literacy–restorative framework and case-to-governance feedback mechanism. Source: Authors’ conceptual synthesis.
Figure 12. Integrated preventive–literacy–restorative framework and case-to-governance feedback mechanism. Source: Authors’ conceptual synthesis.
Education 16 01365 g012
Table 1. Gap analysis and specific contribution of the present framework.
Table 1. Gap analysis and specific contribution of the present framework.
Literature StrandEstablished ContributionPersistent LimitationContribution Developed Here
Academic-integrity cultureShared values, education, prevention, and institutional responsibilityHuman–AI co-production and procedural mediation are not fully specifiedThree-level integrity model and differentiated duties
AI policy and governanceFairness, transparency, human oversight, and risk managementWeak translation into course- and task-level practiceFour-level policy architecture and contestable procedures
AI literacyTechnical, critical, and ethical competenciesKnowledge may be treated as if it reliably produces compliant conductSeparation of capacity, opportunity, incentives, and motivation
Assessment reformProcess evidence, authentic tasks, evaluative judgementAccessibility, workload, reliability, and disciplinary trade-offs may be understatedAI-resilient assessment with explicit safeguards
Restorative justiceParticipation, accountability, repair, and reintegrationSometimes reduced to reflection, retraining, or task revisionEligibility criteria, safeguards, and proportional response pathway
Integrated frameworksCombination of policy, literacy, and response mechanismsLimited account of mechanisms, tensions, and institutional feedbackSeven propositions and a case-to-governance learning loop
Table 2. Review architecture and transparency measures.
Table 2. Review architecture and transparency measures.
ComponentOperational DecisionTransparency Safeguard
PurposeTheory building and normative integration, not effect estimationArticle type and limits stated explicitly
CoveragePurposive academic and policy corpus with foundational and current sourcesComplete 64-source matrix in Appendix A
Search logicIterative combinations across academic integrity, assessment, GenAI, governance, literacy, and restorationSearch concepts, venues, and example strings reported
EligibilityDirect relevance, credible provenance, and substantive conceptual, empirical, policy, or implementation contributionAcademic evidence, policy authority, and author proposals kept distinct
AnalysisComparative coding, negative-case analysis, mechanism mapping, and proposition developmentSource-to-concept audit trail and framework comparison
Quality controlCounter-position inclusion and claim-level calibrationNo claims of validation or universal effectiveness
Table 3. Search concepts, example combinations, and eligibility functions.
Table 3. Search concepts, example combinations, and eligibility functions.
Concept DomainIllustrative Search CombinationEligibility Function
Academic integrity“academic integrity” AND (generative AI OR ChatGPT, version 5.6 OR large language model)Identifies changes in authorship, disclosure, delegation, and misconduct
Assessmentassessment validity AND (generative AI OR AI-assisted assessment)Identifies evidence-of-learning, fairness, accessibility, and verification issues
AI literacy“AI literacy” AND (higher education OR students OR teachers)Identifies technical, critical, ethical, and organisational competencies
Algorithmic governance(human oversight OR explainability OR contestability) AND educationIdentifies procedural safeguards for consequential automated decisions
Restorative justice(restorative justice OR restorative practice) AND higher educationIdentifies harm, participation, repair, facilitation, and implementation conditions
Cross-domain integrationacademic integrity AND policy AND literacy AND restorativeTests whether existing frameworks specify mechanisms and feedback
The searches were iterative and purposive. They are reported for traceability and conceptual reproducibility, not as a claim of exhaustive systematic retrieval.
Table 4. Three levels of academic integrity and their principal questions.
Table 4. Three levels of academic integrity and their principal questions.
LevelPrimary ObjectCore QuestionsTypical Evidence
Integrity of academic workStudent or researcher output and processWhat contribution was made? What was delegated? Was assistance disclosed, verified, and justified?Draft history, AI-use statement, sources, process artefacts, explanation of reasoning
Integrity of assessmentValidity and fairness of the task and judgementDoes the task elicit the intended learning? Are expectations, access, and verification proportionate?Learning-outcome alignment, process evidence, moderation, accessibility review
Institutional procedural integrityRules and decisions affecting academic standingAre procedures transparent, evidence-based, reviewable, documented, and proportionate?Published rules, reasoned decisions, human review, appeal record, audit trail
Table 5. Functional classification of AI use in academic work.
Table 5. Functional classification of AI use in academic work.
DimensionLower InvolvementIntermediate InvolvementHigher InvolvementIntegrity Question
FunctionSpelling, formatting, retrievalIdea generation, feedback, translation, code suggestionsDrafting, problem solving, synthesis, decision generationDoes the function replace a target learning process?
ControlLocal edits accepted selectivelyIterative prompting and substantial revisionOutput adopted with limited human transformationWho made and can justify decisive choices?
DisclosureMinimal under task rulesUse statement and examples of influenceDetailed contribution record, prompts, verification and limitationsIs disclosure proportionate and meaningful?
Evidence of learningIndependent explanation readily availableProcess artefacts and commentary neededOral defence, reconstruction or additional evidence may be requiredCan the learner demonstrate the assessed competence?
Integrity concernLow when aligned with rules and outcomesContext-dependentHigh when use substitutes for the learning outcome or is concealedWhat validity threat or unfair advantage is created?
The categories are analytical judgements, not measured risk scores; institutions must adapt them to disciplinary and task-specific contexts.
Table 6. Epistemic responsibility obligations and proportionate evidence.
Table 6. Epistemic responsibility obligations and proportionate evidence.
ObligationMinimum ExpectationPossible EvidenceBoundary or Caution
VerifyCheck factual, source, computational, and citation claimsSource comparison, calculation trace, reference auditVerification burden should match the stakes and task
JustifyExplain why substantive choices and interpretations were retainedCommentary, oral explanation, decision memoFluent explanation alone does not prove independent authorship
DiscloseIdentify material AI functions when rules require disclosureAI-use statement, examples, prompt excerptsAvoid collecting unnecessary private or sensitive data
InterpretRelate output to disciplinary standards and learning outcomesCritical annotation, comparison with evidenceDo not assess prompt fluency instead of disciplinary knowledge
Accept responsibilityCorrect errors and answer for final claims and consequencesRevision record, acknowledgement, repair actionResponsibility remains human even when provider limitations contributed
Table 7. Allocation of responsibilities and evidence of fulfilment.
Table 7. Allocation of responsibilities and evidence of fulfilment.
ActorPrimary DutiesEvidence of FulfilmentLimits of Responsibility
StudentFollow task rules; disclose; verify; justify; protect dataUse statement, process evidence, explanation, accurate sourcingDoes not bear sole responsibility for ambiguous rules or inaccessible support
EducatorSet task-specific rules; design valid assessment; teach ethical use; review evidenceAssessment brief, examples, rubric, feedback, reasoned decisionCannot guarantee misconduct-free assessment; must avoid inaccessible verification
Programme/institutionCreate coherent policy; train staff; provide appeals; monitor patternsPolicy hierarchy, training records, audits, appeal proceduresShared responsibility does not remove individual attribution
Integrity/quality unitEnsure consistency, proportionality, data quality, and feedbackCase coding, moderation, annual review, anonymised trend reportCase counts are not direct measures of integrity culture
Technology providerDocument limitations; protect data; support review and contestabilityModel cards, contracts, privacy terms, change logs, support channelsProvider duties do not substitute for institutional procurement judgement
Table 8. Safeguards for algorithmically mediated integrity decisions.
Table 8. Safeguards for algorithmically mediated integrity decisions.
Decision StageMinimum SafeguardEvidence RequiredFailure Mode to Avoid
Procurement/deploymentDocumented purpose, legal basis, privacy review, accessibility and bias testingTool register, risk assessment, contract termsOpaque vendor claims or unreviewed model changes
Flag generationValidity evidence and clear limits on inferenceError rates, test conditions, affected groupsTreating probability or similarity as proof
Human reviewCompetence, time, authority, corroboration, and written reasonsReviewer training, evidence log, override recordAutomation bias or rubber-stamping
Student participationIntelligible explanation and opportunity to respondNotice, evidence summary, meeting recordInformation asymmetry and coerced admission
Appeal/auditIndependent review, timelines, outcome monitoring, differential-impact analysisAppeal decisions, reversals, disparity auditNominal appeal routes or unexamined systematic error
Table 9. Policy architecture from institutional principles to task-level rules.
Table 9. Policy architecture from institutional principles to task-level rules.
Policy LevelPrimary FunctionMinimum ContentOwner/Review Cycle
InstitutionalCommon values, definitions, rights, data protections, and due processAI-use principles, automated-tool limits, human review, appeal, privacyGoverning body and integrity/data functions; at least annual review
ProgrammeDisciplinary conventions, progression, and professional obligationsPermitted functions, accreditation constraints, assessment mapProgramme board; curriculum review cycle
CoursePedagogical purpose and recurring expectationsExamples, disclosure norms, support, common rubric languageCourse leader; before each delivery
TaskOperational permission and evidence requirementsAllowed/prohibited functions, disclosure, process evidence, consequencesTask designer/moderator; every assessment iteration
FeedbackRevision based on cases, appeals, access, and technology changeAnonymised trends, differential effects, recurring ambiguitiesQuality/integrity committee; periodic and event-triggered review
Table 10. Assessment strategies, integrity benefits, trade-offs, and safeguards.
Table 10. Assessment strategies, integrity benefits, trade-offs, and safeguards.
StrategyPotential Integrity BenefitPrincipal Trade-OffsSafeguard
Process portfolioShows development and decision pointsWorkload; fabricated artefacts; privacy concernsSelective checkpoints and sampled verification
Reflective AI-use commentaryMakes contribution and judgement visibleFormulaic responses; over-disclosureTask-specific prompts and privacy limits
Oral defence/demonstrationTests understanding and ownershipAnxiety, accessibility, language bias, inter-rater variabilityAdjustments, structured questions, assessor training, moderation
Context-rich projectReduces generic answer substitutionUnequal access to contexts/resources; still outsourceableResource provision and process evidence
In-class/supervised componentProvides complementary evidenceSurveillance, logistics, accessibility, authenticity concernsUse only when justified; minimise intrusive monitoring
AI-output critiqueRequires evaluation of limitations and evidenceMay assess AI familiarity more than disciplinary knowledgeExplicit teaching and alternative pathways
Table 11. Comparative assessment decision matrix.
Table 11. Comparative assessment decision matrix.
Decision CriterionPortfolio/CommentaryOral/DemonstrationContext-Rich ProjectSupervised ComponentAI-Output Critique
Best useTracing development and judgementConfirming understanding and ownershipApplying learning in situated conditionsComplementary evidence for high-stakes decisionsAssessing critical evaluation of AI output
Primary validity riskArtefacts may be manufacturedPerformance anxiety or language effectsContext resources may dominate learningMay measure test conditions rather than authentic competenceMay over-reward prior AI familiarity
Accessibility concernDocumentation burden and digital accessDisability, anxiety, language, schedulingUnequal access to field contexts or toolsMonitoring, space, time, and accommodationAlternative pathways for students with limited tool access
Staff burdenModerate to highHighModerate to highHighModerate
Recommended combinationSampling + oral clarificationPortfolio or project evidenceProcess evidence + structured reflectionUse sparingly with authentic evidenceDisciplinary task + explicit criteria
Table 12. Role-specific AI literacy for academic integrity.
Table 12. Role-specific AI literacy for academic integrity.
RoleMinimum CompetenciesAdvanced CompetenciesIntegrity Application
StudentsRecognise limits; verify output; understand task rules; protect sensitive dataExplain choices; disclose material use; evaluate bias and evidenceResponsible contribution and demonstrable learning
EducatorsUnderstand tool functions and uncertainty; communicate task rulesDesign AI-resilient assessment; interpret indicators; moderate evidenceFair assessment and educational guidance
Programme leadersMap AI literacy to outcomes and disciplinary normsCoordinate assessment and staff development across coursesConsistency and progression
Integrity/administrative staffUnderstand evidence limits, privacy, and due processAudit systems, cases, appeals, and differential impactsProcedural integrity
Procurement/IT teamsEvaluate security, documentation, accessibility, and data termsMonitor model changes, vendor claims, and contestabilityResponsible institutional deployment
Table 13. Four functions of AI literacy and their limits.
Table 13. Four functions of AI literacy and their limits.
FunctionPrimary OutcomeWhat Literacy Can SupportWhat Literacy Cannot Guarantee
Individual capacityInformed and critical tool useVerification, disclosure, error recognition, privacy judgementHonest conduct under pressure or deliberate advantage-seeking
Curricular outcomeDisciplinary and ethical competenceProgressive learning, evaluative judgement, responsible practiceUniform competence without teaching, practice, and assessment
Organisational capacityCompetent policy, oversight, procurement, and case handlingInterpretation of system limitations and procedural risksAdequate authority, time, staffing, or institutional commitment
Preventive supportReduced ambiguity and improved norm comprehensionUnderstanding of permitted functions and consequencesRemoval of incentives, opportunity, competition, or strategic misconduct
Table 14. Response criteria and safeguards.
Table 14. Response criteria and safeguards.
FactorEducational EmphasisRestorative EmphasisDisciplinary/Combined Emphasis
Rule clarityAmbiguous, inaccessible, or poorly taughtClear enough to discuss responsibility while acknowledging contextual failuresClear, accessible, and knowingly breached
IntentMisunderstanding or weak judgementResponsibility accepted; intention may be mixedDeliberate concealment or strategic unfair advantage
Impact/harmLow or readily correctedIdentifiable relational or community harm that can be addressedSerious academic, professional, safety, or equity impact
RepetitionFirst occurrenceMay address recurrence where genuine engagement is possibleRepeated conduct or failed prior interventions
ParticipationCoaching may be required as educationVoluntary, informed, facilitated, non-coerciveFormal procedure protects rights if restoration is declined or unsuitable
OutcomeClarification, rework, learning planAgreement, repair, reintegration, follow-upProportionate sanction, educational conditions, reasons, appeal
Table 15. Restorative process requirements and safeguards.
Table 15. Restorative process requirements and safeguards.
RequirementOperational QuestionMinimum SafeguardUnsuitable or High-Risk Condition
EligibilityAre facts sufficiently established and is the case safe for participation?Screening by trained staff; no presumption from automated evidenceSeriously disputed facts, acute safety risk, or incompatible legal duty
VoluntarinessCan each party decline without losing due process rights?Informed consent; no adverse inference from refusalCoercion, grade leverage, or implied waiver of appeal
ParticipationWho was affected and whose voice is appropriate?Facilitated inclusion, support person, power-sensitive designUnsafe confrontation or token representation
Responsibility and harmCan conduct and consequences be acknowledged?Clear distinction between responsibility, intent, and institutional contributionDenial where facts remain contested or institutional pressure to confess
Repair agreementAre actions specific, feasible, proportionate, and educationally relevant?Written agreement, timelines, privacy limits, monitoringPunitive or humiliating conditions disguised as restoration
Follow-up and recordsHow are completion, recurrence, confidentiality, and reintegration handled?Defined record rules, follow-up, support, anonymised institutional learningPermanent stigma, uncontrolled disclosure, or no implementation capacity
Table 16. Seven provisional propositions for empirical assessment.
Table 16. Seven provisional propositions for empirical assessment.
PropositionMechanismBoundary Condition or Counter-CaseIllustrative Indicator
P1. Task-level policy clarity reduces ambiguity-related breaches.Students can connect principles to permitted functions and disclosure.Clear policy may not deter strategic misconduct.Coverage and comprehension of task-specific guidance
P2. AI literacy supports judgement but does not guarantee conduct.Knowledge improves evaluation, disclosure, and verification.Pressure and perceived advantage may override knowledge.Literacy assessment plus behavioural/process evidence
P3. AI-resilient assessment improves evidential validity when accessible and proportionate.Complementary evidence makes contribution and understanding more interpretable.Added burden may reduce fairness or reliability.Validity, access, workload, moderation, differential effects
P4. Governed verification is more legitimate than automated suspicion.Corroboration, human reasons, and appeal reduce error and procedural harm.Reviewers may exhibit automation bias or lack authority.Override rates, appeal outcomes, differential error
P5. Differentiated responsibility improves prevention when duties remain attributable.Institutional and provider failures become visible without erasing student agency.Vague shared responsibility may diffuse accountability.Role-specific compliance and reason-giving
P6. Restorative processes support learning and repair when substantive safeguards are present.Dialogue and agreed repair address relational consequences and reintegration.Coercion, serious harm, or disputed facts may make restoration inappropriate.Consent quality, completion, recurrence, participant experience
P7. Case-to-governance feedback produces institutional learning when patterns lead to documented change.Cases reveal ambiguity, assessment weakness, access inequity, or procedural problems.Case data may be incomplete, biased, or misinterpreted.Policy/assessment changes linked to anonymised case analysis
Table 17. Stable and adaptable elements of the framework.
Table 17. Stable and adaptable elements of the framework.
ElementExpected Stability Across ContextsRequired Local Adaptation
Human responsibilityA person or institution remains answerable for consequential academic decisions.Role titles, committee structures, professional obligations
Transparency and reasonsMaterial rules and consequential decisions should be intelligible.Disclosure format, language, disciplinary conventions
Assessment validity and fairnessEvidence should match learning outcomes and avoid unjustified disadvantage.Assessment methods, accommodations, resource constraints
ContestabilityStudents need meaningful human review and appeal.Legal procedure, timelines, representation, record retention
Restorative safeguardsVoluntariness, participation, harm, responsibility, and repair are essential.Facilitation model, cultural practices, eligible cases
Continuous reviewPolicies and systems should be revisited using evidence.Indicators, review cycle, institutional capacity
Table 18. Staged implementation and evaluation roadmap.
Table 18. Staged implementation and evaluation roadmap.
StageResponsible ActorsPriority Actions and ResourcesOutputs and IndicatorsInterpretive Caution
0–3 months: minimum safeguardsSenior leadership, integrity, legal/data protection, studentsDefine scope; prohibit automated proof; require reasons and appeal; publish task templateApproved interim policy; named owners; accessible appeal routeFormal publication does not establish comprehension or implementation
3–9 months: curriculum and assessmentProgramme leaders, educators, learning designers, access servicesMap outcomes; identify vulnerable assessment; embed literacy; design accommodations; allocate staff timeProgramme map; revised briefs/rubrics; participation and comprehension dataMore process evidence can increase burden without improving validity
9–18 months: restorative and technology infrastructureIntegrity units, trained facilitators, QA, IT/procurementSet eligibility and record rules; train facilitators; audit tools/vendors; standardise anonymised case codingFacilitator pool; tool register; case-quality rubric; baseline equity and appeal indicatorsSatisfaction or low case counts do not prove effectiveness
Ongoing: assurance and adaptationGovernance committee, QA, programmes, staff and studentsReview appeals, differential impacts, recurrence, workload, policy comprehension, and model changesAnnual public summary; documented policy/assessment changes; action closure rateChanges may be symbolic unless implementation and outcomes are audited
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bădescu, G.; Susinski, M.; Vasile, C.; Săvescu, P.; Constantinescu, E.; Tănasie, G.; Dima, N.; Filip, L.-O.; Savu, A.; Didulescu, C. Preventive and Restorative Academic Integrity in AI-Assisted Higher Education: A Critical Conceptual Synthesis and Integrated Governance Framework. Educ. Sci. 2026, 16, 1365. https://doi.org/10.3390/educsci16091365

AMA Style

Bădescu G, Susinski M, Vasile C, Săvescu P, Constantinescu E, Tănasie G, Dima N, Filip L-O, Savu A, Didulescu C. Preventive and Restorative Academic Integrity in AI-Assisted Higher Education: A Critical Conceptual Synthesis and Integrated Governance Framework. Education Sciences. 2026; 16(9):1365. https://doi.org/10.3390/educsci16091365

Chicago/Turabian Style

Bădescu, Gabriel, Mihail Susinski, Cristian Vasile, Petre Săvescu, Emilia Constantinescu, Gabriel Tănasie, Nicolae Dima, Larisa-Ofelia Filip, Adrian Savu, and Caius Didulescu. 2026. "Preventive and Restorative Academic Integrity in AI-Assisted Higher Education: A Critical Conceptual Synthesis and Integrated Governance Framework" Education Sciences 16, no. 9: 1365. https://doi.org/10.3390/educsci16091365

APA Style

Bădescu, G., Susinski, M., Vasile, C., Săvescu, P., Constantinescu, E., Tănasie, G., Dima, N., Filip, L.-O., Savu, A., & Didulescu, C. (2026). Preventive and Restorative Academic Integrity in AI-Assisted Higher Education: A Critical Conceptual Synthesis and Integrated Governance Framework. Education Sciences, 16(9), 1365. https://doi.org/10.3390/educsci16091365

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop