1. Introduction
A supervisor who receives a well-structured, fluently argued, fully referenced student manuscript now faces a question the finished product cannot answer: what did the student actually practise? Generative AI has made the written artefact a poor proxy for the cognition behind it, and the instructional response has converged on a division of labour—rules specifying which part of the work a student may delegate and which part must remain their own (
Bearman et al., 2024). Such rules are attractive because they are enforceable at the level of the assignment. Whether they are effective at the level of the learner is a separate question, and it is the question this study addresses.
Three strands of literature bear on this question, and they have rarely been integrated empirically.
The first concerns what is offloaded. Cognitive offloading—the transfer of cognitive demand to an external resource—long predates generative AI and carries no fixed valence (
Risko & Gilbert, 2016;
Sparrow et al., 2011); its consequences depend on how the external resource is enrolled in the task (
Grinschgl et al., 2023;
Silva et al., 2026;
Skulmowski, 2023). The two faces can be separated within a single sample:
W. Fan et al. (
2026) found that the cognitive relief students perceive from classroom AI use and the cognitive offloading they perform through it mediate their attitudes in opposite directions, the former complementary and the latter competitive. Work on generative AI has increasingly resolved a single dimension of use into two: engagement in which the learner keeps the cognitive lead and uses the system as scaffolding, and engagement in which the learner transfers the judgement itself (
Y. Fan et al., 2025;
Zhu et al., 2026). The same move has been made independently on engagement outcomes:
Y. Wang et al. (
2026) separated reflective from thoughtless AI use and argued that research has attended to how much students use AI at the expense of how they use it. That distinction predicts downstream outcomes better than frequency of use does. It also produced a finding whose implication has not been tested: the two forms of engagement felt equally satisfying in the moment despite opposite downstream associations, which led to the conjecture—explicitly labelled by its authors as consistent with, but not demonstrating, detection failure—that learners cannot tell from experience which kind of use they are engaged in (
Zhu et al., 2026).
The second strand concerns where in a task the offloading occurs.
Lai et al. (
2026) decomposed writing into drafting, evaluating, and revising and compared three human–AI divisions of labour in a quasi-experiment with 101 EFL undergraduates. Having AI draft while the student evaluated and revised outperformed the more familiar arrangement in which the student drafts and AI evaluates, and both outperformed writing without AI. The peer-review results were more pointed still: students who had been assigned the evaluating role for four rounds were no better at identifying problems than students who had used no AI at all. The stage that is delegated, on this evidence, matters more than the amount delegated. Three gaps remain: the design contained no condition in which AI performed the revision, which the authors identified as a limitation; no individual differences were measured; most consequentially for interpretation, evaluation by students in the evaluating role was prescribed by the assignment rather than established through observed behaviour. Process-tracing evidence has since made that last gap pressing. Synthesising 33 studies of learner interaction with AI-mediated feedback in second-language writing,
Alghamdi and Alghizzi (
2026) reported that higher-regulation learners engage in selective uptake and recursive evaluation of AI feedback, whereas lower-regulation learners more often accept it rapidly and with reduced evaluative engagement. Those profiles are observed, not assigned, which is exactly what leaves open whether prescribing the evaluative role does anything to them.
The third strand explains why the evaluating stage should be the one that matters. Evaluative judgement—the capacity to judge the quality of work, including one’s own and others’—is treated in higher education as a core capability rather than an ancillary skill (
Tai et al., 2018), and feedback literacy research locates the learning value of feedback in the learner’s processing of it rather than in its receipt (
Carless & Boud, 2018;
Molloy et al., 2020). Both traditions have been revisited in light of generative AI, with the concern that a system able to deliver instant authoritative-seeming quality judgements removes the occasion on which a student’s own judgement would otherwise be exercised (
Bearman et al., 2024). This literature supplies the mechanism but almost no experimental evidence. It does not tell us whether a student told to evaluate will evaluate.
Basic learning research suggests why evaluation in particular should resist delegation, while drafting and revising may not. Retrieval and the effortful generation of a response produce retention advantages over passive reception (
Roediger & Karpicke, 2006;
Slamecka & Graf, 1978), and conditions that slow acquisition while improving durable learning—desirable difficulties—are systematically undervalued by learners, who prefer the fluent option (
E. L. Bjork & Bjork, 2011). Diagnosing what is wrong with a text is such a condition: it is slow, effortful, and unrewarding relative to receiving a diagnosis. Cognitive load accounts add a boundary, since removing extraneous load can help while removing the germane processing that constitutes the learning defeats the purpose (
Kirschner et al., 2006;
Sweller, 1988). In this case, the three writing stages are not interchangeable loads. Drafting and revising are largely productive operations that AI can shoulder without eliminating the learner’s cognitive work; evaluation is the stage at which the germane processing lives, so delegating it removes the learning rather than the burden. The prediction is therefore directional and theoretically motivated, not merely a contrast among conditions. It also implies the asymmetry we test in RQ3: if learners systematically prefer fluency and undervalue difficulty, the arrangement that teaches least may be the one that feels best, and self-report will not distinguish them.
Human–computer interaction research provides the closest methodological precedent. Cognitive forcing functions—interface constraints that require a person to commit to their own judgement before seeing the system’s recommendation—reduce over-reliance, at the cost of subjective experience: participants find them effortful and dislike them (
Buçinca et al., 2021). This establishes that compelling engagement is possible, and it hints at the asymmetry central to the present study, since the intervention that improved performance was the one that felt worse. The paradigm, however, is one-shot decision-making with an unambiguous correct answer, not a multi-week learning task, and its outcome is reliance on advice rather than what the person retains.
The Gap and the Present Study
Put side by side, the three strands leave a specific gap. What an instructor can prescribe is the stage; what the student brings is a disposition toward how AI is used. If the two interact, then a single rule produces systematically different outcomes for different students, and its average effect conceals the fact that the students the rule was written for may be the ones it fails. Supervisors observe exactly this: they require students to revise their own drafts, and find that some students develop while others forward the revision to the model.
We therefore randomised undergraduates to four divisions of labour across four rounds of authentic, graded report writing, withdrew AI entirely at post-test, and measured three objective outcomes rather than relying on self-report: unaided transfer performance, peer-review quality, and the rate at which students caught known errors planted in an AI-generated manuscript. Baseline dependent offloading tendency entered as a continuous moderator, and prompt-level logs supplied a behavioural indicator of whether the assigned evaluation was performed or re-delegated.
Three questions organised the design:
RQ1 (stage). Which stage of report writing, when delegated to AI, most affects independent writing once AI is withdrawn? We predicted that arms in which AI drafted (A) or revised (C) would outperform the arm in which AI evaluated (B), and that all AI arms would outperform the no-AI control (D) on transfer (H1a). On peer-review quality, we predicted that arm A would exceed B and C, and—following
Lai et al. (
2026)—that arm B would
not exceed the no-AI control on problem identification (H1b).
RQ2 (interaction; the study’s core). Does the student’s dependent offloading tendency moderate the effect of the assigned stage? We predicted a reliable arm × tendency interaction such that the advantage of arm A weakens as tendency rises (H2), and that within arm A the proportion of prompts that re-delegate judgement rises with tendency, providing behavioural rather than self-report evidence for the mechanism.
RQ3 (detection). Can students detect what AI-generated text gets wrong? We predicted higher catch rates in arm A than in B and C (H3a); that catch rate would be uncorrelated or weakly correlated with immediate satisfaction and self-rated quality (H3b), which would test one behavioural implication of that conjecture—that the subjective experience of the work does not covary with how much of the text’s error the student actually caught; that dependent tendency would negatively predict catch rate after baseline ability was controlled (H3c). H3b concerns the informativeness of these two global judgements, not students’ metacognition about the detection task itself, which this design did not measure.
Figure 1 sets out the resulting conceptual model. The two-dimensional framing of AI use is not new; several research groups have converged on it independently, and a scoping review of generative AI, cognitive offloading, and learner agency in higher education now maps the literature they have produced (
G. Wang et al., 2026). What this study adds is a closed loop: correlational evidence replaced by randomisation, self-report by behaviour, and the question of which stage should be offloaded by the question of for whom the answer holds.
2. Materials and Methods
2.1. Design
The experiment was a four-arm, between-participants randomised design. The writing task was fixed as a three-stage cycle—draft, evaluate, revise—and the manipulation was which stage the AI performed:
Arm A: AI produces the draft; the student evaluates it against the rubric and then revises it.
Arm B: the student drafts; AI evaluates that draft against the rubric and returns feedback; the student revises accordingly.
Arm C: the student drafts; AI returns a revised version directly, without presenting a separate evaluation; the student then adjudicates each change, accepting or rejecting it with a stated reason.
Arm D: the student performs all three stages without AI.
Arms A, B, and D correspond to the three conditions of
Lai et al. (
2026), permitting direct comparison. Arm C supplies the condition those authors identified as missing, in which the revision itself is delegated. All three AI arms used the same model (Qwen 3.0, Alibaba Cloud), accessed at the provider’s default settings with web search disabled, and the same prompt-template pool; the only difference was what the student had to do with the AI’s output.
Two features of this operationalisation determine what the arms identify. First, the arms are workflows rather than an orthogonal decomposition of the three stages. In arm C, the model must appraise the student’s draft in order to revise it, even though that appraisal is never surfaced as feedback and the student never performs it; diagnosing one’s own draft is thus work the student does not do in arm C, as it is work they do not do in arm B. What separates arm C from arm B is the direction of the evaluative act: in arms A and C, the student renders judgement on text the AI produced, whereas in arm B, the AI renders judgement on text the student produced and the student executes it. Second, the object of that judgement differs between arms A and C. Arm A students appraised a whole AI-generated manuscript each round; arm C students appraised a set of localised changes. This difference yields a testable prediction: a task requiring whole-text diagnosis should build a capability that adjudicating discrete edits does not, and the error-detection outcome bears on it directly.
The task in each round was a report of 2500–3000 words containing a research question, a literature basis, an argument, and a conclusion—a miniature of a thesis chapter, and graded as part of the course. Theses themselves were not used. Randomly assigning some students to a less effective division of labour for work that determines whether they graduate places a degree at risk, and no ethics committee should approve it. The report preserved the authentic stakes and the same cognitive structure while removing that risk.
The platform constrained use to the assigned stage. Stages were released in sequence, so one could not begin before the previous had been submitted; the model was reachable only where the arm’s design placed it, the exception being arm A, where the student could consult the model freely once the draft had been returned; every request and response was logged. Within the platform, adherence was complete: every participant in an AI arm produced interactions, every participant in arm D produced none, and self-report agreed with the logs in every case.
Platform logs cover use within the platform. Compliance in arm D was verified against self-report as well as logs, with no case identified; consultation of a public model outside the platform would not appear in either record, as in any field experiment on a tool the participants can reach independently. The estimates are effects of the assigned arrangement under those conditions.
2.2. Participants and Sites
Participants were undergraduates at two institutions in a provincial capital in central China, referred to throughout as Site A and Site B; the institutions are not identified. Site A supplied the primary sample and ran the full four-arm protocol. Site B supplied a reduced three-arm conceptual replication (arms A, B, and D only, without the peer-review task).
The two sites differ in selectivity in a sense specific to the Chinese admissions system, in which undergraduate entry is determined almost entirely by a centrally administered examination and institutions admit within published score bands: Site B admits in the highest of those bands, Site A in a lower one. Both teach and assess in Chinese, and all tasks, rubrics, AI interactions, and instruments were in Chinese; the instruments were translated for reporting only.
Participants were third-year undergraduates majoring in e-commerce, a programme classified under different faculties at the two institutions; we used that classification as a stratum in the randomisation. The course ran across twelve class sections at Site A and four at Site B, taught by six instructors in total, and the study was conducted in the spring semester of the 2024–2025 academic year. The two sites were closely comparable on baseline experience with generative AI. The consent protocol collected the minimum personal information the analysis plan required, which did not extend to age or gender.
Appendix B reports the composition figures.
Students were taught in intact classes, so allocation was nested within teaching units. Assignment was at the individual level within stratum rather than by class, which prevents class from being confounded with each arm but also means students in different arms sat in the same room; any resulting contamination would attenuate differences between arms rather than manufacture them. Class identifiers were not retained in the analysis dataset, so no model represents class- or instructor-level dependence, and faculty classification, which has two categories, cannot stand in for it. Because assignment was individual, a difference shared by all students in a class does not bias the contrasts between arms; residual correlation within classes or instructors could, however, make their standard errors too small.
At Site A, 744 students consented and were randomised, totalling 186 per arm. At Site B, 126 students consented and were randomised, totalling 42 per arm. Retention at the Week 6 post-test was 84.9% to 87.1% across arms at Site A (639 analysed) and 81.0% to 92.9% at Site B (109 analysed).
Figure 2 reports the full participant flow.
The primary sample was placed at the less selective site for two reasons. The interaction test required roughly 520 completers, which the smaller cohort at Site B could not supply; locating the main conclusions in the population where AI dependence is most pronounced increases their practical relevance. Because institution type confounds intake, curriculum, and AI access, all primary tests were conducted within site, and the single pooled analysis reported below treats the site as a fixed-effect covariate. Site B is described as a conceptual, not a direct, replication.
2.3. Randomisation and Blinding
Randomisation was carried out separately at each site, within strata defined by baseline writing tertile and faculty classification. The allocation sequence was generated by the course platform rather than by any member of the research or teaching team; no teaching staff had access to it, and allocation was concealed until a participant had completed baseline measurement. The resulting arms were of equal size and did not differ on any baseline measure (
Table 1). The allocation procedure and the stratum-by-arm counts are presented in
Appendix B.
Participants could not be blinded to their own condition since the workflows visibly differ, but they were not told the direction of any hypothesis. Raters were blind to both condition and phase. All submissions were stripped of metadata, AI traces, and outline residue and shuffled before rating. Raters completed rubric training and a calibration set of 12–15 scripts, and rating proceeded only after ICC(2,
k) ≥ 0.75 was reached on that set. Each script was rated by three raters, with a fourth independent rating whenever two raters differed by two points or more. Inter-rater reliability in the final sample ranged from ICC(2,
k) = 0.787 to 0.846 across outcomes and sites, and rater main effects were not reliable for any outcome; per-outcome coefficients are presented in
Appendix B.
2.4. Procedure
Week 1. Informed consent, baseline measurement, rubric training, and platform training. Baseline writing ability was assessed with an 800-word research note written in 45 min without AI and rated by assessors blinded to condition and phase; it served as both a stratifying variable and a covariate. Trait measures were administered in randomised order. All participants received identical rubric training and completed calibration items afterwards, so that arms began with comparable evaluative knowledge.
Weeks 2–5. Four rounds of report writing, one topic per round, counterbalanced across participants. Within each round, the sequence draft → evaluate → revise → submit was fixed; arms differed only in which stage the AI performed. The platform logged all interactions.
Week 6. Post-test, entirely without AI, comprising the transfer task, the peer-review task, the error-detection task, and post-measures. The post-test followed immediately after the final intervention round rather than after an interval, so it measures performance once AI is withdrawn rather than retention over time.
2.5. Measures
Transfer performance (primary outcome). An independent report on a new topic of equivalent difficulty, written without AI and rated on a six-dimension rubric by assessors blinded to condition and phase, scored 0–24.
Peer-review quality. Within 25 min, participants evaluated a common sample manuscript. Ratings covered problem identification, problem explanation, and constructive suggestion, each scored 0–5, using the same operationalisation as
Lai et al. (
2026) to permit comparison; a score of 0 records a dimension the reviewer did not address at all. This task was administered at Site A only.
Error detection. All participants received the same AI-generated report of approximately 1200 words containing 12 known planted errors, four of each type: fabricated citations (well-formed references to non-existent sources, and real sources whose conclusions were misattributed), logical fallacies (hasty generalisation, reversed causation, circular argument), and evidence–conclusion mismatches (data pointing opposite to the stated conclusion, inferential strength unsupported by sample size). Fabricated citations were included because they are the failure mode supervisors encounter most often in practice (
Walters & Wilder, 2023). Each planted item is verifiable independently of any rater’s judgement: a fabricated citation either does or does not correspond to an existing source, a misattributed one either does or does not represent what its source concluded, and the logical and evidential faults were constructed to instantiate named fallacies rather than to be judged defective on impression. Two researchers confirmed each item against its type definition before piloting. Participants marked problems and stated why within a time limit. Scoring therefore required human judgement rather than string matching: a marked passage counted as a catch only when the accompanying justification identified the planted fault, so that marking the right passage for an unrelated reason did not score, and false positives were counted by the same procedure. Each response was scored by a single trained assistant who was blind to condition, so no inter-scorer agreement estimate is available for this outcome. Scorer latitude was limited instead by a fixed key that specified, for each of the twelve planted errors, what a justification had to identify for the item to count as caught, which reduced every scoring decision to a binary judgement about one item against a stated criterion. That key is provided with the instrument, so any scoring decision can be checked against it. The instrument itself was calibrated in a pilot of 8–10 participants per arm: the target band for each error type was a catch rate between 0.2 and 0.8, and any item whose type-level rate exceeded 0.9 or fell below 0.1 was rewritten before the main study. The catch rate was the number identified out of 12, computed separately by type. False positives—correct passages marked as errors—were recorded and entered the models as a covariate so that indiscriminate marking cannot inflate the catch rate. The item statistics in the final sample confirm that this calibration held. The twelve-item instrument was internally consistent, KR-20 = 0.619; the four-item type subsets are too short to support a confirmatory claim, so all confirmatory detection analyses use the twelve-item total and results by error type are exploratory (
Appendix B).
Interaction logs. Prompts were stored with full context. The behavioural indicator for RQ2 was the re-outsourcing ratio: the proportion of a participant’s prompts that ask the model to make or substitute a judgement (
what is wrong with this paragraph,
just fix it for me) rather than to explain or justify. The unit of coding was the individual prompt rather than the session log, and the ratio is the proportion of a participant’s prompts classified as re-delegating. Two coders independently coded a randomly selected 20% of all prompts; agreement on that subset was Cohen’s
= 0.83, above the threshold set in the protocol, after which disagreements were resolved by discussion and one coder classified the remaining 80% using the agreed scheme. The scheme, with its decision rules and the conventions that settled borderline cases, is presented in
Appendix A. Coders saw the prompt text and arm—arm is inferable from the workflow and could not be masked—but had no access to participants’ offloading scores or to any outcome measure. The ratio is informative in arm A, where the student’s evaluative work was carried out in free interaction with the model and could therefore be handed back to it. It is structurally zero in arms B and C for different reasons: in arm B the student had no evaluative task to delegate, whereas in arm C the adjudication was completed inside a structured platform form—accept or reject each change, with a stated reason—which provided no channel to the model. It is undefined in arm D. This difference between arms A and C is a property of the workflow rather than of the students, and it bears directly on what a design can enforce.
Self-report measures. Dependent and autonomous offloading tendency were measured with the four-item subscales of the instrument developed by
Zhu et al. (
2026), metacognitive monitoring with the same source, along with AI use frequency and faculty classification. Both subscales were internally consistent at both sites, with
and McDonald’s
between 0.79 and 0.83 (
Appendix B). Metacognitive monitoring enters as a composite score, and the two post-measures are single items; internal consistency is not defined for either.
2.6. Analysis
Sample size was fixed in advance by Monte Carlo simulation of the primary test, the arm × tendency interaction, which consumes three degrees of freedom and has substantially lower power than the corresponding main effect. At
= 0.05 an interaction of Cohen’s
f = 0.20 reaches power of 0.82 with 130 completers per arm, and we recruited to 720 against that target.
Appendix B reports the simulation parameters and the power of the remaining tests.
The analysis plan—the three planned contrasts, the smallest effect size of interest, the correction families, and the decision to test equivalence where a null would carry weight—was fixed before data collection and is provided in full as
Supplementary Material. The plan was fixed internally within the research team and carries no third-party timestamp; these analyses are therefore pre-specified rather than preregistered. Every analysis added after the data were seen is identified as such where it is reported. Participants are analysed in the arm to which they were randomised, with no reassignment by compliance or by what they actually did. Because the outcomes are measured only at the post-test, the primary analyses are complete-case analyses conducted according to randomised assignment rather than intention-to-treat in the strict sense: 744 students were randomised at Site A and 639 contributed outcome data, so the 105 who did not complete the post-test are absent from the outcome models. The attrition analysis is reported in full below. Effect sizes are reported as point estimates with 95% confidence intervals, and
p-values are not used as the sole basis for any conclusion. A smallest effect size of interest of
d = 0.30 was set in advance; effects below it are reported on both scales even when statistically reliable.
H1a was tested by ANCOVA of transfer on arm with baseline as a covariate, followed by three pre-specified contrasts under the Holm correction: A versus B, C versus B, and the mean of the AI arms versus D. H1b was tested by MANOVA across the three peer-review dimensions, with univariate follow-ups. H2 was tested as the arm × centred tendency interaction, with autonomous tendency also in the model because the two are not opposite poles of one dimension; tendency was never dichotomised. Simple slopes and a Johnson–Neyman region of significance followed a reliable interaction. H3 comprised the arm effect on catch rate, the association of catch rate with subjective experience, and the individual-difference model. Multiple-comparison correction was applied within families—Holm for the contrast and peer-review families, Benjamini–Hochberg for the detection family—and not applied to the single pre-specified interaction test, a decision fixed in advance and reported here rather than a post hoc choice.
Where a null result carries interpretive weight, equivalence was tested rather than inferred from non-significance, using two one-sided tests against bounds of ±0.30
d (or ±0.10 for correlations). Attrition was analysed before outcomes were examined. Analyses used Python 3.12 (statsmodels, SciPy); the bootstrap used 5000 resamples under a fixed seed. The package websites are
https://www.statsmodels.org/ and
https://scipy.org/ (both accessed on 15 September 2026). The analysis plan and the analysis scripts are provided as
Supplementary Material so that every reported estimate can be traced both to the specification that called for it and to the code that produced it; the participant-level data are subject to the access conditions set out in the Data Availability Statement.
2.7. Ethics
The principal ethical risk in this design is the power asymmetry between instructors and their own students, and the protocol addressed it structurally. Instructors had no access to allocation lists and performed no rating; allocation and data handling were carried out by third-party research assistants. Research ratings and course grades were kept physically separate, produced by different people through different procedures and never shared. The consent form stated that participation, non-participation, and withdrawal at any time without providing a reason would have no bearing on any course evaluation or on the student–instructor relationship. Students who declined were offered equivalent conventional writing instruction.
The four intervention rounds were graded coursework, marked against a single rubric applied irrespective of arm—completeness of content, depth of analysis, structure and expression, quality of visual presentation, and clarity of the research question and its conclusions—and contributed 20% of the course mark. Students in the control arm worked without AI assistance throughout, so allocation could have affected their marks on this component; every participant received the full training materials for the most effective division of labour once the study ended. Course grades were held separately from the research dataset and were never linked to it. Research ratings played no part in any course mark, were produced by different people through a different procedure, and were never shared with the instructor. The study’s own outcomes are unaffected in any case: transfer, peer review and detection were all assessed at the Week 6 post-test, produced without AI by every arm under identical conditions, and carried no course credit.
3. Results
3.1. Randomisation, Attrition, and Compliance
Randomisation produced balanced arms at Site A on every baseline variable: writing ability,
F(3, 740) = 0.04,
p = 0.990; dependent tendency,
F = 0.39,
p = 0.758; autonomous tendency,
F = 0.51,
p = 0.674; metacognitive monitoring,
F = 0.48,
p = 0.699 (
Table 1).
Attrition was 14.1% overall and unrelated to arm, (3) = 0.39, p = 0.943. Completers and non-completers did not differ on baseline writing ability, t = −0.42, p = 0.679; dependent tendency, t = 0.12, p = 0.907; autonomous tendency, t = −0.17, p = 0.862; metacognitive monitoring, t = −1.61, p = 0.107. Compliance in arm D was complete on both records, logs and self-report, agreeing in every case, so the sensitivity analysis the plan specified for non-compliance did not arise.
Completion was unrelated to arm and every baseline measure, and re-estimating the primary models with inverse-probability weights or after multiple imputation moves no estimate by more than 0.08 transfer points (
Table 2;
Appendix B).
3.2. H1a: Which Stage Was Offloaded?
Transfer performance differed by arm,
F(3, 634) = 8.18,
p < 0.001,
= 0.037. Means were 16.01 (
SD = 4.18) in arm A, 14.59 (SD = 3.89) in arm B, 15.92 (SD = 3.56) in arm C, and 14.60 (SD = 3.86) in arm D (
Figure 3a).
The three pre-specified contrasts (
Table 3) were all statistically reliable. Having AI draft outperformed having AI evaluate,
d = 0.35, 95% CI [0.13, 0.57], Holm-adjusted
p = 0.004; likewise, having AI revise outperformed having AI evaluate,
d = 0.36 [0.14, 0.58],
p = 0.004; combining AI arms outperformed the no-AI control,
d = 0.23 [0.05, 0.41],
p = 0.012.
Those three contrasts do not by themselves test every component of H1a, which, as stated, requires each AI arm to exceed the control individually. The aggregate contrast of A, B, and C against D can be reliable while one of its constituents is not, and that is what happened here. All six pairwise contrasts appear in
Table 3, each marked as planned or added, with the Holm correction applied across the six. Arm A exceeded the control,
d = 0.35 [0.13, 0.57], adjusted
p = 0.009, as did arm C,
d = 0.36 [0.13, 0.58], adjusted
p = 0.009. Arm B did not,
d = −0.004 [−0.22, 0.22], adjusted
p = 1.000. Arms A and C did not differ,
d = 0.02 [−0.20, 0.24], adjusted
p = 1.000.
H1a is therefore supported in two of its three components and refuted in the third. Both arrangements in which the student judged AI-produced text outperformed writing without AI at effect sizes above the smallest effect size of interest we set in advance. The arrangement in which AI judged student-produced text did not. The combined-AI advantage over the control, though reliable, is d = 0.23, below that threshold—this an average that, as the component contrasts show, is pulled down by including an arm that conferred no benefit at all.
That failure is the more informative half of the result, so we tested it as a claim rather than treating non-significance as evidence. Arm B was statistically equivalent to the no-AI control on transfer,
d = −0.004 [−0.22, 0.22], TOST against ±0.30
d,
p = 0.004, and arms A and C were statistically equivalent to each other (difference = 0.09 points, TOST
p = 0.007). The equivalence test of B against D was specified in advance, on the ground that a null there would carry interpretive weight; the equivalence of A and C was not, and is reported as such. Four rounds of AI-generated feedback on the student’s own draft—the arrangement most often described in accounts of classroom practice (
Alghamdi & Alghizzi, 2026;
Bearman et al., 2024)—left students no better off at post-test than four rounds of writing with no assistance.
The benefit therefore does not attach to AI involvement as such. What distinguishes the two effective arms from arm B is the direction of the evaluative act: in arms A and C, the student rendered judgement on text the model had produced, whereas in arm B, the model rendered judgement on text the student had produced. Evaluation was not retained by the student in arm C in any absolute sense—appraisal of the student’s own draft was performed by the model there as in arm B—so the transfer data are consistent with the direction of judgement as the difference that matters, although the arms differ in other respects as well and do not isolate it. The detection data reported below show that this is not the whole story.
3.3. H1b: Peer-Review Quality
The multivariate effect of each arm across the three peer-review dimensions was reliable: Wilks’
= 0.967,
F(9, 1538.3) = 2.39,
p = 0.011. Univariate follow-ups (
Table 4) showed effects on problem explanation,
F(3, 634) = 5.96, Holm-adjusted
p = 0.002,
= 0.027, and constructive suggestion,
F = 3.58,
p = 0.027,
= 0.017, with problem identification not reaching the threshold,
F = 2.37,
p = 0.070.
H1b predicted that arm A would exceed both arm B and arm C. Both comparisons are reported across all three dimensions, with the Holm correction within the pre-specified family (
Table 4). Arm A exceeded arm B on problem explanation,
d = 0.41 [0.19, 0.63], adjusted
p = 0.003, on problem identification,
d = 0.24 [0.02, 0.46], adjusted
p = 0.139; on constructive suggestion,
d = 0.26 [0.04, 0.48], adjusted
p = 0.125. Arm A likewise exceeded arm C on problem explanation,
d = 0.35 [0.13, 0.57], adjusted
p = 0.017, on identification,
d = 0.26 [0.04, 0.48], adjusted
p = 0.125; on suggestion,
d = 0.28 [0.06, 0.50], adjusted
p = 0.089.
H1b is therefore partially supported. All six contrasts run in the predicted direction, and the effect sizes are comparable in magnitude, but only the problem-explanation dimension survives correction against either comparison arm. Explaining what is wrong with a text is the dimension on which whole-text diagnosis paid off; identifying that something is wrong and proposing a fix did not separate the arms reliably with this sample size.
The finding that carries the theoretical weight is the comparison of arm B against the no-AI control. On problem identification, the two were indistinguishable:
d = −0.03 [−0.24, 0.19]; the same held for explanation,
d = −0.05, and suggestion,
d = 0.04 (
Figure 3b). Because non-significance alone would not support the claim, we tested equivalence directly: all three dimensions were statistically equivalent against bounds of ±0.30
d (
p = 0.007, 0.014, and 0.011). This replicates the most counter-intuitive result reported by
Lai et al. (
2026) and strengthens it, since equivalence here is demonstrated rather than inferred from a failure to reject. Four rounds in which AI performed the evaluation left students no better at evaluating than four rounds with no AI at all.
3.4. H2: For Whom the Stage Effect Holds
The interaction of arm with dependent offloading tendency was reliable (F(3, 629) = 7.25, p < 0.001, = 0.033) in a model accounting for 24.5% of the variance in transfer. It survived the removal of all covariates (F(3, 631) = 6.77, p < 0.001) and the addition of faculty classification as a fixed covariate (F(3, 628) = 7.22, p < 0.001).
Simple slopes locate the interaction entirely within arm A (
Table 5). Among students who had an AI draft, a one-point difference in dependent tendency was associated with 1.41 fewer transfer points (95% CI [−1.97, −0.85],
p < 0.001). In the other three arms, the slope was indistinguishable from zero: arm B,
b = −0.07,
p = 0.806; arm C,
b = 0.23,
p = 0.367; arm D,
b = −0.10,
p = 0.719 (
Figure 4a). Tendency was measured rather than manipulated, so these slopes describe how the benefit was distributed across students, not what a change in tendency would cause. The slopes above are estimated within each arm, including the specification the analysis plan named. Deriving them instead as conditional effects of the interaction model provides closely similar values and identical conclusions—arm A,
b = −1.32, 95% CI [−1.84, −0.81]; arm B,
b = −0.22,
p = 0.403; arm C,
b = 0.25,
p = 0.326; arm D,
b = −0.05,
p = 0.843—and it is that model’s fit that
Figure 4a plots.
The Johnson–Neyman analysis converts this into the quantity an instructor would want (
Figure 4b). At the low end of the tendency distribution, having AI draft was worth 3.97 points over writing without AI. The advantage declined monotonically and ceased to be distinguishable from zero above a tendency score of 3.45 on the five-point scale—a value 0.48 standard deviations above the sample mean of 2.93, and one exceeded by roughly a third of the sample. Above that threshold, the advantage of arm A was no longer distinguishable from zero at the precision this sample affords, which is not the same as its being zero; the point estimate remains positive, and the confidence interval includes values an instructor would care about. What the analysis supports is that the arrangement producing the largest average benefit produced no
detectable benefit for the students who most habitually delegate judgement to AI.
The logs indicate what those students did with the assignment. Within arm A, where evaluation was the student’s responsibility, the proportion of prompts that re-delegated judgement rose with dependent tendency (b = 0.072 per scale point, 95% CI [0.054, 0.090], p < 0.001), and that proportion was itself negatively associated with transfer (r = −0.31, p < 0.001). Assignment to the evaluative role did not make evaluation happen; students high in dependent tendency forwarded it back to the model. Mean re-outsourcing in arm A was 0.28.
Re-outsourcing does not statistically mediate the tendency effect on transfer. The indirect path was not distinguishable from zero (ab = −0.03, bootstrap 95% CI [−0.40, 0.36]) and the b path was not reliable once tendency was in the model (b = −0.42, p = 0.863), while the direct effect of tendency within arm A remained (b = −1.40, p < 0.001). The prompt ratio therefore indexes the disposition rather than carrying its effect: it establishes that re-outsourcing occurs and scales with tendency, while the transfer shortfall runs through something other than this single measured behaviour.
3.5. H3: Detection and What Subjective Experience Tracked
Across 639 participants, the mean catch rate was 0.539 of the 12 planted errors. Arm affected detection (
F(3, 633) = 10.64,
p < 0.001,
= 0.048), with false positives controlled. Arm A caught 0.612 of the errors, against 0.525 in arm B, 0.508 in arm C, and 0.511 in arm D (
Figure 5a). All three comparisons of arm A against the others were reliable—versus B,
d = 0.42 [0.20, 0.65]; versus C,
d = 0.51 [0.28, 0.73]; versus D,
d = 0.48 [0.26, 0.70]; all were Holm-adjusted (
p ≤ 0.001) (
Table 6). Broken down by error type—an exploratory division, since four binary items cannot support a confirmatory claim—the advantage held for each of the three and was numerically largest for fabricated citations (
d = 0.51 [0.29, 0.74]), the failure mode with the most direct practical consequence (
Walters & Wilder, 2023).
Two further models address the binary structure of the detection outcome. Analysing the catch rate as a count of successes in twelve trials rather than as a continuous score reproduces the arm effect, and entering the error type as a within-participant factor shows no interaction between error type and arm (Wald
(6) = 5.69,
p = 0.459): the arm A advantage is common to fabricated citations, logical fallacies, and evidence–conclusion mismatches rather than driven by one of them (
Appendix B).
This outcome separates arms A and C, whereas the transfer results did not. Arm C was statistically equivalent to the no-AI control on detection (d = −0.015, TOST p = 0.006) and equivalent to arm B as well (d = −0.085, TOST p = 0.027). Students who spent four rounds adjudicating the model’s edits ended the study able to write as well as students who had diagnosed whole AI-generated drafts and no better able than the control group to find what was wrong in one. The two arms differ in the way the design anticipated—arm A required whole-text diagnosis for each round; arm C made judgements about localised changes the model had already identified as worth making—though they differ in other respects as well, considered in the Discussion section. Delegating revision therefore carried no cost on transfer, but it did not build detection capability either, and any claim that revision can be safely offloaded has to be qualified by which outcome is being asked about.
The central test concerns whether either of the global judgements students made about their own work tracked how much error they had caught. Neither did. The catch rate was uncorrelated with immediate satisfaction with their own output (
r = 0.015, 95% CI [−0.063, 0.092]), and equivalence against bounds of ±0.10 was supported (TOST
p = 0.016). Its association with self-rated quality was marginal in the raw correlation (
r = 0.080 [0.002, 0.156],
p = 0.044) and disappeared once the arm and baseline ability were controlled (
b = 0.005,
p = 0.506) (
Figure 5c). Likewise, immediate satisfaction carried no information after adjustment (
b = 0.002,
p = 0.789).
That these self-reports were uninformative about detection is not because they were uninformative in general: self-rated quality was weakly but reliably associated with transfer performance (r = 0.100, p = 0.012). The same judgements that carried some signal about how well students wrote carried none about how much of the planted error they had caught. Whether students could have reported their detection shortfall if asked directly is a separate question; answering it calls for confidence ratings on the detection items themselves, or an estimate of how many errors the student believed they had caught.
Individual differences ran in one direction only (
Table 7). Dependent tendency predicted lower detection (
b = −0.042 per scale point, 95% CI [−0.057, −0.027],
p < 0.001). Neither autonomous tendency (
b = −0.004,
p = 0.610) nor metacognitive monitoring (
b = 0.007,
p = 0.571) predicted it. The students least likely to catch what AI got wrong were those most disposed to delegate judgement to it, and their self-assessed monitoring capacity provided no protection.
3.6. Robustness and Conceptual Replication
The interaction was not an artefact of model specification. It held without covariates and with faculty classification added as a fixed covariate, and there was no evidence that arms differed in the effect of baseline ability (F(3, 631) = 0.26, p = 0.854).
At Site B (n = 109 completers across arms A, B, and D), the stage main effect reproduced in direction and magnitude: arm A scored 16.60 (3.96), arm B 14.67 (3.13), and arm D 14.59 (4.27), providing d = 0.54 [0.08, 1.00] for A versus B and d = 0.49 [0.01, 0.96] for A versus D—estimates were as large as those at Site A, with the omnibus arm test at F(2, 101) = 2.85, p = 0.063. The interaction did not reach significance at Site B (F(2, 101) = 0.98, p = 0.378), and the within-arm slopes were correspondingly imprecise (arm A, b = −0.17, 95% CI [−1.74, 1.41]). With roughly 35 participants per arm, this test had little power to detect an interaction of the observed size; the Site B result is directionally consistent but inconclusive with respect to moderation.
A pooled analysis of the three common arms also supports moderation, with each site entered as a fixed effect rather than a random one, since two sites cannot identify a variance component. The arm × tendency interaction held (
F(2, 579) = 7.00,
p = 0.001), while the site itself and its interaction with the arm contributed nothing (
Appendix B). The analysis was not pre-specified and is exploratory; two sites are not a sample of institutions, so the pooled estimate is not evidence for any wider population.
Finally, dependent and autonomous tendency correlated at r = −0.31, confirming that they are related but distinct rather than opposite poles of one dimension, which is why both entered the models.
4. Discussion
Three findings, taken in order, describe a mechanism, as summarised in
Figure 6.
Which stage is delegated determines what the Week 6 outcomes look like, and the two outcomes do not agree. On transfer, delegating drafting and delegating revision produced equivalent benefits, and both exceeded delegating evaluation, which was itself equivalent to using no AI. What separates the two effective arms from arm B is the direction of the evaluative act: the student judged the model’s text rather than executing the model’s judgement of theirs. This aligns the stage question with what evaluative-judgement research would predict (
Bearman et al., 2024;
Tai et al., 2018), with the qualification that literature has not previously been able to supply, since here the arrangement was randomised rather than observed.
Calling this retaining evaluation would overstate what the arms separate. Appraisal of the student’s own draft was performed by the model in arm C as much as in arm B; the difference lies in what the student had to do afterwards. The design does not orthogonally decompose the three stages, and the transfer result concerns the direction of judgement rather than the possession of a stage.
The detection results then show that even this is incomplete. Arm C matched arm A on transfer but was statistically equivalent to the no-AI control on error detection. This extends
Lai et al. (
2026) in the direction those authors identified—the missing fourth condition—but the answer it returns is conditional rather than clean. Revision can be delegated without cost to how well students subsequently write, and not without cost to whether they can tell that an AI-generated text is wrong.
What produced that difference is not identified by this design, and the candidates are several. Arms A and C differ in at least four respects, any of which could carry the detection effect. Arm A required appraisal of a whole manuscript rather than of a set of localised changes. Arm A therefore presented more diagnostic material per round. The proportion of the text that was AI-generated differed between the arms. And arm A alone gave the student the occasion to go back over material already passed. Arm C is not a control for arm A stripped of one ingredient; it is a different arrangement, and only arm A separated from the control on this outcome. The first is the explanation the design was built to test, and it coheres with the peer-review pattern, where the advantage is concentrated on explaining what is wrong. That is not identification: distinguishing these accounts requires arms that vary the scope of appraisal while holding the rest constant. That is the experiment this result most directly motivates. A single conclusion about whether a stage is safe to offload is in any case not available: the answer depends on which capability is being protected, and a design optimised for writing quality can leave detection untouched.
The practical inversion is nonetheless the sharpest result for teaching. The arrangement most often described in accounts of classroom and supervisory practice (
Alghamdi & Alghizzi, 2026;
Bearman et al., 2024)—the student drafts and the AI provides feedback—was the one arrangement that produced no measurable benefit on any of the three objective outcomes. It is not that this arrangement is mildly less effective. On these outcomes, it is indistinguishable from providing no AI at all.
Assigning the evaluation does not make evaluation occur. The interaction shows that the benefit of the best arrangement was concentrated among students low in dependent offloading tendency and was no longer statistically distinguishable from zero above a tendency score of 3.45, which roughly a third of the sample exceeded. The logs establish that the assignment did not control what it was meant to control: within arm A, the higher a student’s tendency, the more of their prompts asked the model to make the judgement they had been assigned. That the re-delegation occurred is established; that it is what produced the transfer shortfall is not, since the indirect effect was indistinguishable from zero. The point the logs support is that a rule can allocate a task without allocating the cognition, because the same tool that performs the delegated part remains available for the retained one.
This is where the stage literature and the mode literature meet, and the meeting point is not additive. The stage effect is not a property of the arrangement but of the arrangement crossed with the student. An average treatment effect of d = 0.35 conceals a range from a substantial benefit to none, and the students who receive none are precisely those whose habitual practice the intervention was designed to interrupt.
The mechanism goes only so far on this evidence. The prompt-level data establish that re-outsourcing occurred and scaled with tendency, which self-report could not establish. They do not establish that re-outsourcing is the channel: the indirect effect was indistinguishable from zero, and the transfer shortfall persisted with the ratio in the model. A single coded ratio is a coarse proxy for a disposition that plausibly also expresses itself in how long a student attends to feedback, how deeply they revise, and what they do with a judgement once made. The disposition matters and is visible in behaviour, and the pathway remains unidentified.
Neither of the two global judgements students made about their work tracked how many errors they had caught. Participants caught 0.539 of the planted errors on average and 0.612 in arm A, so they were far from unable to detect them. The catch rate was unrelated to immediate satisfaction—it was equivalent to zero, not merely non-significant—and unrelated to self-rated quality once arm and ability were controlled, even though self-rated quality was weakly associated with actual writing performance, r = 0.10. What the design establishes is therefore an asymmetry between two objects of judgement: the same students’ global self-assessments carried a small amount of information about the quality of their writing and none detectable about their detection performance.
What it does not establish is that students were unaware of their detection shortfall. Satisfaction with one’s output and a global quality rating are not metacognitive judgements about the detection task; a direct test would require confidence ratings on the detection items themselves, or an estimate of how many errors the student believed they had caught, and neither was collected. Students asked those questions might well have reported low confidence, and a dissociation between a global judgement and an item-level one is exactly what the metacognition literature would lead one to expect. The present result bears on the conjecture in
Zhu et al. (
2026)—that learners cannot tell from experience which kind of engagement they are in—by showing that one class of experiential signal, satisfaction with the product, does not covary with a performance dimension that matters. It does not settle whether a more targeted signal would. Supplying that test is the most direct extension of this study.
The asymmetry still carries a practical edge, because the judgements that failed to track detection are the ones students and instructors actually have to hand, and the error type where the gap between arms was largest—fabricated citations—is the one most likely to survive into a submitted manuscript (
Walters & Wilder, 2023). It also connects to a broader pattern in how people relate to AI-assisted work, where the subjective experience of assistance and its actual consequences can move independently (
Zhang et al., 2026).
Why the dissociation should take this particular form is answerable from what is known about how such judgements are formed. Metacognitive judgements are inferential rather than direct readouts of a cognitive state; they rest on cues whose validity varies with what is being judged (
Koriat, 1997). A judgement is well calibrated when the cues it happens to recruit predict the criterion and poorly calibrated when they do not. Fluency is the cue most available to a student contemplating a finished text, and fluency is exactly what a competent language model supplies, irrespective of whether the content is sound. That fluency at encoding breeds confidence unwarranted by later performance is among the most robust findings in the study of self-regulated learning (
R. A. Bjork et al., 2013;
Koriat & Bjork, 2005), and the same logic predicts the pattern observed here: the cues that make a text feel satisfying overlap to some degree with those that make writing good, consistent with the weak association between self-rated quality and transfer, and are unrelated to whether a cited source exists, consistent with the absence of any association between either judgement and detection. That metacognitive accuracy is specific to the object being judged rather than a general trait is itself well established (
Kelemen et al., 2000), and the null for trait monitoring reported above is what that literature would lead one to expect. On this reading, the dissociation is not an anomaly of AI-assisted work but an instance of a known regularity, met in a setting where its consequences are unusually costly (
Messeri & Crockett, 2024).
A second question is why four rounds of evaluative work left detection unimproved in three of the four arms. The relevant account here is automation-induced complacency: when a system reliably supplies output of apparent quality, attention to that output declines, and the errors that escape are disproportionately those the system did not itself flag—errors of omission (
Parasuraman & Manzey, 2010;
Skitka et al., 1999). Arms B and C map onto this directly. In arm B, the student received a diagnosis and never searched for one; in arm C, the student ruled on changes the model had already located, so whatever it passed over was never brought into view. Neither arrangement gives an omission error the occasion to be found, which is what detection at post-test demanded. The implication is that the capability will not accrue as a by-product of use, and there is converging reason to expect this: people’s intuitive heuristics for judging AI-generated language are systematically wrong (
Jakesch et al., 2023), and students working with a model on authentic tasks calibrate their confidence to its output rather than to its correctness (
Kaplar et al., 2026). Building detection plausibly requires explicit instruction in critical AI literacy (
Ng et al., 2021)—something no arm of this study provided, and something these results argue is not substitutable by experience in the evaluator’s role.
Three features of this result deserve emphasis. It is not a general miscalibration effect in the sense of
Kruger and Dunning (
1999), since the same participants’ judgements were partially calibrated about their writing. Trait metacognitive monitoring, measured at baseline, predicted neither detection nor the size of the dissociation. Detection was lowest among students high in dependent tendency, so the students who caught the least were also those whose habitual practice the arrangements were meant to interrupt.
4.1. Theoretical Implications
Compared with the learning literature, the pattern is more specific than a general claim that AI harms or helps. On transfer, delegating drafting and delegating revising were equally harmless, which is what a cognitive load account predicts if those stages are largely productive operations whose removal does not eliminate the germane processing (
Sweller, 1988). Delegating evaluation was not harmless on any outcome, and the sharpest expression of this is that four rounds of AI feedback left students no better at diagnosing a text than four rounds without AI. Evaluation behaves as the desirable difficulty of the writing cycle (
E. L. Bjork & Bjork, 2011): the stage whose effortfulness is the point, and whose removal is experienced as relief while the learning it carried disappears with it.
The detection contrast between arms A and C refines this, with the caveat entered above that the two arms differ in more than one respect. Both required the student to judge model-produced text, and both yielded the transfer benefit, but only arm A improved detection. The desirable difficulty plausibly lies in the unaided search for what is wrong rather than in the act of judging as such: adjudicating changes the model has already located is judgement without search, and it did not transfer to finding errors nobody had flagged. This is a hypothesis the present contrast is consistent with rather than a mechanism it isolates. What the contrast does establish is that the granularity at which a stage is defined is a substantive modelling choice: two arrangements that both count as “the student evaluates” behaved differently on one of the two outcomes, and a design treating evaluation as a single stage would not have shown it.
This locates the harm of generative AI in learning at a level below the tool and below the amount of use. What matters is which cognitive operation is displaced, and that operation is not visible in the artefact—which is precisely the supervisor’s problem. It also refines the two-dimensional framings that several groups have converged on (
Y. Fan et al., 2025;
Zhu et al., 2026). Those frameworks treat the mode of engagement as a property of the person. The present results show it is a person-by-situation product: the same disposition was inert in three of four arms and decisive in the fourth. A disposition to offload can only express itself where the design leaves a channel through which to offload. Work on automation trust shows a structurally similar pattern, with individual differences moderating susceptibility to complacency rather than determining it outright (
Tang et al., 2026); that research concerns driving rather than learning, so the parallel is one of form rather than of content.
The comparison of arms A and C makes this concrete, and it is an observation the study did not set out to produce. In arm A, the student’s evaluative work took place in free interaction with the model, so re-delegation was available—and the students disposed to re-delegate did so, and lost the benefit. In arm C, the equivalent work was completed inside a structured form with no channel to the model, and the slope of tendency on transfer in that arm was flat (b = 0.23, p = 0.367). The disposition did not stop existing in arm C; it had nowhere to go. The channel in arm A was also a continuing one: a student’s working relationship with the model carried across the four rounds rather than restarting at each assignment, so a student who had established a habit of asking for verdicts did not have to establish it again. That is how students ordinarily work with such a system over a term, and it is part of what makes an assigned evaluation forwardable in practice rather than only in principle. The two arms differ in more than the presence of a channel and were not designed as a manipulation of it, so this is a lead for direct testing rather than a result. It is nonetheless the pattern to expect if enforcement rather than instruction is what makes the difference.
Finally, the detection result bears on the assumption of learner self-regulation that underwrites much of this field. Self-regulated learning models presume access to a monitoring signal (
Zimmerman, 2002), and interventions from reflective prompting to self-assessment inherit that presumption. Here one such signal was uninformative in a specific and testable sense: satisfaction was statistically equivalent to uncorrelated with detection, while self-rated quality remained weakly informative about writing quality. Whether learners retain access to some other, more targeted signal about detection is precisely what this design leaves open, and it is the question a follow-up should ask. What the present data do show is that the signal most readily available—how satisfied a student is with the text in front of them—does not carry that information. Theories of self-regulated learning developed before generative AI may therefore need to specify which objects of monitoring remain accessible when the work is co-produced with a system whose errors are fluent, rather than treating monitoring as a single general capacity.
4.2. Practical Implications
The design implication follows from the third finding, with its scope stated carefully. Self-regulated correction, reflective prompts, and honour-code-style undertakings all assume the learner can infer from the experience of the work that something went wrong. These data show that one obvious basis for that inference—satisfaction with the product—carries no information about how much error the text still contains. They do not show that no such basis exists. The prudent design posture is therefore not to rely on the student’s global sense of the work having gone well, which is the signal these arrangements most often invoke and the one measured here. This is not to treat monitoring accuracy as fixed. Rewarding accurate monitoring rather than performance alone has been shown to raise both together (
Liu et al., 2025), which indicates that calibration is something a design can aim at directly. Aiming at it, however, is a different intervention from assuming it, and it is the assumption that the present results undercut.
What follows is a preference for structure over exhortation. If the evaluative work is what has to happen, it has to be enforced by the environment rather than requested of the student—which is the logic of cognitive forcing functions (
Buçinca et al., 2021), transplanted from single decisions into a multi-week learning task. Arm C is an instance of such a structure that happened to be in the design: the adjudication had to be entered as a form, there was no channel to the model, and the disposition that cost students the benefit in arm A was inert there. Requiring a committed judgement to be recorded before the model can be consulted, capping consultations during the evaluative stage, or making the evaluative artefact itself the assessed object are all mechanisms of the same kind, and none of them depends on the student noticing anything. The arm C result also carries its own warning, since that arrangement protected transfer without building detection: enforcement has to be aimed at the capability one actually wants. The cost that
Buçinca et al. (
2021) documented—such constraints feel worse—appears here in a sharper form: because neither global self-judgement carried information about detection, student satisfaction is not a usable signal for whether such a design is working, and optimising for it would favour arrangements that feel good over arrangements that produced the outcomes measured here. Student preference for AI involvement in scoring and feedback is itself under active study and varies with context and stakes (
Yildirim-Erbasli et al., 2026); a design selected on that basis would be tracking something these data indicate carries no information about what students failed to detect.
For supervisors specifically, the results argue against the intuitive rule, with two qualifications the data insist on. Telling a student to write the draft themselves and let AI check it is the arrangement that failed on every outcome. Letting AI draft while the student diagnoses it produced gains on all three. Letting AI revise while the student adjudicates the changes produced the writing gain but left detection where it started, so it is the weaker of the two options if the concern is students accepting fabricated material.
The second qualification concerns whom each arrangement serves, and the two effective arms differ here in a way the average effects conceal. The advantage of arm A was concentrated among students low in dependent tendency and was no longer statistically distinguishable from zero above a tendency score of 3.45; in arm C, the slope of tendency on transfer was flat (b = 0.23, p = 0.367), so what benefit that arrangement conferred did not depend on the disposition. In other words, only arm A delivered its benefit selectively. The practical reading is therefore a qualified pair of recommendations rather than a single rule: arm A is the stronger arrangement on average and the only one that built detection, but it is the one whose gain the students it was written for are least likely to receive; arm C was insensitive to the disposition and left detection untouched. Either way, the disposition is worth measuring at intake rather than assumed, since it is what separates the students for whom the better arrangement works from those for whom it does not.
4.3. Limitations
Three limitations bound these claims. The arms are workflows rather than an orthogonal decomposition of the three stages, so the contrasts bear on the direction and object of the student’s evaluative act rather than isolating the offloading of any one stage, and the four respects in which arms A and C differ are not separated either. The inferential base is also narrower than the sample size suggests: the moderation rests on one site, since Site B reproduced the stage main effect at comparable magnitude but had too few participants per arm to test the interaction, and because class identifiers were not retained, residual correlation within classes or instructors is not modelled and may make the reported standard errors too small. And what can be said about mechanism is bounded in two ways. For example, the pathway from tendency to shortfall is unidentified, since re-outsourcing does not mediate it and the cognitive-agency and motivational constructs theorised to carry it (
Zhu et al., 2026) lie outside the measures taken here. The detection findings concern the informativeness of two global self-judgements rather than students’ metacognition about the detection task, which would require item-level confidence ratings or estimates of errors missed (
Fleming & Lau, 2014;
Schraw, 2009).