1. Introduction
Misconceptions in Newtonian mechanics are among the most thoroughly documented obstacles to learning in science education. Despite extensive instruction, a substantial fraction of introductory physics students retain pre-Newtonian beliefs (that motion requires continuous force, that heavier objects fall faster, that action–reaction pairs cancel) long after formal coverage of the relevant material [
1,
2,
3]. The persistence of these misconceptions reflects not a failure of exposure but a feature of cognition: belief revision is gated by dissatisfaction with the prior conception and by the perceived intelligibility, plausibility, and fruitfulness of an alternative [
4,
5,
6]. This makes the design of effective tutoring a
behavioral problem at least as much as a
technological one.
The recent emergence of large language models (LLMs) as general-purpose tutors has reignited interest in how instructional dialogue can be scaled to support conceptual change [
7,
8], building on a long tradition of dialogue-based intelligent tutoring systems in which natural-language interaction was shown to approach the effectiveness of human tutoring [
9,
10]. Retrieval-Augmented Generation (RAG) variants of LLM tutors [
11] enrich generation with externally retrieved content (typically textbook passages, misconception databases, or knowledge-graph traversals) under the hypothesis that grounded responses will produce more accurate, more pedagogically targeted, and ultimately more change-inducing dialogue. To date, however, evaluations of RAG-augmented tutoring have focused predominantly on system-level metrics rather than on learner-level behavioral outcomes (which misconceptions revise, for whom, after how many turns, and with what dialogue dynamics).
The present study addresses three gaps in the existing behavioral evidence on AI tutoring. First, the retrieval-architecture-by-learner-profile interaction has rarely been examined: which simulated learners gain most from retrieval, and whether knowledge-graph-augmented retrieval offers behavioral differences beyond simple chunk-based retrieval. Second, the granularity of effects across specific misconceptions (whether retrieval differentially aids single-concept versus relational misconceptions) has not been systematically mapped. Third, the reproducibility of patterns across LLM providers has rarely been verified.
A practical barrier to behavioral evaluation has been the cost and ethical complexity of running large factorial studies with human participants. Recent methodological work has demonstrated that LLM-based
synthetic learners (computational agents parameterized by behavioral characteristics derived from cognitive theory) can serve as a reproducible, scalable hypothesis-generation tool [
12,
13,
14,
15]. Synthetic learners do not substitute for human-participant research; they enable a level of experimental control (deterministic seeding, factorial coverage, repeated measures) that complements rather than competes with empirical work. The present study adopts this paradigm: we interpret our results as behavioral patterns within a closed simulated system, with all pedagogical implications labeled as hypotheses for human-subject follow-up.
This article makes three contributions. First, it provides an open, deterministically seeded simulation harness in which competing tutoring architectures can be compared under identical conditions, at an API cost of roughly two US dollars for a complete 600-session factorial. Second, it reports the behavioral signature of that comparison across eight outcome measures, including the interaction between retrieval architecture and learner profile that single-condition benchmarks cannot reveal. Third, and unusually for simulation studies, it audits its own primary measure from four directions: two LLM judges drawn from different model families, a re-scoring ablation that replaces the simulator’s trigger vocabulary with alternative and independently authored criteria, a perturbation analysis of the revision rule’s numeric constants, and a generalized linear model with family-wise error control. The audit separates the findings that survive every check (the profile clustering, the interaction and the LLM-versus-template contrast) from the one that does not (the retrieval gradient, which is a property of keyword-based measurement), and places explicit bounds on how the architectural findings should be read.
The remainder of the article is organized as follows.
Section 2 situates the study within work on dialogue tutoring, retrieval augmentation, simulated learners and LLM-based evaluation.
Section 3 describes the profiles, misconceptions, conditions, revision rule and statistical plan, and states which analyses were confirmatory and which exploratory.
Section 4 reports the factorial results, followed by the measurement audit.
Section 5 interprets the findings, states their limitations and lists hypotheses for human-subject follow-up, and
Section 6 concludes with directions for future work.
Research Questions and Hypotheses
Within a 5 (learner profile) × 4 (tutoring condition) × 30 (replications) between-subjects factorial design, we addressed three research questions:
RQ1. How do simulated learner profiles, varying along prior knowledge, motivation, and learning style, differ in misconception-revision trajectories?
RQ2. How does retrieval architecture (KG-RAG vs. RAG vs. Vanilla) shape the behavioral dynamics of simulated misconception revision in LLM-based conditions, relative to a non-LLM Static-Feedback floor?
RQ3. Do the behavioral effects of tutoring condition depend on learner profile (Condition × Profile interaction)?
Drawing on Conceptual Change Theory [
4], Cognitive Load Theory [
16], and the prior FCI literature [
1,
2,
3], we pre-specified the following directional hypotheses: (H1) higher prior knowledge and higher motivation are associated with higher simulated misconception-revision rates; (H2) within the LLM-based conditions, the simulated revision rate follows KG-RAG > RAG > Vanilla; (H3) retrieval benefit is larger for simulated learners with moderate prior knowledge than for the extremes.
A secondary methodological question concerned the
robustness of patterns across LLM providers, addressed via a sensitivity experiment using Anthropic Claude Sonnet 4 (60 sessions targeted;
analyzable,
Section 3.7) alongside the primary 600-session experiment using OpenAI GPT-4o-mini.
The three research questions and hypotheses H1–H3 were fixed before the 600-session run and constitute the
confirmatory part of the study; they are tested on the primary outcome (MRR) with the omnibus two-way ANOVA and, in this revision, with a binomial generalized linear model and Holm-corrected inference (
Section 3.8). All other analyses reported below (the seven secondary outcomes, per-misconception breakdowns, cell-level contrasts, keyword-density analysis, the LLM-judge audits, the re-scoring ablation, the parameter perturbation and the cross-provider sensitivity run) are
exploratory and are labelled as such where they appear.
2. Related Work
Natural-language tutoring has a long empirical record: AutoTutor and its descendants produced learning gains comparable to untrained human tutors [
9,
17], and VanLehn’s meta-analysis [
10] placed step-based tutoring systems close to human tutoring in effect size. In physics, the Force Concept Inventory [
1,
2] established that pre-Newtonian beliefs survive conventional instruction, Hake [
3] tied normalized gains to interactive engagement, and Conceptual Change Theory [
4,
5,
6,
18] together with diSessa’s knowledge-in-pieces account [
19,
20] explains why such beliefs resist revision. Our profiles and revision rule are parameterized from this literature, and
Section 4.10 tests whether the simulator recovers a latency ordering it was not calibrated on.
Instruction-tuned LLMs have produced a rapid series of tutoring deployments [
7], including a physics course in which a prompt-engineered LLM tutor outperformed in-class active learning [
8]. Retrieval-augmented generation [
11] is the standard remedy for hallucination, and in education it has been evaluated mostly on answer quality and groundedness; Levonian et al. [
21] found that retrieval improves groundedness in mathematics question answering while human preference does not always follow. In these evaluations the unit of analysis is the tutor’s utterance, not the learner’s trajectory. We are not aware of prior work that compares retrieval architectures on which misconceptions revise, for which learner profile, and after how many turns; this is the gap the factorial design addresses.
Simulated students predate LLMs [
22], and LLMs have made behaviorally specified agents cheap to instantiate: prompted models reproduce subgroup response distributions [
12], replicate classic human-subject experiments [
13], populate interactive environments [
14], act as economic subjects [
15] and, in education, serve as simulated students for teaching-assistant training [
23]. The caution common to this literature, which we adopt, is that agreement between simulated and human patterns must be established rather than assumed, so that simulation results are hypotheses for human studies rather than substitutes for them. Finally, automatic evaluation of dialogue increasingly relies on an LLM as rater. Strong LLM judges agree with human preferences at rates comparable to inter-human agreement but carry position, verbosity and self-enhancement biases [
24], and LLM evaluators recognise and favour their own generations [
25], so a judge from the same model family as the system under evaluation is not an independent rater. This motivates the two-family judge design of
Section 4.8.
5. Discussion
5.1. Three Behavioral Patterns from the Simulator
Three patterns emerged consistently across the 600-session primary experiment.
The retrieval gradient is a property of the keyword-based measure. Under the simulator’s rule, knowledge-graph-augmented RAG produced approximately 70% higher MRR than the same LLM without retrieval (
vs.
,
), the ordering survives Holm correction and the binomial GLM, and the ablation shows it does not depend on any single phrase, on the richness term, on IDF weighting, on the numeric constants, or even on the particular vocabulary, since an independently authored concept-text vocabulary reproduces it. What it does depend on is the kind of measure: neither LLM judge reproduces it, and the second-provider run did not show it. The most parsimonious reading is that retrieval increases the amount of canonical correct-concept language the tutor emits (
Section 4.7), that any lexical measure of revision will register this, and that a rater reading the learner’s own words does not. We therefore report the gradient as a finding about what retrieval does to tutor output, and as a hypothesis about learners, not as evidence that retrieval improves revision.
Two profile clusters, not five gradations. The five learner profiles partition into two clusters on MRR: a high-revision cluster of Intermediate-Strategic and Advanced-Residual learners (, statistically indistinguishable from one another) and a low-revision cluster of the remaining three profiles (). The defining property of the high cluster within our parameterization is the combination of moderate-to-high prior knowledge and high motivation; learners possessing only one of these characteristics (Naive-Active: high motivation, very low prior knowledge) did not enter the high cluster. The simulator thus predicts a particular interaction between prior knowledge and motivation that warrants empirical testing.
The interaction is reliable and pedagogically meaningful. A Condition × Profile interaction emerged on MRR ( in the ANOVA; , in the binomial GLM), and similar interactions appeared on five of eight outcomes after Holm correction. The most pedagogically meaningful pattern is that no condition was able to produce substantial revision for Naive-Passive learners ( across all conditions). Within the simulator, retrieval-augmented tutoring does not lift learners whose engagement-decay parameter is set high and whose motivation is set low. This is a behavioral hypothesis that, if confirmed empirically, would imply pre-tutoring engagement-raising interventions as a prerequisite for AI-tutoring effectiveness.
5.2. Per-Misconception Heterogeneity and an Alternative Reading of Static-Feedback Competitiveness
The per-misconception breakdown revealed two patterns worth highlighting. The Impetus misconception showed the largest within-condition retrieval gradient (KG-RAG vs. Vanilla ). For this cognitively entrenched belief, structured retrieval, in our simulator, surfaces multiple mutually-reinforcing keyword-bearing passages that the LLM weaves into multi-turn explanations.
Static-Feedback was unexpectedly competitive with Vanilla-LLM on MC03 (Motion-Implies-Force) and MC04 (Action-Reaction-Cancel). We considered two readings. The first is substantive: hand-curated templates whose phrasings closely match the canonical correct framings of certain misconceptions are competitive with parametric LLM generation. The second, and in our view more parsimonious, reading is a methodological artifact: the Static-Feedback templates and the revision-trigger keywords used by the simulator were both authored against the same misconception ontology. Static templates therefore enjoy a built-in advantage on the simulator’s keyword-driven MRR rule that does not reflect a pedagogical mechanism observable in human learners. We caution against interpreting the Static-Feedback competitiveness as evidence that templates can replace LLMs for these misconceptions; the result more likely reflects the geometry of the simulator’s evaluation criterion.
5.3. Construct Fidelity: What the Simulator Recovers and What It Does Not
The simulator reproduces several qualitative patterns that were not built directly into the calibration rules. Out-of-sample, Heavier-Faster emerges as the misconception with the shortest revision latency, consistent with FCI documentation that vacuum-based demonstrations efficiently support its revision [
20]. The profile clustering (P4 ≈ P5 ≫ others) is consistent with the documented role of prior knowledge in conceptual change [
5,
18] and with the Hake-gain observation that interactive engagement is more effective for students with sufficient initial scaffold [
3].
However, the simulator does not reproduce the FCI prediction that Impetus should be the most persistent misconception under naive instruction. Impetus latency in our simulation is near the median, not the maximum. We attribute this to the design of the keyword-trigger rules: the Impetus trigger keywords (“inertia,” “first law,” “no force needed,” “keeps moving”) happen to be high-frequency phrases in the retrieved physics corpora, giving the simulator a structural advantage on Impetus that does not exist in human cognition. This is the kind of out-of-sample mismatch that deliberate construct-fidelity probing is designed to surface. An obvious next calibration step, deferred to future work, would be to downweight or threshold trigger keywords by their inverse document frequency in the retrieval corpus, so that the revision rule rewards corpus-rare keyword matches more strongly than corpus-common ones; we expect this would partially correct the Impetus mismatch and yield more conservative MRR estimates for misconceptions whose trigger phrases are central to the host corpus.
5.4. Limitations
Several limitations are important to acknowledge.
Synthetic learners are not human learners. Our findings characterize the behavioral dynamics of a closed simulator with calibrated parameters; they do not directly establish that real human learners would show the same patterns under the same conditions. We position this work as hypothesis-generating; the patterns reported here warrant empirical testing with human participants before being interpreted as evidence about learning.
Misconception revision is operationalized as keyword-triggered state flipping. This is a substantial simplification of the cognitive process underlying conceptual change.
Section 4.7 quantifies the mechanistic limitations;
Section 4.8 shows only slight agreement between the rule and two LLM judges from different model families (
and
), both of which find far more revisions than the rule does. The reported MRR magnitudes should therefore be read as simulator-internal quantities, not as estimates of true revision rates, and the retrieval gradient as a property of the keyword-based measure. The judges have limitations of their own: each observed only the final three learner utterances, the two agree with each other only moderately (
) and differ substantially in leniency, and neither has been validated against human raters on this task. A future calibration round should either adopt a validated judge as the primary measure or tighten the keyword rule until its positive rate matches a validated judge’s.
Static-Feedback evaluation is structurally favorable, and the trigger vocabulary is author-declared. The Static-Feedback templates and the revision-trigger keywords were authored against the same misconception ontology (
Section 5.2);
Section 4.9 shows that this inflates Static-Feedback under strict phrase-pair criteria to the level of KG-RAG, while under the original rule, under an independent vocabulary and under both judges Static-Feedback is clearly lowest. The claim that the trigger vocabulary was fixed before the tutors were run rests on the authors’ declaration and on the structure of the code, not on an auditable version history (
Section 3.6); we mitigate this with the vocabulary-independent criteria and the judges, and recommend that future simulation studies commit the evaluation rule to a public repository before generating data.
Second-provider evidence is preliminary. The 58-session Claude Sonnet 4 run is a sensitivity check with –15 per condition; it can show that a pattern is not automatic, which it does for the retrieval gradient, but it cannot establish or rule out a provider difference. A full 600-session run on a second provider, projected to cost approximately $53 in API charges and roughly 5 hours of sequential compute, is the most important next step and is not provided here. Resumable infrastructure for it is released alongside the manuscript.
Cell-level subsamples have limited power, and most analyses are exploratory. The 5 × 4 × 30 design was powered for the omnibus test on MRR, not for cell-level contrasts (
). Cell-level effect sizes (e.g., the Intermediate-Strategic × KG-RAG peak) are reported with bootstrap 95% CIs but are exploratory, as are the seven secondary outcomes, the judge, ablation and perturbation analyses and the second-provider run (
Section 3.8). The confirmatory results are the three MRR effects, which survive Holm correction and the binomial GLM.
Misconception set coverage. The five-misconception set excludes other prevalent FCI misconceptions, notably centrifugal-force beliefs and the role of medium in gravity. Generalization to a broader FCI-style misconception space remains to be tested.
Single-domain coverage. Newtonian mechanics was chosen for its uniquely well-documented misconception literature; whether the patterns generalize to other STEM domains is an empirical question.
5.5. Hypotheses for Human-Subject Follow-Up
Within the synthetic-learner system, several patterns are sufficiently robust to warrant pre-registered empirical testing:
Hemp 1. For real introductory physics students, LLM tutoring of any of the three architectures will produce higher pre-/post-FCI gains than template-based feedback; whether knowledge-graph-augmented RAG additionally outperforms non-RAG LLM tutoring is the measure-dependent part of our results and should be tested as a secondary, directional hypothesis.
Hemp 2. The interaction between prior knowledge and motivation predicts FCI-gain such that students with high motivation but low prior knowledge gain less than students with moderate-to-high prior knowledge and high motivation, even when receiving the same tutoring.
Hemp 3. Engagement raised pre-tutoring (e.g., via goal-setting or interest interventions) will substantially lift FCI-gain for students who would otherwise fall into the low-engagement, low-prior-knowledge phenotype.
These hypotheses are offered as testable predictions; they are not claims about how human learning works.
6. Conclusions and Future Work
This article asked whether a deterministically seeded simulation harness with LLM-based synthetic learners can compare retrieval architectures for LLM tutoring at a learner-level outcome, and how far its answers can be trusted. Across 600 sessions on GPT-4o-mini, the main findings are the following. First, the learner-profile effect is the largest in the design () and resolves into two clusters rather than a ranked gradient: simulated learners with moderate-to-high prior knowledge and high motivation revise, and the other three profiles largely do not; this structure survives Holm correction, a binomial GLM, every re-scoring criterion, every perturbation of the rule’s constants, and is visible under both LLM judges. Second, a Condition × Profile interaction (; GLM ) shows that no architecture lifts the low-engagement, low-prior-knowledge profile, whose MRR stays at or below in every condition and under every criterion. Third, every LLM-based tutor exceeds the non-LLM template baseline under the simulator’s rule, under both judges, and in the second-provider run. Fourth, and stated with its qualification, the simulator’s measure orders the three LLM tutors as KG-RAG > RAG > Vanilla (); this ordering is stable across 17 leave-one-phrase-out vocabularies, an independently authored vocabulary and 224 perturbations of the rule’s constants, but it is not reproduced by either LLM judge ( with the simulator and ) and was absent in the 58-session Claude Sonnet 4 run. It is therefore a finding about what retrieval does to tutor language, and a hypothesis about learners, not evidence that retrieval improves revision.
The methodological contribution is the audit itself. Judging a simulator’s primary measure against two raters from different model families, replaying the measure over fixed logs under alternative criteria, and perturbing its constants are all cheap once the dialogue corpus exists, and together they separate the results that are properties of the simulated learners from those that are properties of the measurement rule. We recommend this practice for simulation-based evaluations of educational technology in general.
Future work follows directly from the limitations. (1) A full 600-session run on a second provider, for which resumable infrastructure is released, would test whether the retrieval gradient is provider-specific. (2) Validating the LLM judges against human raters on a subset of dialogues, and then adopting a validated judge as the primary revision measure, would remove the keyword rule as the point of failure; the judges’ own moderate agreement (
) shows that this validation is needed rather than optional. (3) The single-phrase dependence identified for the Advanced-Residual profile (“centripetal”) should be removed by enlarging the MC05 vocabulary before the harness is reused. (4) Extending the misconception set to centrifugal-force beliefs and the role of medium in gravity, and the domain beyond Newtonian mechanics, would test the generality of the two-cluster structure. (5) The three empirical hypotheses of
Section 5.5 will be pre-registered on the Open Science Framework before any human-subject follow-up, with the simulator-derived effect sizes as the basis for the registered power analyses. The full pipeline, all 600 dialogue logs, both judges’ verdicts, the ablation and perturbation scripts and the analysis code are released so that every number in this article can be recomputed.