Next Article in Journal
Excess Properties, CO2 Absorption, and FTIR Spectra of Monoethanolamine with Dimethyl Sulfoxide or N,N–Dimethylformamide Binary Solutions
Previous Article in Journal
A Fast Shell-Based Framework for Predicting Chucking-Induced In-Plane Distortion in Silicon Wafers from Measured Out-of-Plane Geometry
Previous Article in Special Issue
Predicting Student Engagement Characteristics Using a Multi-Instance Localization Approach with a Gradient-Boosted Deep LSTM Classifier
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evaluating Retrieval-Augmented LLM Tutoring with Synthetic Learners: A Factorial Simulation of Newtonian Misconception Revision

by
Semih Yumuşak
1,* and
Güngör Yumuşak
2
1
Department of Computer Engineering, TOBB University of Economics and Technology, Söğütözü Caddesi No. 43, 06560 Ankara, Türkiye
2
Department of Curriculum and Instruction, Ahmet Keleşoğlu Faculty of Education, Necmettin Erbakan University, 42090 Konya, Türkiye
*
Author to whom correspondence should be addressed.
Appl. Sci. 2026, 16(19), 9938; https://doi.org/10.3390/app16199938 (registering DOI)
Submission received: 11 September 2026 / Revised: 26 September 2026 / Accepted: 1 October 2026 / Published: 8 October 2026
(This article belongs to the Special Issue Artificial Intelligence in Education: Latest Advances and Prospects)

Featured Application

The simulation harness released with this article allows developers of retrieval-augmented tutoring systems to benchmark competing retrieval architectures, and to identify which learner profiles each architecture favours, at a cost of a few US dollars per full factorial and before any classroom deployment.

Abstract

Retrieval-augmented generation (RAG) is increasingly used to ground large language model (LLM) tutors, but comparing tutoring architectures is expensive, because every design decision ordinarily requires a fresh study with students. We present a reproducible simulation harness that evaluates tutoring architectures against LLM-based synthetic learners, and we use it to run a full factorial comparison. Five learner profiles, parameterized from Conceptual Change Theory and the Force Concept Inventory literature, were crossed with four tutoring conditions (knowledge-graph-augmented RAG, BM25 RAG, a vanilla LLM without retrieval, and a non-LLM template baseline) across 30 deterministically seeded replications, giving 600 fully logged sessions on GPT-4o-mini, plus a 58-session cross-provider sensitivity run on Claude Sonnet 4. Two findings were robust to every check we applied. First, the learner-profile effect ( η p 2 = 0.315 ) resolved into two clusters rather than a ranked gradient, and a Condition × Profile interaction ( η p 2 = 0.097 ) showed that no architecture lifted the low-engagement, low-prior-knowledge profile. Second, every LLM-based condition exceeded the non-LLM floor under the simulator’s rule, under two independent LLM judges from different model families, and under a binomial generalized linear model with Holm-corrected inference. A third finding is measure-dependent and we report it as such: under the simulator’s keyword-triggered revision rule, the Misconception Revision Rate followed a monotonic retrieval gradient across the LLM-based conditions ( M = 0.40 for KG-RAG, 0.33 for RAG and 0.24 for the vanilla LLM; η p 2 = 0.099 ), but this gradient was not reproduced by either LLM judge (Cohen’s κ with the simulator = 0.12 and 0.07 ) and was not reproduced in the Claude Sonnet 4 sensitivity run. A re-scoring ablation over the fixed dialogue logs shows that the gradient survives leave-one-keyword-out, IDF-weighted and independently authored vocabularies, and that the simulator’s numeric constants can be perturbed by ± 50 % without changing any ordering, so the gradient is a stable property of keyword-based measurement rather than of one keyword list, while its absence under the judges bounds how far it should be read as evidence about learners. All architectural findings are reported as simulator-internal results and as pre-registerable hypotheses for human-subject testing. The full pipeline, all session logs, both judges’ outputs and the analysis code are released.

1. Introduction

Misconceptions in Newtonian mechanics are among the most thoroughly documented obstacles to learning in science education. Despite extensive instruction, a substantial fraction of introductory physics students retain pre-Newtonian beliefs (that motion requires continuous force, that heavier objects fall faster, that action–reaction pairs cancel) long after formal coverage of the relevant material [1,2,3]. The persistence of these misconceptions reflects not a failure of exposure but a feature of cognition: belief revision is gated by dissatisfaction with the prior conception and by the perceived intelligibility, plausibility, and fruitfulness of an alternative [4,5,6]. This makes the design of effective tutoring a behavioral problem at least as much as a technological one.
The recent emergence of large language models (LLMs) as general-purpose tutors has reignited interest in how instructional dialogue can be scaled to support conceptual change [7,8], building on a long tradition of dialogue-based intelligent tutoring systems in which natural-language interaction was shown to approach the effectiveness of human tutoring [9,10]. Retrieval-Augmented Generation (RAG) variants of LLM tutors [11] enrich generation with externally retrieved content (typically textbook passages, misconception databases, or knowledge-graph traversals) under the hypothesis that grounded responses will produce more accurate, more pedagogically targeted, and ultimately more change-inducing dialogue. To date, however, evaluations of RAG-augmented tutoring have focused predominantly on system-level metrics rather than on learner-level behavioral outcomes (which misconceptions revise, for whom, after how many turns, and with what dialogue dynamics).
The present study addresses three gaps in the existing behavioral evidence on AI tutoring. First, the retrieval-architecture-by-learner-profile interaction has rarely been examined: which simulated learners gain most from retrieval, and whether knowledge-graph-augmented retrieval offers behavioral differences beyond simple chunk-based retrieval. Second, the granularity of effects across specific misconceptions (whether retrieval differentially aids single-concept versus relational misconceptions) has not been systematically mapped. Third, the reproducibility of patterns across LLM providers has rarely been verified.
A practical barrier to behavioral evaluation has been the cost and ethical complexity of running large factorial studies with human participants. Recent methodological work has demonstrated that LLM-based synthetic learners (computational agents parameterized by behavioral characteristics derived from cognitive theory) can serve as a reproducible, scalable hypothesis-generation tool [12,13,14,15]. Synthetic learners do not substitute for human-participant research; they enable a level of experimental control (deterministic seeding, factorial coverage, repeated measures) that complements rather than competes with empirical work. The present study adopts this paradigm: we interpret our results as behavioral patterns within a closed simulated system, with all pedagogical implications labeled as hypotheses for human-subject follow-up.
This article makes three contributions. First, it provides an open, deterministically seeded simulation harness in which competing tutoring architectures can be compared under identical conditions, at an API cost of roughly two US dollars for a complete 600-session factorial. Second, it reports the behavioral signature of that comparison across eight outcome measures, including the interaction between retrieval architecture and learner profile that single-condition benchmarks cannot reveal. Third, and unusually for simulation studies, it audits its own primary measure from four directions: two LLM judges drawn from different model families, a re-scoring ablation that replaces the simulator’s trigger vocabulary with alternative and independently authored criteria, a perturbation analysis of the revision rule’s numeric constants, and a generalized linear model with family-wise error control. The audit separates the findings that survive every check (the profile clustering, the interaction and the LLM-versus-template contrast) from the one that does not (the retrieval gradient, which is a property of keyword-based measurement), and places explicit bounds on how the architectural findings should be read.
The remainder of the article is organized as follows. Section 2 situates the study within work on dialogue tutoring, retrieval augmentation, simulated learners and LLM-based evaluation. Section 3 describes the profiles, misconceptions, conditions, revision rule and statistical plan, and states which analyses were confirmatory and which exploratory. Section 4 reports the factorial results, followed by the measurement audit. Section 5 interprets the findings, states their limitations and lists hypotheses for human-subject follow-up, and Section 6 concludes with directions for future work.

Research Questions and Hypotheses

Within a 5 (learner profile) × 4 (tutoring condition) × 30 (replications) between-subjects factorial design, we addressed three research questions:
  • RQ1. How do simulated learner profiles, varying along prior knowledge, motivation, and learning style, differ in misconception-revision trajectories?
  • RQ2. How does retrieval architecture (KG-RAG vs. RAG vs. Vanilla) shape the behavioral dynamics of simulated misconception revision in LLM-based conditions, relative to a non-LLM Static-Feedback floor?
  • RQ3. Do the behavioral effects of tutoring condition depend on learner profile (Condition × Profile interaction)?
Drawing on Conceptual Change Theory [4], Cognitive Load Theory [16], and the prior FCI literature [1,2,3], we pre-specified the following directional hypotheses: (H1) higher prior knowledge and higher motivation are associated with higher simulated misconception-revision rates; (H2) within the LLM-based conditions, the simulated revision rate follows KG-RAG > RAG > Vanilla; (H3) retrieval benefit is larger for simulated learners with moderate prior knowledge than for the extremes.
A secondary methodological question concerned the robustness of patterns across LLM providers, addressed via a sensitivity experiment using Anthropic Claude Sonnet 4 (60 sessions targeted; N = 58 analyzable, Section 3.7) alongside the primary 600-session experiment using OpenAI GPT-4o-mini.
The three research questions and hypotheses H1–H3 were fixed before the 600-session run and constitute the confirmatory part of the study; they are tested on the primary outcome (MRR) with the omnibus two-way ANOVA and, in this revision, with a binomial generalized linear model and Holm-corrected inference (Section 3.8). All other analyses reported below (the seven secondary outcomes, per-misconception breakdowns, cell-level contrasts, keyword-density analysis, the LLM-judge audits, the re-scoring ablation, the parameter perturbation and the cross-provider sensitivity run) are exploratory and are labelled as such where they appear.

2. Related Work

Natural-language tutoring has a long empirical record: AutoTutor and its descendants produced learning gains comparable to untrained human tutors [9,17], and VanLehn’s meta-analysis [10] placed step-based tutoring systems close to human tutoring in effect size. In physics, the Force Concept Inventory [1,2] established that pre-Newtonian beliefs survive conventional instruction, Hake [3] tied normalized gains to interactive engagement, and Conceptual Change Theory [4,5,6,18] together with diSessa’s knowledge-in-pieces account [19,20] explains why such beliefs resist revision. Our profiles and revision rule are parameterized from this literature, and Section 4.10 tests whether the simulator recovers a latency ordering it was not calibrated on.
Instruction-tuned LLMs have produced a rapid series of tutoring deployments [7], including a physics course in which a prompt-engineered LLM tutor outperformed in-class active learning [8]. Retrieval-augmented generation [11] is the standard remedy for hallucination, and in education it has been evaluated mostly on answer quality and groundedness; Levonian et al. [21] found that retrieval improves groundedness in mathematics question answering while human preference does not always follow. In these evaluations the unit of analysis is the tutor’s utterance, not the learner’s trajectory. We are not aware of prior work that compares retrieval architectures on which misconceptions revise, for which learner profile, and after how many turns; this is the gap the factorial design addresses.
Simulated students predate LLMs [22], and LLMs have made behaviorally specified agents cheap to instantiate: prompted models reproduce subgroup response distributions [12], replicate classic human-subject experiments [13], populate interactive environments [14], act as economic subjects [15] and, in education, serve as simulated students for teaching-assistant training [23]. The caution common to this literature, which we adopt, is that agreement between simulated and human patterns must be established rather than assumed, so that simulation results are hypotheses for human studies rather than substitutes for them. Finally, automatic evaluation of dialogue increasingly relies on an LLM as rater. Strong LLM judges agree with human preferences at rates comparable to inter-human agreement but carry position, verbosity and self-enhancement biases [24], and LLM evaluators recognise and favour their own generations [25], so a judge from the same model family as the system under evaluation is not an independent rater. This motivates the two-family judge design of Section 4.8.

3. Materials and Methods

3.1. Theoretical Framing of the Simulation

We model misconception revision as a probabilistic state transition modulated by simulated learner characteristics, taking inspiration from Conceptual Change Theory [4,6]. We are explicit that our operationalization (Section 3.6) collapses the intelligibility–plausibility–fruitfulness sequence onto a deterministic keyword-exposure rule; this is a substantial simplification whose consequences are quantified in Section 4.7 and Section 5.4.
Our use of LLM-based synthetic learners follows a methodological tradition in which language models simulate behaviorally specified agents whose response patterns are analyzed using the same statistical tools applied to human-participant data [12,13,14,15]. The validity of this approach hinges on the construct fidelity of the simulation (whether the synthetic learners exhibit qualitative patterns consistent with the literature they were inspired by) and on out-of-sample predictions whose direction was not built into the calibration. We address both in Section 4.10.
The complete experimental design is summarized in Figure 1.

3.2. Synthetic Learner Profiles

Five learner profiles were defined a priori based on the cross-product of prior knowledge, motivation, and learning style (Table 1). Each profile was operationalized through six behavioral parameters in [ 0 , 1 ] : question-asking initiative, help-seeking threshold, resistance to conceptual change, elaboration depth, engagement decay rate, and metacognitive awareness [26]. Parameter values were chosen to be consistent with documented patterns in the FCI literature [1,2].
The synthetic learner is implemented as an LLM agent prompted with (i) a persona description, (ii) the current cognitive state, and (iii) behavioral guidelines derived from the parameters. The agent’s verbal output is generated by the LLM; misconception-state transitions are governed by a separate deterministic update rule, decoupling expressed dialogue from internal cognitive state.

3.3. Misconception Set

We focus on five canonical Newtonian-mechanics misconceptions (Table 2), selected for their prevalence in the FCI literature [1]. Following diSessa [19], we type misconceptions by what cognitive coordination they require for correct understanding. MC04 (Action-Reaction-Cancel) is uniquely cross-object relational: correctly resolving it requires coordinating forces acting on two distinct objects. The other four are single-concept misconceptions, although MC03 (Motion-Implies-Force) is best characterized as a velocity–force relation error within a single object. We note that the five-misconception set excludes other prevalent FCI misconceptions, notably centrifugal-force beliefs and the role of medium in gravity (see Section 5.4).

3.4. Tutoring Conditions

Four tutoring conditions were compared (Table 3; Figure 1C) in a between-subjects design intended to isolate the contribution of retrieval mechanisms. All LLM-using conditions (C1–C3) shared the same base LLM and Socratic-tutoring system prompt; they differed only in information access at generation time. Retrieval in C1 and C2 used BM25 ranking [27].
C4 (Static-Feedback) is a non-LLM lower-bound control that confounds retrieval with LLM presence relative to C3. Throughout this manuscript we therefore describe the retrieval gradient over the three LLM-based conditions (KG-RAG vs. RAG vs. Vanilla) and use C4 as a non-LLM floor for effect-size anchoring, rather than as a retrieval comparator.

3.5. Experimental Design and Procedure

We employed a 5 × 4 × 30 between-subjects factorial design (600 sessions). A G*Power analysis [28] confirmed that 30 replications per cell yield power = 0.80 to detect medium effects ( f = 0.25 ) at α = 0.05 . Replications were seeded deterministically (master seed = 42; per-session seed = master_seed + profile_idx × 1000 + condition_idx × 100 + replication). The full corpus comprised 6650 LLM API calls. We note that cell-level analyses (e.g., individual Condition × Profile cells with n = 30 ) operate below the power level for which the design was sized; we therefore report 95% bootstrap CIs for cell-level effect sizes.

3.6. Behavioral Outcome Measures and the Revision Rule

Eight outcomes were operationalized at the session level. The primary outcome was the simulated Misconception Revision Rate (MRR): the proportion of initially-active misconceptions revised by session end. We provide the explicit functional form of the revision rule below to support direct reproduction from the manuscript.
For each tutor turn t and active misconception m, the synthetic learner scans the tutor’s text for a small set of literature-grounded revision-trigger keywords associated with m. Let k t , m ∈ { 0 , 1 , 2 , … } denote the count of distinct keyword matches in turn t for misconception m, and let c t , m denote the cumulative exposure count (the number of prior tutor turns containing ≥ 1 keyword for m). The revision probability on turn t is then
p revision ( t , m ) = min 0.90 , p base ( m ) + p exposure ( t , m ) + p richness ( t , m ) ,
with
p base ( m ) = ( 1 − r ) · 0.15 , p exposure ( t , m ) = ( 1 − r ) · c t , m · 0.12 , p richness ( t , m ) = min ( 0.20 , k t , m · 0.08 ) ,
where r is the simulated learner’s profile-specific resistance to conceptual change parameter as set in the released configuration file (P1: r = 0.70 ; P2: 0.60 ; P3: 0.50 ; P4: 0.30 ; P5: 0.20 ). If a stochastic draw succeeds, misconception m is flagged as revised and its confidence value drops to 0.10 ; otherwise confidence is reduced by 0.15 . The rule is evaluated only on tutor turns to which the learner actually replies: a tutor turn that is immediately followed by a scenario introduction is not scored, and a turn answered by a disengaged learner (engagement below 0.30 ) is skipped, since the learner agent returns before the update. This formulation makes explicit that MRR is, by construction, a keyword-density-driven simulator, and we quantify its consequences in Section 4.7 and Section 4.9.

Provenance of the Trigger Vocabulary and Separation from the Tutoring Conditions

The 17 trigger phrases (listed with their corpus inverse document frequencies in Table S1 of the Supplementary Materials; MC01 “inertia”, “first law”, “no force needed”, “keeps moving”; MC02 “same rate”, “same acceleration”, “regardless of mass”, “vacuum”; MC03 “zero net force”, “constant velocity”, “balanced forces”; MC04 “different objects”, “free body”, “third law pairs”; MC05 “acceleration direction”, “centripetal”, “perpendicular”) were written by the authors from the canonical correct formulations in the misconception literature (“inertia”, “first law”, “regardless of mass”, “balanced forces”, “third law pairs”, “centripetal” and so on) as part of the learner agent, before any tutor was run; they are hard-coded in the learner module, and the mock LLM used for offline pipeline tests was written against the same list. They were not derived from tutor output, and the tutor prompts, the retrieval corpus and the knowledge graph do not reference them. We state this as an author declaration: the released repository was committed as a single snapshot after the experiments, so its version history cannot independently attest to the order of authorship, which is one reason we add the vocabulary-independent checks below. The one place where construction and evaluation are not independent is the Static-Feedback condition: its hand-written templates were authored against the same misconception ontology as the trigger list, so they share phrasing with it. We treat this as a known structural advantage for Static-Feedback on the simulator’s measure (Section 5.2) and, in this revision, test the rule’s dependence on the specific vocabulary directly by re-scoring all 600 logged dialogues under alternative criteria (Section 4.9).

3.7. LLM Implementation

The primary experiment used OpenAI gpt-4o-mini-2024-07-18, selected to enable the full 5 × 4 × 30 factorial at low marginal cost (total API cost: $1.93; 9.10M input + 0.95M output tokens). A sensitivity experiment was conducted with Anthropic claude-sonnet-4-20250514 using a 3 (profiles: P1, P3, P5) × 4 (conditions) × 5 (replications) design, targeting 60 sessions (cost: $5.34). Two sessions from an earlier pipeline-validation mini-test were preserved on disk under an older metric-aggregation schema and were excluded from the present statistical analysis, yielding N = 58 sessions available for cross-provider comparison. The two LLM judges of Section 4.8 were gpt-4o-mini-2024-07-18 and claude-sonnet-4-5-20250929, each called at temperature 0 with the same rubric prompt; the second judge was added in revision after the first run, over the identical 174-session sample selected by a fixed random seed. We adopt the convention of typesetting model identifiers in monospaced font throughout. Generation parameters were temperature = 0.3 for tutors and 0.7 for learners. Each session’s full dialogue, learner final state, and per-turn metadata were persisted as JSON files.

3.8. Statistical Analysis

For each continuous outcome we conducted two-way factorial ANOVA with Condition (4 levels) and Profile (5 levels) as between-subjects factors. The design is balanced ( n = 30 per cell). Effect sizes are reported as partial eta-squared ( η p 2 ), with conventional thresholds: small ( η p 2 ≥ 0.01 ), medium ( η p 2 ≥ 0.06 ), large ( η p 2 ≥ 0.14 ) [29]. Significant ANOVA main effects were followed by Welch pairwise comparisons with Bonferroni correction. Effect sizes for pairwise comparisons are reported as Cohen’s d with bootstrap 95% confidence intervals (10,000 resamples). The two top profiles (P4 and P5) were also subjected to a two one-sided test (TOST) for equivalence at a smallest-effect-of-interest of d = ± 0.10 .

3.8.1. Family-Wise Error and Model Checks

Because eight outcomes are tested, the omnibus p-values for each factor are additionally reported after Holm’s step-down correction across the eight outcomes [30]; the confirmatory conclusions rest on MRR, and the secondary outcomes are exploratory. MRR is a bounded proportion with a Bernoulli-like distribution for the single-misconception profiles (Table 4), so ANOVA assumptions were checked with Levene’s test on the 20 cells and a Shapiro–Wilk test on the residuals, and the MRR analysis was repeated with a binomial generalized linear model (logit link) in which the response for each session is the pair (misconceptions revised, misconceptions not revised) [31]. Main effects and the interaction were tested with likelihood-ratio tests, overdispersion was assessed with the Pearson χ 2 to residual degrees of freedom ratio, and the tests were repeated with dispersion-scaled statistics (quasi-binomial).

3.8.2. Measurement Audit

Four exploratory analyses probe the revision rule itself. (i) Two LLM judges. A stratified random sample of 174 sessions (506 misconception-level items) was rated independently by two LLM judges shown the final three learner utterances of each session: gpt-4o-mini, from the same model family as the tutor and learner agents, and claude-sonnet-4-5, from a different family. Each judge classified each initially active misconception as still held, revised, or indeterminate (disengaged or not addressed; indeterminate is never counted as revision). Cohen’s κ [32] is reported between the simulator and each judge and between the judges, with bootstrap 95% CIs, together with the direction of every disagreement. (ii) Re-scoring ablation. Holding all 600 logged dialogues fixed, the stochastic revision rule was replayed 200 times per session under alternative operationalizations of “the tutor addressed misconception m”: the original vocabulary; each of the 17 leave-one-phrase-out vocabularies; a strict criterion requiring at least two distinct phrases in one turn; a version with the richness term removed; IDF-weighted matching, with and without a gate on total IDF weight; and a vocabulary authored independently of the trigger list, built from the correct_concept prose descriptions in the released misconception configuration (content terms and bigrams with corpus document frequency at most 5%, at least two matches required). The expected MRR by condition and profile is reported under each criterion. Because the tutor agents adapt to the learner’s current misconception state, a replay over fixed logs cannot reproduce the generative process exactly (after a logged revision the tutor stops targeting that misconception, so a replay in which the draw fails receives no further exposures); replayed levels are therefore conservative, and the analysis is read for orderings, not for absolute levels. (iii) Parameter perturbation. The same replay was repeated with each numeric constant of Equation (1) varied one at a time by ± 25 % and ± 50 % , the resistance parameters shifted by ± 0.075 and ± 0.15 for all profiles, and 200 joint random draws within ± 50 % of all constants. (iv) Keyword density. Section 4.7 relates condition-level keyword density to MRR. All analysis scripts are released with the data.

3.9. Reproducibility

Configuration files, retrieval corpora, scenarios, and per-session seeds are SHA-256-hashed at run start. All 600 session dialogue logs are released as JSON files. The full codebase, including the multi-provider LLM client supporting both the Anthropic and OpenAI APIs, is archived at Zenodo (https://doi.org/10.5281/zenodo.22713409). Data are released under the CC BY 4.0 license and code under the MIT license.

4. Results

4.1. Two Profile Clusters, Not a Ranked Five-Profile Gradient

The five profiles partition cleanly into two clusters on MRR (Table 4). A high-revision cluster comprising Intermediate-Strategic and Advanced-Residual ( M = 0.525 and 0.517 respectively) was statistically indistinguishable on a pairwise Welch t-test (mean difference + 0.008 , 95% CI [ − 0.119 , + 0.135 ] ; t ( 238 ) = 0.13 , p = 0.898 ; Cohen’s d = 0.017 , 95% bootstrap CI [ − 0.251 , + 0.269 ] ). A TOST equivalence test at SESOI d = ± 0.10 supported the conclusion that the two top profiles are statistically equivalent. A low-revision cluster comprising Intermediate-Confused, Naive-Active, and Naive-Passive ( M = 0.175 , 0.145 , 0.018 ) was separated from the high cluster by a large between-cluster Cohen’s d = 1.20 (95% bootstrap CI [ + 1.005 , + 1.408 ] ; Welch t = 12.19 , p < 0.001 ). We therefore frame the profile main effect as a two-cluster rather than a fully ranked five-profile gradient.

4.2. Retrieval Gradient Across LLM-Based Conditions Under the Simulator’s Measure

Under the simulator’s keyword-triggered revision rule, MRR follows a monotonic retrieval gradient among the three LLM-based conditions: KG-RAG-Tutor ( M = 0.40 ) > RAG-Tutor ( M = 0.33 ) > Vanilla-LLM ( M = 0.24 ); Static-Feedback ( M = 0.13 ) anchors the non-LLM floor (Figure 2, Table 5). The KG-RAG vs. Vanilla contrast among LLM conditions corresponds to Cohen’s d ≈ 0.42 . Because the LLM-vs-Static contrast confounds retrieval with LLM presence, we describe the retrieval gradient strictly within the LLM-based conditions. Two qualifications, established in Section 4.8, Section 4.9, Section 4.10, Section 4.11 and Section 4.12, apply to this gradient and are carried through the Discussion and Conclusions: it is a property of the keyword-based measure (neither LLM judge reproduces it, although both reproduce the LLM-vs-Static contrast), and it was not reproduced in the small Claude Sonnet 4 sensitivity run. The re-scoring ablation (Section 4.9) shows that, within keyword-based measurement, it does not depend on any single phrase or on the particular list used.

4.3. Main Effects and Interaction

A two-way factorial ANOVA on MRR (Table 6) yielded significant main effects of Condition, F ( 3 , 580 ) = 21.20 , p < 0.001 , η p 2 = 0.099 , and Profile, F ( 4 , 580 ) = 66.54 , p < 0.001 , η p 2 = 0.315 , as well as a significant Condition × Profile interaction, F ( 12 , 580 ) = 5.17 , p < 0.001 , η p 2 = 0.097 . Across the eight outcomes, the profile main effect was significant in every case; the condition main effect in six of eight; and the interaction in five of eight.

4.4. Family-Wise Error Control, Assumption Checks and a Binomial GLM

After Holm correction across the eight outcomes (Table 6, p H columns), every effect that reaches p < 0.001 uncorrected remains significant, including all three MRR effects; the only change is that the Scaffolding Responsiveness condition effect ( p = 0.034 ) no longer reaches the 5% level ( p H = 0.067 ). The condition main effect therefore holds for five of eight outcomes after correction, and the interaction for five of eight.
The ANOVA assumptions do not hold for MRR. Levene’s test rejects homogeneity of variance across the 20 cells ( W = 6.22 , p < 0.001 ; cell SDs range from 0.04 to 0.50 ), the Shapiro–Wilk test rejects normality of the residuals ( W = 0.962 , p < 0.001 ), and 78% of sessions have MRR exactly 0 or 1. The binomial GLM described in Section 3.8 is the appropriate model for such data, and it reproduces the ANOVA conclusions (Table 7): likelihood-ratio tests give Condition χ 2 ( 3 ) = 60.4 , p < 0.001 , Profile χ 2 ( 4 ) = 314.7 , p < 0.001 , and the Condition × Profile interaction χ 2 ( 12 ) = 29.5 , p = 0.003 , with the full model preferred over the additive model by AIC ( 828.4 vs. 833.9 ). The Pearson dispersion of the full model is 0.96 , so there is no overdispersion to correct, and the dispersion-scaled tests are essentially unchanged. The GLM’s fitted condition-level revision probabilities equal the observed session means ( 0.403 , 0.329 , 0.240 , 0.132 ), and the pooled misconception-level proportions ( 0.240 , 0.187 , 0.127 , 0.082 ) preserve the same ordering. The interaction is weaker in the GLM ( p = 0.003 ) than in the ANOVA ( p < 0.001 ), which is the expected consequence of modelling the binary sessions correctly; we therefore describe the interaction as reliable but of moderate strength.

4.5. Condition × Profile Interaction

Figure 3 visualizes the interaction. In the high-revision cluster (Intermediate-Strategic, Advanced-Residual), the retrieval gradient is fully expressed: KG-RAG-Tutor reaches MRR ≥ 0.70 , with Cohen’s d = 2.40 (95% bootstrap CI [ + 1.63 , + 4.05 ] ) between the peak cell (Intermediate-Strategic × KG-RAG, M = 0.77 , n = 30 ) and the floor cell (Naive-Passive × KG-RAG, M = 0.03 , n = 30 ). In the low-revision cluster (Naive-Active, Naive-Passive), all conditions converge near zero, with no condition producing substantial revision; for Naive-Passive, MRR is ≤ 0.03 in every condition.
The complete cell-level heatmap is shown in Figure 4, on a 0–1.0 scale.

4.6. Per-Misconception Revision

Per-misconception MRR by tutoring condition (Figure 5) reveals heterogeneity; this breakdown is exploratory. Three observations stand out. First, the Impetus misconception (MC01) showed the largest within-condition retrieval gradient (KG-RAG 0.40 vs. Vanilla 0.14 , Δ = 0.26 ). Second, the Heavier-Faster misconception (MC02) was revised at moderate rates across all LLM conditions, and was the only misconception in our set where the full retrieval gradient was preserved across all conditions including Static-Feedback at lower levels. Third, Static-Feedback was competitive with Vanilla-LLM on the two relational/cross-object misconceptions (MC03 Motion-Implies-Force and MC04 Action-Reaction-Cancel). We return to this third observation in Section 5.2.

4.7. Keyword Density and the Mechanism of the Condition Effect

Because the revision rule (Section 3.6) is by construction keyword-driven, we quantified, as an exploratory analysis, the extent to which the condition effect on MRR is attributable to keyword density in tutor turns rather than to any non-keyword pedagogical feature. Across the 600 sessions, tutor-turn keyword density (revision-trigger keywords per tutor turn) varied by condition: KG-RAG M = 0.91 , RAG M = 0.67 , Static M = 0.69 , Vanilla M = 0.55 . At the condition level, mean keyword density correlated with mean MRR at r = 0.61 , indicating that keyword density accounts for a substantial share of the between-condition variance.
However, two findings complicate a purely keyword-density account. First, within each condition, the session-level correlation between keyword density and MRR was weak or null in the two LLM-based RAG conditions (KG-RAG r = − 0.12 , p = 0.16 ; RAG r = +0.12, p = 0.16 ), and moderate only in Static-Feedback ( r = + 0.30 , p < 0.001 ) and Vanilla-LLM ( r = + 0.22 , p < 0.01 ). Second, Static-Feedback exhibits the second-highest keyword density (0.69, near RAG-Tutor’s 0.67) but the lowest MRR (0.13). This non-monotonicity rules out a purely keyword-density-mediated account: keyword density is a necessary but not sufficient input. We interpret this as evidence that retrieval contributes additional structure beyond keyword density (for instance, contextual coherence of multi-keyword passages or sequencing of explanations across turns), although the present simulation does not isolate these mechanisms further.

4.8. Two LLM Judges from Different Model Families

To address the construct-validity concern that MRR is computed by a keyword-trigger rule rather than by direct assessment of the learner’s expressed beliefs (Section 3.6 and Section 4.7), we had the same stratified random sample of 174 GPT-4o-mini sessions (29% of the primary corpus; 506 misconception-level items) rated by two LLM judges, each a separate instance from the tutor and learner stack and each shown only the last three learner utterances of a session. Judge 1 (gpt-4o-mini) belongs to the same model family as the agents; Judge 2 (claude-sonnet-4-5) does not. Each judge classified every initially active misconception as (i) still held, (ii) revised (the learner explicitly expresses the correct view), or (iii) indeterminate (disengaged, or the misconception was not addressed). The third category is critical: absence of misconception expression is never counted as revision. This analysis is exploratory.
Judge 1 returned 204 indeterminate items (40%) and Judge 2 returned 244 (48%). Table 8 summarizes agreement on the determinate items. Against Judge 1 the simulator’s rule agrees in 65.1% of 301 items, Cohen’s κ = 0.12 (bootstrap 95% CI [ 0.01 , 0.23 ] ); against Judge 2 it agrees in 41.6% of 262 items, κ = 0.07 (95% CI [ 0.00 , 0.13 ] ). Both values are in the “slight” band of the Landis and Koch scale [32]. Per-condition kappas are uniformly low for both judges (Judge 1: KG-RAG 0.16 , RAG 0.12 , Vanilla 0.09 , Static 0.06 ; Judge 2: 0.10 , 0.02 , 0.02 , 0.18 ), per-misconception kappas range from − 0.01 to 0.30 , with Motion-Implies-Force at chance under both, and on the 155 items on which the two judges are both determinate and agree with each other the simulator’s κ with that consensus is 0.10 (95% CI [ − 0.01 , 0.21 ] ). We treat this low agreement as a central result rather than a caveat: the simulator’s MRR and an LLM rater’s reading of the learner’s own words are largely independent operationalizations of revision. The disagreement runs in the same direction for both judges, which find revisions the rule does not: Judge 1 classified 33% of its determinate items as revised against the simulator’s 20% (73 judge-revised/simulator-not items versus 32 in the reverse cell), Judge 2 76% against 26% (142 versus 11). The rule therefore under-detects revision relative to either judge, and the shared-family judge is the stricter of the two. Between the judges agreement is moderate ( κ = 0.41 , 95% CI [ 0.32 , 0.50 ] , on 228 items determinate under both; three-category κ = 0.48 over all 506 items) and entirely one-sided: no item was called revised by Judge 1 and still held by Judge 2, whereas 73 items ran the other way. Whatever self-preference bias the same-family judge may carry [25], it did not take the form of inflated revision credit for GPT-generated dialogue, so the low simulator–judge agreement is not an artefact of a lenient in-family rater: a stricter in-family rater and a more lenient out-of-family rater both disagree with the rule in the same direction.
Condition-level revision proportions under each judge (determinate items; per-condition n = 56 –101; descriptive because the subsets are unbalanced) preserve the LLM-versus-Static-Feedback contrast: pooled over the three LLM conditions, Judge 1 rates 38.8% of items revised against 14.5% for Static-Feedback (Fisher’s exact p < 0.001 ) and Judge 2 80.6% against 57.1% ( p < 0.001 ). Neither judge preserves the retrieval gradient (Judge 1: KG-RAG 0.34 , RAG 0.44 , Vanilla 0.38 ; Judge 2: 0.77 , 0.78 , 0.85 ; χ 2 across the three LLM conditions p = 0.49 and p = 0.40 ); the same picture holds when indeterminate items are counted as not revised (Figure 6B). The profile clustering is visible under both judges: Naive-Passive items are never judged revised (0 of 29 and 0 of 9), Naive-Active items rarely (14% and 61%), and the three higher profiles most often (55–79% and 88–92%). These results are direct evidence that the rule systematically under-detects revisions that either judge recognizes from explicit endorsement of the correct concept, and that the retrieval gradient is specific to the keyword-based measure.

4.9. Robustness of the Revision Measure: Re-Scoring Ablation and Parameter Perturbation

Dependence on the trigger vocabulary can be tested directly (exploratory analysis), because the dialogue logs are fixed and the rule is cheap to replay. Table 9 reports the expected MRR by condition when all 600 logged sessions are re-scored under the alternative criteria of Section 3.8 (200 stochastic replays per session; the complete table including all 17 leave-one-phrase-out variants and profile means is Table S2). Replayed levels under the original vocabulary ( 0.36 , 0.31 , 0.27 , 0.17 ) are lower and more compressed than the logged means ( 0.40 , 0.33 , 0.24 , 0.13 ) for the reason given in Section 3.8: the tutors adapt to the learner’s state, so a replay over fixed text loses the exposures a live session would have generated after a failed draw. The profile means, which do not depend on that adaptation, are reproduced almost exactly ( 0.01 , 0.15 , 0.17 , 0.51 , 0.53 against logged 0.02 , 0.15 , 0.18 , 0.53 , 0.52 ), so the table is read for orderings.
Three results follow. First, the ordering KG-RAG > RAG > Vanilla is preserved under every one of the 17 leave-one-phrase-out vocabularies, under IDF weighting, under removal of the richness term, under the two-phrase gate, and under the independently authored concept-text vocabulary: the gradient is a stable consequence of how much correct-concept language the three LLM tutors emit, not of any single phrase, although its size depends on the corpus-frequent phrases (removing “centripetal”, “different objects” or “inertia” shrinks all three LLM conditions together). Second, the LLM-versus-Static contrast is the part that depends on the criterion. Under the original vocabulary and its perturbations Static-Feedback is lowest, but under the two strict gates (at least two phrases, or IDF weight at least one) it rises to the level of KG-RAG ( 0.16 vs. 0.11 ), because its templates were authored against the same ontology and reliably contain the canonical phrase pairs, whereas LLM tutors more often use one phrase at a time; this quantifies the structural advantage anticipated in Section 5.2. Under the independent concept-text vocabulary, which shares no phrase list with the rule or the templates, Static-Feedback is again clearly lowest ( 0.05 vs. 0.12 – 0.27 ), matching both LLM judges. Third, the Naive-Passive floor holds under every criterion (best cell ≤ 0.03 ) and the two-cluster profile structure in 22 of 23 variants; the exception is removal of “centripetal”, the phrase through which the Advanced-Residual profile’s sole misconception (MC05) is most often addressed (its other two triggers are corpus-rare), which collapses that profile from 0.53 to 0.12 and with it the high-cluster gap and the LLM-versus-Static contrast in that one variant. This is the one place where the measure rests on a single phrase, and it concerns a profile-level result rather than the condition ordering.
The same replay was used, again as an exploratory analysis, to ask whether the orderings depend on the particular numeric constants of Equation (1). Varying 0.15 , 0.12 , 0.08 , the richness cap 0.20 and the probability cap 0.90 one at a time by ± 25 % and ± 50 % (20 variants), shifting all five resistance values by ± 0.075 and ± 0.15 (4 variants), and drawing all constants jointly at random within ± 50 % (200 variants) preserved the KG-RAG > RAG > Vanilla > Static ordering, the positive high-cluster gap and the Naive-Passive floor in all 224 variants. Across the joint draws the KG-RAG mean ranged from 0.23 to 0.46 , the KG-RAG minus RAG difference from 0.04 to 0.07 , the RAG minus Vanilla difference from 0.02 to 0.05 , the high-cluster gap from 0.23 to 0.42 , the largest Naive-Passive cell never exceeded 0.04 and the two top profiles never differed by more than 0.04 (Table S3). The constants therefore set the level of MRR, which we do not interpret, and not the orderings, which we do.

4.10. Out-of-Sample Construct Fidelity

A central concern in synthetic-learner methodology is that calibration and validation can become circular: agents tuned to reproduce FCI patterns will, trivially, reproduce them. To partially address this, we examined an out-of-sample prediction whose direction was not used in calibration: per-misconception revision latency (the number of turns to first revision, conditional on revision occurring). The FCI literature places Heavier-Faster among the easiest misconceptions to revise via classroom demonstrations [20], while Impetus is among the most persistent under traditional instruction [1,19].
Observed latencies (sessions where the given misconception was revised): Heavier-Faster M = 3.95 turns (shortest, consistent with FCI prediction); Velocity-Force-Confusion M = 5.29 turns (longest); Impetus M = 4.08 , Action-Reaction-Cancel M = 4.25 , Motion-Implies-Force M = 4.33 (intermediate). The simulation thus correctly predicts the easiest-to-revise misconception (Heavier-Faster) without that ordering being built into the keyword-trigger rules, but does not reproduce the FCI prediction that Impetus should require the most exposure. We interpret this as partial construct fidelity: the simulation captures some FCI-documented patterns out-of-sample, but not all. The Impetus mismatch is a calibration limitation that warrants caution when generalizing simulation predictions to human learners.

4.11. Behavioral Side-Channels

Beyond MRR, the seven secondary outcomes provide a richer, exploratory picture (Table 6). Engagement was overwhelmingly profile-determined ( η p 2 = 0.979 ) because engagement decay is a profile parameter. Inquiry Depth (Figure 7) showed a very large profile main effect ( η p 2 = 0.715 ) but no condition effect ( p = 0.72 ): who is asking dominates how they are scaffolded for the depth of inquiry in this simulator.

4.12. Sensitivity Run on a Second Provider

The Claude Sonnet 4 run (3 profiles × 4 conditions × 5 replications; 60 sessions targeted, N = 58 analyzable; Section 3.7) is a small, exploratory sensitivity check, not a replication, and we report it as preliminary sensitivity evidence only (Figure 8). Two of the primary patterns were also present in this run: every LLM condition exceeded Static-Feedback, and the high-revision profiles were separated from the low-revision ones. The retrieval gradient among LLM-based conditions was not reproduced: Vanilla-LLM had the highest MRR ( M = 0.357 ) of the three LLM conditions. With n = 14 –15 sessions per condition the 95% CIs of the three LLM conditions overlap heavily, so the run can neither confirm nor reject the gradient, and it cannot distinguish a provider difference from sampling noise. What it does establish is that the gradient is not a pattern that appears automatically whenever the harness is run; combined with the two judges (Section 4.8), this is why we describe the gradient as measure-dependent throughout. A full 600-session run on a second provider is listed as future work (Section 6).

5. Discussion

5.1. Three Behavioral Patterns from the Simulator

Three patterns emerged consistently across the 600-session primary experiment.
The retrieval gradient is a property of the keyword-based measure. Under the simulator’s rule, knowledge-graph-augmented RAG produced approximately 70% higher MRR than the same LLM without retrieval ( 0.40 vs. 0.24 , d ≈ 0.42 ), the ordering survives Holm correction and the binomial GLM, and the ablation shows it does not depend on any single phrase, on the richness term, on IDF weighting, on the numeric constants, or even on the particular vocabulary, since an independently authored concept-text vocabulary reproduces it. What it does depend on is the kind of measure: neither LLM judge reproduces it, and the second-provider run did not show it. The most parsimonious reading is that retrieval increases the amount of canonical correct-concept language the tutor emits (Section 4.7), that any lexical measure of revision will register this, and that a rater reading the learner’s own words does not. We therefore report the gradient as a finding about what retrieval does to tutor output, and as a hypothesis about learners, not as evidence that retrieval improves revision.
Two profile clusters, not five gradations. The five learner profiles partition into two clusters on MRR: a high-revision cluster of Intermediate-Strategic and Advanced-Residual learners ( M ≈ 0.52 , statistically indistinguishable from one another) and a low-revision cluster of the remaining three profiles ( M < 0.18 ). The defining property of the high cluster within our parameterization is the combination of moderate-to-high prior knowledge and high motivation; learners possessing only one of these characteristics (Naive-Active: high motivation, very low prior knowledge) did not enter the high cluster. The simulator thus predicts a particular interaction between prior knowledge and motivation that warrants empirical testing.
The interaction is reliable and pedagogically meaningful. A Condition × Profile interaction emerged on MRR ( η p 2 = 0.097 in the ANOVA; χ 2 ( 12 ) = 29.5 , p = 0.003 in the binomial GLM), and similar interactions appeared on five of eight outcomes after Holm correction. The most pedagogically meaningful pattern is that no condition was able to produce substantial revision for Naive-Passive learners ( ≤ 0.03 across all conditions). Within the simulator, retrieval-augmented tutoring does not lift learners whose engagement-decay parameter is set high and whose motivation is set low. This is a behavioral hypothesis that, if confirmed empirically, would imply pre-tutoring engagement-raising interventions as a prerequisite for AI-tutoring effectiveness.

5.2. Per-Misconception Heterogeneity and an Alternative Reading of Static-Feedback Competitiveness

The per-misconception breakdown revealed two patterns worth highlighting. The Impetus misconception showed the largest within-condition retrieval gradient (KG-RAG 0.40 vs. Vanilla 0.14 ). For this cognitively entrenched belief, structured retrieval, in our simulator, surfaces multiple mutually-reinforcing keyword-bearing passages that the LLM weaves into multi-turn explanations.
Static-Feedback was unexpectedly competitive with Vanilla-LLM on MC03 (Motion-Implies-Force) and MC04 (Action-Reaction-Cancel). We considered two readings. The first is substantive: hand-curated templates whose phrasings closely match the canonical correct framings of certain misconceptions are competitive with parametric LLM generation. The second, and in our view more parsimonious, reading is a methodological artifact: the Static-Feedback templates and the revision-trigger keywords used by the simulator were both authored against the same misconception ontology. Static templates therefore enjoy a built-in advantage on the simulator’s keyword-driven MRR rule that does not reflect a pedagogical mechanism observable in human learners. We caution against interpreting the Static-Feedback competitiveness as evidence that templates can replace LLMs for these misconceptions; the result more likely reflects the geometry of the simulator’s evaluation criterion.

5.3. Construct Fidelity: What the Simulator Recovers and What It Does Not

The simulator reproduces several qualitative patterns that were not built directly into the calibration rules. Out-of-sample, Heavier-Faster emerges as the misconception with the shortest revision latency, consistent with FCI documentation that vacuum-based demonstrations efficiently support its revision [20]. The profile clustering (P4 ≈ P5 ≫ others) is consistent with the documented role of prior knowledge in conceptual change [5,18] and with the Hake-gain observation that interactive engagement is more effective for students with sufficient initial scaffold [3].
However, the simulator does not reproduce the FCI prediction that Impetus should be the most persistent misconception under naive instruction. Impetus latency in our simulation is near the median, not the maximum. We attribute this to the design of the keyword-trigger rules: the Impetus trigger keywords (“inertia,” “first law,” “no force needed,” “keeps moving”) happen to be high-frequency phrases in the retrieved physics corpora, giving the simulator a structural advantage on Impetus that does not exist in human cognition. This is the kind of out-of-sample mismatch that deliberate construct-fidelity probing is designed to surface. An obvious next calibration step, deferred to future work, would be to downweight or threshold trigger keywords by their inverse document frequency in the retrieval corpus, so that the revision rule rewards corpus-rare keyword matches more strongly than corpus-common ones; we expect this would partially correct the Impetus mismatch and yield more conservative MRR estimates for misconceptions whose trigger phrases are central to the host corpus.

5.4. Limitations

Several limitations are important to acknowledge.
Synthetic learners are not human learners. Our findings characterize the behavioral dynamics of a closed simulator with calibrated parameters; they do not directly establish that real human learners would show the same patterns under the same conditions. We position this work as hypothesis-generating; the patterns reported here warrant empirical testing with human participants before being interpreted as evidence about learning.
Misconception revision is operationalized as keyword-triggered state flipping. This is a substantial simplification of the cognitive process underlying conceptual change. Section 4.7 quantifies the mechanistic limitations; Section 4.8 shows only slight agreement between the rule and two LLM judges from different model families ( κ = 0.12 and 0.07 ), both of which find far more revisions than the rule does. The reported MRR magnitudes should therefore be read as simulator-internal quantities, not as estimates of true revision rates, and the retrieval gradient as a property of the keyword-based measure. The judges have limitations of their own: each observed only the final three learner utterances, the two agree with each other only moderately ( κ = 0.41 ) and differ substantially in leniency, and neither has been validated against human raters on this task. A future calibration round should either adopt a validated judge as the primary measure or tighten the keyword rule until its positive rate matches a validated judge’s.
Static-Feedback evaluation is structurally favorable, and the trigger vocabulary is author-declared. The Static-Feedback templates and the revision-trigger keywords were authored against the same misconception ontology (Section 5.2); Section 4.9 shows that this inflates Static-Feedback under strict phrase-pair criteria to the level of KG-RAG, while under the original rule, under an independent vocabulary and under both judges Static-Feedback is clearly lowest. The claim that the trigger vocabulary was fixed before the tutors were run rests on the authors’ declaration and on the structure of the code, not on an auditable version history (Section 3.6); we mitigate this with the vocabulary-independent criteria and the judges, and recommend that future simulation studies commit the evaluation rule to a public repository before generating data.
Second-provider evidence is preliminary. The 58-session Claude Sonnet 4 run is a sensitivity check with n = 14 –15 per condition; it can show that a pattern is not automatic, which it does for the retrieval gradient, but it cannot establish or rule out a provider difference. A full 600-session run on a second provider, projected to cost approximately $53 in API charges and roughly 5 hours of sequential compute, is the most important next step and is not provided here. Resumable infrastructure for it is released alongside the manuscript.
Cell-level subsamples have limited power, and most analyses are exploratory. The 5 × 4 × 30 design was powered for the omnibus test on MRR, not for cell-level contrasts ( n = 30 ). Cell-level effect sizes (e.g., the Intermediate-Strategic × KG-RAG peak) are reported with bootstrap 95% CIs but are exploratory, as are the seven secondary outcomes, the judge, ablation and perturbation analyses and the second-provider run (Section 3.8). The confirmatory results are the three MRR effects, which survive Holm correction and the binomial GLM.
Misconception set coverage. The five-misconception set excludes other prevalent FCI misconceptions, notably centrifugal-force beliefs and the role of medium in gravity. Generalization to a broader FCI-style misconception space remains to be tested.
Single-domain coverage. Newtonian mechanics was chosen for its uniquely well-documented misconception literature; whether the patterns generalize to other STEM domains is an empirical question.

5.5. Hypotheses for Human-Subject Follow-Up

Within the synthetic-learner system, several patterns are sufficiently robust to warrant pre-registered empirical testing:
  • Hemp 1. For real introductory physics students, LLM tutoring of any of the three architectures will produce higher pre-/post-FCI gains than template-based feedback; whether knowledge-graph-augmented RAG additionally outperforms non-RAG LLM tutoring is the measure-dependent part of our results and should be tested as a secondary, directional hypothesis.
  • Hemp 2. The interaction between prior knowledge and motivation predicts FCI-gain such that students with high motivation but low prior knowledge gain less than students with moderate-to-high prior knowledge and high motivation, even when receiving the same tutoring.
  • Hemp 3. Engagement raised pre-tutoring (e.g., via goal-setting or interest interventions) will substantially lift FCI-gain for students who would otherwise fall into the low-engagement, low-prior-knowledge phenotype.
These hypotheses are offered as testable predictions; they are not claims about how human learning works.

6. Conclusions and Future Work

This article asked whether a deterministically seeded simulation harness with LLM-based synthetic learners can compare retrieval architectures for LLM tutoring at a learner-level outcome, and how far its answers can be trusted. Across 600 sessions on GPT-4o-mini, the main findings are the following. First, the learner-profile effect is the largest in the design ( η p 2 = 0.315 ) and resolves into two clusters rather than a ranked gradient: simulated learners with moderate-to-high prior knowledge and high motivation revise, and the other three profiles largely do not; this structure survives Holm correction, a binomial GLM, every re-scoring criterion, every perturbation of the rule’s constants, and is visible under both LLM judges. Second, a Condition × Profile interaction ( η p 2 = 0.097 ; GLM p = 0.003 ) shows that no architecture lifts the low-engagement, low-prior-knowledge profile, whose MRR stays at or below 0.03 in every condition and under every criterion. Third, every LLM-based tutor exceeds the non-LLM template baseline under the simulator’s rule, under both judges, and in the second-provider run. Fourth, and stated with its qualification, the simulator’s measure orders the three LLM tutors as KG-RAG > RAG > Vanilla ( η p 2 = 0.099 ); this ordering is stable across 17 leave-one-phrase-out vocabularies, an independently authored vocabulary and 224 perturbations of the rule’s constants, but it is not reproduced by either LLM judge ( κ with the simulator = 0.12 and 0.07 ) and was absent in the 58-session Claude Sonnet 4 run. It is therefore a finding about what retrieval does to tutor language, and a hypothesis about learners, not evidence that retrieval improves revision.
The methodological contribution is the audit itself. Judging a simulator’s primary measure against two raters from different model families, replaying the measure over fixed logs under alternative criteria, and perturbing its constants are all cheap once the dialogue corpus exists, and together they separate the results that are properties of the simulated learners from those that are properties of the measurement rule. We recommend this practice for simulation-based evaluations of educational technology in general.
Future work follows directly from the limitations. (1) A full 600-session run on a second provider, for which resumable infrastructure is released, would test whether the retrieval gradient is provider-specific. (2) Validating the LLM judges against human raters on a subset of dialogues, and then adopting a validated judge as the primary revision measure, would remove the keyword rule as the point of failure; the judges’ own moderate agreement ( κ = 0.41 ) shows that this validation is needed rather than optional. (3) The single-phrase dependence identified for the Advanced-Residual profile (“centripetal”) should be removed by enlarging the MC05 vocabulary before the harness is reused. (4) Extending the misconception set to centrifugal-force beliefs and the role of medium in gravity, and the domain beyond Newtonian mechanics, would test the generality of the two-cluster structure. (5) The three empirical hypotheses of Section 5.5 will be pre-registered on the Open Science Framework before any human-subject follow-up, with the simulator-derived effect sizes as the basis for the registered power analyses. The full pipeline, all 600 dialogue logs, both judges’ verdicts, the ablation and perturbation scripts and the analysis code are released so that every number in this article can be recomputed.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/app16199938/s1, Table S1: revision-trigger vocabulary with corpus IDF; Table S2: complete re-scoring ablation including all leave-one-phrase-out variants and profile means; Table S3: parameter perturbation results.

Author Contributions

Conceptualization, S.Y. and G.Y.; methodology, S.Y. and G.Y.; software, S.Y.; validation, S.Y. and G.Y.; formal analysis, S.Y.; investigation, S.Y. and G.Y.; data curation, S.Y.; writing–original draft preparation, S.Y.; writing–review and editing, S.Y. and G.Y.; visualization, S.Y.; supervision, G.Y.; project administration, G.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. This study used computationally simulated agents only; no human participants were involved.

Informed Consent Statement

Not applicable.

Data Availability Statement

All session dialogue logs, computed metrics, configuration files, the verdicts of both LLM judges, the re-scoring ablation and perturbation scripts, and the analysis code are openly available at Zenodo at https://doi.org/10.5281/zenodo.22713409, under the CC BY 4.0 license for data and the MIT license for code.

Acknowledgments

Experimental use of LLMs. The synthetic-learner and tutor agents in this study were implemented by querying OpenAI gpt-4o-mini-2024-07-18 and Anthropic claude-sonnet-4-20250514 via their respective public APIs; the two LLM judges were gpt-4o-mini-2024-07-18 and claude-sonnet-4-5-20250929. The dialogue dataset analyzed in this manuscript is the output of these queries. The authors take full responsibility for the experimental design, prompting strategies, and post-hoc analysis. Manuscript-preparation use of LLMs. During the preparation of this manuscript, the authors used Claude (Anthropic) to assist with copy-editing of draft sections, generation of LaTeX boilerplate from a Markdown outline, and to suggest formulations of statistical descriptions. All ideas, hypotheses, theoretical framings, interpretations, and final wording are the authors’ own. The authors reviewed and edited all LLM-assisted output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ANOVAAnalysis of Variance
BM25Best Matching 25 (retrieval algorithm)
CIConfidence Interval
FCIForce Concept Inventory
IDSInquiry Depth Score
KG-RAGKnowledge-Graph-augmented Retrieval-Augmented Generation
LLMLarge Language Model
MRRMisconception Revision Rate
RAGRetrieval-Augmented Generation
SESOISmallest Effect Size Of Interest
TOSTTwo One-Sided Tests (equivalence)

References

  1. Hestenes, D.; Wells, M.; Swackhamer, G. Force concept inventory. Phys. Teach. 1992, 30, 141–158. [Google Scholar] [CrossRef] [Scilit]
  2. Halloun, I.A.; Hestenes, D. The initial knowledge state of college physics students. Am. J. Phys. 1985, 53, 1043–1055. [Google Scholar] [CrossRef] [Scilit]
  3. Hake, R.R. Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses. Am. J. Phys. 1998, 66, 64–74. [Google Scholar] [CrossRef] [Scilit]
  4. Posner, G.J.; Strike, K.A.; Hewson, P.W.; Gertzog, W.A. Accommodation of a scientific conception: Toward a theory of conceptual change. Sci. Educ. 1982, 66, 211–227. [Google Scholar] [CrossRef] [Scilit]
  5. Vosniadou, S. The cognitive-situative divide and the problem of conceptual change. Educ. Psychol. 2007, 42, 55–66. [Google Scholar] [CrossRef] [Scilit]
  6. Strike, K.A.; Posner, G.J. A revisionist theory of conceptual change. In Philosophy of Science, Cognitive Psychology, and Educational Theory and Practice; Duschl, R.A., Hamilton, R.J., Eds.; State University of New York Press: Albany, NY, USA, 1992; pp. 147–176. [Google Scholar]
  7. Kasneci, E.; Sessler, K.; Küchemann, S.; Bannert, M.; Dementieva, D.; Fischer, F.; Gasser, U.; Groh, G.; Günnemann, S.; Hüllermeier, E.; et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individ. Differ. 2023, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
  8. Kestin, G.; Miller, K.; Klales, A.; Milbourne, T.; Ponti, G. AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Sci. Rep. 2025, 15, 17458. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. Graesser, A.C.; Lu, S.; Jackson, G.T.; Mitchell, H.H.; Ventura, M.; Olney, A.; Louwerse, M.M. AutoTutor: A tutor with dialogue in natural language. Behav. Res. Methods Instrum. Comput. 2004, 36, 180–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  10. VanLehn, K. The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educ. Psychol. 2011, 46, 197–221. [Google Scholar] [CrossRef] [Scilit]
  11. Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2020; Volume 33, pp. 9459–9474. [Google Scholar]
  12. Argyle, L.P.; Busby, E.C.; Fulda, N.; Gubler, J.R.; Rytting, C.; Wingate, D. Out of one, many: Using language models to simulate human samples. Polit. Anal. 2023, 31, 337–351. [Google Scholar] [CrossRef] [Scilit]
  13. Aher, G.V.; Arriaga, R.I.; Kalai, A.T. Using large language models to simulate multiple humans and replicate human subject studies. In Proceedings of the 40th International Conference on Machine Learning (ICML), Honolulu, HI, USA, 23–29 July 2023; Volume 202, pp. 337–371. [Google Scholar]
  14. Park, J.S.; O’Brien, J.C.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), San Francisco, CA, USA, 29 October–1 November 2023; pp. 1–22. [Google Scholar]
  15. Horton, J.J.; Filippas, A.; Manning, B.S. Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? NBER Working Paper No. 31122; National Bureau of Economic Research: Cambridge, MA, USA, 2023. [Google Scholar]
  16. Sweller, J. Cognitive load theory. Psychol. Learn. Motiv. 2011, 55, 37–76. [Google Scholar] [CrossRef] [Scilit]
  17. Nye, B.D.; Graesser, A.C.; Hu, X. AutoTutor and family: A review of 17 years of natural language tutoring. Int. J. Artif. Intell. Educ. 2014, 24, 427–469. [Google Scholar] [CrossRef] [Scilit]
  18. Chinn, C.A.; Brewer, W.F. The role of anomalous data in knowledge acquisition: A theoretical framework and implications for science instruction. Rev. Educ. Res. 1993, 63, 1–49. [Google Scholar] [CrossRef]
  19. diSessa, A.A. Toward an epistemology of physics. Cogn. Instr. 1993, 10, 105–225. [Google Scholar] [CrossRef] [Scilit]
  20. Hammer, D. Misconceptions or p-prims: How may alternative perspectives of cognitive structure influence instructional perceptions and intentions? J. Learn. Sci. 1996, 5, 97–127. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Levonian, Z.; Li, C.; Zhu, W.; Gade, A.; Henkel, O.; Postle, M.-E.; Xing, W. Retrieval-augmented generation to improve math question-answering: Trade-offs between groundedness and human preference. In Proceedings of the NeurIPS 2023 Workshop on Generative AI for Education (GAIED), New Orleans, LA, USA, 15 December 2023. [Google Scholar]
  22. Käser, T.; Alexandron, G. Simulated learners in educational technology: A systematic literature review and a Turing-like test. Int. J. Artif. Intell. Educ. 2024, 34, 545–585. [Google Scholar] [CrossRef] [Scilit]
  23. Markel, J.M.; Opferman, S.G.; Landay, J.A.; Piech, C. GPTeach: Interactive TA training with GPT-based students. In Proceedings of the Tenth ACM Conference on Learning @ Scale (L@S), Copenhagen, Denmark, 20–22 July 2023; pp. 226–236. [Google Scholar]
  24. Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2023; Volume 36, pp. 46595–46623. [Google Scholar]
  25. Panickssery, A.; Bowman, S.R.; Feng, S. LLM evaluators recognize and favor their own generations. In Advances in Neural Information Processing Systems; Curran Associates: Red Hook, NY, USA, 2024; Volume 37, pp. 68772–68802. [Google Scholar]
  26. Zimmerman, B.J. Becoming a self-regulated learner: An overview. Theory Into Pract. 2002, 41, 64–70. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Robertson, S.; Zaragoza, H. The probabilistic relevance framework: BM25 and beyond. Found. Trends Inf. Retr. 2009, 3, 333–389. [Google Scholar]
  28. Faul, F.; Erdfelder, E.; Lang, A.-G.; Buchner, A. G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behav. Res. Methods 2007, 39, 175–191. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Cohen, J. Statistical Power Analysis for the Behavioral Sciences, 2nd ed.; Lawrence Erlbaum Associates: Hillsdale, NJ, USA, 1988. [Google Scholar]
  30. Holm, S. A simple sequentially rejective multiple test procedure. Scand. J. Stat. 1979, 6, 65–70. [Google Scholar]
  31. Warton, D.I.; Hui, F.K.C. The arcsine is asinine: The analysis of proportions in ecology. Ecology 2011, 92, 3–10. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Study design as one integrated pipeline. (A) Between-subjects factorial: 5 simulated learner profiles × 4 tutoring conditions × 30 seeded replications = 600 sessions on GPT-4o-mini, plus a 58-session sensitivity run on Claude Sonnet 4; C1 RAG-Tutor (BM25 over textbook chunks and a misconception knowledge base), C2 KG-RAG-Tutor (C1 plus 2-hop knowledge-graph traversal), C3 Vanilla-LLM (parametric knowledge only), C4 Static-Feedback (regex-matched templates, no LLM). (B) One session: warm-up, five scenario blocks and a wrap-up within 20 turns; the tutor agent and the synthetic learner alternate, and the learner’s misconception state is updated by the revision rule of Equation (1) on every answered tutor turn, independently of the LLM-generated utterance. (C) Measurement: eight session-level outcomes analysed by ANOVA and a binomial GLM with Holm correction, and the four-part audit of the revision measure (two LLM judges, re-scoring ablation, parameter perturbation, second-provider run).
Figure 1. Study design as one integrated pipeline. (A) Between-subjects factorial: 5 simulated learner profiles × 4 tutoring conditions × 30 seeded replications = 600 sessions on GPT-4o-mini, plus a 58-session sensitivity run on Claude Sonnet 4; C1 RAG-Tutor (BM25 over textbook chunks and a misconception knowledge base), C2 KG-RAG-Tutor (C1 plus 2-hop knowledge-graph traversal), C3 Vanilla-LLM (parametric knowledge only), C4 Static-Feedback (regex-matched templates, no LLM). (B) One session: warm-up, five scenario blocks and a wrap-up within 20 turns; the tutor agent and the synthetic learner alternate, and the learner’s misconception state is updated by the revision rule of Equation (1) on every answered tutor turn, independently of the LLM-generated utterance. (C) Measurement: eight session-level outcomes analysed by ANOVA and a binomial GLM with Holm correction, and the four-part audit of the revision measure (two LLM judges, re-scoring ablation, parameter perturbation, second-provider run).
Applsci 16 09938 g001
Figure 2. Condition effect on Misconception Revision Rate with 95% confidence intervals. Among the three LLM-based conditions, MRR is monotonic in retrieval richness; Static-Feedback ( n = 150 ) anchors a non-LLM floor and is not part of the retrieval contrast.
Figure 2. Condition effect on Misconception Revision Rate with 95% confidence intervals. Among the three LLM-based conditions, MRR is monotonic in retrieval richness; Static-Feedback ( n = 150 ) anchors a non-LLM floor and is not part of the retrieval contrast.
Applsci 16 09938 g002
Figure 3. Condition × Profile interaction on MRR (mean ± 95% CI). The peak cell (Intermediate-Strategic × KG-RAG, M = 0.77 ) is statistically equivalent to the corresponding Advanced-Residual cell. At the Naive-Passive end, all conditions converge near zero.
Figure 3. Condition × Profile interaction on MRR (mean ± 95% CI). The peak cell (Intermediate-Strategic × KG-RAG, M = 0.77 ) is statistically equivalent to the corresponding Advanced-Residual cell. At the Naive-Passive end, all conditions converge near zero.
Applsci 16 09938 g003
Figure 4. MRR heatmap on a 0–1.0 scale. Each cell shows the mean MRR for one Condition × Profile cell ( n = 30 ).
Figure 4. MRR heatmap on a 0–1.0 scale. Each cell shows the mean MRR for one Condition × Profile cell ( n = 30 ).
Applsci 16 09938 g004
Figure 5. Per-misconception revision rate by tutoring condition. Impetus shows the largest within-condition retrieval gradient.
Figure 5. Per-misconception revision rate by tutoring condition. Impetus shows the largest within-condition retrieval gradient.
Applsci 16 09938 g005
Figure 6. Two-judge construct-validity check (174 sessions, 506 items). (A) Cohen’s κ with bootstrap 95% CIs for the simulator against each judge, between the judges, and for the simulator against the judges’ consensus. (B) Share of items classified as revised by condition under the simulator’s rule and under each judge, counting indeterminate items as not revised so that all 506 items enter. Both judges separate the non-LLM baseline from the LLM conditions; neither reproduces the KG-RAG > RAG > Vanilla ordering.
Figure 6. Two-judge construct-validity check (174 sessions, 506 items). (A) Cohen’s κ with bootstrap 95% CIs for the simulator against each judge, between the judges, and for the simulator against the judges’ consensus. (B) Share of items classified as revised by condition under the simulator’s rule and under each judge, counting indeterminate items as not revised so that all 506 items enter. Both judges separate the non-LLM baseline from the LLM conditions; neither reproduces the KG-RAG > RAG > Vanilla ordering.
Applsci 16 09938 g006
Figure 7. Inquiry Depth Score (taxonomy-weighted) by learner profile. Naive-Active and Advanced-Residual learners reach the Conceptual–Metacognitive band, but only the latter translate this depth into measurable misconception revision.
Figure 7. Inquiry Depth Score (taxonomy-weighted) by learner profile. Naive-Active and Advanced-Residual learners reach the Conceptual–Metacognitive band, but only the latter translate this depth into measurable misconception revision.
Applsci 16 09938 g007
Figure 8. Sensitivity run on a second provider. (A) GPT-4o-mini primary experiment ( N = 600 ): retrieval gradient among LLM-based conditions under the simulator’s measure, Static floor. (B) Claude Sonnet 4 sensitivity run ( N = 58 ): LLM-vs-Static contrast present, retrieval gradient not reproduced; CIs overlap heavily at this sample size. Panel B is preliminary sensitivity evidence, not a replication.
Figure 8. Sensitivity run on a second provider. (A) GPT-4o-mini primary experiment ( N = 600 ): retrieval gradient among LLM-based conditions under the simulator’s measure, Static floor. (B) Claude Sonnet 4 sensitivity run ( N = 58 ): LLM-vs-Static contrast present, retrieval gradient not reproduced; CIs overlap heavily at this sample size. Panel B is preliminary sensitivity evidence, not a replication.
Applsci 16 09938 g008
Table 1. Synthetic learner profiles.
Table 1. Synthetic learner profiles.
IDLabelPrior KnowledgeInitial MisconceptionsMotivation & Style
P1Naive-PassiveVery Low5 of 5Low; passive/receptive
P2Naive-ActiveVery Low5 of 5High; active/exploratory
P3Intermediate-ConfusedMedium3 of 5Medium; mixed
P4Intermediate-StrategicMedium1 of 5High; strategic/goal-oriented
P5Advanced-ResidualHigh1 of 5 (specific)High; reflective/analytical
Table 2. Targeted Newtonian misconceptions and typing.
Table 2. Targeted Newtonian misconceptions and typing.
IDLabelDescriptionType
MC01ImpetusObjects need continuous force to keep movingSingle-concept
MC02Heavier-FasterHeavier objects fall faster than lighter onesSingle-concept
MC03Motion-Implies-ForceA moving object has a net force in the direction of motionSingle-concept (velocity–force)
MC04Action-Reaction-CancelNewton’s third-law pairs cancel outCross-object relational
MC05Velocity-Force-ConfusionNet force and velocity always point in the same directionSingle-concept
Table 3. Tutoring conditions.
Table 3. Tutoring conditions.
CondLabelInformation AccessLLM
C1RAG-TutorBM25 over 10 textbook chunks + 5 misconception entries; top-5 retrievedYes
C2KG-RAG-TutorC1 + 2-hop BFS over 25-node, 31-edge concept graphYes
C3Vanilla-LLMParametric knowledge onlyYes
C4Static-FeedbackRegex keyword matching to hand-written templatesNo
Table 4. MRR by learner profile ( n = 120 per profile). The top two profiles are statistically indistinguishable.
Table 4. MRR by learner profile ( n = 120 per profile). The top two profiles are statistically indistinguishable.
ProfileMSDCluster
Intermediate-Strategic0.5250.502High-revision
Advanced-Residual0.5170.502High-revision
Intermediate-Confused0.1750.198Low-revision
Naive-Active0.1450.158Low-revision
Naive-Passive0.0180.064Low-revision
Note: For P4 and P5, only one misconception is initially active, so session-level MRR is effectively binary (0 or 1); standard deviations near 0.50 for these profiles reflect this Bernoulli-like distribution rather than measurement anomalies.
Table 5. MRR by tutoring condition ( n = 150 per condition).
Table 5. MRR by tutoring condition ( n = 150 per condition).
ConditionMSD95% CIRole
KG-RAG-Tutor0.4030.418[0.335, 0.470]LLM with KG-augmented RAG
RAG-Tutor0.3290.414[0.262, 0.395]LLM with BM25 RAG
Vanilla-LLM0.2400.387[0.178, 0.302]LLM, no retrieval
Static-Feedback0.1320.303[0.083, 0.181]Non-LLM floor
Table 6. Two-way ANOVA for the eight behavioral outcomes, with omnibus p-values after Holm correction across the eight outcomes ( p H ; step-down, α = 0.05 ). ns: not significant at p < 0.05 after correction.
Table 6. Two-way ANOVA for the eight behavioral outcomes, with omnibus p-values after Holm correction across the eight outcomes ( p H ; step-down, α = 0.05 ). ns: not significant at p < 0.05 after correction.
Condition Profile Interaction
Outcome F(3, 580) η p 2 p H F(4, 580) η p 2 p H F(12, 580) η p 2 p H
MRR21.200.099<0.001 66.540.315<0.001 5.170.097<0.001
Revision Latency22.230.103<0.001 23.610.140<0.001 3.270.063<0.001
Help-Seeking Freq.17.070.081<0.001 188.520.565<0.001 3.430.066<0.001
Inquiry Depth (IDS)0.450.0020.951 ns 349.880.715<0.001 1.170.0250.723 ns
Question-Asking Rate4.180.0210.024 531.760.786<0.001 4.510.085<0.001
Resistance Episodes3.640.0180.038 31.780.180<0.001 0.740.0151.000 ns
Scaffolding Resp.2.000.0250.067 ns 9.630.140<0.001 0.780.0381.000 ns
Final Engagement9.660.048<0.001 6751.530.979<0.001 3.780.073<0.001
Notes: Uncorrected p < 0.001 for all F values except Question-Asking Rate (Condition, p = 0.006), Resistance Episodes (Condition, p = 0.013), Scaffolding Responsiveness (Condition, p = 0.034) and the entries marked ns. Degrees of freedom for outcomes computed on conditional subsamples (IDS, n = 578; Scaffolding Responsiveness, n = 257) are adjusted accordingly; Holm p-values were recomputed from the released session metrics with listwise deletion of undefined values, which changes no conclusion.
Table 7. Binomial GLM (logit link) for MRR with the pair (revised, not revised) per session as the response. Likelihood-ratio tests; dispersion-scaled rows divide the statistic by the Pearson dispersion ( 0.96 ).
Table 7. Binomial GLM (logit link) for MRR with the pair (revised, not revised) per session as the response. Likelihood-ratio tests; dispersion-scaled rows divide the statistic by the Pearson dispersion ( 0.96 ).
EffectLR χ 2 dfp
Condition (given Profile)60.43<0.001
Profile (given Condition)314.74<0.001
Condition × Profile29.5120.003
Condition, dispersion-scaled63.03<0.001
Profile, dispersion-scaled328.14<0.001
Condition × Profile, dispersion-scaled30.7120.002
Table 8. Agreement between the simulator’s keyword rule and two LLM judges on the 174-session sample (determinate items only; indeterminate items excluded pairwise). Positive rate is the share of items classified as revised. Bootstrap 95% CIs for κ (5000 resamples).
Table 8. Agreement between the simulator’s keyword rule and two LLM judges on the 174-session sample (determinate items only; indeterminate items excluded pairwise). Positive rate is the share of items classified as revised. Bootstrap 95% CIs for κ (5000 resamples).
PairnAgreement κ [95% CI]Positive Rate APositive Rate B
Simulator (A) vs. Judge 1, GPT (B)30165.1%0.12 [0.01, 0.23]0.200.33
Simulator (A) vs. Judge 2, Claude (B)26241.6%0.07 [0.00, 0.13]0.260.76
Judge 1, GPT (A) vs. Judge 2, Claude (B)22868.0%0.41 [0.32, 0.50]0.410.73
Simulator (A) vs. consensus of both judges (B)15549.7%0.10 [−0.01, 0.21]0.220.61
Table 9. Re-scoring ablation: expected MRR by condition when all 600 logged dialogues are re-scored under alternative revision criteria (200 replays per session). Gradient: KG-RAG > RAG > Vanilla. LLM > Static: every LLM condition above Static-Feedback. LOO row: ranges over the 17 leave-one-phrase-out vocabularies. Full table with profile means in Table S2.
Table 9. Re-scoring ablation: expected MRR by condition when all 600 logged dialogues are re-scored under alternative revision criteria (200 replays per session). Gradient: KG-RAG > RAG > Vanilla. LLM > Static: every LLM condition above Static-Feedback. LOO row: ranges over the 17 leave-one-phrase-out vocabularies. Full table with profile means in Table S2.
CriterionKG-RAGRAGVanillaStaticGradientLLM > Static
Logged (live simulator)0.4030.3290.2400.132yesyes
Original vocabulary, replayed0.3620.3050.2650.167yesyes
Leave-one-phrase-out (17 variants)0.249–0.3620.204–0.3060.153–0.2680.154–0.16717/1716/17
Exposure only (no richness term)0.2790.2310.2070.092yesyes
IDF-weighted richness0.3450.2870.2510.165yesyes
Strict: ≥ 2 phrases or IDF gate0.110.060.040.16yesno
Independent concept-text vocabulary0.2710.2310.1230.047yesyes
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Yumuşak, S.; Yumuşak, G. Evaluating Retrieval-Augmented LLM Tutoring with Synthetic Learners: A Factorial Simulation of Newtonian Misconception Revision. Appl. Sci. 2026, 16, 9938. https://doi.org/10.3390/app16199938

AMA Style

Yumuşak S, Yumuşak G. Evaluating Retrieval-Augmented LLM Tutoring with Synthetic Learners: A Factorial Simulation of Newtonian Misconception Revision. Applied Sciences. 2026; 16(19):9938. https://doi.org/10.3390/app16199938

Chicago/Turabian Style

Yumuşak, Semih, and Güngör Yumuşak. 2026. "Evaluating Retrieval-Augmented LLM Tutoring with Synthetic Learners: A Factorial Simulation of Newtonian Misconception Revision" Applied Sciences 16, no. 19: 9938. https://doi.org/10.3390/app16199938

APA Style

Yumuşak, S., & Yumuşak, G. (2026). Evaluating Retrieval-Augmented LLM Tutoring with Synthetic Learners: A Factorial Simulation of Newtonian Misconception Revision. Applied Sciences, 16(19), 9938. https://doi.org/10.3390/app16199938

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop