3.1. Research Design and Analytical Framework
This study employed a cluster-randomized comparative classroom intervention in which intact English for Academic Purposes (EAP) class sections were randomly assigned to one of three instructional conditions: (a) AI-adaptive instruction, (b) teacher-differentiated instruction, and (c) non-differentiated instruction. Cluster randomization was required because students were enrolled in fixed class sections that could not be reorganized. Four intact sections participated (N = 87): two were allocated to the AI-adaptive condition (N = 44), one to the teacher-differentiated condition (N = 21), and one to the non-differentiated condition (N = 22). All groups were taught by the same instructor using scripted lesson plans and condition-specific protocols to minimize instructor effects and cross-condition contamination.
Because only four classroom clusters were available, this study provides exploratory classroom-level evidence rather than definitive causal estimates. Moreover, each teacher-differentiated and non-differentiated condition was represented by a single cluster, leaving instructional condition partially confounded with cluster membership. Consequently, observed differences may reflect both instructional characteristics and unmeasured class-level factors (e.g., classroom climate, peer composition, group dynamics, or motivation). Although the sections were broadly comparable in demographic characteristics and CEFR distributions, baseline equivalence was not established on the continuous reading measure, and residual cluster-specific influences cannot be excluded. To assess the robustness and clustering sensitivity of the observed patterns, mixed-effects modeling, intraclass correlation estimation, design effect adjustment, and sensitivity analyses were conducted. These procedures improve transparency but cannot overcome the limited number and unequal allocation of clusters or separate instructional condition from cluster-specific influences.
Cluster Randomization Procedure
Randomization was conducted at the classroom level because institutional scheduling precluded assigning students individually. Before baseline assessment, a research assistant not involved in instruction or outcome assessment assigned anonymous identification codes to the four intact EAP sections. A computer-generated random-number sequence then allocated the sections to the three instructional conditions, resulting in two AI-adaptive classes and one class each in the teacher-differentiated and non-differentiated conditions. Given the availability of only four clusters, stratified or constrained randomization was not feasible.
Allocation occurred before the intervention and prior to baseline instructional activities. Because of the nature of the intervention, allocation concealment was not possible; the instructor needed to know group assignments to prepare condition-specific materials. Nevertheless, baseline assessments were administered under standardized conditions, and scoring, psychometric calibration, and statistical analyses followed prespecified protocols to minimize investigator bias.
3.4. Instructional Materials and Content Control
All instructional conditions addressed the same reading topics, learning objectives, and underlying propositional content to ensure curricular comparability. Materials were drawn from a common bank of CEFR-aligned passages based on the institutional curriculum. While content remained constant, its linguistic realization varied by instructional architecture. In the AI-adaptive condition, lexical support, syntactic complexity, scaffolding, and question difficulty were adjusted dynamically. Comparable adaptations were implemented through teacher-designed tiered materials in the differentiated condition, whereas the non-differentiated condition used a single unmodified version of each passage.
To preserve internal validity, the following elements were held constant across conditions:
Reading topics and thematic content.
Total instructional time and session length.
Assessment tasks and evaluation criteria.
The principal source of variation was the instructional architecture, including adaptation source, responsiveness, continuity, feedback density, and personalization granularity (
Section 3.5.1). This design minimized content-related confounding and improved curricular comparability across conditions; however, observed differences cannot be attributed uniquely to instructional condition because classroom-level influences remain inseparable from condition in the four-cluster design.
3.5. Instructional Conditions
3.5.1. AI-Adaptive Personalized Reading Instruction (G1)
The AI-adaptive condition preserved the propositional meaning and learning objectives of each reading passage while dynamically modifying selected linguistic and instructional features according to learner performance. Rather than assigning learners to fixed proficiency groups, a researcher-configured adaptive platform continuously calibrated text complexity, lexical support, syntactic complexity, scaffolding, comprehension-question difficulty, and feedback specificity to align instructional demands with learners’ evolving reading ability.
The platform employed a transparent, deterministic, rule-based adaptation engine rather than machine learning or generative AI. Although the platform employed deterministic rule-based adaptation rather than machine learning, it is referred to as an AI-adaptive system because it continuously analyzed learner performance and autonomously selected instructional responses without instructor intervention, consistent with the operational definition of AI-adaptive instruction adopted in this study. It continuously monitored response accuracy, response latency, and error patterns, then applied predefined CEFR-aligned decision rules informed by Cognitive Load Theory. For example, sustained high performance (≥80% accuracy) triggered progression to more inferential tasks with reduced scaffolding, whereas repeated errors or prolonged response times prompted simplified syntax, additional lexical support, and more explicit guidance. Adaptation was therefore performance-responsive rather than predictive, ensuring transparency, instructional consistency, and reproducibility. Reading passages were selected from a researcher-developed, pre-reviewed CEFR-aligned text bank. The system did not generate or rewrite texts; instead, it selected appropriate passages and adjusted scaffolding, lexical support, question difficulty, and instructional prompts. All materials were reviewed before implementation to ensure semantic equivalence, curricular alignment, and consistency of learning objectives.
In addition, the platform recorded learner identification codes, response accuracy, response latency, adaptation levels, completed activities, and instructional progression. No personally identifiable information was transmitted to external providers.
Appendix A provides technical specifications of the adaptive architecture, decision rules, and logging procedures. This comparison therefore examined two adaptive instructional architectures rather than AI versus non-AI instruction. The AI-adaptive condition implemented continuous algorithm-mediated personalization, whereas teacher differentiation relied on intermittent instructor-mediated adaptation constrained by classroom pacing and teacher attention. Because these architectures differed simultaneously in adaptation source, responsiveness, continuity, feedback density, and personalization granularity, the study was not designed to isolate the independent effect of AI. Accordingly, the findings should be interpreted as reflecting the combined characteristics of the instructional systems rather than AI alone.
3.5.2. Teacher-Designed Differentiated Reading Instruction (G2)
In the teacher-differentiated condition, instruction was delivered through pedagogically grounded, teacher-created tiered tasks aligned with learners’ proficiency levels. All learners worked with the same reading texts to ensure content equivalence; however, task demands and the depth of scaffolding varied by proficiency level. Differentiation strategies included the following:
- (a)
Tiered comprehension questions (literal vs. inferential focus);
- (b)
Adjusted scaffolding prompts (explicit guidance vs. strategic hints);
- (c)
Variable depth of inferential and critical reading tasks.
This condition reflects established practices of differentiation in EFL pedagogy, drawing on teachers’ professional judgment, contextual awareness, and instructional experience.
3.5.3. Traditional Non-Differentiated Reading Instruction (G3)
The control group followed the standard institutional reading curriculum without adaptive or differentiated mechanisms. All learners received identical texts, tasks, and instructional support, regardless of proficiency level. Instruction consisted of uniform pre-reading explanation, guided reading, and whole-class review of comprehension tasks. (see
Table 3). No systematic adjustments were made to task difficulty, scaffolding intensity, or feedback specificity.
3.6. Study Procedure and Treatment Implementation
The instructional intervention was implemented over twelve weeks following a standardized, phase-based protocol designed to ensure treatment fidelity, comparability across groups, and replicability.
Phase 1: Orientation and Pre-Intervention Assessment (Week 1)
Phase 1 spanned two 90-min sessions and focused on orientation, ethical briefing, and baseline assessment. During the orientation session, participants were informed of the study’s objectives, procedures, and assessment schedule. Ethical assurances, including voluntary participation, confidentiality, and non-impact on course grades, were clearly communicated. Learners also received a brief orientation to the instructional format specific to their assigned condition, without disclosure of comparative hypotheses to prevent expectancy effects.
In the pre-test session, all participants completed a CEFR-aligned reading comprehension test under standardized examination conditions. The assessment measured literal comprehension, inferential understanding, and the construction of interpretive meaning. Pre-test scores were used for proficiency stratification, baseline assessment, and the longitudinal pre-to-post analyses; they were not used for evaluative grading.
Phase 2: Instructional Treatment (Weeks 2–11)
The instructional treatment phase consisted of three 60-min sessions per week, totaling 3 h per group. All groups covered the same reading topics and texts during this phase. To control for extraneous instructional variables, instructional time was identical across groups, the same instructor taught all conditions, and assessment expectations were held constant. The study compared two adaptive instructional architectures and one non-adaptive instructional architecture rather than isolating a single AI component (see
Table 4).
Open-ended reflections were obtained from all 87 participants immediately after completing the instructional intervention. Responses were collected across all instructional conditions (AI-adaptive: N = 44; teacher-differentiated: N = 21; non-differentiated: N = 22). Participants completed the reflections anonymously using identification codes assigned for research purposes. These identifiers are used when presenting illustrative quotations to preserve participant confidentiality while allowing readers to identify the instructional condition from which each quotation originated.
Phase 3: Post-Intervention Assessment (Week 12)
Participants completed a learner perception questionnaire assessing cognitive load, engagement, instructional comfort, and perceived ownership of reading achievement, followed by open-ended written reflections on their learning experiences, perceived benefits, challenges, and sense of agency. Participants completed the reflections immediately after the intervention under classroom conditions and wrote them in English as part of their regular English for Academic Purposes coursework; therefore, no translation was required. Qualitative data were analyzed using
Braun and Clarke’s (
2006) six-phase inductive thematic analysis. Analysis involved repeated familiarization with the data, inductive coding, theme development and refinement, and selection of representative quotations. No predetermined codebook was used. Additionally, both authors independently familiarized themselves with the data and developed preliminary inductive codes. The authors then discussed the resulting interpretations to refine theme definitions and ensure consistent application of the analytical framework. Given the exploratory purpose of the qualitative component, the analysis aimed to identify recurring patterns of meaning rather than quantify coding agreement or estimate inter-coder reliability.
3.6.1. Structure of Each Instructional Session (60 Minutes)
All instructional sessions followed a standardized four-stage structure to control instructional time and time-on-task effects across conditions. Each 60 min session consisted of (a) pre-reading activation (10 min), involving schema activation and topical orientation; (b) guided reading and task engagement (30 min), during which learners completed comprehension-focused reading activities; (c) comprehension consolidation (10 min), focused on reinforcing understanding and clarifying ambiguities; and (d) reflection and feedback (10 min), during which learners reviewed performance and learning strategies. All temporal allocations, instructional stages, and task objectives were held constant across conditions (see
Figure 3).
3.6.2. AI-Adaptive Personalized Reading Treatment (G1)
The AI-adaptive condition operationalized personalization as a continuous performance-responsive process. During pre-reading activation, learners received system-generated synthesis tasks and topic prompts calibrated to their proficiency levels. In the guided reading stage, the AI system dynamically adjusted instructional input in real time based on learner performance, including text complexity, lexical support, task difficulty, and feedback specificity. Comprehension tasks progressed from literal to inferential and evaluative levels according to demonstrated performance. During consolidation, learners completed AI-generated synthesis tasks and received immediate automated feedback targeting comprehension accuracy and strategic reading behaviors. In the reflection stage, learners responded to AI-delivered prompts addressing perceived difficulty, instructional comfort, and task clarity.
3.6.3. Teacher-Designed Differentiated Reading Treatment (G2)
The teacher-differentiated condition followed the same session structure while implementing differentiation through instructor-mediated task tiering and scaffolding. Pre-reading activities included proficiency-aligned vocabulary instruction and guided background activation. During guided reading, all learners engaged with the same core texts, but task demands varied by proficiency level: A2 learners completed more literal comprehension tasks, whereas B1 learners engaged in inferential and evaluative activities. The instructor provided contingent scaffolding and individualized support throughout the session. Consolidation activities involved proficiency-appropriate summarization and response tasks, followed by oral and written feedback targeting comprehension and strategy use. Reflection activities focused on task difficulty and reading strategies.
3.6.4. Traditional Non-Differentiated Reading Treatment (G3)
The control condition followed the same session structure but implemented uniform instruction without adaptive or differentiated mechanisms. All learners received identical explanations, reading tasks, scaffolding, and feedback regardless of proficiency level. Sessions consisted of whole-class instruction, identical comprehension activities, and general feedback led by the instructor. Reflection activities were limited to brief content-focused discussions without structured evaluation of learning strategies or instructional experience.
3.6.5. Treatment Fidelity and Contamination Control
Treatment fidelity was monitored throughout the 12-week intervention using calibration sessions, structured classroom observations, and a standardized treatment fidelity checklist with condition-specific sections completed independently by two trained observers. The complete fidelity instrument and scoring procedures are provided in the
Supplementary Materials. Inter-rater reliability was high (κ = 0.82), and adherence exceeded 90% across all instructional conditions. Contamination was minimized through separate LMS environments, restricted platform access, and AI activity-log verification. Post-test screening indicated minimal cross-condition exposure (<5%).
Figure 4 summarizes the fidelity and contamination control procedures.
3.7. Reading Comprehension Assessment, Equating, and CEFR Classification
Reading comprehension was assessed using two researcher-constructed parallel-form academic reading tests modeled on the IELTS Academic Reading specification. Because secure operational IELTS materials cannot be reproduced for research dissemination, both forms were developed using publicly available IELTS descriptors, item formats, and construct specifications. Source texts were selected from open-access academic and journalistic materials appropriate for tertiary EFL learners. To enhance content validity, all passages and items were independently reviewed by two experienced EAP instructors and one applied linguistics researcher with IELTS examiner training experience. Reviewers evaluated alignment with the IELTS Academic Reading descriptors, the CEFR reading scales, and the curricular objectives.
Psychometric comparability between forms was established through KR-20 reliability estimation, Rasch modeling, and common-item equating using 12 anchor items. Item functioning, score distributions, and cognitive-domain coverage demonstrated strong comparability across forms, supporting valid measurement of instructional gains.
Table 5 summarizes the psychometric and structural characteristics of the two parallel forms.
A detailed assessment blueprint and a parallel-form comparability framework are provided in
Appendix B. Because adaptive instructional systems are highly sensitive to measurement quality, the present study treated assessment validity as a foundational design requirement rather than a procedural necessity. Parallel-form equating, Rasch calibration, and CEFR-linked interpretation were implemented to ensure that observed developmental patterns reflected meaningful differences in reading proficiency rather than measurement artifacts. Consequently, this study contributes to ongoing discussions on interpreting assessments in AI-mediated learning environments.
3.7.1. Scoring Procedures
Each test consisted of 40 dichotomously scored items (1 = correct, 0 = incorrect), yielding raw scores ranging from 0 to 40. Raw scores were converted to Rasch-scaled ability estimates to permit interval-level analysis and to support CEFR classification. Two trained raters independently verified the accuracy of the scoring for 20% of randomly selected scripts. Inter-rater agreement exceeded 99%, and all discrepancies were resolved through discussion.
3.7.2. CEFR Mapping and Proficiency Classification
CEFR categories were used as an interpretive framework rather than as independently validated proficiency certifications. Following Rasch calibration and common-item equating, learner ability estimates were placed on a common logit scale. Published IELTS–CEFR correspondence documents (e.g., the British Council, Cambridge Assessment English, and the Council of Europe) informed the development of locally defined Rasch cut-score ranges approximating CEFR proficiency within the present assessment system. These thresholds were used solely for descriptive interpretation and should not be interpreted as equivalent to official CEFR certification or externally validated IELTS scores. Preliminary local validation compared Rasch-based classifications with institutional placement records for a random subsample (
N = 20), yielding substantial agreement (Cohen’s κ = 0.81). Because institutional placement targets the university’s A2–B1 instructional range, this comparison provides only preliminary support for the local classification framework and does not constitute external validation across the full CEFR continuum (see
Table 6).
3.7.3. Interpretive Cautions Regarding Vertical Progression
Although the Rasch-scaled score distribution extended beyond the instructional target range, interpret the resulting B2 and C1 classifications as locally derived Rasch-based categories rather than externally validated CEFR proficiency. Because the assessments were designed primarily to measure development within the institutional A2–B1 range, classifications above this range represent descriptive indicators of relative performance on the local Rasch scale. As no formal CEFR standard-setting study or external CEFR-calibrated assessment was conducted, broader proficiency claims require external validation.
3.7.4. Learner Experience Measures and Psychometric Validation
Learner experience was assessed using an 18-item questionnaire comprising four theoretically defined constructs: Cognitive Load (5 items), Engagement (5 items), Instructional Comfort (4 items), and Perceived Ownership of Reading Achievement (4 items). Responses were recorded on a five-point Likert scale ranging from 1 (Strongly Disagree) to 5 (Strongly Agree). The questionnaire was administered once immediately after the intervention; consequently, it supports comparisons of post-intervention learner perceptions across instructional conditions but not within-participant change over time. The questionnaire was adapted from established educational psychology and computer-assisted language learning measures and reviewed by three experts for content relevance, clarity, and contextual appropriateness for tertiary EFL learners. A test–retest assessment demonstrated satisfactory temporal stability (Spearman’s r = 0.878, p = 0.001). Internal consistency was acceptable to excellent across all four constructs (Cronbach’s α = 0.81–0.90). The instrument’s internal structure was evaluated using a four-factor confirmatory factor analysis (CFA), with each item loading on its prespecified construct. Because responses were measured on five-point ordinal Likert scales, the model was estimated using the weighted least squares mean- and variance-adjusted (WLSMV) estimator in the lavaan package (v0.6-18) in R 4.4.1. Missing data were minimal (<2%), and all participants (N = 87) were retained.
The final four-factor model demonstrated acceptable-to-good fit: χ
2(129) = 187.42,
p < 0.001, χ
2/df = 1.45, CFI = 0.963, TLI = 0.955, RMSEA = 0.071, 90% CI [0.048, 0.092], and SRMR = 0.061. Standardized factor loadings ranged from 0.68 to 0.85 (all
p < 0.001), providing preliminary support for the intended measurement structure, pending independent replication (see
Table 7). Latent factor correlations were moderate and theoretically consistent, with Cognitive Load negatively associated with Engagement (
r = −0.42), Instructional Comfort (
r = −0.55), and Perceived Ownership of Reading Achievement (r = −0.31), whereas Engagement was positively associated with Instructional Comfort (r = 0.61) and Perceived Ownership of Reading Achievement (r = 0.58); Instructional Comfort was also positively correlated with Perceived Ownership of Reading Achievement (r = 0.49), supporting satisfactory discriminant validity.
Inspection of modification indices identified one theoretically justified correlated residual between items CL2 and CL4, both reflecting perceived time pressure during reading. Allowing for this residual covariance significantly improved model fit (Δχ2 = 12.8, Δdf = 1, p < 0.001) without altering the factor structure. No additional modifications were introduced because the remaining indices lacked theoretical justification.
A four-factor confirmatory factor analysis (CFA), conducted on the same sample of 87 participants included in the principal analyses with a modest case-to-parameter ratio, provided preliminary support for the hypothesized measurement structure. The model demonstrated acceptable-to-good fit, χ
2(129) = 187.42,
p < 0.001, CFI = 0.963, TLI = 0.955, RMSEA = 0.071 (90% CI [0.048, 0.092]), and SRMR = 0.061. Standardized factor loadings ranged from 0.68 to 0.85, and all loadings were statistically significant (
p < 0.001), providing preliminary evidence of structural validity that requires independent replication in a larger, independent sample. Detailed CFA results are provided in the
Supplementary Materials.
To our knowledge, no previously validated reading-specific measure of AI-mediated perceived ownership of reading achievement was available for tertiary EFL reading contexts. Accordingly, the four-item Perceived Ownership of Reading Achievement subscale was developed through theory-informed adaptation of constructs related to learner agency, ownership, feedback literacy, and academic identity (e.g.,
Norton, 2013;
Carless & Boud, 2018). Item wording was adapted to the context of reading comprehension rather than writing while retaining the theoretical focus on learners’ perceived ownership of their reading processes. Three specialists in applied linguistics and educational measurement independently evaluated the items for conceptual relevance, clarity, and construct alignment. Minor wording revisions were made following expert feedback before pilot testing.
Although conceptually related, perceived ownership of reading achievement differs from several adjacent constructs. Learner autonomy concerns independent regulation of learning decisions; agency refers more broadly to perceived capacity for intentional action; ownership emphasizes psychological investment in learning outcomes; and self-efficacy reflects confidence in one’s capability to perform successfully. By contrast, perceived ownership of reading achievement refers specifically to learners’ perception that successful comprehension remains fundamentally their own intellectual accomplishment despite adaptive instructional support. The construct therefore integrates elements of ownership, agency, and attribution, but is operationalized narrowly as learners’ perceived ownership of their own reading achievement rather than as a broader claim about authorship or identity.
3.8. Statistical Analysis Plan
For RQ1, the primary analysis is a descriptive, exploratory comparison of the recoverable pre-to-post trajectories: model-implied pre-test and post-test values for each of the three reported instructional groups (with the two AI-adaptive classrooms analyzed jointly, not as two individually demonstrated trajectories), and model-implied pre-to-post change is reported directly, without treating the classrooms as interchangeable replicates of their assigned instructional condition. Classroom section identifiers (G1a, G1b, G2, G3) were retained throughout the analytic dataset and were used as the clustering factor in the mixed-effects analysis; simple, unmodeled pre-test and post-test descriptive statistics were also computed separately for G1a and G1b. However, instructional condition was entered into the model as a fixed effect shared by both AI-adaptive classrooms, so the reported AI-adaptive coefficients and corresponding model-implied trajectory represent G1a and G1b jointly rather than as two separately estimated classroom effects. Because only two clusters represent the AI-adaptive condition, classroom-specific model-based estimates were not treated as stable inferential estimates or as independent replication of an instructional effect. Accordingly, the primary presentation for RQ1 combines raw classroom-specific descriptives with the pooled model-implied pre-to-post pattern. This presentation does not constitute independent classroom-level replication and cannot provide separable, generalizable estimates of instructional condition effects.
A participant-level 2 × 3 mixed-design ANOVA, with Time (pre-test, post-test) as the within-subject factor and Instructional Condition (AI-adaptive, teacher-differentiated, and non-differentiated) as the between-subject factor, is additionally reported as a secondary, supplementary analysis. Because this analysis treats the 87 learners as independent units despite instructional condition having been assigned at the classroom level, and because it does not model classroom clustering, its Time × Instructional Condition F test, p-value, and partial η2 should not be treated as primary evidence of an instructional condition effect. The study reports them instead for transparency and comparability with prior participant-level conventions. The exploratory cluster-adjusted mixed-effects model described below serves the same secondary, robustness-checking role and, with only four clusters, cannot overcome the unit-of-analysis limitation.
To assess the robustness of these patterns, an exploratory cluster-adjusted linear mixed-effects model was fitted with repeated observations nested within learners and learners nested within classroom sections. Because only four classroom clusters were available, the model cannot separate instructional effects from cluster membership or provide stable treatment effect estimates. Accordingly, coefficients, confidence intervals, and p-values are reported for transparency as sample-specific indicators of direction and uncertainty, not as confirmatory evidence of treatment effectiveness. Statistical significance for the mixed-design ANOVA was evaluated at α = 0.05 (two-tailed), with multiplicity adjustments applied to pairwise comparisons where appropriate. Model assumptions were examined using standardized residuals and Q–Q plots (or Shapiro–Wilk tests) for residual normality and Levene’s test for homogeneity of variance. Because the within-subject factor contained only two levels, sphericity was satisfied automatically, making Mauchly’s test and Greenhouse–Geisser/Huynh–Feldt corrections unnecessary. Effect sizes are reported as partial eta squared (η2p), with 95% confidence intervals where available.
Preliminary analyses included descriptive statistics, baseline-equivalence testing, outcome-specific intraclass correlation coefficients (ICCs), and calculation of design effects and nominal effective sample sizes. Design effects were calculated using both the conventional formula, DE = 1 + (
− 1) ρ, and the unequal-cluster correction, DE = 1 + {[(1 + CV
2)
] − 1} ρ. Because classroom sizes were nearly identical (22, 22, 21, and 22 learners;
= 21.75; CV ≈ 0.02), both methods produced virtually identical estimates. Effective sample sizes (N/DE) are presented descriptively to illustrate the potential reduction in independent information associated with clustering and should not be interpreted as overcoming the inferential limitations imposed by the four-cluster design. Baseline reading ICCs were estimated from pre-intervention scores, whereas ICCs for cognitive load, engagement, instructional comfort, and perceived ownership of reading achievement were estimated from post-intervention questionnaire data because these outcomes were measured only after the intervention, as shown in
Table 8.
Appendix C provides full calculation steps for each design effect and effective sample size, including verification using the unequal-cluster-size correction formula.
For RQ2 and the quantitative component of RQ3, learner experience variables were available only at post-test. Because participant-level one-way ANOVAs assume independence and this study included only four classroom clusters (with single clusters representing the teacher-differentiated and non-differentiated conditions), inferential comparisons were not considered sufficiently robust. These outcomes are therefore reported descriptively by instructional condition using means and standard deviations, and are interpreted as classroom-specific patterns rather than causal treatment effects.
The non-differentiated classroom showed an estimated pre-to-post increase of b = 0.89 (p < 0.001). Relative to this trajectory, the AI-adaptive classrooms showed an additional estimated increase of b = 0.44 (p = 0.002), whereas the teacher-differentiated classroom showed an additional increase of b = 0.29 (p = 0.057). The AI classrooms also showed a higher observed baseline score than the control classroom (b = 0.27, p = 0.027). Because only four classroom clusters were available, these coefficients describe model-implied baseline contrasts and pre-to-post trajectories and should not be interpreted as separable or generalizable estimates of instructional effectiveness.
Table 9 presents the exploratory cluster-adjusted mixed-effects robustness analysis. Pre-test and the non-differentiated condition served as the reference categories; consequently, the condition main effects represent estimated baseline contrasts, and the Time × Condition terms represent differential pre-to-post trajectories relative to the non-differentiated class. The analysis followed the mixed-design ANOVA to examine whether the observed within-sample trajectory pattern remained evident after accounting for learner and classroom membership. Given the four-cluster design, the coefficients, confidence intervals, and
p-values are sample-specific indicators of direction and uncertainty and should not be interpreted as separable or generalizable estimates of instructional condition effects.
The class assigned to the non-differentiated condition showed an estimated pre-to-post increase of b = 0.89 (SE = 0.09,
p < 0.001). Relative to this trajectory, the two classes assigned to the AI-adaptive condition showed an additional estimated increase of b = 0.44 (SE = 0.14,
p = 0.002), whereas the teacher-differentiated class showed an additional estimated increase of b = 0.29 (SE = 0.15,
p = 0.057). The AI-versus-control baseline coefficient indicated that the AI-adaptive classes began with a higher estimated reading score than the control class, b = 0.27 (SE = 0.12,
p = 0.027). These coefficients describe sample-specific contrasts among the model-implied pre-to-post trajectories, distinguishing three instructional groups (with the two AI-adaptive classrooms analyzed jointly) drawn from the four participating clusters, and should not be interpreted as separable or generalizable estimates of instructional condition effects, because instructional condition was partially confounded with classroom membership and the estimates were derived from only four clusters. Sensitivity analyses testing this Time × AI coefficient across alternative model specifications, along with pairwise effect sizes for reading gains, are reported in
Appendix D.