1. Introduction
AI-powered educational applications, from conversational assistants to adaptive practice platforms, are increasingly embedded in the daily study routines of secondary school students. They provide personalized explanations, worked examples, practice recommendations, and writing feedback [
1,
2,
3]. A recent meta-analysis of experimental studies found that ChatGPT-based interventions improved academic outcomes, learner motivation, and higher-order cognitive skills [
2]. In China’s K-12 sector, locally developed AI tools, including Doubao, ERNIE Bot, and iFlytek Spark, are now used in both classroom teaching and after-school independent study, supported by national policies that promote AI literacy and the adoption of intelligent technologies in education [
4,
5].
As AI tools spread through secondary classrooms and self-study, a question that remains under-explored is whether students from low-SES families gain as much from them as their higher-SES peers. The question sits at the heart of sustainable educational equity: whether emerging technologies act as a lever for inclusive learning or as another channel through which existing advantages reproduce themselves. Two perspectives offer opposite predictions. The Matthew effect hypothesis [
6,
7] holds that students from higher-SES families, who already have more digital resources, parental guidance, and supplementary education, are better positioned to extract value from new technologies, so existing achievement gaps widen. Recent evidence from a digital divide perspective points the same way: SES is positively associated with AI literacy, and conventional digital inequalities persist and evolve in the era of generative AI [
8]. The resource substitution hypothesis [
9] and the equalizer perspective on digital technology [
10,
11] predict the opposite: technology can compensate for the resource gaps low-SES students face and narrow achievement gaps, because students with fewer existing supports stand to gain more when high-quality learning resources finally become accessible.
Prior work on this question has been limited in three ways. Prior research on digital divides in education has largely concentrated on whether students can access technology and how often they use it, with less attention to how students actually use digital tools to learn [
11,
12]. Existing AI-in-education research has mostly examined university students; high school students, who operate in a more structured and exam-oriented environment, have received less attention [
1,
13]. And studies linking SES to academic outcomes rarely look at the mechanisms through which technology use might moderate that link, especially in the context of emerging AI tools [
14].
The present study addresses these gaps. We test whether and how family SES moderates the relationship between AI use quality and deep learning approach in Chinese high school students—that is, whether high-quality AI use is more strongly associated with the deep learning approach among students from lower-SES backgrounds. Deep learning approach, defined as a learning approach oriented toward understanding and meaning-seeking strategies [
15], is consistently associated with higher academic achievement and stronger learning transfer [
16,
17]. We chose it as the outcome because it is theoretically sensitive to both resource availability and instructional support, which makes it well suited for capturing differences in learning processes across SES groups. Rather than treating AI use as a binary or unidimensional variable, we measured AI use quality as a three-part construct: seeking AI help, evaluating AI responses, and applying AI output. The measure was adapted from [
18] the Self-Assessment Practice Scale and indexes the cognitive depth of students’ engagement with AI tools.
By bringing digital divide theory [
11] together with the resource substitution hypothesis [
9], we tested the prediction that AI use quality moderates the SES–deep learning approach association. The resource substitution argument is straightforward: a resource is more valuable to people who lack alternatives. Low-SES students typically have less access to private tutoring, parental academic guidance, and other supplementary resources, so AI tools may carry greater marginal value for them. If that logic holds, the positive association between AI use quality and deep learning approach should be stronger for low-SES students than for their higher-SES peers. Confirming this would provide direct evidence that high-quality AI use can partly offset the resource disadvantages associated with lower family SES.
This study offers three contributions. First, it extends the resource substitution hypothesis into research on AI-assisted learning, which gives a reason to expect uneven gains from AI tools across SES groups. Second, it measures AI use quality as a three-part construct (seeking, evaluating, applying) rather than as access or frequency. Third, it tests the equalizer prediction in a population that has received little attention so far: Chinese high school students.
3. Method
3.1. Participants and Procedures
Participants were students from three public high schools in Hangzhou, China. All three schools were classified as provincial-level standardized high schools and were selected to ensure variation in student socioeconomic composition. A stratified approach was used to recruit participants from Grades 10 and 11, targeting a roughly even split between the two grade levels. Grade 12 students were excluded because their curriculum was almost entirely devoted to preparation for the gaokao, China’s high-stakes National College Entrance Examination, which largely determines students’ higher education placement and is a dominant source of academic pressure in secondary schooling. This intensive preparation schedule leaves little room for exploratory use of AI learning tools. All three schools had voluntarily incorporated locally developed AI tools (e.g., Doubao, ERNIE Bot, SparkDesk) into classroom teaching or after-school activities, without any compulsory curriculum mandate. This selection criterion ensured that respondents had sufficient experience with AI tools to answer the survey items meaningfully, though it also limits generalizability to students with prior AI exposure.
We distributed 580 questionnaires. After removing incomplete questionnaires (N = 14), responses showing uniform answer patterns across extended item sequences (N = 11), and responses from students who reported never having used any AI learning tool (N = 7), we retained 548 valid questionnaires, an effective response rate of 94.5%. All retained questionnaires were complete with no missing data. Excluding non-users was necessary because the AI-use quality items require actual experience with AI tools to be meaningfully answered; that exclusion also restricts generalization to students with at least some AI exposure. Because China’s education system follows a fixed age-grade structure, grade level serves as a reliable proxy for age; we therefore recorded grade rather than collecting age separately.
Table 1 summarizes the demographic profile of the final sample.
Participants ranged from 15 to 17 years of age. Data were gathered over a two-week window within a regular academic term. The research team organized students to complete an online questionnaire on the Wenjuanxing platform in school computer rooms. Average completion time was about 10 to 15 min. To reduce common method bias, we put several procedural remedies in place: anonymity and confidentiality were emphasized on the first page, predictor and outcome measures were placed on different pages, and response directions were varied across measures. All valid respondents received stationery items as a token of appreciation. Ethical approval was granted by the researchers’ institutional review board and the administrations of the participating schools. Written informed consent was collected from each student and their parent or guardian.
3.2. Measures
All Likert-type items were rated on a 5-point scale; the FAS-III uses its own scoring system. To ensure cross-cultural appropriateness, all adapted scales went through a standard translation–back-translation procedure followed by expert panel review. The adaptation procedure was systematic. In the first phase, two bilingual researchers with training in educational psychology each produced an independent Chinese translation of all items. The two drafts were then compared, and any inconsistencies were resolved by joint discussion. Next, a third translator with no prior contact with the original instruments rendered the Chinese version back into English. The research team reviewed this back-translation alongside the source items and refined the phrasing where needed to ensure semantic equivalence. Finally, a six-member expert panel, comprising three scholars in psychometrics and educational measurement and three practicing high school teachers experienced in AI-integrated instruction, assessed every item for linguistic clarity, developmental fit for high school students, and alignment with the target construct. Wording was adjusted on the basis of their recommendations. The finalized questionnaire items are provided in the
Supplementary Materials.
To verify that the adapted items retained their intended factor structure in the target population, the questionnaire was pilot-tested with 56 students at a public high school comparable in type and academic level to the main-study schools but not included in the final sample. Confirmatory factor analysis was performed on the pilot data. The three-factor model for AI use quality (seeking, evaluating, and applying) yielded acceptable fit indices (χ2/df = 1.78, CFI = 0.951, TLI = 0.940, RMSEA = 0.058, SRMR = 0.048), and the single-factor model for deep learning approach likewise met conventional thresholds (χ2/df = 1.92, CFI = 0.962, TLI = 0.948, RMSEA = 0.063, SRMR = 0.042).
3.2.1. Family Socioeconomic Status
Family SES was measured with the Family Affluence Scale III [
30,
31]. The FAS-III is a six-item self-report measure designed for adolescents and assesses family material wealth through questions about household possessions and resources. Total scores range from 0 to 13, with higher scores indicating greater family affluence. Full item content and scoring rules are reported in the Supporting Materials. We chose the FAS-III because it sidesteps the methodological problems of asking adolescents to report parental income or education directly, and it has been validated across multiple cultural contexts, including Chinese samples [
30,
31]. The FAS-III captures only the material dimension of SES, but its psychometric properties have been stable across diverse populations. For regression analyses, FAS-III total scores were standardized. The FAS-III uses a binary and ordinal categorical response format different from Likert-type scales, so we did not include it in the confirmatory factor analysis reported below.
3.2.2. AI Use Quality
AI use quality was measured with twelve items adapted from [
18] Self-Assessment Practice Scale (SaPS). The original self-assessment context was modified to reflect AI learning tool use. The scale comprised three dimensions with four items each: seeking AI help, evaluating AI responses, and applying AI output. The original SaPS contains more items per dimension. We selected four items per dimension based on expert panel ratings of content relevance and factor loading strength from pilot testing, prioritizing items most applicable to the AI context while keeping respondent burden low. The full item set is reported in the Supporting Materials. For the main moderation analysis, we used the mean of all twelve items as a composite AI use quality score. Cronbach’s α was 0.832 for seeking, 0.824 for evaluating, and 0.830 for applying; for the overall twelve-item scale, α = 0.903. The CFA fit indices for the three-factor AI use quality model were χ
2/df = 1.92, CFI = 0.976, TLI = 0.970, RMSEA = 0.041, and SRMR = 0.033.
3.2.3. Deep Learning Approach
Deep learning approach was measured with six items adapted from the deep approach subscale of the Revised Two-Factor Study Process Questionnaire [
15]. The original deep approach subscale contains ten items spanning deep motive and deep strategy; the present study selected three motive items and three strategy items based on factor loading strength and content relevance established in prior Chinese adaptations [
32]. The full item set is reported in the Supporting Materials. Cronbach’s alpha was 0.851. The single-factor CFA fit indices were χ
2/df = 3.18, CFI = 0.0986, TLI = 0.972, RMSEA = 0.065, and SRMR = 0.024.
3.2.4. Control Variables
Gender (0 = female, 1 = male) and grade (0 = Grade 10, 1 = Grade 11) were included as control variables. Because all three participating schools had introduced AI learning tools on a voluntary basis and provided students with comparable access through school facilities, AI access was not included as a separate control variable.
3.3. Data Analysis
A priori power analysis was conducted using G*Power 3.1 [
33] to determine the minimum sample size required to detect a small-to-medium moderation effect. For hierarchical linear regression with five predictors, a significance level of 0.05, power of 0.80, and an anticipated effect size of f
2 = 0.03 (corresponding to ΔR
2 ≈ 0.025 for the interaction term, based on typical interaction effect sizes in field studies [
34]; the minimum required sample size was approximately 377. The present sample of 548 exceeded this threshold, providing adequate statistical power.
The analysis proceeded in four stages. First, we computed descriptive statistics, reliability (Cronbach’s α), and bivariate correlations. Confirmatory factor analysis (CFA) on the 18 Likert-type items (12 AI use quality + 6 deep learning approach) was used to examine the measurement model. We computed composite reliability (CR) and average variance extracted (AVE) following [
35], and discriminant validity was examined by comparing the square root of AVE for each construct with its inter-construct correlations. Second, hierarchical regression analysis was used to test the main effects and moderation effect. In Step 1, control variables (gender and grade) were entered. In Step 2, the main predictors (SES and AI use quality) were added. In Step 3, the SES × AI use quality interaction term was added. All continuous predictors were standardized prior to creating the interaction term to reduce multicollinearity [
36]. Variance inflation factor (VIF) values were examined for all models to verify the absence of problematic multicollinearity. Third, simple slope analysis was conducted to probe the significant interaction by estimating the effect of AI use quality on deep learning approach at ±1 SD of SES. Fourth, bias-corrected bootstrap estimation with 5000 resamples was used to construct 95% confidence intervals for the interaction effect. Supplementary dimension-specific moderation analyses were conducted to examine whether the moderation pattern held across the three individual dimensions of AI use quality.
Because all Likert-type variables were collected through self-report at a single time point, common method bias was examined by comparing a single-factor CFA model with the hypothesized measurement model [
37].
4. Results
4.1. Common Method Bias Test
Because all Likert-type variables were collected through self-report at a single time point, we examined common method bias first. We compared a single-factor CFA model in which all 18 Likert-type items (12 AI use quality + 6 deep learning approach) loaded on one latent factor with the hypothesized four-factor model (seeking, evaluating, applying, deep learning approach). The FAS-III items were excluded from this analysis because they use a different response format. The single-factor model showed poor fit (χ
2 = 3412.56, df = 135, χ
2/df = 25.28, CFI = 0.498, TLI = 0.462, RMSEA = 0.136, SRMR = 0.119), whereas the four-factor model fit the data well (χ
2 = 218.52, df = 129, χ
2/df = 1.69, CFI = 0.978, TLI = 0.973, RMSEA = 0.036, SRMR = 0.031). The CFI difference of 0.480 far exceeded the conventional benchmark of 0.10 [
38], suggesting that a single common factor is unlikely to account for the observed covariance structure. Together with the procedural remedies during data collection, these results indicate that common method bias is not a serious threat, although shared-method variance cannot be entirely ruled out.
4.2. Measurement Model
We estimated the four-factor CFA model (seeking, evaluating, applying, deep learning approach) to evaluate the measurement model. The model fit the data well (χ
2 = 218.52, df = 129, χ
2/df = 1.69, CFI = 0.978, TLI = 0.973, RMSEA = 0.036, SRMR = 0.031).
Table 2 reports the standardized factor loadings, squared multiple correlations (SMC), composite reliability (CR), and average variance extracted (AVE) for each construct. All standardized loadings ranged from 0.69 to 0.84, exceeding the recommended threshold of 0.50; CR values ranged from 0.85 to 0.90, exceeding the 0.70 cutoff; and AVE values ranged from 0.59 to 0.61, exceeding the 0.50 cutoff [
35]. As shown in
Table 3 the square root of AVE for each construct exceeded all of its inter-construct correlations, supporting discriminant validity.
4.3. Descriptive Statistics and Correlations
Table 3 presents descriptive statistics and bivariate correlations for all study variables. All continuous variables had acceptable skewness (|skew| < 1) and kurtosis (|kurt| < 1), indicating approximate normality.
Family SES (FAS-III) was significantly and positively correlated with the deep learning approach (r = 0.290, p < 0.01), and AI use quality was also significantly and positively correlated with deep learning approach (r = 0.482, p < 0.01). The three AI use dimensions were highly intercorrelated (r = 0.0599–0.614), supporting their conceptualization as components of an overarching AI use quality construct while also suggesting that the dimensions share substantial common variance.
4.4. Hierarchical Regression Analysis
Table 4 presents the hierarchical regression results predicting deep learning approach. In Step 1, control variables (gender and grade) explained essentially no variance in deep learning approach (R
2 = 0.001, F(2, 545) = 0.292,
p = 0.747), so these demographic factors were not meaningfully associated with deep learning approach in this sample.
In Step 2, adding SES and AI use quality explained a substantial additional proportion of variance (ΔR2 = 0.266, p < 0.001). Both SES (B = 0.085, SE = 0.017, β = 0.188, t = 4.964, p < 0.001, 95% CI [0.051, 0.119]) and AI use quality (B = 0.198, SE = 0.017, β = 0.439, t = 11.589, p < 0.001, 95% CI [0.164, 0.232]) significantly predicted deep learning approach, supporting H1 and H2. All VIF values in Step 2 were below 1.07, indicating no multicollinearity concerns.
In Step 3, adding the interaction between family SES and AI use quality explained significant additional variance (ΔR
2 = 0.031,
p < 0.001). The interaction coefficient was negative and significant (B = −0.083, SE = 0.017, β = −0.177, t = −4.892,
p < 0.001, 95% CI [−0.116, −0.049]), supporting the resource substitution hypothesis (H3): the association between AI use quality and deep learning approach was stronger among low-SES students (H3). The full model (Step 3) explained 29.8% of the variance in deep learning approach (Adjusted R
2 = 0.291). All VIF values in Step 3 were below 1.07. The significance of the interaction was further confirmed by bias-corrected bootstrap estimation (95% CI [−0.118, −0.048]), with the confidence interval clearly excluding zero. The f
2 effect size for the interaction was 0.044 (ΔR
2/(1 − R
2full) = 0.031/.702), which corresponds to a small-to-medium effect by conventional benchmarks (small = 0.02, medium = 0.15) [
34]. The effect is modest in absolute terms, but interaction effects of this magnitude are typical in field-based educational research [
34], and the practical significance of narrowing SES-based learning gaps makes the effect meaningful.
4.5. Simple Slope Analysis
To probe the significant interaction, simple slope analysis was conducted at ±1 SD of SES (
Figure 2). Both slopes were significant. The association between AI use quality and deep learning approach was substantially stronger for low-SES students (B = 0.280, SE = 0.024, t = 11.67,
p < 0.001) than for high-SES students (B = 0.114, SE = 0.024, t = 4.75,
p < 0.001). The slope for low-SES students was about 2.46 times that of their high-SES peers. AI use quality was therefore more strongly associated with deep learning approach among students from less affluent families.
This pattern is consistent with the equalizer hypothesis. When AI use quality was low, the gap in deep learning approach between high and low SES students was relatively large. As AI use quality increased, this gap narrowed because low-SES students showed a steeper association between AI use quality and deep learning approach.
4.6. Supplementary Analyses: Dimension-Specific Moderation
To check whether the moderation effect was consistent across the three dimensions of AI use quality, we estimated separate regression models using each dimension as the moderator.
Table 5 presents the results.
All three interaction terms were statistically significant (p < 0.001), indicating that the equalizer pattern was not limited to a single dimension of AI use quality. The strongest moderation effect was for the applying dimension (β = −0.168, t = −4.407, p < 0.001), followed by seeking (β = −0.152, t = −4.053, p < 0.001) and evaluating (β = −0.143, t = −3.864, p < 0.001). We had initially expected seeking to be the dominant moderator, on the assumption that low-SES students would benefit most from simply having a willing tutor; the actual ranking pushed us toward a different reading. Active application of AI output to one’s own learning and proactive help-seeking appear to be the more important compensatory mechanisms for low-SES students.
6. Conclusions, Limitations, and Implications
6.1. Conclusions
We examined whether family SES moderates the association between AI use quality and deep learning approach among Chinese high school students, with implications for sustainable educational equity. The results supported the resource substitution hypothesis: the association between AI use quality and deep learning approach was stronger for low-SES students, and consequently, when AI use quality was high, the SES-based gap in deep learning approach was narrower. The effect was consistent across all three dimensions of AI use quality and was confirmed by bootstrap confidence intervals. AI learning tools have the potential to serve as compensatory resources for students from economically disadvantaged backgrounds, narrowing the SES-based gap in deep learning approach, provided that students engage with these tools in a deliberate and sophisticated way. In doing so, AI literacy education can become a sustainable lever for educational equity rather than another channel through which existing advantages are reproduced.
6.2. Limitations
This study has several limitations. The cross-sectional design precludes causal inference. Although the theoretical framework and the hypothesized direction of moderation are well grounded, the observed associations could reflect bidirectional influence or the operation of omitted variables; students who already adopt a deep approach to learning may also be more inclined to use AI tools in elaborate ways. Longitudinal or experimental designs are needed to confirm whether changes in AI use quality lead to changes in deep learning approach over time and whether this effect varies across SES groups. Relatedly, some interpretations offered in the Discussion, such as the suggestion that AI tools carry greater marginal value for low-SES students because these students lack alternative resources, go beyond what the cross-sectional data can establish. These interpretations are speculative and should be tested with longitudinal or quasi-experimental designs that can disentangle the proposed mechanisms.
A related concern is that all constructs were measured through self-report. Although procedural remedies were implemented and the common method bias test suggested that a dominant method factor was unlikely, self-report measures remain vulnerable to social desirability and retrospective bias. Future studies could supplement self-report data with objective indicators from AI platforms, such as query logs, session duration, and revision patterns.
The SES measure also has its own boundary. While well validated for adolescent populations, the FAS-III captures only the material dimension of SES and does not directly assess parental education or occupation. A more comprehensive SES measure incorporating cultural and social capital dimensions might reveal additional nuances in the moderation pattern.
Sampling further constrains generalizability. Participants came from three public high schools in Hangzhou that had already adopted AI tools, and students who had never used AI tools were excluded. These design choices ensured ecological validity but limited the populations to which the findings can be extended. Students in rural areas, private schools, or schools with less institutional support for AI integration, as well as students with no AI experience at all, may show different patterns, and replication across more diverse educational contexts is warranted.
Beyond these methodological caveats, the present design also left several analytical extensions for future work. Potential mediating mechanisms, such as academic self-efficacy or learning engagement, were not examined; moderated mediation models could identify the psychological pathways through which AI use quality differentially benefits low-SES students. Because only three schools were included, the data also did not support formal multilevel modeling, so school-level variables that may influence the SES × AI use quality interaction remain unexplored.
6.3. Concluding Remark
AI is now widespread in secondary schools, and whether it widens or narrows educational inequality has become one of the more pressing questions for sustainable educational equity. The Matthew and equalizer hypotheses have offered competing predictions for decades, but evidence on which pattern emerges with generative AI in secondary education has been thin. This study adds evidence from Chinese high school students, an under-studied group, and finds support for the equalizer reading. Among students who engaged with AI at higher cognitive depth, the SES-based gap in deep learning approach was meaningfully smaller, and low-SES students gained more from quality engagement than their higher-SES peers did. The pattern aligns with the spirit of SDG 4: equitable, inclusive, and quality education for all.
The result extends the resource substitution hypothesis into AI-mediated learning. When students engage with AI critically and apply its output deliberately, the tool can substitute for the kind of academic support that more affluent families pay for: tutoring, parental coaching, and enrichment activities. The multidimensional measure of AI use (seeking, evaluating, applying) was central to detecting this moderation. If the same data had been summarized only by access or frequency of use, the equalizer effect could easily have been missed.
The buffering documented here is conditional, not automatic. The relevant divide has shifted from access alone to the cognitive depth of engagement: who proactively seeks AI help, who evaluates AI responses critically, and who applies AI output back into their own learning. For policymakers and schools, this reframes the policy question. Closing the access gap is necessary but not sufficient. What schools also owe their students, especially those whose families cannot supplement classroom instruction, is explicit guidance in using AI as a deliberate cognitive partner rather than a shortcut to answers. Without such guidance, even universal AI access could leave existing inequalities intact or relocate them to a second-level digital divide.
In a system marked by intense academic competition and uneven family resources, AI may turn out to be one of the few educational resources whose marginal value is highest for students with the fewest alternatives. Whether that holds in real classrooms depends on whether schools treat AI literacy as an instructional priority across subjects, instead of leaving it to whatever students happen to absorb at home. For sustainable educational equity to take root in the AI era, this priority cannot remain rhetorical. The stakes are high enough to warrant serious, equity-focused attention.