1. Introduction
Formative assessment (FA) has been increasingly recognized as a central component of effective teaching and learning, particularly when conceptualized as a process enacted with students rather than merely for them (
Andrade et al., 2021). In contrast to summative assessment, which primarily serves to certify learning outcomes, FA aims to generate information to adapt instruction and support students’ learning processes in real time (
Black & Wiliam, 2009). Within this perspective, alignment between teachers’ and students’ perceptions of assessment practices has emerged as a critical condition for assessment quality and instructional effectiveness, as it shapes how assessment information is interpreted and used for learning (
Balbi et al., 2025;
Veugen et al., 2024).
Formative assessment has been widely recognized as a powerful approach for supporting student learning and improving classroom instruction (
Black & Wiliam, 1998;
Wiliam, 2011).
Beyond its instructional function, FA has been consistently linked to the development of students’ self-regulated learning (SRL). SRL refers to learners’ capacity to plan, monitor, and adapt their cognitive, motivational, and behavioral processes in order to achieve academic goals (
Brandmo et al., 2020;
Zimmerman, 2002). Prominent theoretical models conceptualize SRL as a multidimensional process involving cognitive and metacognitive strategies, motivational beliefs, and behavioral engagement (
Wolters et al., 2005). A growing body of research highlights the close conceptual and empirical relationship between FA and SRL, suggesting that formative practices create conditions that enable students to regulate their learning more effectively through clearer goal orientation, the proactive use of feedback, and structured opportunities to evaluate and adjust their study strategies (
Panadero et al., 2018).
This relationship is particularly salient in mathematics education, a domain characterized by cumulative knowledge structures, high cognitive demands, and persistent student difficulties. Successful learning in mathematics requires sustained engagement in problem-solving, strategic planning, monitoring solution processes, and persistence in the face of errors. Empirical studies indicate that FA practices can support these processes by promoting meaningful feedback, instructional adaptation, and student engagement (
Rakoczy et al., 2019). In parallel, SRL has been consistently associated with improved mathematical learning and achievement, underscoring its importance in this subject area (
Dignath & Büttner, 2008;
Özcan, 2016).
In the Portuguese educational context, data from PISA 2022 (
Duarte et al., 2023) indicate a decline in students’ average performance in mathematics. Moreover, over the past decade, mathematics has consistently had the highest failure rate among students in the second and third cycles of basic education (ages 10–15) (
DGEEC, 2023a,
2023b). From a motivational perspective, the TIMSS 2023 (
Duarte et al., 2024) report further shows that a substantial proportion of Portuguese eighth-grade students (ages 13–14) report disliking learning mathematics. These indicators persist despite educational policy initiatives aimed at improving pedagogical practices and promoting formative assessment, most notably the MAIA Project (Monitoring, Support, and Research in Pedagogical Assessment) (
Fernandes et al., 2021). Implemented nationwide by the Directorate-General for Education in 2019, the MAIA Project sought to integrate the formative function of assessment more coherently within the curriculum. However, the mismatch between efforts to improve teaching practices and students’ learning outcomes suggests that changes in teachers’ practices alone are insufficient, highlighting the need for students’ active involvement in the teaching and learning process.
Recent research suggests that the effectiveness of FA does not depend solely on the presence or frequency of formative practices, but critically on how students perceive and interpret these practices. Several studies have documented systematic discrepancies between teachers’ self-reports of FA practices and students’ perceptions, with teachers often perceiving their practices as more formative than students do (
Pat-El et al., 2015;
van der Kleij, 2019). Such perceptual incongruence may weaken the regulatory function of assessment, limiting students’ ability to use feedback to guide their learning and self-regulation (
J. Hattie & Timperley, 2007;
Sadler, 1989). These findings are consistent with contemporary views of FA as a relational, dialogic, and interpretive practice, whose impact depends on shared teacher–student sense-making rather than on technical implementation alone (
Carless & Winstone, 2023).
Despite these advances, few studies have jointly examined formative assessment, multiple dimensions of self-regulated learning, and teacher–student perceptual congruence within the specific context of lower secondary mathematics. Existing research has often focused on either students’ perceptions or teachers’ practices in isolation, or on FA-SRL relations without explicitly considering the alignment between teachers’ intentions and students’ interpretations. Addressing this gap, the present study integrates students’ and teachers’ perspectives to examine how perceived formative assessment practices relate to different dimensions of SRL in mathematics, and whether congruence between teachers’ and students’ perceptions is associated with students’ regulatory functioning.
This body of research suggests that formative assessment may support self-regulated learning not only through instructional design, but through students’ perceptions and shared sense-making with teachers, an assumption tested in the present study.
Conceptually, the present study advances the literature by integrating three strands of research that have often been examined separately: (a) the multidimensional structure of self-regulated learning, (b) students’ and teachers’ parallel perceptions of formative assessment practices, and (c) multilevel modeling of teacher–student perceptual congruence in a mathematics-specific context. By simultaneously examining global formative assessment effects, dimensional patterns, and congruence indices at the classroom level, this study moves beyond simple association models and offers a more relational, system-oriented account of how formative assessment operates in authentic instructional settings. In doing so, it contributes empirical evidence to ongoing debates about the role of perceptual alignment in activating the regulatory potential of formative assessment.
3. Materials and Methods
3.1. Participants
Participants were 305 students enrolled in lower secondary education (Grades 5–9). This level was selected because it encompasses the transition to upper secondary education, during which students are expected to demonstrate more advanced self-regulated learning competencies (
Meusen-Beekman et al., 2016), and because few studies cover this level of schooling.
The sample included 161 girls (52.8%) and 144 boys (47.2%). Students were distributed across Grades 5 to 9, with 7.5% in Grade 5 (n = 23), 40.1% in Grade 6 (n = 123), 26.6% in Grade 7 (n = 81), 14.1% in Grade 8 (n = 43), and 11.5% in Grade 9 (n = 35). It is important to note that the higher concentration of Grade 6 students was not intentional but resulted from the availability of classes in the schools where data collection took place. This asymmetry will be duly considered in the study’s discussion and limitations.
Regarding grade retention, 91.1% of students (n = 278) reported no history of grade repetition, while 8.9% (n = 27) indicated having repeated at least one school year.
Data were collected in two public schools in the Greater Lisbon region (NUTS III sub-region), in a predominantly low- to middle-socioeconomic context. The questionnaires were administered during regular school hours. Depending on the school organization, in one school, data collection took place in the library without teachers present; in the other, it occurred in the classroom, with teachers present but remaining at their desks and not intervening in the completion of the questionnaires.
The teacher sample consisted of 39 mathematics teachers. Regarding sex, most participants were female (n = 29), while ten were male (n = 10). Teachers’ ages ranged from 27 to 62 years, with a mean age of 48.21 years (SD = 7.55), indicating a predominantly mid- to late-career teaching workforce. Regarding teaching experience, participants reported 3 to 36 years of professional experience. On average, teachers had 22.41 years of experience (SD = 7.67), reflecting a sample composed mainly of experienced professionals. The teacher sample was evenly distributed across age and experience groups and was predominantly female and experienced.
3.2. Instruments
Detailed psychometric analyses of all instruments used in this study, including exploratory and confirmatory factor analyses and reliability indices, are reported in the
Supplementary Materials.
The instruments were adapted from previous research, with specific adjustments, including adding new items to the formative assessment scale and reorganizing items in the self-regulated learning scales to optimize their psychometric properties. In this section, we present the final versions of the scales used in the data analysis, along with psychometric evidence supporting the reliability of the findings. A comprehensive description of the original scales as administered to participants, including the full item pool and the step-by-step validation process, is provided in the
Supplementary Material.
Instruments that had not been previously adapted or validated were first subjected to exploratory factor analyses (EFA) to examine their latent structure, followed by confirmatory factor analyses (CFA) to test the adequacy of the proposed measurement models (
Brown, 2015;
Kline, 2016). Model fit was evaluated using commonly reported fit indices and established decision guidelines. Internal consistency reliability was assessed using both Cronbach’s alpha and McDonald’s omega. Omega was reported alongside alpha because it provides a less restrictive and often more accurate estimate of reliability when the assumption of tau-equivalence is violated (
Dunn et al., 2014).
The internal consistency of the scales was evaluated using the guidelines proposed by
George and Mallery (
2003), which provide conventional benchmarks for interpreting Cronbach’s alpha coefficients. The interpretation of the CFA models was conducted in accordance with the recommendations and fit index guidelines proposed by
Hair et al. (
2019).
1. Self-Regulated Learning in Mathematics
The scale used in this study to assess students’ self-regulated learning is a composite of three scales, each reflecting how the individual regulates the specific dimensions proposed by
Wolters et al. (
2005): Strategies for the Regulation of Academic Cognition, Strategies for the Regulation of Academic Behavior, and Strategies for the Regulation of Academic Motivation.
1.1. Strategies for the Regulation of Academic Cognition
The scale evaluating strategies for regulating academic cognition included items that measured students’ use of cognitive and metacognitive strategies in mathematics learning. Cognitive self-regulation consisted of 14 items reflecting rehearsal, elaboration, and organization strategies, while metacognitive self-regulation comprised 12 items assessing students’ planning, monitoring, and regulation of their learning processes. All items were rated on a five-point Likert scale, with higher scores indicating more frequent use of self-regulated learning strategies.
A confirmatory factor analysis supported a two-factor structure distinguishing Cognitive Self-Regulation (14 items) and Metacognitive Self-Regulation (12 items), with acceptable model fit (CFI = 0.919, TLI = 0.906, SRMR = 0.044, RMSEA = 0.046). Internal consistency was good for both dimensions (Cognitive Self-Regulation: α = 0.857, ω = 0.859; Metacognitive Self-Regulation: α = 0.817, ω = 0.822), and the overall scale demonstrated excellent reliability (α = 0.911, ω = 0.912).
1.2. Strategies for the Regulation of Academic Behavior
Behavioral self-regulation strategies were assessed using a unidimensional model adapted from the original two-factor self-report scales developed by
Wolters et al. (
2005), as detailed in the
Supplementary Material. The unidimensional model combines all 12 items into a single Behavioral Self-Regulation dimension (e.g., “I try to do my math assignments on my own, even if I need help”), reflecting an integrated process in which students manage both their persistence and their study conditions simultaneously. The behavioral regulation instrument was rated on a five-point Likert scale, with higher scores indicating greater use of behavioral self-regulation strategies. Scores for each dimension were computed by averaging the corresponding item responses, with higher values reflecting more effective regulation of time, study environment, and effort.
The unidimensional scale was subjected to confirmatory factor analysis. This model showed acceptable fit, χ2(48) = 106, p < 0.001, CFI = 0.932, TLI = 0.907, SRMR = 0.049, RMSEA = 0.063, 90% CI [0.047, 0.080]. Standardized loadings ranged from 0.218 to 0.709 and were statistically significant.
Internal consistency was adequate (α = 0.809; ω = 0.817), supporting retention of the unidimensional structure for subsequent analyses.
1.3. Strategies for the Regulation of Academic Motivation
Motivational self-regulation strategies were evaluated using the Academic Self-Regulation Questionnaire (SRQ-A), the Portuguese validated version developed by
Gomes et al. (
2019), based on Self-Determination Theory (
Ryan & Deci, 2000). The instrument measures students’ reasons for participating in academic activities and captures various forms of motivational regulation along a spectrum of self-determination.
The version used in this study included 16 items divided into four categories (four items each): External Regulation, which reflects motivation driven by external demands or rewards; Introjected Regulation, representing behavior influenced by internal pressures such as guilt or contingent self-worth; Identified Regulation, referring to engagement based on the personal significance of the activity; and Intrinsic Regulation, indicating motivation driven by interest and enjoyment. All items were rated on a five-point Likert scale, with higher scores showing a stronger endorsement of each regulatory style.
Internal consistency coefficients ranged from acceptable to good across the four dimensions (α = 0.699–0.856; ω = 0.707–0.859).
In addition, a Relative Autonomy Index (RAI) was computed to provide a global indicator of motivational self-regulation, reflecting the extent to which students’ motivation was relatively autonomous rather than controlled (
Grolnick & Ryan, 1989).
2. Formative assessment scale
2.1. Students’ version
Students’ and teachers’ perceptions of formative assessment practices were evaluated using a 27-item self-report scale. Although the instrument was initially designed to align with
Wiliam and Thompson’s (
2007) five formative assessment strategies, the psychometric analyses presented in the
Supplementary Materials supported a simpler three-dimensional structure: Clarification and Elicitation (11 items), Feedback (12 items), and Shared Regulation (4 items).
Clarification and Elicitation evaluate how clearly learning goals and criteria are communicated and how teachers use activities and classroom interactions to gather evidence of students’ understanding. Feedback measures how effectively teachers offer improvement-focused feedback that enhances students’ learning. Shared Regulation involves practices that encourage student involvement in managing their learning, including self-assessment and peer assessment.
All items were rated on a five-point Likert scale, with higher scores indicating a stronger perceived presence of formative assessment practices. A confirmatory factor analysis supported the three-factor structure with good model fit (CFI = 0.933, TLI = 0.925, SRMR = 0.049, RMSEA = 0.050). Internal consistency was good to excellent across the dimensions (α = 0.706–0.925; ω = 0.712–0.925), with excellent reliability for the overall formative assessment score (α = 0.931; ω = 0.931).
2.2. Teachers’ version
To examine the psychometric properties of the teacher version of the instrument, a confirmatory factor analysis was conducted using the same three-dimensional structure identified in the student sample. The results indicated weak model fit, likely due to the small sample size of teachers (N = 39), which can affect the stability of CFA estimates. Nonetheless, internal consistency estimates were acceptable across dimensions (α = 0.70–0.77; ω = 0.75–0.76).
Given the model’s theoretical coherence and the acceptable reliability of the dimensions, the teacher version of the scale was kept to allow comparisons between students’ and teachers’ perceptions of formative assessment practices.
3.3. Data Analysis Procedures
The data analysis proceeded in three consecutive stages aligned with the study hypotheses: (1) student-level correlational analyses (H1), (2) student-level multiple regression analyses evaluating the predictive contributions of formative assessment dimensions (H2), and (3) classroom-level analyses examining teacher–student perceptual discrepancies and agreement (H3).
Perceived formative assessment was operationalized using both a global score and dimensional scores. The global score reflected students’ overall perceptions of formative assessment practices and was used in correlational analyses, while dimensional scores, clarification and elicitation, feedback, and shared regulation were used to examine the specific contributions of these practices to different components of self-regulated learning.
Hypothesis H1 was tested using Pearson correlation analyses at the student level to explore links between students’ perceptions of formative assessment and indicators of self-regulated learning.
Hypothesis H2 explored whether different dimensions of formative assessment predicted specific aspects of self-regulated learning. In these analyses, the three formative assessment dimensions were entered simultaneously as predictors in multiple regression models. Four models were estimated, with cognitive self-regulation, metacognitive self-regulation, behavioral self-regulation, and motivational autonomy (Relative Autonomy Index) designated as dependent variables. Standardized regression coefficients (β), coefficients of determination (R2), and 95% confidence intervals were reported. Multicollinearity among predictors was evaluated using Variance Inflation Factors (VIFs) and tolerance statistics.
Because teacher–student congruence was calculated using class-aggregated student perceptions, the adequacy of this aggregation was assessed. ICC (1) values for the student-reported formative assessment ranged from 0.17 to 0.32, indicating significant between-class variability. With an average class size of (k = 7), ICC (2) values ranged from 0.59 to 0.76, supporting the reliability of the classroom means used as Level 2 predictors.
Hypothesis H3a was tested at the classroom level using matched teacher–class pairs. Student perception scores were combined at the class level and compared with teacher self-reports using paired-sample t-tests.
Hypothesis H3b was examined using multilevel linear models with random classroom intercepts to account for the hierarchical data structure, with students (Level 1) nested within classrooms (Level 2). For each self-regulated learning outcome (cognitive, metacognitive, behavioral, and motivational regulation), an unconditional (intercept-only) model was first fitted to calculate the intraclass correlation coefficient (ICC), which measures the proportion of variance due to differences between classrooms. ICC values ranged from 0.068 to 0.145, indicating significant between-class variance and justifying the use of multilevel modeling.
Teacher–student perceptual agreement in formative assessment practices was measured using classroom-level discrepancy scores. Absolute differences were calculated between teachers’ self-reported formative assessment scores and the corresponding class-aggregated student perception scores for each formative assessment dimension (clarification and elicitation, feedback, and shared regulation). Lower values indicate closer teacher–student perceptual alignment, while higher values reflect greater perceptual divergence.
These discrepancy indices were entered as Level 2 fixed predictors in the multilevel models, while students’ self-regulated learning outcomes were specified at Level 1. This modeling approach allowed examination of whether variation in teacher–student perceptual alignment across classrooms was linked to differences in students’ self-regulated learning, while accounting for the non-independence of observations within classrooms.
For interpretative purposes, approximate standardized coefficients were calculated by rescaling the unstandardized fixed effects using the ratio of the predictor and outcome standard deviations (β_std ≈ b × SD_X/SD_Y). This transformation does not influence statistical inference but makes it easier to compare effect sizes across different outcomes.
Before interpreting the regression and multilevel models, diagnostic procedures were performed to evaluate underlying assumptions. For the single-level regression models, linearity and homoscedasticity were checked through visual inspection of standardized residuals plotted against predicted values, and normality was assessed using histograms and normal P–P plots of standardized residuals. No significant violations were detected. Influential observations were examined using Cook’s distance and standardized residuals; Cook’s distance values stayed below 0.14, and standardized residuals did not exceed |2.07|. Multicollinearity diagnostics showed no problematic collinearity (VIF < 2.12; tolerance > 0.47).
For the multilevel models, examining Level 1 residuals plotted against fitted values showed no serious violations of linearity or homoscedasticity. Considering the Level 1 sample size (N = 305), the estimation methods are deemed robust to small deviations from normality. Overall, no significant violations of model assumptions were found that would threaten the validity of the reported results. All statistical analyses were carried out using Jamovi software (version 2.5.3) and IBM SPSS Statistics (version 31.01.0 (49)).
4. Results
4.1. Associations Between Formative Assessment and Self-Regulated Learning
To test H1, Pearson correlation analyses were conducted at the student level to examine the relationships between students’ perceptions of formative assessment practices and various indicators of self-regulated learning (SRL). As shown in
Table 1, results indicated that Global FA was positively and significantly associated with all major dimensions of self-regulated learning. Higher perceived levels of formative assessment were moderately associated with cognitive SRL, metacognitive SRL, and behavioral SRL, as well as with students’ Relative Autonomy Index (RAI).
As reported in
Table 1, some correlations among self-regulated learning dimensions were relatively high, particularly between shared regulation and metacognitive self-regulation. This pattern likely reflects the conceptual proximity between these constructs, as both involve processes of monitoring, reflection, and regulation during learning activities. Because both variables were assessed via student self-report, shared method variance may also have contributed to the magnitude of these associations. Accordingly, these correlations were interpreted as indicative of strong alignment rather than as evidence of construct redundancy.
All three dimensions of formative assessment were positively and significantly associated with cognitive, metacognitive, behavioral, and motivational aspects of SRL. The strongest associations were observed for Shared Regulation, especially with metacognitive (r = 0.83, p < 0.001) and cognitive self-regulation (r = 0.70, p < 0.001). Theoretically, this pattern is coherent: shared regulation practices involve dialogue, joint monitoring, and collaborative evaluation, processes that externalize regulatory thinking and may support the development of individual metacognitive control.
Clarification and Elicitation also showed consistent positive associations across SRL dimensions, particularly with behavioral regulation (e.g., effort and time management). Feedback demonstrated positive, though comparatively smaller, relationships with SRL indicators. In addition, all three formative assessment dimensions were positively related to students’ Relative Autonomy Index, suggesting links between formative classroom practices and motivational autonomy.
These results indicate that while formative assessment operates as a coherent global construct, its components relate differentially to specific self-regulatory processes.
4.2. Predictors of Self-Regulated Learning Dimensions
To test Hypothesis 2, a series of multiple linear regression analyses was performed at the student level to evaluate the specific contributions of different formative assessment dimensions to students’ self-regulated learning (SRL). In these analyses, three formative assessment dimensions, clarification and elicitation, feedback, and shared regulation, were included simultaneously as predictors, and each SRL dimension was analyzed as a separate outcome variable.
Results indicated that the regression model predicting cognitive self-regulation was statistically significant, F(3, 301) = 17.55, p < 0.001, accounting for 14.9% of the variance (R2 = 0.149). Both clarification and elicitation (β = 0.16, p = 0.022) and feedback (β = 0.17, p = 0.030) emerged as significant positive predictors, while shared regulation showed a marginal association (β = 0.13, p = 0.057).
The overall model for metacognitive self-regulation was also significant, F(3, 301) = 22.30, p < 0.001, accounting for 18.2% of the variance (R2 = 0.182). In this model, clarification and elicitation was the only significant predictor (β = 0.38, p < 0.001), while feedback and shared regulation did not show significant effects.
The model predicting behavioral self-regulation was significant, F(3, 300) = 43.48, p < 0.001, explaining 30.3% of the variance (R2 = 0.303). Clarification and elicitation again emerged as the only significant predictor (β = 0.54, p < 0.001), while feedback and shared regulation were not significantly linked to behavioral regulation.
Finally, the regression model predicting motivational autonomy (Relative Autonomy Index) was also significant, F(3, 297) = 15.24, p < 0.001, explaining 13.3% of the variance (R2 = 0.133). Clarification and elicitation significantly predicted higher levels of autonomous motivation (β = 0.33, p < 0.001), while feedback and shared regulation were not significant predictors.
Across models, multicollinearity diagnostics showed no issues, as Variance Inflation Factor (VIF) values stayed well below standard thresholds.
These findings suggest that different formative assessment dimensions display distinct predictive patterns across SRL outcomes. Specifically, practices related to clarifying learning goals and gathering evidence of student understanding stand out as the most reliable predictors across cognitive, metacognitive, behavioral, and motivational aspects of self-regulated learning.
4.3. Discrepancies Between Teachers’ and Students’ Perceptions of Formative Assessment
To test Hypothesis H3a, paired-samples
t-tests were conducted at the teacher/class level (
N = 39) to compare teachers’ self-reported use of formative assessment (FA) practices with students’ class-aggregated perceptions. In addition to dimensional analyses, a paired-samples
t-test was also conducted using a global formative assessment score (Global FA) to provide an overall synthesis of teacher–student perceptual differences. Because the teacher version of the FA measure did not show ideal psychometric properties (see
Section 3.2), these comparisons should be interpreted cautiously as contrasts between parallel observed perceptions rather than between psychometrically equivalent latent constructs.
Results based on the global score indicated that teachers reported significantly higher overall levels of formative assessment than students perceived, t(38) = −10.45, p < 0.001, with a very large effect size (Cohen’s d = −1.67). Dimensional analyses revealed the same pattern across all FA components. Teachers reported significantly higher levels of Clarification and Elicitation (M = 4.49, SD = 0.34) than students perceived (M = 3.65, SD = 0.43), t(38) = −11.32, p < 0.001, d = −1.81. Teachers’ reports of Feedback (M = 3.68, SD = 0.35) also exceeded students’ perceptions (M = 2.67, SD = 0.63), t(38) = −10.10, p < 0.001, d = −1.62. Finally, teachers reported higher levels of Shared Regulation (M = 2.53, SD = 0.69) than students did (M = 1.90, SD = 0.53), t(38) = −5.04, p < 0.001, d = −0.81. Overall, these findings provide strong support for H3a, indicating large and systematic discrepancies between teachers’ self-reports and students’ perceptions, both globally and across dimensions.
In addition to mean differences, Pearson correlations between teachers’ and class-aggregated students’ perceptions were examined. These correlations were small to moderate across dimensions (rs ranging from 0.19 to 0.30; r = 0.23 for the global score), indicating partial but limited convergence between teachers’ and students’ views of formative assessment practices.
4.4. Perceptual Congruence and Student Self-Regulation
To investigate whether teacher–student perceptual consistency across formative assessment areas relates to students’ self-regulated learning (SRL), a series of multilevel models was used to account for the nested structure of students within classrooms. Null (intercept-only) models showed meaningful variability between classes for SRL outcomes, with intraclass correlation coefficients (ICC) ranging from 0.068 (motivational autonomy; RAI) to 0.145 (cognitive regulation), indicating that between 6.8% and 14.5% of the variance in students’ self-regulated learning was due to classroom membership. These results demonstrate significant between-class differences and justify using multilevel modeling.
Perceptual congruence was operationalized as absolute difference scores between teachers’ self-reports and class-aggregated student perceptions for each FA dimension. Because both measures were assessed on five-point Likert scales, discrepancy scores could theoretically range from 0 (perfect alignment) to 4 (maximum divergence). In the present sample, observed discrepancy scores were substantially lower, with standard deviations ranging approximately from 0.65 to 0.75, indicating moderate variability across classrooms. Accordingly, lower discrepancy values indicate greater teacher–student alignment, whereas higher values reflect greater perceptual divergence. Because congruence was modeled using absolute difference scores, negative regression coefficients indicate that smaller discrepancies (i.e., greater alignment) are associated with higher SRL outcomes.
Multilevel models were estimated to examine the association between teacher–student perceptual congruence and students’ self-regulated learning outcomes.
To improve transparency and enable a comprehensive evaluation of the analyses,
Table 2 displays the full set of multilevel regression models, including both significant and non-significant effects, along with standard errors, standardized coefficients (where applicable), and
p-values.
For cognitive self-regulation, congruence in Clarification and Elicitation significantly predicted outcomes (β = −0.34, SE = 0.06, t = −5.80, p < 0.001; β_std ≈ −0.32). Congruence in Feedback was also significant (β = −0.29, SE = 0.06, t = −5.20, p < 0.001; β_std ≈ −0.30). In this outcome only, congruence in Shared Regulation showed a small but statistically significant association (β = −0.17, SE = 0.06, t = −2.69, p = 0.008; β_std ≈ −0.16).
For metacognitive self-regulation, congruence in Clarification and Elicitation was a significant predictor (β = −0.42, SE = 0.05, t = −7.96, p < 0.001; β_std ≈ −0.43), as was congruence in Feedback (β = −0.28, SE = 0.05, t = −5.17, p < 0.001; β_std ≈ −0.30). Congruence in Shared Regulation was not statistically significant (β = −0.08, SE = 0.06, p = 0.178; β_std ≈ −0.08).
A similar pattern was observed for behavioral self-regulation. Congruence in Clarification and Elicitation showed a strong association (β = −0.50, SE = 0.05, t = −10.40, p < 0.001; β_std ≈ −0.53), representing the largest standardized effect across models. Congruence in Feedback was also significant (β = −0.24, SE = 0.05, t = −4.69, p < 0.001; β_std ≈ −0.27), whereas congruence in Shared Regulation was not significant (β = −0.06, SE = 0.06, p = 0.311; β_std ≈ −0.06).
For motivational autonomy (RAI), congruence in Clarification and Elicitation emerged as a significant predictor (β = −2.01, SE = 0.28, t = −7.09, p < 0.001; β_std ≈ −0.42). Congruence in Feedback was also significant (β = −1.21, SE = 0.29, t = −4.21, p < 0.001; β_std ≈ −0.25). In contrast, congruence in Shared Regulation did not significantly predict RAI (β = −0.22, SE = 0.31, p = 0.484; β_std ≈ −0.04).
In addition to the fixed effects reported in
Table 2, the marginal and conditional
R2 values in
Table 3 indicate the proportion of variance explained by the models. Marginal
R2 values, reflecting variance explained by the fixed effects alone, ranged from 0.002 to 0.272 across outcomes, indicating modest to moderate explanatory power, depending on the formative assessment dimension considered. The largest marginal
R2 values were observed for behavioral self-regulation predicted by Clarification and Elicitation (
R2m = 0.272) and for metacognitive regulation predicted by Clarification and Elicitation (
R2m = 0.178). Conditional
R2 values, which incorporate both fixed and random effects, ranged from 0.074 to 0.339, suggesting that between-class variability accounted for an additional proportion of variance across outcomes. Overall, these indices indicate that perceptual congruence in core formative assessment practices explains a meaningful, though not exhaustive, share of classroom-level differences in self-regulated learning.
These findings suggest that although some effects operate primarily at the individual level, task–objective incongruence exhibits a more stable class-level pattern. Importantly, the multilevel mixed-effects models reported above already accounted for clustering by class, reinforcing the robustness of the main results.
The pattern of findings provides strong but dimension-specific support for H3b. Across outcomes, greater teacher–student alignment in Clarification and Elicitation showed the most consistent and strongest associations with cognitive, metacognitive, behavioral, and motivational regulation, followed by moderate and stable associations for Feedback. Congruence in Shared Regulation showed generally weak or nonsignificant associations, except for a small association with cognitive regulation.
5. Discussion
The present findings indicate that students’ perceptions of formative assessment (FA) practices in mathematics are positively associated with cognitive, metacognitive, behavioral, and motivational dimensions of self-regulated learning (SRL). At a global level, FA showed moderate correlations with all SRL components, suggesting that students who perceive their classroom as more formative also report engaging more frequently in regulatory processes.
Rather than merely replicating accounts that position formative assessment as inherently linked to self-regulated learning (
Black & Wiliam, 2009;
Nicol & MacFarlane-Dick, 2006;
Panadero et al., 2018), these findings align with a growing body of work conceptualizing FA as a contextual mechanism that structures and sustains regulatory processes (
Brandmo et al., 2020;
van der Linden et al., 2023;
Li & Gu, 2024). Within this perspective, formative assessment operates not as an isolated feedback event but as an integrated instructional ecology in which goals, evidence, and feedback are continuously aligned. When students experience mathematics classrooms as environments where expectations are transparent, evidence of understanding is actively elicited, and feedback informs subsequent action, regulatory processes are embedded within instructional interaction (
Clark, 2012;
Heritage & Wylie, 2023).
Clarification and elicitation practices showed moderate-to-strong associations with cognitive, metacognitive, and behavioral regulation. This pattern is theoretically coherent. In
Zimmerman’s (
2000) cyclical model, the forethought phase involves task analysis and strategic planning. When teachers make learning intentions and success criteria explicit, students gain access to the standards against which performance is evaluated (
Andrade & Heritage, 2017). Clarity reduces ambiguity about quality and supports anticipatory alignment between task demands and strategy selection. In mathematics, where effective problem solving requires coordination between procedural execution and conceptual justification, transparent criteria may facilitate deliberate strategy choice and ongoing monitoring of reasoning processes. Empirical research examining students’ perceptions of assessment transparency further indicates that clarity of goals and criteria is associated with stronger engagement and regulatory behavior (
Wolterinck-Broekhuis et al., 2024;
Šimić Šašić & Atlaga, 2024).
Exposure to criteria also supports the development of evaluative judgment, defined as the capacity to appraise the quality of one’s own work (
Panadero & Broadbent, 2018). Evaluative judgment inherently activates metacognitive processes, particularly comparisons between current performance and explicit standards. Evidence from studies on self- and co-regulated learning suggests that evaluative processes serve as a bridge between formative assessment practices and metacognitive activation (
Panadero et al., 2019;
van der Linden et al., 2023). Accordingly, the observed association between clarification practices and metacognitive SRL reflects theoretical alignment rather than incidental covariance. These findings do not imply causal direction but suggest that classrooms characterized by transparent goal communication are also those in which students report greater engagement in planning and monitoring processes.
Formative feedback was positively associated with all SRL dimensions, although effect sizes were smaller. Feedback has long been conceptualized as information that reduces the discrepancy between current and desired performance (
Nicol & MacFarlane-Dick, 2006). However, its regulatory potential depends on learners’ capacity to interpret, internalize, and act on it. Research on feedback literacy emphasizes that feedback functions formatively only when it is meaningfully processed and translated into strategic adjustment (
Carless & Winstone, 2023;
He et al., 2023). The more moderate correlations observed may therefore reflect the conditional nature of feedback use. Although feedback may be present in classroom practice, its impact varies depending on whether students perceive it as dialogic, actionable, and relevant to improvement. In mathematics contexts, where feedback often targets procedural correctness, its influence on deeper metacognitive processes may be attenuated if it is experienced as corrective rather than developmental.
The strongest correlations emerged between shared regulation of learning and both metacognitive and cognitive self-regulation. This finding aligns with frameworks that distinguish individual, co-regulated, and socially shared regulation (
Hadwin et al., 2017). Socially shared regulation involves joint planning, monitoring, and evaluation enacted through dialogue. In assessment contexts, regulatory processes may be externalized in classroom discourse before being internalized at the individual level (
Allal, 2013;
Andrade et al., 2021). Classroom-based research further indicates that co-regulated and socially shared practices are associated with increased engagement and strategic participation in secondary education settings (
Veugen et al., 2024;
Fernández-Ferrer et al., 2025). In shared regulation environments, monitoring and evaluation are made visible and collectively negotiated, which may normalize metacognitive engagement. At the same time, given that both constructs were assessed via self-report and involve conceptually overlapping activities, part of the magnitude of these correlations may reflect shared method variance. The results, therefore, suggest strong functional alignment rather than construct equivalence.
All FA dimensions were positively associated with motivational autonomy (RAI), though effect sizes were more moderate. Formative assessment may support adaptive motivation by shifting the evaluative focus from normative comparison to personal improvement (
Panadero et al., 2018). When goals are transparent and feedback emphasizes progress, students may experience enhanced perceived competence and clearer expectations. Empirical evidence indicates that students who perceive assessment as constructive and improvement-oriented report higher levels of intrinsic motivation and engagement (
Pat-El et al., 2024;
van der Linden et al., 2023).
Nevertheless, motivational regulation is shaped by multiple interacting influences. Situated Expectancy-Value Theory posits that prior achievement contributes recursively to the development of competence beliefs and task values (
Eccles & Wigfield, 2020;
Marsh et al., 2012). Self-Determination Theory similarly conceptualizes autonomy as the outcome of progressive internalization processes supported by satisfaction of competence, autonomy, and relatedness needs (
Ryan & Deci, 2000,
2017,
2020). Evidence further suggests that perceptions of feedback influence regulatory engagement, in part, through motivational mediators such as self-efficacy and goal orientation (
He et al., 2023). The comparatively modest effect sizes observed are therefore theoretically coherent: formative assessment may facilitate motivational internalization, but it operates within a broader motivational ecology shaped by prior experiences, identity-related meanings, and classroom climate.
We also examined whether different dimensions of students’ perceived formative assessment practices predict variance in various self-regulated learning (SRL) dimensions (H2). Conceptualizing formative assessment as a system of interconnected practices suggests that its different elements may contribute in distinct ways to students’ regulatory processes rather than functioning as a single undifferentiated construct (
Wiliam & Thompson, 2007). Therefore, the analyses considered the simultaneous contribution of three formative assessment dimensions, clarification and elicitation, feedback, and shared regulation, to multiple SRL outcomes.
The results showed different predictive patterns across SRL dimensions. In particular, practices involving clarifying learning goals and gathering evidence of student understanding emerged as the most consistent predictors across regulatory domains. This dimension significantly predicted cognitive, metacognitive, behavioral, and motivational aspects of SRL, suggesting that classrooms where expectations are made clear and evidence of understanding is regularly collected may offer stronger support for students’ learning processes. This pattern aligns with theories that see clarifying learning goals and collecting evidence as key parts of formative assessment (
Wiliam & Thompson, 2007;
Panadero et al., 2018). When students know what quality work looks like and are often asked to show their understanding through tasks and discussions, they can better monitor their progress, evaluate their strategies, and adjust their learning accordingly.
Clarification and elicitation practices were especially strongly linked to behavioral self-regulation, which accounted for the largest portion of variance among the SRL outcomes. Behavioral regulation reflects the active side of self-regulation, including effort, persistence, time management, and strategic follow-through. In math learning, where tasks often demand sustained reasoning and repeated problem-solving, these active regulatory processes are particularly important. Practices that clarify next steps, create opportunities for revision, and make progress visible may thus directly impact students’ behavioral engagement with math tasks.
This pattern is especially meaningful in mathematics education. Mathematical understanding develops through cumulative, conceptually structured processes that integrate procedural fluency, conceptual understanding, and strategic competence (
Kilpatrick & Swafford, 2002). Empirical research also shows that successful mathematical problem solving depends on metacognitive monitoring and persistence when facing difficulties (
Özcan, 2016;
Wafubwa & Csíkos, 2022). In such contexts, formative practices that make learning goals explicit and regularly gather evidence of understanding can help students regulate effort, revise strategies, and stay engaged during cognitively demanding tasks.
In contrast, formative feedback demonstrated more limited predictive effects once the other formative dimensions were considered together. Although feedback has long been recognized as a key mechanism linking assessment and learning, its regulatory influence may depend on the extent to which it is integrated into broader cycles of clarification, evidence collection, and revision (
Nicol & MacFarlane-Dick, 2006;
Panadero et al., 2018). When combined with practices that clarify expectations and gather evidence of understanding, feedback alone might only cover part of the regulatory processes through which formative assessment supports learning.
Clarification and elicitation practices also predicted motivational autonomy, although the proportion of explained variance was smaller than that for cognitive, metacognitive, and behavioral regulation. Within Self-Determination Theory, autonomy refers to the extent to which regulatory processes have been internalized and integrated into the self (
Ryan & Deci, 2020). Formative assessment practices that clarify expectations and emphasize improvement may support students’ perceptions of competence and reduce controlling evaluative pressures, thereby fostering autonomous regulation (
Ryan & Deci, 2000,
2017,
2020;
Jang et al., 2010). At the same time, motivational internalization typically develops gradually within sustained need-supportive environments rather than as an immediate result of individual instructional practices (
Ryan & Deci, 2017,
2020). From a Situated Expectancy-Value perspective, prior achievement experiences and evolving beliefs about competence also significantly influence motivational pathways (
Eccles & Wigfield, 2020;
Marsh et al., 2012). Therefore, formative assessment may serve as one of several contextual factors affecting students’ motivational regulation.
These findings highlight that formative assessment practices do not contribute equally to students’ regulatory processes. Instead, practices related to clarifying learning goals and eliciting evidence of understanding seem to play a particularly central role in supporting multiple aspects of self-regulated learning. This varied pattern emphasizes the importance of viewing formative assessment not only as a unified instructional approach but also as a collection of practices that can influence different facets of students’ regulatory engagement with learning.
Interestingly, the pattern observed in the congruence analyses partially mirrors the student-level results reported for Hypothesis 2. In the regression models examining students’ perceptions of formative assessment practices, clarification and elicitation emerged as the most consistent predictors across several self-regulated learning dimensions. A similar pattern is observed in the congruence analyses, where alignment between teachers and students in this dimension shows the strongest and most stable associations with classroom-level self-regulated learning outcomes.
Hypothesis 3a was tested using both the global Formative Assessment (FA) score and its specific dimensions (Clarification and Elicitation, Feedback, and Shared Regulation), enabling both a synthetic and a differentiated analysis of teacher–student discrepancies. The results revealed a clear and consistent pattern: teachers reported significantly higher levels of formative assessment practices than students perceived, both at the global level and across all assessed dimensions. Effect sizes were large to very large, indicating that the discrepancy is not marginal but structural and systematic.
This pattern aligns with the literature on perceptual incongruence in formative assessment. Studies by
Pat-El et al. (
2013,
2015) show that teachers report higher levels of Assessment for Learning practices than students recognize, highlighting systematic discrepancies between pedagogical intention and student experience. These authors distinguish between implementation incongruence (a practice is reported by the teacher but not clearly perceived by students) and interpretative incongruence (both parties recognize the practice but attribute different meanings to it). While the present study does not allow us to determine which of these mechanisms may be operating, the observed discrepancies are compatible with the types of incongruence described in prior research.
The particularly large discrepancies observed in the dimensions of Clarification, Elicitation, and Feedback suggest that practices teachers consider central to formative assessment may not be experienced by students as regulatory tools. As argued by
Panadero et al. (
2018), mere exposure to criteria does not guarantee the development of evaluative judgment; students must appropriate those criteria as tools for monitoring and self-regulation. Similarly, the feedback literature emphasizes that its formative function depends on how learners interpret and use it (
Nicol & MacFarlane-Dick, 2006;
Carless & Winstone, 2023). Thus, the observed discrepancy may reflect differences in how the purpose and nature of feedback are perceived, particularly when feedback is experienced as corrective rather than developmental.
Although smaller in magnitude, the discrepancy in Shared Regulation remained statistically significant. Given that this dimension involves interactive, socially visible practices (
Hadwin et al., 2017), it may be more readily recognized by students than less explicitly collaborative practices. Nevertheless, the results suggest that opportunities teachers interpret as co-regulatory may not always be experienced by students as genuine moments of joint monitoring and negotiation. Beyond mean differences, correlations between teachers’ perceptions and aggregated student perceptions were small to moderate, indicating partial but limited convergence. This pattern suggests that both informants capture related aspects of the formative classroom climate, but not in equivalent ways. The literature consistently shows that multiple informants provide complementary perspectives and that the absence of strong convergence does not invalidate either source but rather reflects the relational and interpretative nature of assessment practices (
Pat-El et al., 2015).
The interpretation of these findings should be considered, considering the psychometric limitations of the teacher-reported scale. As described in the instruments section, the relatively small number of teacher participants constrained the stability of psychometric estimates and limited the robustness of factor-analytic procedures at the teacher level. Although the pattern of discrepancies is consistent and theoretically plausible, the magnitude of the differences should be interpreted with caution, acknowledging potential measurement-related constraints.
The results of H3a confirm the presence of large, systematic discrepancies between teachers’ self-reported formative assessment practices and students’ perceptions, both globally and across dimensions. Hypothesis 3b examined whether teacher–student perceptual congruence across specific formative assessment (FA) dimensions was associated with student-level SRL outcomes, estimated using multilevel models that accounted for classroom-level clustering. Congruence was operationalized as the absolute difference between teachers’ reports and class-level student perceptions for each FA dimension. Importantly, analyses were conducted at a dimensional level rather than using a global FA score. This decision was theoretically motivated, as prior research suggests that discrepancies between teachers’ and students’ perceptions may vary across specific formative practices (
Pat-El et al., 2013,
2015), and aggregating across dimensions could obscure differentiated patterns of alignment.
Importantly, the current analyses were conducted using multilevel models with student-level SRL outcomes nested within classrooms. Therefore, the results should be viewed as cross-level associations between classroom-level perceptual congruence and students’ self-regulated learning, rather than as solely classroom-aggregated effects.
The results provided dimension-specific support for H3b. Across outcomes, greater teacher–student alignment in Clarification and Elicitation showed the most consistent and substantial associations, followed by moderate associations for Feedback. Congruence in Shared Regulation was generally weak or non-significant, except for a small association with cognitive regulation.
At the cognitive level, congruence in Clarification and Elicitation and in Feedback was significantly associated with higher class-level cognitive self-regulation, with a smaller but significant effect also emerging for Shared Regulation. Cognitive regulation refers to the strategic processing and organization of information; thus, alignment in how learning goals, criteria, and evidence-eliciting tasks are perceived may strengthen students’ understanding of what constitutes adequate mathematical reasoning and solution strategies. Formative assessment frameworks emphasize that learning intentions and criteria serve as regulatory anchors when students meaningfully understand and use them (
Panadero et al., 2018). When teachers and students share similar interpretations of these practices, the transparency of expectations may facilitate strategic engagement with mathematical tasks. Importantly, however, given the correlational and multilevel nature of the analyses, these findings indicate association rather than directional influence.
A similar pattern emerged for metacognitive regulation, with congruence in Clarification and Elicitation, and in Feedback, showing stable, substantial associations, whereas Shared Regulation did not reach significance. Metacognitive regulation involves monitoring and evaluating one’s understanding and strategies. If teachers perceive that criteria and questioning practices are clearly articulated, but students do not recognize or interpret them similarly, the metacognitive potential of these practices may be attenuated. The present findings are consistent with the idea that a shared understanding of learning goals and success criteria may enhance the extent to which students engage in reflective monitoring processes at the classroom level. At the same time, caution is warranted: these data do not allow conclusions about whether perceptual alignment enhances metacognitive engagement or whether classrooms characterized by stronger metacognitive climates foster more aligned perceptions.
For behavioral self-regulation, congruence in Clarification and Elicitation again showed the strongest association, followed by Feedback, while Shared Regulation was non-significant. Behavioral regulation reflects effort management, persistence, and sustained engagement, processes particularly salient in mathematics learning. When teachers and students converge on their understanding of what is expected and how evidence of learning is elicited, students may be better positioned to translate goals and feedback into observable effort and persistence on tasks. Feedback, when perceived similarly by teachers and students, may function as actionable guidance that supports strategic adjustment. These findings align with prior work suggest that discrepancies in how feedback practices are perceived may weaken their formative function (
Pat-El et al., 2015). Nonetheless, as with clarification practices, the present results remain associational.
Finally, for motivational autonomy (RAI), congruence in Clarification and Elicitation and in Feedback showed significant associations, whereas Shared Regulation did not. This pattern suggests that alignment in core formative practices, such as transparent expectations and improvement-oriented feedback, may coexist with stronger classroom-level motivational internalization. When students recognize assessment practices as coherent and aligned with teachers’ intentions, they may experience greater clarity and perceived competence, which are theoretically linked to autonomous motivation. However, the present design does not permit claims about motivational mechanisms, and these associations should not be interpreted causally.
In contrast, perceptual congruence in Shared Regulation was associated with SRL outcomes only to a limited extent. At first glance, this may appear inconsistent with H1, which predicted that students’ perceptions of shared regulation would be strongly correlated with individual SRL. However, the constructs and levels of analysis differ. H1 examined student-level associations between perceived shared regulation and individual SRL, whereas H3b examined whether teacher–student alignment in that perception predicted class-level outcomes. The weaker congruence effects observed here, therefore, do not contradict earlier findings; rather, they suggest that alignment between teacher and student perceptions of shared regulation may be less strongly related to aggregated regulatory outcomes than alignment in clarification and feedback practices.
One possible interpretation, consistent with frameworks that distinguish individual, co-regulated, and socially shared regulation (
Hadwin et al., 2017), is that students may experience socially shared regulation in ways that do not necessarily depend on symmetrical perceptions between teachers and students. Shared regulation is inherently distributed and interactional; thus, students may experience collaborative regulation even when teachers conceptualize those practices differently. However, the present design does not allow conclusions regarding underlying mechanisms. The weaker associations may also reflect measurement characteristics, contextual variability in collaborative practices, or the inherently distributed nature of shared regulation processes. Further research using observational or longitudinal designs would be needed to clarify how perceptual alignment interacts with socially mediated regulation in mathematics classrooms.
Measurement constraints must also be acknowledged. As noted in the Instruments section, the teacher-reported formative assessment scale did not demonstrate optimal psychometric stability, largely due to the relatively small number of teacher participants. Difference scores are inherently sensitive to measurement error in both component variables, and instability in the teacher measure may attenuate the effects of congruence. Consequently, although the overall pattern of results is theoretically interpretable and statistically robust for certain dimensions, the magnitude of associations should be interpreted with caution.
Despite these limitations, the findings extend prior work on perceptual discrepancies in Assessment for Learning (
Pat-El et al., 2013,
2015) by examining, through multilevel modeling, whether perceptual congruence is associated with classroom-level SRL in mathematics. While few studies have applied similar multilevel difference-score approaches in subject-specific contexts, the present results suggest that teacher–student alignment in core formative practices, particularly those related to goal clarification and feedback, may be meaningfully associated with collective regulatory functioning. At the same time, the differentiated pattern across dimensions underscores that not all formative practices rely equally on perceptual alignment to relate to self-regulated learning outcomes.
6. Conclusions
The present study examined how students’ and teachers’ perceptions of formative assessment relate to students’ self-regulated learning in lower secondary mathematics. This study shows that formative assessment plays an important role in supporting students’ self-regulated learning in lower secondary mathematics. Students who perceived a stronger formative assessment environment reported higher levels of cognitive, metacognitive, behavioral, and motivational regulation. More specifically, among the formative assessment dimensions, clarification and elicitation emerged as the most consistent predictor across all SRL domains.
Simultaneously, the findings showed consistent differences between teachers’ reported practices and students’ perceptions, emphasizing that the effectiveness of formative assessment relies on how students interpret and apply assessment information. Significantly, greater teacher–student perceptual alignment, especially in clarifying learning goals and gathering evidence of learning, was linked to improved behavioral regulation and increased motivational autonomy. These findings support viewing formative assessment as a relational and co-regulatory process with important implications for both instructional practice and assessment policy.
From a pedagogical perspective, the findings carry important implications for mathematics teaching. First, the particularly strong associations between formative assessment and behavioral regulation suggest that formative practices in mathematics may be especially effective when they structure visible cycles of revision, persistence, and strategic adjustment. Mathematics learning often requires sustained engagement with complex, multi-step tasks in which errors serve as opportunities for refinement rather than as final evaluations. Teachers may therefore enhance students’ self-regulated engagement by explicitly framing feedback and success criteria as tools for iterative improvement, making the revision process visible and expected rather than optional.
Second, the robust effects of teacher–student congruence in Clarification and Elicitation underscore the importance of a shared understanding of learning goals and success criteria. In mathematics classrooms, where conceptual coherence and procedural justification are central, clarifying what counts as a valid strategy, a complete explanation, or a mathematically sound argument may strengthen students’ capacity to plan, monitor, and evaluate their work. Importantly, the findings suggest that clarity must not only be provided but also perceived as meaningful by students. This underscores the need for dialogic negotiation of criteria, the use of exemplars, and opportunities for students to articulate their understanding of expectations.
Third, the weaker and less consistent effects observed for shared regulation congruence suggest that collaborative practices alone may not guarantee regulatory impact unless their purpose and function are made explicit. Teachers may benefit from making the regulatory intentions behind peer discussion and group problem-solving visible, helping students recognize these moments as opportunities for joint monitoring and strategy development rather than as routine classroom interaction.
The findings suggest that in mathematics education, formative assessment practices are most likely to foster self-regulated learning when they are coherent, transparent, and explicitly linked to students’ regulatory processes. Professional development initiatives may therefore prioritize not only implementing formative techniques but also strategies that enhance perceptual alignment between teachers and students regarding the purpose and function of assessment practices.
Despite its contributions, the present study has several limitations that should be considered when interpreting the findings and that point to directions for future research.
First, the study relied on self-reported measures of formative assessment and self-regulated learning. Although students’ perceptions are theoretically central to understanding the regulatory function of formative assessment (
Pat-El et al., 2013), self-report data may be influenced by social desirability and shared method variance. Future studies would benefit from combining self-reports with classroom observations, analysis of instructional artifacts, or process-oriented measures of self-regulation (
Panadero et al., 2018).
Second, the cross-sectional design limits causal inference. While the findings are consistent with theoretical models proposing that formative assessment supports the development of self-regulated learning (
Clark, 2012;
Meusen-Beekman et al., 2016), reciprocal relations are also plausible. Longitudinal and intervention studies are needed to examine how changes in formative assessment practices and perceptual alignment over time influence students’ self-regulation and motivation.
Third, interpreting teacher–student discrepancies and perceptual congruence requires caution. Although the student version of the formative assessment scale demonstrated a robust factor structure and strong reliability, the teacher version showed limited structural validity and was therefore treated as a descriptive observed measure. Consequently, the discrepancy and congruence indices represent associations between parallel perceptions rather than comparisons between psychometrically equivalent constructs. This limitation may have attenuated the magnitude of congruence effects and highlights the need for further refinement and validation of teacher-reported measures of formative assessment. Accordingly, findings related to H3a and H3b should be interpreted as indicative patterns of perceptual alignment and misalignment rather than precise estimates of latent agreement.
Fourth, because the study relied on self-reported measures, common method bias cannot be fully ruled out. Future studies could complement questionnaire data with classroom observations or performance-based measures.
Another limitation involves the relatively small number of classrooms included in the multilevel analyses (N = 39). Simulation studies show that cluster-level sample sizes of this size have limited statistical power to detect small effects at Level 2. Therefore, some non-significant results should be interpreted carefully, as the analyses may have been underpowered to identify small but potentially meaningful classroom-level relationships.
Finally, future studies could extend the present work by examining mediating and moderating mechanisms, such as whether the effects of formative assessment on motivational autonomy are mediated by cognitive or behavioral self-regulation, or whether perceptual congruence operates differently across subject domains, age groups, or instructional contexts. Such research would further refine the understanding of how formative assessment can be effectively leveraged to support the development of self-regulated learners (
Pat-El et al., 2024).
These findings contribute to the growing body of research examining how formative assessment practices operate within subject-specific contexts, particularly in mathematics education. They highlight the importance of fostering shared understandings of assessment processes in classrooms, suggesting that the effectiveness of formative assessment depends not only on teachers’ instructional intentions but also on how students perceive and engage with these practices during learning.