2. Methods
2.1. Participants
To determine the required sample size, an a priori power analysis was conducted using G*Power 3.1.8 (
Faul et al., 2009) for an ANCOVA with four groups. ANCOVA was selected because adjusting for relevant covariates can reduce error variance and thereby improve statistical power in experimental studies (
Shieh, 2020). Based on an anticipated effect size of
f = 0.40, an alpha level (
α) of 0.05, and power of 0.85, the analysis indicated a minimum sample of 81 participants (
Cohen, 1988). Because the learning-outcome analyses were based on an analytical sample of 40 participants, an additional sensitivity analysis was conducted for the effective analytical sample. With four groups, one covariate,
α = 0.05, and power = 0.85, the analysis indicated a minimum detectable effect of approximately
f = 0.59. Thus, the ANCOVA analyses were primarily sensitive to large effects, whereas smaller between-condition effects may have remained undetected (
Lakens, 2022;
Shieh, 2020).
The study was conducted at a private primary school in Turkey, where English is taught as a foreign language (EFL). Students’ first language (L1) was Turkish, and English was taught as a compulsory subject within the national curriculum. Students typically received three to four hours of English instruction per week, following a communicative, textbook-based syllabus. Prior to the study, all participants had been learning English for approximately two years, focusing mainly on basic vocabulary, everyday expressions, and simple grammatical structures such as the present simple and subject–verb agreement. They had not received explicit instruction on the Simple Past Tense or Superlative forms before this research. Students’ daily exposure to English occurred primarily through formal classroom instruction, as English is not used for communication in the broader community.
A total of 88 fifth-grade students participated in the study. They included 38 males and 50 females, all aged between 11 and 12, which is considered the young learner stage (
Lyster & Mori, 2006). All learners were assessed at the A1 (Breakthrough) proficiency level, following the Common European Framework of Reference for Languages (CEFR). Students were already divided into four intact classes (22 students each) by school administration. This study employed a cluster-randomized experimental design. As all classes had been established before the study, random assignment was conducted at the class rather than the individual-student level. The four intact classes were randomly assigned to the recast, explicit correction, metalinguistic feedback, and control conditions. Thus, the intact class constituted the unit of assignment, whereas the individual student constituted the unit of analysis for the learning-outcome measures, consistent with previous classroom-based corrective-feedback studies (
Li et al., 2016;
Lira-Gonzales et al., 2024).
All four participating teachers held degrees in English Language Teaching (ELT) and had at least ten years of experience teaching English at the primary school level. To ensure consistency in instructional delivery, teachers were provided with common lesson plans. Prior to the study, permission to conduct the research was obtained from the school. In accordance with the school’s established procedure, parents or legal guardians provided general consent at the beginning of each academic year for their children’s participation in school-based research activities. The school confirmed that this standing parental consent covered participation in the present study. Students were also informed about the study and provided their assent to participate. In addition, the study was approved by the university ethics committee (Approval No. 77082166-604.01.02/96).
2.2. Grammar Forms and Oral CF Types
This study focused on grammar instruction targeting the acquisition of the Simple Past Tense and Superlative forms. These grammatical structures were selected based on their pedagogical and linguistic relevance. First, the learners were at the A1 (Breakthrough) level according to the CEFR, at which these forms are typically introduced. In addition, the study aimed to include grammar structures of differing complexity. The Simple Past Tense was chosen as a morphologically regular and perceptually salient form, consistent with prior corrective feedback research (
Ellis et al., 2006;
Yang & Lyster, 2010). In contrast, comparative and superlative structures are more complex than the simple past because they involve both morphological and syntactic elements and occur less frequently in input (
Ellis, 2007;
Lyster et al., 2013). Moreover, all participants encountered both grammatical structures for the first time during the study.
To provide oral CF during instruction, three feedback types were selected: recasts, explicit correction, and metalinguistic feedback. Each was chosen based on its theoretical and pedagogical relevance to young learners. Recasts were included because they represent the most commonly used type of feedback in classroom interaction (
Zhang et al., 2025) and offer an implicit way of correcting learners’ errors without interrupting the communicative flow. For example, when a learner said, “
He go to school yesterday,” the teacher might respond, “
Yes, he went to school yesterday,” by providing the correct form indirectly while maintaining focus on meaning. Explicit correction was selected because previous research suggests that younger learners benefit more from clear, rule-based instruction than from implicit feedback (
Ellis et al., 2006;
Lichtman, 2016). This feedback type overtly identifies the error and supplies the correct form, often accompanied by a brief grammatical explanation. For instance, if a learner produced “
Mount Everest is more high than other mountains”, the teacher might respond, “
You should say ‘the highest,’ not ‘more high’; we use the superlative form here”. Finally, metalinguistic feedback was selected for its potential to prompt learners to reflect on their errors and engage in self-correction, thereby promoting deeper grammatical awareness (
Ellis et al., 2006). Rather than providing the correct form directly, teachers offered hints or questions that guided learners toward discovering the correct structure themselves. For example, if a student said, “
She go to the park yesterday” the teacher might respond by asking, “
Remember, what happens to the verb in the past tense?” thereby prompting the learner to produce the correct form: “
went”.
2.3. Materials
2.3.1. English Diagnostic Tests (EDTs)
The data were collected through the English Diagnostic Tests (EDTs), which included two distinct tests for each grammar form. The English Simple Past Tense Diagnostic Test (ESPTDT) and the English Superlative Diagnostic Test (ESDT) were developed to assess learners’ performance in the target grammar forms. Each ESPTDT and ESDT included two different tasks; the Oral Picture (OP) description task and the Scenario Interaction (SI) task. These task types were selected as they elicit focused production of specific linguistic targets, consistent with experimental research procedures in CF studies (
Vuono & Li, 2021). All tasks for both ESPTDT and ESDT were developed by the authors. To ensure comparability across test versions, all versions followed the same task format, targeted the same grammatical forms, provided the same number of elicitation opportunities, and used the same scoring criteria. The different versions were reviewed by the three experts who teach “Teaching English to Young Learners” at the undergraduate level and were subsequently piloted with 20 young learners to ensure comparable content, difficulty, and age appropriateness.
Oral Picture (OP) Description Task
Oral Picture (OP) description task was selected, as this type of task is one of the most commonly used methods in previous research to assess the acquisition of grammatical forms in EFL and SLA contexts (
Koizumi & In’nami, 2024;
Lee & Lyster, 2023). Both the ESPTDT and ESDT included two OP tasks: one administered during the pre- and post-tests, and another during the delayed post-test (see
Appendix A). In each OP task, participants were shown a visual prompt depicting a situation and asked to provide an oral description. They were given 30 s to prepare and were expected to produce five sentences incorporating the target grammatical forms. Each task provided five elicitation opportunities, corresponding to five target sentences, with a maximum score of 5. Each grammatically correct use of the target form received 1 point, whereas incorrect, missing, or partially correct target forms received 0 points.
Scenario Interaction (SI) Task
Scenario Interaction (SI) tasks, which depict sequences of actions and events to immerse learners in realistic situations, allow participants to interact using the target language and apply grammatical structures in context. Such scenario-based tasks have been frequently employed in EFL and SLA studies to evaluate learners’ acquisition of grammatical forms (
Rosson & Carroll, 2002;
Wu & Roever, 2025). Considering participants’ proficiency levels, ages, and interests, two different scenarios were developed for ESPTDT and ESDT: one used for the pre- and post-tests, and the other for the delayed post-test (see
Appendix B). Both scenarios for each target grammar form were designed as interactive tasks to effectively engage learners. During the SI tasks, participants were asked to describe the scenario using the target grammar forms in a natural language context. They were given two minutes to prepare and were expected to produce five sentences related to the target structures. The SI task also provided five elicitation opportunities and had a maximum score of 5. Each correct use of the target grammatical form received 1 point, whereas incorrect, missing, or partially correct responses received 0 points.
2.4. Procedure
A preliminary needs analysis was conducted to evaluate the teachers’ knowledge and practices regarding oral CF and FonF instruction. Using an observation form and a semi-structured interview developed by the researchers and validated by experts, each teacher was observed for two class hours and interviewed to assess their awareness of FonF and CF strategies. The results indicated that each teacher primarily relied on one type of feedback in their teaching. Based on the results of the classroom observations and semi-structured interviews, the authors conducted a brief training session that covered the principles and implementation of oral CF and FonF instruction, supported by sample activities and practice tasks. Following the training, the teachers in the experimental groups were randomly assigned to apply one specific oral feedback type during the treatment phase.
During the instructional treatment phase, three teachers in the experimental groups taught the target grammar structures—Simple Past Tense and Superlative—through FonF instruction with one assigned type of oral feedback (recasts, explicit correction, or metalinguistic feedback) in response to learners’ oral grammatical errors. The control group teacher used only FonF instruction but did not provide any oral feedback during the treatment process. Each teacher delivered instruction across 80 min (two course hours) for each grammar structure, totaling 640 min (10.7 h) of classroom instruction across all groups. The same lesson plans, target structures, communicative activities, and instructional time were used across the four conditions to provide comparable opportunities. Thus, the intended difference among the conditions was the type of oral CF provided, or its absence in the control condition. The instructional content and activities were standardized across all classrooms to ensure consistency.
FonF instruction was integrated into meaning-focused lessons that included communicative tasks, such as information gap activities, guided storytelling, and pair or group discussions. For the Simple Past Tense, learners were introduced to regular and common irregular verb forms (e.g., played, watched, went, ate), along with the basic rule of adding -ed for regular verbs. For the Superlative, instruction focused on the use of the + adjective + -est (e.g., the tallest, the fastest) and irregular forms such as the best, the worst. Across all four conditions, FonF instruction included the same planned explanations, examples, and modeling of the target grammatical forms within the communicative activities. These instructional moves were provided as part of the lesson and were not contingent on an individual learner producing an error. Oral CF was operationally distinguished from these instructional activities as a teacher response occurring specifically after a learner produced an erroneous target form. In the three experimental conditions, such errors were followed by the assigned feedback type (recast, explicit correction, or metalinguistic feedback). For example, during a storytelling task, if a student said, “She go to the park yesterday” the teacher providing recasts reformulated the sentence as “Yes, she went to the park yesterday” Similarly, in the explicit correction condition, the teacher directly identified the error and supplied the correct form, while in the metalinguistic condition, the teacher prompted learners to reflect on the rule with questions such as “What happens to the verb in the past tense?”. In contrast, the control-group teacher provided the same planned FonF instruction and modeling but did not correct, reformulate, or prompt learners in response to their individual errors. All instructional sessions were audio-recorded and transcribed for analysis. Implementation fidelity was also examined using the complete classroom recordings and transcripts. The review confirmed that the experimental-group teachers applied their assigned feedback type and that no oral CF was provided in the control condition.
In the evaluation phase, ten students from each class were randomly selected, resulting in a subsample of forty students. This assessment subsample was used because each selected participant completed multiple individual oral assessments covering two grammatical structures at three testing occasions (pre-test, post-test, and delayed test), requiring individual administration, recording, transcription, and scoring. This procedure is consistent with several comparable CF studies that have used relatively modest samples with repeated oral assessments across pre-test, post-test, and delayed-test occasions (
Ellis et al., 2006;
Lyster & Izquierdo, 2009;
Saito & Lyster, 2012). The remaining forty-eight students participated in the classroom instruction but did not complete the individual pre-test, post-test, or delayed-test measures.
All EDT tests were conducted orally to assess learners’ accuracy in using the target grammatical forms (Simple Past Tense and Superlative). Testing sessions were administered at three points: before the treatment (pre-test), immediately after (post-test), and three weeks later (delayed test), allowing for measurement of both short-term and three-week delayed learning effects. The total testing time varied depending on participants’ responses; approximately 320 min for pre-tests (M = 8 min), 280 min for post-tests (M = 7 min), and 300 min for delayed tests (M = 7.5 min), resulting in about 15 h (900 min) of recorded testing data across all sessions.
2.5. Data Analysis
A total of 10.7 h of classroom treatment data, including young learners’ oral responses during feedback episodes, were transcribed and coded using
Lyster and Ranta’s (
1997) oral CF framework. The classroom interaction data were divided among five external researchers, who independently coded their assigned data using the same coding framework. Each feedback episode was coded for the learner error, feedback type (recast, explicit correction, or metalinguistic feedback), and repair outcome (repair or no repair). The first author subsequently reviewed all coded data to ensure consistent application of the coding categories. Any uncertainties or discrepancies identified during this review were re-examined against the coding framework and resolved through discussion and consensus. This multi-researcher coding and review procedure was used to enhance the consistency and transparency of the coding process (
Campbell et al., 2013;
O’Connor & Joffe, 2020).
When a learner failed to modify their utterance following feedback, the response was coded as “no repair”, whereas “repair” referred to a successful self-correction of the target form. For instance, in the metalinguistic feedback group, when a student said, “
Saturday is usually busiest day of the week” the teacher prompted, “
Do we say busiest?”, leading the learner to self-correct to “
the busiest”. This episode was coded as metalinguistic feedback and repair. In the recast group, when a learner said, “
Molly didn’t tidied her room” the teacher reformulated the utterance as “
Molly didn’t tidy her room” providing the correct form implicitly within the flow of conversation. The learner then repeated the corrected version, which was coded as recast and repair. In contrast, in the explicit correction group, the same error was addressed more directly. When a student said, “
Molly didn’t tidied her room” the teacher explicitly corrected both the grammatical and pronunciation error by saying, “
Be careful—not ‘didn’t tidied,’ but ‘didn’t tidy”. Sample coding examples are provided in
Appendix C.
Statistical analyses were conducted using SPSS 31. Fifteen hours of recorded data from the pre-tests, post-tests, and delayed tests were transcribed and analyzed. For RQ1 and RQ2, separate ANCOVAs were conducted for each grammatical structure and assessment task (OP and SI) at the post-test and delayed-test occasions. Feedback condition was entered as the between-subjects factor, and the corresponding pre-test score was included as the covariate. Prior to the analyses, the ANCOVA assumptions of normality of residuals, homogeneity of variances, linearity between the covariate and outcome, and homogeneity of regression slopes were examined. Homogeneity of variances was assessed using Levene’s test, and homogeneity of regression slopes was evaluated through the interaction between the pre-test covariate and feedback condition. No missing data were present in the analyzed assessment sample, and potential outliers were screened before the analyses. When significant group effects were identified, Bonferroni-adjusted pairwise comparisons were conducted to control for multiple comparisons. Statistical significance was set at α = 0.05. Partial eta squared (ηp2) was reported as the effect size for omnibus tests.
To address RQ3, separate mixed-design repeated-measures ANOVAs were conducted for the OP and SI tasks, with grammatical structure and time as within-subject factors and feedback condition as the between-subjects factor. The Grammatical Structure × Feedback interaction was used to test whether feedback effects differed across the two grammatical structures, while the Grammatical Structure × Time × Feedback interaction examined whether these differences varied across measurement occasions. Pillai’s trace was reported for multivariate effects, with α = 0.05. Partial eta squared (ηp2) and 95% confidence intervals were reported for effect sizes and relevant Bonferroni-adjusted comparisons.
3. Results
Table 1 presents the distribution of learner errors, teacher feedback moves, and repair outcomes across all conditions for both the Superlative and Simple Past tense instructional sessions. Because oral CF was provided contingently in response to naturally occurring learner errors, the number of errors and feedback moves was not experimentally equated across conditions. Each learner error was counted once, whereas each teacher corrective response was counted as a separate feedback move. Therefore, when a teacher used more than one corrective move in response to the same error, the number of feedback moves could exceed the number of learner errors. Repair or no-repair was determined at the level of the learner error episode based on the learner’s response following the feedback sequence. The results showed that the number of feedback moves exceeded the number of total learner errors, indicating that teachers frequently provided more than one corrective response for individual errors. Across both grammatical targets, young learners in the explicit correction group consistently produced the fewest total errors, while the recast and metalinguistic feedback groups showed similar error frequencies. The results also indicated that recasts were delivered more frequently than explicit or metalinguistic feedback. Metalinguistic feedback was the second most frequently employed feedback type and also produced strong repair outcomes. In contrast, explicit correction moves were less than recasts or metalinguistic feedback but still led to substantial, though slightly lower, rates of learner repair. These differences reflect variation in naturally occurring learner errors and in the number of corrective moves required within individual feedback episodes. For the Superlative form, recasts led to the highest proportion of successful repairs (98%), followed by explicit correction (89%) and metalinguistic feedback (88%) during the instructional sessions. A similar pattern was observed in the Simple Past Tense instructional sessions, where recasts again produced the highest repair rate (100%), followed by metalinguistic feedback (96%), while explicit correction resulted in a lower but still substantial repair rate (78%).
Table 2 presents the mean scores and standard deviations (
SDs) for learners’ performance on the Superlative and Simple Past Tense forms across the OP and SI tasks at the pre-test, post-test, and delayed-test phases. Preliminary analyses indicated no statistically significant differences among the groups on either grammatical structure at the pre-test stage (
p > 0.05), suggesting that the groups were comparable prior to the instructional intervention. Regarding RQ1, learners in all three feedback conditions demonstrated substantial gains from pre-test to post-test and maintained higher levels of performance at the delayed test than the control group (
p < 0.05). Across both grammatical structures and tasks, the feedback groups consistently outperformed the control group, indicating that the provision of oral CF enhanced learners’ grammar acquisition beyond the effects of FonF instruction alone.
In relation to RQ2, statistically significant effects of feedback type were observed for both grammatical structures. For the Superlative form, ANCOVA results revealed statistically significant differences across the conditions in both tasks. In the OP task, the effect of feedback condition was statistically significant at the post-test, F(3, 35) = 2.95, p = 0.046, ηp2 = 0.20, and delayed test, F(3, 35) = 13.60, p < 0.001, ηp2 = 0.54. Similarly, significant group differences were found in the SI task at the post-test, F(3, 35) = 3.30, p = 0.032, ηp2 = 0.22, and delayed test, F(3, 35) = 3.46, p = 0.027, ηp2 = 0.23. For the Simple Past Tense, the feedback conditions also differed statistically significantly in both the OP and SI tasks at the post-test and delayed-test phases. In the OP task, significant effects were found at the post-test, F(3, 35) = 21.93, p < 0.001, ηp2 = 0.65, and delayed test, F(3, 35) = 39.16, p < 0.001, ηp2 = 0.77. Likewise, significant effects were found in the SI task at the post-test, F(3, 35) = 11.98, p < 0.001, ηp2 = 0.51, and delayed test, F(3, 35) = 5.82, p = 0.002, ηp2 = 0.33. Across both grammatical structures, explicit correction consistently produced the highest post-test and delayed-test scores, followed by recasts and metalinguistic feedback.
Bonferroni post-hoc comparisons further clarified these differences. For the Superlative form in the OP task, the explicit correction group significantly outperformed the control group on the post-test (p = 0.040). At the delayed test, explicit correction yielded significantly higher scores than both the recast group (p = 0.001) and the control group (p < 0.001), while the metalinguistic feedback group also performed significantly better than the control group (p = 0.003). In the SI task, the metalinguistic feedback group outperformed the control group on the post-test (p = 0.047), whereas explicit correction again demonstrated a significant advantage over the control group at the delayed test (p = 0.025). For the Simple Past Tense, all three feedback groups significantly outperformed the control group in both the OP and SI tasks at the post-test and delayed test (all p’s < 0.05). Across the RQ2 Bonferroni comparisons, explicit correction showed a significant advantage over the other feedback conditions (all ps < 0.05).
Regarding RQ3, for the OP task, repeated-measures ANOVA results indicated no significant main effect of grammatical structure, Pillai’s trace = 0.061, F(1, 36) = 2.36, p = 0.133, ηp2 = 0.061. However, the Grammatical Structure × Feedback interaction was significant, Pillai’s trace = 0.198, F(3, 36) = 2.97, p = 0.045, ηp2 = 0.198, indicating that the effects of the feedback conditions differed between the Simple Past Tense and Superlative forms. The Grammatical Structure × Time × Feedback interaction was also significant, Pillai’s trace = 0.372, F(6, 72) = 2.72, p = 0.019, ηp2 = 0.185, indicating that these differences varied across measurement occasions. A significant main effect of feedback condition was also observed, F(3, 36) = 30.44, p < 0.001, ηp2 = 0.717. Bonferroni comparisons showed that all three feedback groups significantly outperformed the control condition: recasts (MD = 0.95, 95% CI [0.44, 1.46], p < 0.001), explicit correction (MD = 1.67, 95% CI [1.16, 2.17], p < 0.001), and metalinguistic feedback (MD = 1.23, 95% CI [0.73, 1.74], p < 0.001). Explicit correction also significantly outperformed recasts (MD = 0.72, 95% CI [0.21, 1.22], p = 0.002). No significant differences were found between explicit correction and metalinguistic feedback (MD = 0.43, 95% CI [−0.07, 0.94], p = 0.132) or between recasts and metalinguistic feedback (MD = −0.28, 95% CI [−0.79, 0.22], p = 0.758).
For the SI task, there was no significant main effect of grammatical structure, Pillai’s trace = 0.001, F(1, 36) = 0.02, p = 0.890, ηp2 = 0.001. The Grammatical Structure × Feedback interaction was also not significant, Pillai’s trace = 0.174, F(3, 36) = 2.52, p = 0.073, ηp2 = 0.174, indicating that the effects of the feedback conditions did not differ statistically between the Simple Past Tense and Superlative forms. Similarly, the Grammatical Structure × Time × Feedback interaction was not significant, Pillai’s trace = 0.110, F(6, 72) = 0.70, p = 0.654, ηp2 = 0.055, indicating that this pattern did not significantly vary across measurement occasions. In contrast, the main effect of feedback condition was statistically significant, F(3, 36) = 12.25, p < 0.001, ηp2 = 0.505. Bonferroni comparisons showed that all three feedback groups significantly outperformed the control condition: recasts (MD = 0.83, 95% CI [0.21, 1.45], p = 0.004), explicit correction (MD = 1.28, 95% CI [0.66, 1.90], p < 0.001), and metalinguistic feedback (MD = 0.98, 95% CI [0.36, 1.60], p < 0.001). No significant differences were found among the three feedback conditions (all ps ≥ 0.301).
4. Discussion
This study investigated the effects of different types of oral CF on young learners’ grammar acquisition within FonF instruction in an EFL classroom context. Regarding RQ1, the results showed that learners who received oral CF outperformed those in the control group at both the post-test and delayed-test phases. This finding suggests a potential benefit of combining oral CF with FonF instruction in young learner classrooms. One possible explanation is that oral CF helps learners notice gaps between their own production and target-like forms (
Li, 2018;
Lyster et al., 2013;
Nassaji, 2016;
Parlak, 2024). Although the control group also improved from pre-test to post-test, these gains were less consistent and less sustained than those observed in the feedback conditions. This finding suggests that oral CF may further strengthen and sustain grammar learning by providing immediate, interactionally embedded responses to learners’ errors while only FonF instruction may support learners’ short-term attention to grammatical form. However, given that each condition was represented by a single teacher and intact class, these differences cannot be attributed exclusively to oral CF. Teacher characteristics, classroom dynamics, and group composition may also have contributed to the observed outcomes. The findings also support the view that oral CF can function as a classroom-based instructional strategy for promoting language development during communicative interaction (
Lira-Gonzales et al., 2024;
Tan et al., 2024;
Zhang et al., 2025). Furthermore, the findings suggest that integrating oral CF into FonF instruction may be beneficial in primary EFL classrooms, as corrective feedback can reinforce learners’ attention to grammatical forms during meaningful communication and provide timely support for the development of grammatical accuracy (
Afitska, 2015;
Ellis, 2016;
Lyster, 2015;
Saito & Lyster, 2012).
With respect to RQ2, the findings revealed differences among the three feedback types. The results showed that learners in the explicit-correction condition generally achieved higher scores than those in the other feedback conditions. This finding may reflect the greater salience and clarity of explicit correction for young learners. Such explicitness may facilitate noticing by helping learners recognize the gap between their own production and the target language more readily. In young learner classrooms, this clarity may be especially important, as learners may not always perceive the corrective intent of more implicit feedback moves such as recasts. Accordingly, explicit correction may reduce ambiguity and provide young learners with more accessible information about both the location and nature of their errors (
Lyster et al., 2013;
Lichtman, 2016). These findings further support previous research suggesting that more explicit forms of oral CF promote noticing, form awareness, learner uptake, and grammatical development (
Lichtman, 2016;
Yilmaz, 2013;
Zhang et al., 2025). However, this interpretation should remain tentative, as teacher characteristics and classroom dynamics may also have contributed to the observed differences among the feedback conditions.
Metalinguistic feedback also produced strong learning outcomes, particularly in the delayed tests. This suggests that prompting learners to think about a grammatical rule may support deeper processing and longer-term retention. Unlike explicit correction, metalinguistic feedback does not immediately provide the correct form; instead, it encourages learners to reflect on the rules and generate the correction themselves. This may help learners develop stronger form–meaning connections and more durable grammatical knowledge. This finding is consistent with previous research on output-prompting feedback, which suggests that prompts can promote learner-generated repair and deeper cognitive engagement with the target form (
Saito & Lyster, 2012;
Sheen & Ellis, 2011). In the context of young learners, however, the effectiveness of metalinguistic feedback may depend on whether the prompts are simple, age-appropriate, and closely connected to the communicative activity.
Recasts were also effective during the instructional phase, where they produced high rates of learner repair. These repairs reflected successful immediate uptake, indicating that learners were able to notice the corrective information and produce the target form during interaction. This result suggests that recasts may support immediate correction while maintaining the flow of communication in young learner classrooms. The high repair rates further indicate that recasts might be beneficial when target forms are sufficiently salient and learners are developmentally ready to notice the corrective information. However, despite their effectiveness in promoting immediate repair, the delayed-test results showed that the delayed-test performance of recasts were generally weaker than those of explicit correction. This may be because recasts are often implicit and may not always be perceived by young learners as corrective, especially when they are embedded naturally in ongoing communication. This finding is also consistent with previous research showing that recasts are frequently used in classrooms, but their implicit nature may reduce their salience and limit their delayed impact on acquisition (
Brown, 2016;
Lyster & Saito, 2010;
Tan et al., 2024;
Vuono & Li, 2021;
Zhang et al., 2025). However, the observed differences among the oral feedback types should be interpreted with caution, as each condition was implemented by a different teacher in a single intact class. Therefore, teacher-related factors, classroom dynamics, and group composition may also have contributed to the differences observed across the feedback conditions.
Despite the advantage of explicit correction over the other feedback types in the RQ2 analyses, the mixed-design analyses for RQ3 did not reveal a consistent advantage of explicit correction over both recasts and metalinguistic feedback across the two assessment tasks. The findings for RQ3 showed that the effectiveness of oral CF varied according to the target grammatical structure and assessment task. For the OP task, the significant Grammatical Structure × Feedback interaction indicated that the relative effects of the feedback conditions differed between the Simple Past Tense and Superlative forms, while the significant Grammatical Structure × Time × Feedback interaction showed that these differences also varied across measurement occasions. In contrast, for the SI task, neither interaction was statistically significant, indicating that the effects of the feedback conditions did not differ between the Simple Past Tense and Superlative forms across the measurement occasions. The findings did not indicate a general advantage of oral CF for one grammatical structure over the other; rather, structure-related differences emerged only in the OP task and not in the SI task. This finding is consistent with previous research showing that CF effectiveness could vary according to the linguistic feature being targeted and its characteristics (
Yilmaz & Granena, 2021;
Sato & Loewen, 2018), as well as evidence that feedback-related outcomes may differ across target structures and outcome measures (
Yilmaz & Granena, 2021). Moreover, the significant interaction observed in the OP task may reflect differences in the form–meaning characteristics of the two target structures. Previous research has also identified grammatical complexity and form–meaning transparency as factors that can shape responsiveness to CF (
Yilmaz & Granena, 2021;
Sato & Loewen, 2018). Similarly,
Yang and Lyster (
2010) found that feedback effects differed for regular and irregular English past-tense forms. In the present study, the Simple Past Tense instruction included regular -
ed forms and common irregular verbs, whereas the Superlative required learners to coordinate the definite article with -
est marking and irregular forms such as the best and the worst. These linguistic differences may have influenced how readily learners noticed and used corrective information during the relatively structured OP task. The absence of corresponding interactions in the SI task, however, suggests that such structure-related differences may become less apparent when learners produce the target forms in a more contextualized interaction. This task-dependent pattern is consistent with previous CF research showing that observed feedback effects can vary according to the type of outcome measure used (
Lira-Gonzales et al., 2024;
Lyster & Saito, 2010).
6. Implications and Limitations
This study offers several implications for grammar teaching in young learner EFL classrooms. First, the results suggest that integrating oral CF into FonF instruction could enhance the effectiveness of communicative grammar lessons. Teachers may benefit from incorporating systematic feedback into communicative classroom activities rather than relying solely on exposure to target forms. Second, the results indicate that teachers could carefully consider the type of feedback they provide. Explicit correction appeared to be particularly effective in promoting immediate gains and maintaining these gains at the three-week delayed test, suggesting that young learners may benefit from clear and direct feedback that makes target forms highly salient. Explicit correction showed a significant advantage over the other feedback types, suggesting that young learners may benefit from clear and direct feedback. However, this advantage was not consistent across the RQ3 mixed analyses and should therefore be interpreted in relation to the assessment task. At the same time, metalinguistic feedback may encourage learners to reflect on grammatical rules and develop greater awareness of language forms, whereas recasts may be especially useful for maintaining communicative flow and supporting immediate uptake during interaction. In addition, feedback strategies could be selected not only according to learners’ age and proficiency level but also according to the grammatical structure being taught. Different linguistic forms may place different cognitive demands on learners and therefore may respond differently to particular feedback strategies. Teachers may benefit from adopting a flexible approach to oral CF, adjusting their feedback practices according to instructional goals, learner needs, and the characteristics of the target structure. In addition, teacher- and classroom-level factors might have influenced the observed differences among the feedback.
The study also has several limitations. First, although the four intact classes were randomly assigned to the instructional conditions, only one class and one teacher represented each condition. Consequently, treatment effects could not be separated from possible teacher- and class-level influences, such as teacher characteristics, classroom dynamics, or group composition. This limits the strength of causal conclusions regarding differences among the feedback types. Future studies could include multiple teachers and classes within each condition. Second, although 88 students participated in the instructional phase, the individual learning-outcome analyses were based on a randomly selected subsample of 40 students. The sensitivity analysis indicated that the ANCOVAs were primarily sensitive to relatively large effects, suggesting that smaller between-condition differences may have gone undetected. The limited sample size may also have reduced sensitivity for detecting the interaction effects examined in the mixed-design repeated-measures analyses; therefore, particularly nonsignificant interaction effects should be interpreted cautiously. Future studies could include larger assessment samples and account for the classroom-level structure when planning sample size. Third, another limitation concerns the assessment of inter-rater agreement for the classroom-interaction coding. Although the coded data were reviewed and discrepancies were resolved through discussion and consensus, a formal agreement coefficient was not reported. Future studies could independently double-code a proportion of the data and report an appropriate inter-rater agreement coefficient. Fourth, the participants were 11–12-year-old learners, representing the upper end of the young learner age range. As developmental differences may influence how learners perceive and respond to oral CF, the findings may not be generalizable to younger children. Fifth, the delayed post-test was administered three weeks after the intervention, limiting conclusions about the long-term durability of the observed effects. Finally, the study focused on two grammatical structures, the Simple Past Tense and Superlative forms. As the effectiveness of oral corrective feedback may vary according to the linguistic characteristics of target forms, caution is warranted when generalizing the findings to other grammatical features.