The Paradox of Enjoyment: Unpacking the Relationships Among Enjoyment, Engagement, Burnout, and L2 Achievement Among Left-Behind Students
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe present study is a noteworthy undertaking, looking into a complex set of relationships among foreign language enjoyment (FLE), learning engagement, learning burnout, and academic achievement of left-behind children. The study is both conceptually and methodologically robust, yielding unexpected findings likely to interest the journal’s readers. The manuscript is coherent and logical, grounded in relevant literature, and supported by a large sample of students. The research questions are clear, and the conceptual model is compelling.
Although the study is overall strong, the authors may consider a few suggestions for improvement. One suggestion concerns the adequacy of the measurement and structural models. Although the SEM approach is appropriate for the proposed research questions, the reported model-fit indices are problematic.
e.g., Lines 461-470: The final structural model reports CFI = .944, TLI = .903, RMSEA = .127, and SRMR = .088. Although the CFI and TLI are acceptable, the RMSEA and SRMR are problematic. The authors could either improve the model specification or present a more cautious interpretation of the findings.
Lines 434–440: The CFA results for FLE are problematic. The reported fit indices, CFI = .891, TLI = .860, and RMSEA = .145, do not support a strong measurement model. The FLE scale fit should not be described as acceptable without additional justification. They should report factor loadings and inspect problematic items.
The conclusions are aligned with the study’s findings, but some claims could be expressed more cautiously given that the study is cross-sectional. The authors claim that FLE is a “double-edged emotional force,” but the evidence is based on a model with problematic fit indices. The authors should acknowledge the limitations of the data.
Overall, I find the manuscript timely, with a relevant topic and an interesting theoretical contribution.
Comments for author File:
Comments.pdf
Author Response
1) Lines 461-470: The final structural model reports CFI = .944, TLI = .903, RMSEA = .127, and SRMR = .088. Although the CFI and TLI are acceptable, the RMSEA and SRMR are problematic. The authors could either improve the model specification or present a more cautious interpretation of the findings.
We thank the reviewer for this insightful concern and suggestion. We totally agree that the initial model fits were not ideal, and we have conducted additional analyses to address these concerns.
Based on refined measurement model of FLE, we re-estimated the structural model while retaining the same residual covariances. The revised SEM demonstrated considerably improved model fit: CFI = .962, TLI = .929, SRMR = .067, RMSEA = .096, 90% CI [0.830, 0.109], with all indices now meeting or approaching the recommended thresholds. All hypothesized path coefficients remained stable in direction and significance, confirming the robustness of our conclusions.
We have revised the Methods (Pages 11) and Results (Pages 11–13) sections accordingly. We believe these revisions have considerably enhanced the methodological rigor and transparency of our psychometric reporting
2) Lines 434–440: The CFA results for FLE are problematic. The reported fit indices, CFI = .891, TLI = .860, and RMSEA = .145, do not support a strong measurement model. The FLE scale fit should not be described as acceptable without additional justification. They should report factor loadings and inspect problematic items.
We sincerely thank the reviewer for raising this important point. We acknowledge that the RMSEA of the initial model was indeed high, and we have carefully addressed this concern in the revised manuscript.
We re-examined the FLE measurement model. The initial model indeed showed poor fit (CFI = .891, TLI = .860, RMSEA = .145). Inspection of modification indices indicated local dependencies between FLE3–FLE4 and FLE8–FLE9 due to similar item wording. Following standard CFA practices (Kline, 2023), we allowed their residual covariances to correlate. This theoretically justified adjustment substantially improved model fit: CFI = .950, TLI = .941, SRMR = .040, RMSEA = .099, 90% CI [.089, .109]. All standardized factor loadings were statistically significant (p < .001) and ranged from .62 to .90, and the scale demonstrated excellent internal consistency (Cronbach’s α = .926).
While the RMSEA (0.099) remains slightly above the strict .08 cutoff, we note that the CFI and TLI well exceed conventional thresholds, the SRMR is well within the acceptable range (<.08), and the lower bound of the 90% confidence interval for RMSEA (0.089) is very close to the recommended cutoff. Taken together, these indices suggest that the overall model fit is acceptable.
Specifically addressing the reviewer's suggestions regarding factor loadings and problematic items:
- Factor loadings: All standardized factor loadings were statistically significant (p < .001) and ranged from .62 to .90, with the majority exceeding .70, indicating adequate item quality.
- Problematic items: We examined modification indices and item residuals. No items exhibited excessively low loadings (< .40) that would warrant deletion. The residual covariances between FLE3–FLE4 and FLE8–FLE9 were released based on both statistical and content considerations (similar wording).
We have revised the Measures section accordingly (Page 9). We believe these revisions have substantially strengthened the transparency and rigor of our psychometric analyses.
3) The conclusions are aligned with the study’s findings, but some claims could be expressed more cautiously given that the study is cross-sectional. The authors claim that FLE is a “double-edged emotional force,” but the evidence is based on a model with problematic fit indices. The authors should acknowledge the limitations of the data.
We thank the reviewer for this insightful concern. We totally agree that our cross-sectional design does not support causal inferences. Accordingly, we have systematically revised our manuscript to replace causal language with associational terminology throughout the manuscript. In addition, we have explicitly acknowledged the limitations of the cross-sectional design in the discussion section (Page. 17)
Reviewer 2 Report
Comments and Suggestions for AuthorsWhat I like about this manuscript is that it addresses an important and underexplored population, offering a potentially original examination of enjoyment, engagement, burnout, and L2 achievement. However, I noticed several major issues (before it can be published):
1) The measurement and structural model fit raise serious concerns. The FLE CFA shows poor fit (CFI = .891, TLI = .860, RMSEA = .145), and the final SEM also shows problematic fit (RMSEA = .127; SRMR = .088), yet the manuscript describes these models as acceptable or fitting well.
2) The manuscript uses causal language such as “predicts,” “pathway,” and “inhibitory effect” despite a cross-sectional design. The mediation findings should be interpreted as associational.
3) The abstract contains an apparent inconsistency: it initially states that LE positively predicts achievement, whereas the SEM later reports a nonsignificant LE → achievement path.
4) The theoretical explanation of the positive FLE–burnout association is interesting but remains speculative. Alternative explanations, measurement overlap, omitted variables, and the small effect size should be discussed more cautiously.
Overall, this is a valuable study, with several notable strengths, including a clear research focus, a substantial sample, and strong engagement with recent scholarship. The examination of enjoyment, engagement, and burnout within a unified model is particularly interesting, and the findings concerning the coexistence of enjoyment and burnout offer a potentially original contribution to the literature. I think this manuscript has strong potential and could make a useful contribution to research on second language achievement.
Comments on the Quality of English Language/
Author Response
1) The measurement and structural model fit raise serious concerns. The FLE CFA shows poor fit (CFI = .891, TLI = .860, RMSEA = .145), and the final SEM also shows problematic fit (RMSEA = .127; SRMR = .088), yet the manuscript describes these models as acceptable or fitting well.
We sincerely thank the reviewer for this critical and constructive comment. We totally agree that the initial model fits were not ideal, and we have conducted additional analyses to address these concerns.
First, we re-examined the FLE measurement model. The initial model indeed showed poor fit (CFI = .891, TLI = .860, RMSEA = .145). Inspection of modification indices indicated local dependencies between FLE3–FLE4 and FLE8–FLE9 due to similar item wording. Following standard CFA practices (Kline, 2023), we allowed their residual covariances to correlate. This theoretically justified adjustment substantially improved model fit: CFI = .950, TLI = .941, SRMR = .040, RMSEA = .099, 90% CI [.089, .109]. All standardized factor loadings were statistically significant (p < .001) and ranged from .62 to .90, and the scale demonstrated excellent internal consistency (Cronbach’s α = .926).
While the RMSEA (0.099) remains slightly above the strict .08 cutoff, we note that the CFI and TLI well exceed conventional thresholds, the SRMR is well within the acceptable range (<.08), and the lower bound of the 90% confidence interval for RMSEA (0.089) is very close to the recommended cutoff. Taken together, these indices suggest that the overall model fit is acceptable.
Second, based on this refined measurement model, we re-estimated the structural model while retaining the same residual covariances. The revised SEM demonstrated considerably improved model fit: CFI = .962, TLI = .929, SRMR = .067, RMSEA = .096, 90% CI [0.830, 0.109], with all indices now meeting or approaching the recommended thresholds. All hypothesized path coefficients remained stable in direction and significance, confirming the robustness of our conclusions.
We have revised the Methods (Pages 9 and 11) and Results (Pages 11–13) sections accordingly. We believe these revisions have substantially strengthened the transparency and rigor of our psychometric analyses.
2) The manuscript uses causal language such as “predicts,” “pathway,” and “inhibitory effect” despite a cross-sectional design. The mediation findings should be interpreted as associational.
We thank the reviewer for this valuable and insightful suggestion. We fully agree that our cross-sectional design does not support causal inferences. Accordingly, we have systematically revised our manuscript to replace causal language with associational terminology throughout the manuscript. Specifically: (1) predicts has been changed to associational terms such as is associated with, is linked to, and is connected to; (2) pathway has been revised to indirect association to reflect the associational nature of our mediation analysis; and (3) inhibitory effect has been rephrased as negative indirect effect to accurately describe the direction of the association without implying causality. In addition, we have explicitly acknowledged the limitations of the cross-sectional design in the discussion section (Page. 17)
3) The abstract contains an apparent inconsistency: it initially states that LE positively predicts achievement, whereas the SEM later reports a nonsignificant LE → achievement path.
We apologize for our misleading wording and thank the reviewer for pointing out this inconsistency. In the original abstract, 'predicts' inappropriately referred to the bivariate correlation, which contradicted the non-significant direct LE→achievement path. We have revised it to 'positively correlated' to accurately reflect the correlational matrix. In addition, following the reviewer's suggestion, we have replaced causal language with associational language throughout the abstract (Page. 1).
4) The theoretical explanation of the positive FLE–burnout association is interesting but remains speculative. Alternative explanations, measurement overlap, omitted variables, and the small effect size should be discussed more cautiously.
We sincerely thank the reviewer for this constructive comment. We fully agree that our initial interpretation of the positive FLE–burnout association was somewhat speculative and that we failed to adequately discuss the methodological caveats. To address this, we have substantially revised the Discussion to offer a more balanced and cautious interpretation. Specifically, we have added a new paragraph (Page 16, third paragraph) to more carefully discuss other possible explanations, including: (1) explicitly addressing the potential issue of measurement overlap; (2) acknowledging the threat of omitted variable bias (e.g., personality traits, academic motivation, academic resilience, or perceived teacher support); and (3) emphasizing the practical implications of the small effect size, clarifying that statistical significance does not equate to practical salience. We believe these revisions make our theoretical claims more grounded and our conclusions more transparent
Reviewer 3 Report
Comments and Suggestions for AuthorsThis study investigates the relationships among foreign language enjoyment (FLE), learning engagement (LE), learning burnout (LB), and L2 academic achievement among 816 Chinese left-behind children (LBC) in a rural boarding school. Using structural equation modeling, the authors found that FLE positively predicted both LE and LB, while LE did not significantly predict achievement and LB negatively predicted achievement. The results challenge the unidirectional positive view of FLE in L2 learning, revealing a "paradox of enjoyment" where positive emotions may act as a double-edged sword for this vulnerable population. The authors argue that fostering enjoyment alone is insufficient without targeted interventions to alleviate burnout and translate engagement into actual success. This is a well-conceptualised and methodologically sound study that addresses a genuinely important and under-researched population—left-behind children in Chinese EFL contexts. The theoretical grounding is robust, the sample size is impressive, and the findings offer meaningful contributions to the literature on emotions, engagement, and burnout in L2 learning. The identification of the "paradox of enjoyment" is both novel and theoretically significant, challenging the prevailing positive psychology narrative. The manuscript is well-written and already of high quality. I recommend minor revision with a few suggestions to further strengthen the work.
1#
Method – The fit indices for the FLE scale require attention. The RMSEA for the FLE scale is 0.145, which exceeds the recommended cutoff of 0.08. While the authors note that this may be attributable to the large sample size, RMSEA is actually less sensitive to sample size than χ². Consider reporting alternative fit indices or exploring whether model modifications (e.g., correlating error terms or removing problematic items) could improve fit. A brief explanation of why this elevated RMSEA is acceptable would also help readers.
2#
Method – The Cronbach’s α for the “sense of inadequacy” subscale of LB is low. The α coefficient for this dimension is 0.561, which is below the generally accepted threshold of 0.70. While the authors note that this is acceptable when the total scale demonstrates good reliability, it would be prudent to acknowledge this as a limitation more explicitly and consider whether this subscale should be interpreted with caution. Adding a sentence in the limitations section would strengthen the manuscript’s transparency.
3#
Participants – The age range of participants is not clearly reported. The average age is reported (14.10 years, SD = 0.85), but the minimum and maximum ages are not provided. Given that middle school students typically range from 12 to 15 years, it would be helpful to confirm this range. Please add the age range to Table 1 or the participant description.
4#
Participants – The grade levels of participants are not specified. The manuscript states that participants were from a “rural private boarding school” and that English exam scores were “standardized within each grade level,” but it does not specify which grades were included. Please clarify whether participants were from Grades 7, 8, 9, or a combination, as grade level may affect English proficiency and emotional experiences.
5#
Discussion – The interpretation of the non-significant LE → AA path could be strengthened. The authors suggest that FLE may foster “affective engagement” rather than “deep cognitive engagement,” which is plausible. However, the discussion could be further enriched by considering alternative explanations, such as the possibility that LE was not measured in a way that captures the specific forms of engagement most relevant to test performance, or that the relationship is moderated by other variables (e.g., self-regulation, teacher support). Acknowledging these possibilities would add nuance.
Comments on the Quality of English LanguageEnglish is great.
Author Response
1#
Method – The fit indices for the FLE scale require attention. The RMSEA for the FLE scale is 0.145, which exceeds the recommended cutoff of 0.08. While the authors note that this may be attributable to the large sample size, RMSEA is actually less sensitive to sample size than χ². Consider reporting alternative fit indices or exploring whether model modifications (e.g., correlating error terms or removing problematic items) could improve fit. A brief explanation of why this elevated RMSEA is acceptable would also help readers.
We sincerely thank the reviewer for this expert and constructive comment. We acknowledge that the RMSEA of the initial model was indeed high, and we have carefully addressed this concern in the revised manuscript.
We re-examined the FLE measurement model. The initial model indeed showed poor fit (CFI = .891, TLI = .860, RMSEA = .145). Inspection of modification indices indicated local dependencies between FLE3–FLE4 and FLE8–FLE9 due to similar item wording. Following standard CFA practices (Kline, 2023), we allowed their residual covariances to correlate. This theoretically justified adjustment substantially improved model fit: CFI = .950, TLI = .941, SRMR = .040, RMSEA = .099, 90% CI [.089, .109].
While the RMSEA (0.099) remains slightly above the strict .08 cutoff, we note that the CFI and TLI well exceed conventional thresholds, the SRMR is well within the acceptable range (<.08), and the lower bound of the 90% confidence interval for RMSEA (0.089) is very close to the recommended cutoff. Taken together, these indices suggest that the overall model fit is acceptable.
We have revised the Measures section accordingly (Page 9). We believe this revision has significantly strengthened the psychometric presentation.
2#
Method – The Cronbach’s α for the “sense of inadequacy” subscale of LB is low. The α coefficient for this dimension is 0.561, which is below the generally accepted threshold of 0.70. While the authors note that this is acceptable when the total scale demonstrates good reliability, it would be prudent to acknowledge this as a limitation more explicitly and consider whether this subscale should be interpreted with caution. Adding a sentence in the limitations section would strengthen the manuscript’s transparency.
We appreciate the reviewer's careful observation and suggestion. We acknowledge that the Cronbach's α of .561 for the "sense of inadequacy" subscale is below the conventional .70 threshold.
Following the reviewer's recommendation, we have explicitly stated this limitation in the Discussion section and noted that interpretations involving this specific subscale should be made with caution (Page 17). We believe this addition improves the transparency of the manuscript.
3#
Participants – The age range of participants is not clearly reported. The average age is reported (14.10 years, SD = 0.85), but the minimum and maximum ages are not provided. Given that middle school students typically range from 12 to 15 years, it would be helpful to confirm this range. Please add the age range to Table 1 or the participant description.
We thank the reviewer for this careful and constructive suggestion. Based on this reviewer’s suggestion, we have added the age range of the participants to the manuscript (Page 8).
4#
Participants – The grade levels of participants are not specified. The manuscript states that participants were from a “rural private boarding school” and that English exam scores were “standardized within each grade level,” but it does not specify which grades were included. Please clarify whether participants were from Grades 7, 8, 9, or a combination, as grade level may affect English proficiency and emotional experiences.
We thank the reviewer for this valuable and insightful observation. We agree that grade level is an important factor that may relate to English proficiency and emotional experiences.
Accordingly, we have added the grade distribution to the Participants section (Page. 8). Specifically, participants were recruited from Grades 7 (31.2%, n = 255), 8 (27.3%, n = 223), and 9 (41.5%, n = 338).
5#
Discussion – The interpretation of the non-significant LE → AA path could be strengthened. The authors suggest that FLE may foster “affective engagement” rather than “deep cognitive engagement,” which is plausible. However, the discussion could be further enriched by considering alternative explanations, such as the possibility that LE was not measured in a way that captures the specific forms of engagement most relevant to test performance, or that the relationship is moderated by other variables (e.g., self-regulation, teacher support). Acknowledging these possibilities would add nuance.
We thank the reviewer for this insightful, valuable and constructive suggestion. We totally agree that the interpretation of the non-significant LE → AA path can be further enriched by considering alternative explanations, and we have revised the Discussion section accordingly.
Specifically, we have added discussion of two additional possibilities:
Measurement-related explanation: The measurement of LE may not capture the specific facets of engagement most relevant to test performance (e.g., deep cognitive strategy use vs. general behavioral engagement). The multidimensional nature of engagement and its differential links to achievement outcomes are now acknowledged.
Potential moderating mechanisms: The relationship between LE and achievement is moderated by other variables, such as self-regulation strategies or perceived teacher support. Future research could explore these potential moderators to clarify under what conditions LE contributes more strongly to academic outcomes.
These additions have been incorporated into the Discussion section (Page 15, second paragraph). We believe they add valuable nuance to our interpretation of the findings.
Round 2
Reviewer 2 Report
Comments and Suggestions for Authors/
