Middle School Students’ Interest and Self-Efficacy in a One-Day Informal STEM Learning Experience
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThis paper reports on the NSF STEM Day event at UNLV, a one-day informal learning experience for middle school students that combined hands-on robotics and engineering demonstrations on campus with a visit to the Museum of Illusions. The paper makes a useful contribution to the informal STEM learning literature, particularly in documenting a university outreach model that integrates near-peer mentorship and real-world applications. Nevertheless, here are serious methodological and reporting issues that need to be addressed before publication.
The most significant problem is the treatment of the pre- and post-survey data as if they constitute a matched pre-post design when they do not. The pre-survey yielded N=58 responses from 150 attendees (a 39% response rate), and the post-survey yielded N=34 responses (23% of attendees). These are two different, unmatched samples drawn from the same event. The authors acknowledge this in the Data Analysis section, but the problem runs deeper than they indicate. Throughout the abstract, results, and discussion, the paper makes claims about "increases" in student interest and self-efficacy that the design simply cannot support. When you compare mean STEM interest scores of 4.43 (pre) and 4.40 (post) in two unmatched groups of different sizes, you have no basis for claiming interest increased or decreased. The authors should reframe all comparative language as descriptive observations about two samples rather than evidence of change.
Relatedly, the three self-efficacy items (Q7, Q8, Q9) appear only in the post-survey. There is no pre-event baseline for self-efficacy at all. The abstract states that "participants reported enhanced self-efficacy in STEM," but there is no comparison point that would allow this claim. What the data actually show is that post-event self-efficacy scores were generally positive. The authors should correct this framing throughout, especially in the abstract and conclusion.
A ceiling effect is also present in the STEM interest data that the paper does not adequately address. Pre-event interest was already at 90.7% "interested," with a mean of 4.43 on a 5-point scale. Detecting meaningful change in a group that is already near the top of the scale is statistically unlikely, and this matters for interpreting the essentially flat post-event scores. The paper would benefit from discussing this limitation more directly.
The interaction effect reported in Section 4.2 raises concern: F = 10.07, p = .011, partial eta squared = .93 for the Grade Level x Ethnicity interaction on composite self-efficacy scores. A partial eta squared of .93 is extraordinarily large, essentially indicating that 93% of variance in self-efficacy scores is explained by this interaction. With a post-survey sample of N=34 distributed across multiple grade and ethnicity cells, many cells almost certainly contain two or three participants. Factorial ANOVA results under these conditions are unreliable and should not be reported without a clear caveat about cell sizes. The authors should provide cell-level n counts for all reported interactions. If cell sizes are as small as suspected, these interaction effects should either be removed or discussed with substantially more caution.
The significant Gender x Grade Level interaction (p = .031) reported in Section 5.3 does not appear in the Results section at all and lacks any supporting statistics. If this is a finding, it needs to be in the results with the corresponding F value, degrees of freedom, and effect size.
The paper also conflates the ANOVA p-value (p = .023 for grade level on Q7) with the Tukey post-hoc p-value (p = .048 for the 6th vs. 8th grade comparison). The abstract reports p = .048 as the primary significance indicator for grade-level differences, but this is the post-hoc pairwise value, not the overall ANOVA result. The abstract should report the ANOVA result (F = 3.76, p = .023) or clarify that the .048 figure is from the Tukey procedure.
The paper states two research questions in the Introduction but references a "Research Question 3" in Section 5.3. This is a clear error. Either a third research question needs to be added to the Introduction, or the discussion section needs to be revised to eliminate the reference to a nonexistent question.
There is also a structural mismatch between the Results and Discussion. Section 4 (Results) begins with three pages of photographs before presenting any data. Figures 3, 4, and 5 are photographs of students at the event. While these provide context, they belong in the Methods or Event Description sections, not the Results. The reader expects data in a Results section.
The mentor survey data in Table 2 is interesting and adds texture to the paper, but its relationship to the two stated research questions is unclear. The research questions are about student outcomes. If the mentor development component is presented as a finding, it should either be tied to a stated research question or repositioned as a supplementary observation. Some of the mentor responses in the table also contain informal language that reads more like raw data than analyzed results ("I found out I'm really good at working with kids"; "the freestyler software is pretty annoying"). These may be authentic quotes, but the paper does not frame them as quotations from a qualitative analysis. If they are direct quotes, label them as such and provide some analytical commentary. If they are paraphrases, revise accordingly.
The abstract and conclusion describe the event as boosting STEM engagement and confidence, but Section 4.1 shows that willingness to pursue a STEM major or career actually declined slightly in "yes" responses (41.8% to 37.9%). The paper addresses this by noting that "maybe" responses increased, but the framing in the abstract does not reflect this nuance. The abstract should be revised to accurately represent what the data show.
The paper presents the Museum of Illusions as a "field trip science learning experience" in the abstract and throughout. The museum is a commercial entertainment venue. The connection to science learning is asserted but not demonstrated. Post-event survey results show that students could distinguish scientific principles from optical illusions, but this does not clearly establish what learning specifically occurred at the museum versus at the UNLV demonstrations. The claims about museum-based learning should be more carefully scoped.
The instrument used for the study has no discussion of validity or reliability. The surveys appear to have been developed by the authors. At minimum, the paper should note whether these items were adapted from validated instruments in the field, and if not, acknowledge instrument validation as a limitation.
Comments on the Quality of English LanguageThere are several proofreading errors that should be corrected. On page 11, the sentence reads "certain projects had more than pone mentor" (should be "one"). On page 17, "analytical pparatus" should be "analytical apparatus." The manuscript would benefit from another careful proofread.
The abstract uses the phrase "combining participation in demonstration and hands-on learning of captivating robotics and IoT projects" which is awkward. The language throughout is generally clear but could be edited further.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThank you for the opportunity to review this manuscript. The study addresses a potentially worthwhile educational initiative and raises several interesting issues. However, I consider that a number of fundamental concerns relating to the study's significance, conceptual framing, measurement, and interpretation of findings should be carefully addressed before the manuscript is suitable for publication.
First, the significance of the study and the research gap are not sufficiently articulated. It remains unclear how the findings extend existing knowledge beyond evaluating the impact of a specific STEM outreach activity. Without this, the manuscript currently reads more as an evaluation of a specific educational initiative than as a research study that makes a clear contribution to the scholarly literature. In this sense, the authors are suggested clearly identify why these findings are important to a body of knowledge (for example, addressing what is missing or what is under-researched and explain why these missing parts are important to address).
Second, the conceptualisation of key constructs, particularly STEM interest and STEM self-efficacy, requires greater clarity. The manuscript would benefit from a clearer explanation of how these constructs are defined and distinguished from related concepts in the literature.
Third, the measurement of these constructs raises concerns. The survey instruments appear to have been developed by the authors; however, limited information is provided regarding their development process, validity, and reliability. In particular, additional evidence is needed to demonstrate how the survey items measure the intended constructs. For example, if specific items (i.e. survey questions) are used to measure STEM self-efficacy or STEM interest, the manuscript should provide a clearer rationale for why these items are considered valid measures of those variables, as well as evidence supporting their reliability and validity. As the study's conclusions rely heavily on these measures, additional evidence supporting the quality and appropriateness of the instruments is necessary.
Finally, the interpretation of the findings should be more closely aligned with the limitations of the dataset. While the analyses provide some interesting preliminary patterns, the relatively small sample size, particularly for subgroup analyses involving grade level, gender, and ethnicity, limits the strength of the conclusions that can be drawn. In several cases, the findings appear to be interpreted in a manner that exceeds the evidential strength of the data. Given the limited sample sizes within some demographic categories, the results should be presented more cautiously as preliminary or exploratory findings rather than as strong evidence of demographic differences or broader educational implications. The authors are suggested that modify the manuscript with a more exploratory and cautious tone. The findings are best interpreted as preliminary evidence that may inform future research, rather than as strong or broadly generalisable conclusions.
Author Response
Please see the attachment.
Author Response File:
Author Response.docx
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe authors have done substantial and responsive work on this revision, and the most serious problems from the first round have been resolved. The implausible Grade Level by Ethnicity interaction (partial eta squared = .93) has been removed, the unsupported Gender by Grade Level interaction is gone, the self-efficacy claims have been reframed as post-event perceptions rather than gains, the ceiling effect is acknowledged, and the abstract now correctly reports the ANOVA result (F = 3.76, p = .023) rather than the post-hoc value. The photographs have been moved out of the Results section, a third research question and a qualitative methods subsection have been added for the mentor data, and the museum description is more carefully scoped.
A few issues remain.
There is a contradiction between the Results and the Discussion regarding gender and ethnicity. Section 4.2 states that the ANOVAs showed no statistically significant differences by gender or ethnicity on the self-efficacy items, yet Section 5.2 opens by saying the results indicated several statistically significant differences in STEM self-efficacy across grade levels, gender, and ethnic subgroups. Section 5.2 then describes gender differences, with female students being less likely to see themselves as future engineers or scientists and reporting lower peer inspiration, that do not appear anywhere in the Results with supporting statistics. This is the same kind of mismatch flagged in the first round. Please either report these gender comparisons in the Results with the relevant statistics, or remove the claims from the Discussion so that the two sections agree.
The research question numbering is now inconsistent across the paper. Section 2.5 lists RQ1 as STEM interests and career aspirations, RQ2 as self-efficacy, and RQ3 as mentor experiences. However, the qualitative methods subsection says it addresses Research Question 1 for the mentor reflections, which should be RQ3. Section 5.1 is headed RQ1 but its text refers to Research Question 2, and Section 5.2 is headed RQ2 but its text refers to Research Question 3. Please align the in-text references and the section headings with the numbering established in Section 2.5.
Some change-oriented language remains even though the unmatched design supports only descriptive comparison. For example, Section 5.1 says engineering interest was increasing from 25.6% to 38.0%, and Section 4.1 uses "up from," "dropped by," and "slightly declined." Please carry the descriptive reframing into these passages as well. Relatedly, the title still uses "Boosting," which asserts a causal effect the study does not test. Consider softening it to reflect the descriptive, exploratory nature of the findings.
The significant grade-level result is for a single item (Q7, applying scientific principles), but the abstract and Section 5.2 describe it as STEM self-efficacy and overall self-efficacy. Please specify that the difference was found on that one item rather than on overall self-efficacy, so the claim matches the analysis.
Then there are a few smaller items. The Bandura (1997) sentence in Section 5.2 about confidence developing through repeated success, growing maturity, and opportunities for mastery is duplicated and should appear once. Two subsections are both numbered 3.4 (Limitations and Qualitative Methods), so the second should be renumbered. The ceiling effect, now well discussed in the Results, could also be noted briefly in the Limitations, as the response letter indicates was intended.
The mentor reflection analysis in Table 2 and Section 5.3 is improved and reads more cleanly than the earlier raw responses, though it still sits closer to a structured summary than a full thematic analysis. A sentence or two naming the themes that organize Table 2 would tie it more clearly to the qualitative methods described in Section 3.
Comments on the Quality of English LanguageThe manuscript is generally clear, but it needs one more proofreading pass. Section 5.2 contains a duplicated sentence about Bandura's sources of self-efficacy. A few sentences remain awkward, for example in Section 5.3 ("meaningful aspects of their participation to be creating the projects themselves" and "their training and implementation turned to be a rich learning experience"). The research question cross-references and the repeated 3.4 section number should also be corrected during editing.
Author Response
- The authors have done substantial and responsive work on this revision, and the most serious problems from the first round have been resolved. The implausible Grade Level by Ethnicity interaction (partial eta squared = .93) has been removed, the unsupported Gender by Grade Level interaction is gone, the self-efficacy claims have been reframed as post-event perceptions rather than gains, the ceiling effect is acknowledged, and the abstract now correctly reports the ANOVA result (F = 3.76, p = .023) rather than the post-hoc value. The photographs have been moved out of the Results section, a third research question and a qualitative methods subsection have been added for the mentor data, and the museum description is more carefully scoped.
Response: Thank you very much for the opportunity to revise our manuscript. The reviewer’s comments have been extremely helpful in improving flaws observed in the contents and the presentation of our work. Below, please find our detail response to the comments. We have revised the manuscript accordingly, and included most of the explanation below in the manuscript in its new edition.
- There is a contradiction between the Results and the Discussion regarding gender and ethnicity. Section 4.2 states that the ANOVAs showed no statistically significant differences by gender or ethnicity on the self-efficacy items, yet Section 5.2 opens by saying the results indicated several statistically significant differences in STEM self-efficacy across grade levels, gender, and ethnic subgroups. Section 5.2 then describes gender differences, with female students being less likely to see themselves as future engineers or scientists and reporting lower peer inspiration, that do not appear anywhere in the Results with supporting statistics. This is the same kind of mismatch flagged in the first round. Please either report these gender comparisons in the Results with the relevant statistics, or remove the claims from the Discussion so that the two sections agree.
Response: We sincerely thank the reviewer for identifying this inconsistency. We agree that the Discussion should reflect only the statistically supported findings presented in the Results. Accordingly, we removed all unsupported statements regarding gender and ethnicity from Section 5.2, including references to female students' perceptions of future engineering or science careers and peer inspiration, because these findings were not supported by the statistical analyses reported in the Results.
The revised Discussion now focuses exclusively on the statistically significant finding reported in Section 4.2: a grade-level difference on STEM self-efficacy Question 7, which assessed students' confidence in applying scientific principles. We also clarified in both the Abstract and the Discussion that this significant difference pertains only to this individual self-efficacy item rather than to overall STEM self-efficacy. These revisions ensure complete consistency among the Results, Discussion, and Abstract.
In addition, during this revision we removed career aspirations as a study construct throughout the manuscript because our paper focused directly on STEM interests and STEM self-efficacy but did not explicitly assess career aspirations. This conceptual refinement further improved the alignment among the research questions, methods, results, and discussion.
- The research question numbering is now inconsistent across the paper. Section 2.5 lists RQ1 as STEM interests and career aspirations, RQ2 as self-efficacy, and RQ3 as mentor experiences. However, the qualitative methods subsection says it addresses Research Question 1 for the mentor reflections, which should be RQ3. Section 5.1 is headed RQ1 but its text refers to Research Question 2, and Section 5.2 is headed RQ2 but its text refers to Research Question 3. Please align the in-text references and the section headings with the numbering established in Section 2.5.
Response: We sincerely thank the reviewer for identifying these inconsistencies. We carefully reviewed the entire manuscript and corrected all research question numbering and corresponding in-text references to ensure consistency throughout the manuscript. Specifically, the qualitative methods section, Results, Discussion, and all section headings were revised so that they correctly correspond to the revised research questions.
Again, after carefully reconsidering the alignment between the research questions and our survey, we removed career aspirations as a study construct because the survey directly measured students' STEM interests and STEM self-efficacy but did not explicitly assess career aspirations. Consequently, the research questions were revised as follows:
- RQ1: What STEM interests do middle school students report following participation in NSF STEM Day?
- RQ2: What STEM self-efficacy perceptions do middle school students report following participation in NSF STEM Day, and do these perceptions differ by grade level, gender, or ethnicity?
- RQ3: How do college student mentors describe their learning experiences and mentoring activities during NSF STEM Day?
These revisions further improved the alignment among the conceptual framework, research questions, survey instrument, methods, results, and discussion, ensuring that each research question is directly addressed by the corresponding analyses and findings.
- Some change-oriented language remains even though the unmatched design supports only descriptive comparison. For example, Section 5.1 says engineering interest was increasing from 25.6% to 38.0%, and Section 4.1 uses "up from," "dropped by," and "slightly declined." Please carry the descriptive reframing into these passages as well. Relatedly, the title still uses "Boosting," which asserts a causal effect the study does not test. Consider softening it to reflect the descriptive, exploratory nature of the findings.
Response: We carefully reviewed the entire manuscript and further revised the language to ensure that the findings are presented descriptively and do not imply causal effects that cannot be supported by our unmatched pre- and post-survey design.
Specifically, we replaced remaining change-oriented expressions such as "increased," "boosted," "up from," "dropped by," and "slightly declined" with neutral descriptive language, including "reported," "compared with," "most frequently reported," and "selected in the pre- and post-surveys." We also revised the Discussion to emphasize descriptive comparisons rather than changes attributable to participation in NSF STEM Day.
In addition, we revised the manuscript title by removing the term "Boosting," which could imply a causal effect. The revised title more accurately reflects the descriptive and exploratory nature of the study and is consistent with the study design and the interpretation of the findings.
These revisions further align the manuscript with the study's methodology and strengthen the consistency between the analyses, results, and conclusions.
- The significant grade-level result is for a single item (Q7, applying scientific principles), but the abstract and Section 5.2 describe it as STEM self-efficacy and overall self-efficacy. Please specify that the difference was found on that one item rather than on overall self-efficacy, so the claim matches the analysis.
Response: We agree that our interpretation should accurately reflect the scope of the statistical finding. Accordingly, we revised both the Abstract and Section 5.2 (Discussion) to clarify that the statistically significant grade-level difference was observed only for STEM self-efficacy Question 7, which assessed students' confidence in applying scientific principles, rather than for overall STEM self-efficacy.
Specifically, the revised manuscript now states that a significant grade-level difference was observed on the STEM self-efficacy item assessing confidence in applying scientific principles, with eighth-grade students reporting higher confidence than sixth-grade students (F = 3.76, p = .023). We also revised the Discussion to ensure that the interpretation focuses on this individual survey item and does not imply statistically significant differences in overall STEM self-efficacy.
These revisions ensure that the interpretation of the findings is fully consistent with the statistical analyses and accurately reflects the scope of the observed result.
- Then there are a few smaller items. The Bandura (1997) sentence in Section 5.2 about confidence developing through repeated success, growing maturity, and opportunities for mastery is duplicated and should appear once. Two subsections are both numbered 3.4 (Limitations and Qualitative Methods), so the second should be renumbered. The ceiling effect, now well discussed in the Results, could also be noted briefly in the Limitations, as the response letter indicates was intended.
Response: We have addressed each of these points in the revised manuscript.
First, the duplicated sentence describing Bandura's (1997) sources of self-efficacy in Section 5.2 has been removed so that the discussion appears only once. Second, the duplicate subsection numbering in Section 3 has been corrected by renumbering the qualitative methods subsection to ensure consistent organization throughout the manuscript. Finally, as suggested, we added a brief statement to the Limitations section acknowledging the potential ceiling effect observed in the STEM interest measure. Specifically, we note that the high level of pre-event STEM interest may have reduced the opportunity to observe larger descriptive differences in the post-survey responses.
- The mentor reflection analysis in Table 2 and Section 5.3 is improved and reads more cleanly than the earlier raw responses, though it still sits closer to a structured summary than a full thematic analysis. A sentence or two naming the themes that organize Table 2 would tie it more clearly to the qualitative methods described in Section 3.
Response:
we revised Section 5.3 by explicitly identifying the major themes that emerged from the qualitative content analysis and that organize Table 2. Specifically, we summarized the mentors' reflections into four overarching themes: (1) engaging middle school students through hands-on STEM learning, (2) developing communication and leadership skills, (3) applying STEM knowledge in authentic teaching contexts, and (4) fostering personal and professional growth. We also added a brief introductory statement to Section 5.3 to explain that these themes were derived through qualitative content analysis, thereby strengthening the connection between the qualitative methods described in Section 3 and the presentation of the findings in Table 2.
- Comments on the Quality of English Language
The manuscript is generally clear, but it needs one more proofreading pass. Section 5.2 contains a duplicated sentence about Bandura's sources of self-efficacy. A few sentences remain awkward, for example in Section 5.3 ("meaningful aspects of their participation to be creating the projects themselves" and "their training and implementation turned to be a rich learning experience"). The research question cross-references and the repeated 3.4 section number should also be corrected during editing.
Responses: we carefully proofread the entire manuscript to improve its clarity, readability, and overall quality of English. The duplicated Bandura (1997) sentence in Section 5.2 was removed, and the awkward sentences identified in Section 5.3 were revised for improved clarity and readability. We also reviewed the manuscript to ensure consistency throughout, correcting all research question cross-references, renumbering the duplicated Section 3.4 subsection, and making additional editorial revisions where appropriate. We believe these revisions have further improved the overall presentation and readability of the manuscript. We appreciate the reviewer's careful attention to these editorial details, which have helped improve the overall quality of the manuscript.

