6.1. Section 1
Section 1 provides a thorough look at how participants responded to the task of finding research articles, showing how students found, understood, and sorted scholarly sources both with and without help from AI-based tools. This section illustrates the accuracy and variability in their selections, the criteria they used to distinguish research articles from other academic texts, and the degree to which they demonstrated awareness of research design, methodological features, and publication conventions.
By looking at the trends in their choices and reasons, along with how AI prompts affected their decisions, the results show important strengths, common misunderstandings, and new skills in students’ growing ability to critically assess academic research. Ultimately, this section demonstrates how AI-assisted reasoning shaped students’ performance and where human judgment remained essential.
Figure 3 and
Table 4 demonstrated the descriptive statistics for survey questions.
Figure 3 summarizes students’ survey responses on AI use, showing that AI is used most frequently for search assistance and prompt identification, while more advanced evaluative and integrative practices receive lower average ratings. The error bars indicate noticeable variability across students, suggesting uneven adoption of higher-order AI-supported strategies. Overall, the figure highlights that AI is commonly used for initial support but less consistently for deeper critical evaluation and refinement.
Prompt Strategy Use and Effects:
Here are the four prompt strategies we compared:
Summarizing abstracts
Generating search keywords
Ranking articles by relevance or citation count
Suggesting related articles
Table 5 demonstrated the descriptive statistics for the prompt strategies.
ANOVA Test:
To see if the perceived efficiency gains varied between different prompt strategies, an ANOVA was performed using the numerical values of students’ survey answers. The item was, “Did prompt engineering improve efficiency?” was measured on a three-level ordinal scale (“No change,” “Somewhat improved,” “Significantly improved”). For analytical purposes, these categories were mapped to numerical values (1 = No change, 2 = Somewhat improved, 3 = Significantly improved), allowing for exploratory comparison across groups.
Prompt strategies were treated as the grouping variable, with mean efficiency ratings compared across all four strategies using a one-way ANOVA.
The ANOVA did not yield a statistically significant effect (p ≥ 0.05), indicating that differences in mean efficiency ratings across prompt strategies were not large or consistent enough to support a strong association. Although slight mean differences were observed, these variations did not exceed expected within-group variability.
Importantly, this result does not suggest that prompt strategies had no instructional value. Rather, it indicates that perceived efficiency gains were relatively stable across prompt types, implying that students engaged in revision in a broadly consistent manner regardless of the specific prompt strategy employed. This finding aligns with the broader pattern observed across tasks, where revision behavior appeared to be driven more by iterative engagement than by prompt category alone.
Table 6 presents the full ANOVA results and descriptive statistics supporting this interpretation.
Chi-Square and Cramer’s V Test
To assess whether participants’ answers to one question were statistically associated with another, we computed chi-square tests between all pairs of multiple-choice questions.
Table 7 shows the results.
All associations are statistically significant, with p < 0.001 for nine of the ten pairings and p = 0.016 for the efficiency ↔ quality pair. Severe truncation of any text has been avoided here, so you can fully read each question and interpretation.
Participants who used AI prompts more often tended to report:
(Both results are statistically significant at p < 0.05)
However, the frequency of use did not significantly correlate with the perceived quality improvement of selected articles or the specific prompt strategy considered most useful.
Here’s the textual summary table of Cramér’s V effect sizes among all multiple-choice question pairs.
This table shows the strength of association (based on Cramér’s V), chi-square value, and p-value. The chi-square test only tells us if a relationship between two categorical variables is statistically significant. It doesn’t tell us how strong that relationship is. Cramer’s V is an effect size measure that provides a standardized value between 0 and 1 to quantify the strength of this association, helping to interpret the practical significance of the findings.
Cramér’s V Effect Size Summary Table (Top Associations):
Cramer’s V was used in Section 1 to evaluate the strength of associations identified by statistically significant chi-square tests. Whereas the chi-square statistic indicates whether an association exists, Cramer’s V provides an estimate of its practical magnitude.
The observed effect sizes varied from weak to moderate, suggesting that while the relationships were statistically detectable, their practical impact was limited. The notably small effect sizes indicate that prompt type had a negligible impact on students’ revision behavior. Students demonstrated similar levels of revision regardless of whether they employed summarizing, explaining, paraphrasing, or other prompt strategies.
This pattern indicates that prompt selection alone did not meaningfully drive revision depth. Rather, broader engagement practices that transcend prompt categories appear to shape revision behavior.
The full Cramer’s V results are reported in
Table 8.
The top two relationships (Cramér’s V = 0.45 and 0.37) are statistically significant (p < 0.05), indicating meaningful behavioral trends:
Frequent AI prompt users report higher efficiency and more active revision practices.
The remaining pairs show weaker relationships (Cramér’s V < 0.3), suggesting mostly independent response patterns among questions.
Text Similarity:
Having used different prompt strategies, we now wanted to compare the AI-generated article suggestions, summaries, or rankings and the revised evaluation/final selection of articles.
How to Interpret Together:
Cosine similarity captures idea retention → semantic fidelity.
The Jaccard Index captures vocabulary reuse → lexical overlap.
Levenshtein Similarity captures surface rewording → editing effort.
Table 9 provides a conceptual explanation.
Together, these show how deeply students revised AI-generated suggestions:
A high cosine score combined with a low Jaccard or Levenshtein score indicates that students demonstrated conceptual understanding through paraphrasing.
High across all = minimal revision (copying).
Low cosine = conceptual change or misunderstanding.
Table 10 reports three complementary text-similarity measures used to characterize students’ revision behavior. Cosine similarity captures
semantic overlap using TF–IDF weighting, the Jaccard Index measures
lexical overlap based on shared unique words, and Levenshtein similarity reflects
character-level similarity based on edit distance.
As illustrated in
Table 10, the three metrics display distinct but convergent patterns. Average cosine similarity falls in the mid-range (≈0.25), indicating partial conceptual alignment between AI-generated and student-revised texts. This suggests that students retained core ideas while allowing meaning to diverge through paraphrasing or restructuring.
In contrast, the Jaccard Index shows low average overlap (≈0.11), indicating that students frequently replaced vocabulary rather than reusing AI-generated wording. Levenshtein similarity is also low (≈0.14), reflecting substantial character-level editing and further supporting the presence of extensive textual transformation rather than surface modification.
Taken together, these patterns indicate active engagement with AI-generated content. Rather than copying or lightly editing AI output, students systematically reformulated text while preserving elements of semantic content.
The similarity in both meaning and wording suggests that the way students revised the text was a thoughtful change instead of just reusing it without much thought.
6.2. Section 2
Section 2 examines participants’ performance on the AI summary task for videos, focusing on how effectively students used AI tools to generate concise, coherent, and accurate summaries of academic video content. This section looks at the ways that students interacted with AI, how well they could improve or refine drafts made by AI, and how well they could identify key arguments, themes, and supporting evidence in the source material. By looking at the trends in their summaries and how they changed their prompts, the results show how students combined help from AI with their understanding, where they had confusion, and how AI affected the quality and detail of their final work. Ultimately, this section demonstrates the role of AI in shaping students’ summarization practices and highlights the areas in which human oversight and critical interpretation remained essential.
The descriptive statistics for word count (
Figure 4 and
Table 11) revealed a consistent pattern of student revision behavior. On average, the AI-generated summaries were considerably longer (M = 190.13 words, SD = 174.75) than the students’ revised summaries (M = 127.06 words, SD = 115.76). This reduction in length suggests that students tended to condense and refine the AI output rather than expand upon it. The paired comparison confirmed that this difference was statistically significant, indicating that the shortening was not due to random variation but reflected a meaningful shift in how students reworked the text. Taken together, these results suggest that students generally used the AI-generated summary as a starting point and then selectively trimmed, reorganized, or simplified the material to produce more concise summaries in their words.
Chi-Square and Cramer’s V Test:
The chi-square analyses reported in
Table 12 were conducted to examine whether associations existed among prompt choice, revision level, perceived usefulness of rephrasing, and self-reported improvements in listening or note-taking.
Across most comparisons, the chi-square tests did not reach statistical significance, indicating no strong associations among these variables. The link between the type of prompt and how much students revised was not significant, with a small Cramer’s V, meaning that the choice of prompt didn’t really affect how much students changed the AI-generated summaries. Students tended to revise AI output to a similar degree regardless of prompt style, adapting the text to fit their voice and intentions.
Similarly, the association between revision level and reported improvements in listening or note-taking was non-significant with a small effect size. This pattern implies that engagement with the task itself, rather than the amount of textual revision performed, more closely correlates with perceived listening benefits.
One comparison—between prompt type and perceived usefulness of rephrasing—approached statistical significance (p = 0.053) and was associated with a moderate effect size. While this result does not support a definitive conclusion, it suggests that certain prompting strategies (e.g., summarizing versus extracting key ideas) may offer modest advantages for comprehension. This trend warrants further investigation in instructional design and future studies.
Overall, the chi-square and effect-size patterns indicate that while students varied in prompting strategies and revision practices, the educational benefits of the activities were broadly distributed across groups. Structured engagement with AI-supported tasks appears to drive learning gains, not a single prompt or revision variable.
Interpretation Summary:
Cramer’s V values < 0.20 → negligible association
0.20–0.35 → small to moderate association
No comparison here shows a strong relationship.
The only borderline case is Prompt Type × Understanding, suggesting students’ prompting strategy may slightly influence how much rewriting helps comprehension—potentially worth exploring pedagogically.
Figure 5,
Figure 6 and
Figure 7 provided some intriguing data on listening benefits on summarization with AI prompts and the effect of prompt type on the revision process with paraphrasing.
AI-Generated and Revised Text Comparison—
Table 13 presents a comparison of AI-generated summaries and student-revised texts using complementary semantic, lexical, and surface-level similarity measures. Together, these metrics were used to characterize how students transformed AI output while revising.
The cosine similarity between AI-generated and revised texts was relatively high on average (M = 0.61, SD = 0.23), indicating substantial retention of core ideas and conceptual structure. This suggests that students largely preserved meaning even as they revised wording and organization.
In contrast, the Jaccard Index showed moderate lexical overlap (M = 0.44, SD = 0.26), indicating that fewer than half of the original words were reused. This level of overlap is consistent with paraphrasing rather than direct copying. Levenshtein similarity was also moderate (M = 0.48, SD = 0.25), reflecting substantial sentence-level editing and further confirming that revisions extended beyond minor surface changes.
Taken together, the combined pattern across semantic, lexical, and surface measures indicates that students actively reformulated AI-generated summaries while maintaining the underlying message. This behavior aligns with productive AI-supported learning, in which students use AI for initial structure or content guidance while retaining ownership of the final language.
6.3. Section 3
Section 3 analyzes participants’ responses to the task requiring them to transform casual English into academically appropriate language with the support of AI tools. This section explores how effectively students identified informal features, how they used AI to revise or elevate the tone and clarity of their writing, and the degree to which they could recognize, and correct inaccuracies or inappropriate stylistic shifts introduced by the model. By examining the patterns in their rewritten texts, prompt adjustments, and justifications, the results highlight students’ developing awareness of academic discourse conventions and their ability to collaborate with AI to improve linguistic precision. Ultimately, this section demonstrates how AI shaped students’ writing choices and reveals the balance between automated assistance and human editorial judgment in achieving academically sound output.
Interpretation
Figure 8 shows a comparison between the types of prompts students used at the start of the task and the prompt strategies they later judged as most effective.
Specifically, it illustrates that:
Initially, most students relied on simple, direct prompts (e.g., “rewrite academically,” “improve this text”).
After engaging with the task, students increasingly identified structured prompts—such as step-by-step or role-based prompts—as more effective for producing academically appropriate writing.
Taken together,
Figure 8 visualizes a shift in prompting strategy over time, indicating that students learned through experience which types of prompts better supported tone control, clarity, and academic style. The figure highlights iterative learning and prompt refinement, rather than suggesting that prompt type alone determines writing quality.
Chi-Square Test + Cramer’s V:
Research question: Does the type of prompt used influence meaning consistency across rewrites? (
Table 14)
Interpretation:
The association between prompt type and meaning preservation did not reach statistical significance. However, the small-to-moderate Cramer’s V suggests a weak and inconsistent tendency for more structured prompts to support meaning stability, indicating that prompt structure may play a limited supportive role rather than serving as a primary driver of meaning preservation in this dataset.
ANOVA/Correlation:
We tested whether trying more prompt versions relates to meaning consistency and is shown in
Table 15 as a combined ANOVA summary table.
Participants’ perceived consistency was mapped as
Correlation: r = 0.26
Interpretation:
There is a positive but weak relationship:
Students who experimented with multiple prompts tended to preserve meaning more effectively.
In other words, iteration helps, but the improvement is not dramatic.
Interpretation
Differences in prompt strategy at the start did not produce statistically significant differences in either academic rewrite length or casual reversion length. This suggests that students generally produced text of similar length regardless of the initial prompting method used.
Table 16 reports the results of one-way ANOVAs examining whether text similarity differed across prompt strategy conditions at three stages of rewriting.
Across all comparisons, the ANOVAs did not yield statistically significant differences in similarity scores (all p > 0.05). Effect sizes were small (η2 = 0.04–0.08), indicating that prompt strategy accounted for only a limited proportion of variance in meaning preservation or transformation.
For the original-to-academic rewrite (sim_OA), the small effect size suggests that prompt strategy did not meaningfully influence how much meaning changed when students shifted to an academic register. Similarly, in the academic-to-final casual rewrite (sim_AF), similarity levels were comparable across groups, indicating that students reverted to a casual tone in a consistent manner regardless of initial prompting. The original-to-final casual comparison (sim_OF) shows a similar pattern, with final casual texts converging toward original meaning across all prompt strategies.
Taken together, these findings suggest that while prompt structure may shape how students approach writing tasks, it does not strongly determine how much meaning is preserved or altered across revisions. Students demonstrated consistent control over tone shifting and meaning recovery, pointing to stable cognitive-linguistic strategies that operate independently of prompt framing.
Figure 9 visualizes these patterns, illustrating the relative stability of meaning consistency across prompt strategies.
Tukey HSD Post Hoc Comparison: Academic Word Count by Prompt Strategy:
Even though the overall ANOVA did not show statistically significant differences, Tukey HSD allows us to examine pairwise comparisons between prompt strategies.
Table 17 showed the test results.
Interpretation:
A Tukey HSD post hoc analysis was conducted following the ANOVA to examine pairwise differences in academic word count across prompt strategies. The analysis did not identify any statistically significant differences between prompt pairs, indicating that academic response length remained stable regardless of the prompting style used. This pattern suggests that while students employed different prompting approaches, output length was not meaningfully affected. Prompt strategy appears to shape how students structured or refined academic writing rather than how much they wrote.
Semantic retention was further examined using cosine similarity between the original casual paragraph and the first academic rewrite. The mean cosine similarity was 0.256. This mid-to-low similarity indicates substantial transformation in expression when students shifted from casual to academic tone, with wording and structure diverging noticeably while core meaning was partially retained.
Taken together, these results align with the instructional objective of the task. Students demonstrated that academic rewriting involves rhetorical reorganization rather than surface-level embellishment, reflecting meaningful engagement with register and discourse conventions rather than simple lexical substitution.
Overall Findings (Synthesis): The overall findings have been reported in
Table 18.
6.4. Section 4
Section 4 examines participants’ performance on the integrated concept-mapping task, in which students synthesize information from both a reading and a video with the support of AI tools. This section explores how effectively students identified key ideas, established logical connections between concepts, and organized multimodal This study focuses on transforming information into a coherent visual or structured representation. It also considers how students used AI to clarify relationships, generate initial map structures, or refine their drafts, as well as the extent to which they recognized and corrected AI-generated inaccuracies or oversimplifications. By analyzing the patterns in their conceptual linkages, the depth of their integration, and their prompt-refinement strategies, the results highlight students’ emerging ability to merge multiple sources while leveraging AI as a cognitive aid. Ultimately, this section demonstrates the role of AI in shaping students’ integrative reasoning and underscores the continued importance of human interpretation in constructing accurate and meaningful concept maps.
Descriptive Statistics (Effectiveness of AI Prompts, as Shown in
Table 19):
Interpretation:
Students rated AI prompts as highly effective on average, clustering mostly around “Very effective” (3).
The following
Table 20 shows whether the effectiveness ratings depend on the specific computer science topics considered by the students.
Chi-Square Test of Independence (
Table 21):
A chi-square test of independence was conducted to examine whether the topic was associated with students’ perceived effectiveness of AI prompts (
Table 21). The chi-square result did not reach conventional statistical significance (χ
2 = 17.71, df = 10,
p = 0.060). While the
p-value approaches the 0.05 threshold, it does not provide sufficient evidence to confirm a reliable association within this dataset. However, Cramer’s V indicates a moderate effect size (V = 0.391), suggesting that topic choice accounted for a meaningful proportion of variation in perceived AI effectiveness. Taken together, these results point to a
possible topic-related influence that is present but not consistently strong enough to reach statistical confirmation.
This pattern suggests that topic familiarity or domain complexity, rather than prompt design alone, may shape students’ evaluations of AI prompt effectiveness. Importantly, the finding is best interpreted as indicating variability across topics rather than a definitive topic effect, consistent with the exploratory nature of the analysis.
A Kruskal–Wallis test was conducted to examine whether students’ effectiveness ratings of AI prompts differed across topic groups (
Table 22). This non-parametric test assesses whether at least one topic exhibits a different distribution of ratings compared to others.
The test did not yield a statistically significant result (H = 8.061, p = 0.153). As the p-value exceeds the 0.05 threshold, there is insufficient evidence to conclude that perceived prompt effectiveness varied systematically by topic. This result suggests that, despite some variation in ratings across topics, these differences were not large or consistent enough to indicate a reliable topic-based effect within this dataset. The Kruskal–Wallis results agree with the chi-square findings, showing that while the topic might affect how students see the prompts, it doesn’t strongly influence their ratings of AI prompt effectiveness.
Conclusions
There is no statistically significant difference in how effective students rated AI prompts across different topics. Students from various fields, such as AI, cybersecurity, and other data science fields, have similar perceptions of the AI prompts. perceived the AI prompts similarly.
Use of Prompts for Organizing Information and Idea Coordination:
The violin plot emphasizes distribution shape, showing how responses cluster.
Violin Plots—Interpretation
“How effective were AI prompts…?”
- o
A tall, narrow peak around 4–5 → Strong agreement clustered at the positive end.
- o
Thin distribution on lower numbers → Few students rated AI prompts weak.
“Did rephrasing help you understand relationships?”
- o
A slightly wider violin → More varied experiences.
- o
Peaks near 3–4 suggest general usefulness but non-uniform adoption.
“Own synthesis vs. AI text?”
- o
Possible two bumps (bimodal) → Some students relied heavily on AI; others relied more on their own ideas.
- o
Wider shape overall → Indicates variation in learning styles and comfort with AI.