Next Article in Journal
Extra-Curricular Activities and Children’s Bilingual Language Learning in Singapore
Previous Article in Journal
Relationships Between Problematic Internet Use, Physical Activity, and Mental Health in University Students
Previous Article in Special Issue
Integrating Generative AI in Engineering Education: Enhancing Learning and Attendance in a Vehicle Theory Course
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

From Expectations to Measured Pragmatism: A Pre- and Post-Experience Study of Student Engagement in AI-Supported Academic Exams

1
Faculty of Instructional Technologies, Holon Institute of Technology, Holon 5810201, Israel
2
Faculty of Engineering, Ruppin Academic Center, Kfar Monash 4025000, Israel
3
The Center for Research in Technological and Engineering Education, Ruppin Academic Center, Kfar Monash 4025000, Israel
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(4), 642; https://doi.org/10.3390/educsci16040642
Submission received: 2 March 2026 / Revised: 8 April 2026 / Accepted: 14 April 2026 / Published: 17 April 2026

Abstract

Generative AI (GenAI) is transforming higher education assessments, yet empirical research on students’ lived experiences with GenAI during graded, time-constrained classroom assessments remains scarce. This study investigates how direct experience with GenAI in examinations shapes student perceptions of learning, metacognition, and engagement. Drawing on self-regulated learning research and cognitive load theory, we employed a retrospective pre–post design to analyze qualitative reflections and quantitative data from 90 undergraduate computer science and engineering students. Our qualitative analysis suggests a complex recalibration from idealized expectations of efficiency toward what may be described as a state of measured pragmatism. Interpretive analysis of Post-experience reflections indicates that direct practical engagement appeared to make students more conscious of the need for metacognitive engagement, with a focus on real-time output verification and the restrictive role of time pressure. Concerns regarding assessment authenticity and fairness emerged only after direct engagement. Quantitative results show that although 68.5% preferred the GenAI format, this preference did not correlate significantly with academic performance (r = 0.014, p = 0.89). Those findings suggest that student engagement is driven by pedagogical and professional relevance rather than grade improvement alone. Overall, the findings underscore the need for assessment designs that balance cognitive support with active student monitoring and responsibility.

1. Introduction

The advent of generative artificial intelligence (GenAI), particularly large language models (LLMs) such as ChatGPT, has fundamentally disrupted traditional practices in higher education, including assessments (Xia et al., 2024). Since the emergence of this new technology in late 2022, academic institutions have moved rapidly from initial prohibition and uncertainty toward a more cautious and intentional integration of these tools into pedagogical frameworks (Amzalag & Kurtz, 2025; Kasneci et al., 2023; Zawacki-Richter et al., 2024). Yet while GenAI is increasingly recognized as an essential component of contemporary digital literacy, there is as yet no consensus regarding its appropriate role in educational environments (Bauer et al., 2025; Kasneci et al., 2023). Some scholars frame GenAI as a robust cognitive scaffold capable of reducing barriers to learning, while others emphasize the risks of cognitive atrophy, over-reliance, and the erosion of foundational problem-solving skills (Bates et al., 2020; Kasneci et al., 2023).
Despite the growing body of literature on GenAI in education, most existing studies focus on formative learning, general student attitudes, or institutional policy at a theoretical level (Deng et al., 2025; Wang & Fan, 2025). The present study examines how students actually engage with GenAI during classroom-based assessments, an area marked by a notable scarcity of empirical research (Chiu et al., 2023; Xia et al., 2024). In assessment environments, conditions such as strict time pressure and evaluative consequences fundamentally shape students’ cognitive strategies and perceptions (Yusuf et al., 2024).
The present study was conducted in the context of computer science and engineering examinations, where traditional assumptions about individual knowledge production and algorithmic reasoning are being directly challenged (Cotton et al., 2024; Xia et al., 2024). We explore undergraduate students’ perceptions through the lens of self-regulated learning (SRL) theory (Zimmerman, 2002) and cognitive load theory (Sweller, 1988, 2011). Using a retrospective pre–post design, we examine how direct practical experience with GenAI during a classroom assessment affected students’ attitudes. In this respect, our study is one of the few to employ longitudinal or pre–post designs to examine how students’ conceptualization of AI changes following direct, practical application. The study integrates qualitative thematic analysis with quantitative measures to provide novel empirical evidence on the evolving role of GenAI in academic evaluations.
This study contributes to the field by offering a process-oriented exploration of how direct experience with GenAI-supported assessments reshapes student perceptions. Theoretically, it bridges the gap between GenAI literacy and self-regulated learning (SRL) frameworks; practically, it provides empirical evidence to guide the design of authentic, fair, and cognitively engaging assessment environments in the AI era.
The remainder of this paper is organized as follows. Section 2 reviews the relevant literature on GenAI in higher education and self-regulated learning. Section 2 presents research questions. Section 3 presents research objectives. Section 4 details the research methodology, including data collection and analysis procedures. Section 5 presents the study’s findings, while Section 6 discusses these results in the context of broader educational implications and concluding remarks. Finally, Section 7 and Section 8 offer limitations, and suggestions for future research.

2. Background and Related Works

2.1. Generative AI in Higher Education

Following their rapid emergence, GenAI tools, particularly LLMs, were initially met with uncertainty and restrictive policies. Yet over the past few years, GenAI has increasingly been reframed as an integral component of contemporary digital literacy, prompting universities to move from prohibition toward cautious integration. Recent reviews show that students now use GenAI to explain or summarize material, to help with programming, and to generate feedback across disciplines (Amzalag & Kurtz, 2025; Kasneci et al., 2023; Kurtz et al., 2024; Qian, 2025).
A growing body of literature highlights the pedagogical potential of GenAI as a cognitive support tool. It can help learners structure information, generate ideas, explore alternative perspectives, and receive immediate support when working through complex tasks, thereby functioning as a form of academic scaffolding (Holmes et al., 2019; Chen et al., 2020; Zawacki-Richter et al., 2019). At the same time, the literature reflects ongoing debate regarding how GenAI should be used effectively in pedagogy and assessment (Kasneci et al., 2023; Xia et al., 2024; Yusuf et al., 2024). While some scholars emphasize its potential to enhance understanding, motivation, and access to feedback, others warn that unreflective reliance on AI may reduce epistemic effort, encourage superficial task completion, and erode foundational skills (Bates et al., 2020; Kasneci et al., 2023). These tensions are particularly salient in assessment contexts, where traditional assumptions about individual knowledge production, authorship, and students’ reasoning skills are increasingly challenged (Cotton et al., 2024; Eke, 2023; Xia et al., 2024).
Crucially, the relationship between GenAI use and learning success is not uniform. Rather than being inherently beneficial or detrimental, GenAI is better understood as a dual-use technology whose effects depend on context, task design, and the quality of learner engagement. Students with stronger prior knowledge and better prompting and evaluation skills are more likely to use GenAI as support for explanation, reflection, and revision, whereas less critical use may foster cognitive offloading and superficial acceptance of AI-generated output (Chen et al., 2020; Hou et al., 2025; Kasneci et al., 2023). Higher education therefore appears to be navigating a transitional phase in which institutional policy, pedagogical practice, and student behavior evolve at different rates: universities are moving from blanket bans toward AI literacy guidelines, educators are redesigning assessments to integrate or restrict GenAI more deliberately, and students are often already using these tools autonomously for tasks ranging from brainstorming to algorithmic debugging (Amzalag & Kurtz, 2025; Cotton et al., 2024; Hou et al., 2025; Kurtz et al., 2024; Xia et al., 2024).

2.2. GenAI and Student Learning: Enhancement, Dependence, or Dual-Use Technology

Empirical research on GenAI-supported learning presents a divided picture. On the one hand, generative AI is increasingly framed as a powerful educational support tool, offering unprecedented opportunities for personalized learning and assessment (Chiu et al., 2023; Deng et al., 2025). Prior studies emphasize improvements in efficiency, access to information, and learner confidence, with students frequently using GenAI to clarify complex concepts, generate examples, and receive immediate feedback—particularly in technical and programming-related tasks (Denny et al., 2024; Wang & Fan, 2025; Zviel-Girshin, 2024). From this perspective, GenAI is conceptualized as an adaptive scaffold that can enhance motivation and learning outcomes when used reflectively and in combination with human instruction. Supporting this view, the Beimel et al. (2025) demonstrate that GenAI-supported instruction can positively influence students’ learning experiences by enhancing autonomy and perceived competence, while maintaining academic performance comparable to that of traditional teaching approaches.
At the same time, emerging scholarship cautions that the educational impact of GenAI is highly contingent on how learners engage with these tools. Bauer et al. (2025) emphasize that artificial intelligence can either support deeper learning or reduce tasks to a superficial level of completion, depending on the extent to which learners purposefully integrate its use and critically verify the information and conclusions it supplies. Other studies also document potential learning risks associated with unreflective use of GenAI. Reliance on AI-generated outputs may promote cognitive offloading and uncritical acceptance of information, leading to superficial engagement that undermines deeper understanding (Denny et al., 2024; Eke, 2023). Experimental and observational research suggests that students who rely heavily on GenAI may exhibit reduced depth of problem-solving and weaker knowledge transfer than peers who engage more actively with learning tasks (Bastani et al., 2024).
Recent empirical work also complicates claims that position GenAI primarily as a tool for educational equity. Wecks et al. (2024), in a post hoc analysis examining GenAI use in examination contexts, find that it is associated with diminished academic outcomes, including lower exam scores. Notably, the most pronounced performance declines were observed among high-potential students, suggesting that when misused, reliance on these tools may fundamentally impede rather than support the learning process.
Complementing these findings, in a study comparing traditional instruction with GenAI-supported self-regulated learning, the Amzalag et al. (2025) report that while students experienced increased autonomy and flexibility when using GenAI, they also expressed heightened concerns about reliability and critical thinking. Similarly, Zviel-Girshin (2024) finds that although AI tools can support novice programmers in tasks such as debugging, code comprehension, and enhancing real-world relevance, their use also introduces risks related to over-reliance, academic integrity, and weakened understanding of core programming concepts.
Taken together, this body of research indicates that the educational impact of GenAI is neither uniformly beneficial nor inherently detrimental, but instead reflects a dual-use dynamic shaped by context, task design, and learner regulation. Notably, much of the existing literature focuses on the use of GenAI in formative learning activities, coursework, or informal study contexts. Empirical investigations of how students engage with GenAI within assessment environments remain notably scarce, despite the likelihood that exam conditions—such as time pressure, evaluative consequences, and constrained decision-making—fundamentally shape students’ strategies, perceptions, and learning processes (Fan et al., 2025; Wecks et al., 2024; Xia et al., 2024).

2.3. GenAI, Metacognition, and Self-Regulated Learning

Self-regulated learning (SRL) theory provides a valuable framework for understanding the nuanced role of GenAI in learning contexts. SRL theory emphasizes learners’ active involvement in planning, monitoring, and evaluating their own cognitive processes, positioning learners as agents who strategically regulate both internal cognition and external resources (Zimmerman, 2002). From this perspective, learning effectiveness depends not merely on access to informational support, but also on students’ capacity to critically reflect on their understanding and to deliberately manage the tools they employ.
Two divergent theoretical streams can explain the mechanisms by which GenAI supports or hinders this process. First, cognitive load theory posits that the limited capacity of working memory constrains learning, and that instructional supports are most effective when they reduce extraneous cognitive load while preserving or enhancing the germane load associated with meaningful learning (Sweller, 1988, 2011). In this context, GenAI tools can serve as cognitive aids to simplify or scaffold information processing, thereby freeing cognitive resources for higher-order reasoning. However, when GenAI substitutes for rather than supports cognitive processing, it may also suppress germane load, limiting opportunities for deep conceptual engagement.
In contrast, constructivist theory (Bada & Olusegun, 2015; Piaget, 1954) posits that learners must actively construct their own knowledge through experience and reflection rather than passively absorbing information. This theory holds that a deeper understanding is built upon the active reconciliation of new data with prior schemas. In this model, the use of GenAI presents a unique risk: it allows students to bypass the essential cognitive processes of comprehension, analysis, and summarization (Bastani et al., 2024; Fan et al., 2025). That is, by providing complete or highly polished outputs, GenAI may prevent the productive cognitive effort that constructivist theory identifies as necessary for building durable knowledge (Bada & Olusegun, 2015).
In keeping with the idea that the value of GenAI for learning is contingent on learners’ engagement with these tools, empirical evidence can be brought to support both theories. In line with cognitive load theory, GenAI may function as a metacognitive catalyst that prompts students to verify information and reflect on their understanding (Tankelevitch et al., 2024); at the same time, in line with constructivist theory, it risks turning an active learning process into a passive consumption experience when used as an authoritative shortcut. This dual potential aligns with broader conceptions of AI literacy, which extend beyond technical proficiency to include critical awareness, epistemic responsibility, and reflective use (Ng et al., 2021). From an SRL perspective, AI literacy can therefore be conceptualized not only as technical competence, but as a metacognitive capacity that involves epistemic judgment, responsibility for final decisions, and the regulation of AI-supported learning processes (Ng et al., 2021; Panadero, 2017).
In assessment contexts, as in all learning settings, the educational value of GenAI depends on metacognitive engagement with the technology—that is, whether learners treat AI as a cognitive partner that supports monitoring, or as an authoritative shortcut that replaces thinking (Annapureddy et al., 2025). This distinction takes on an extra dimension in assessments, where time pressure and evaluative stakes may further shape students’ regulatory strategies. In these settings, time pressure functions not merely as an external constraint but as a central regulatory cue that may either activate heightened metacognitive monitoring or encourage uncritical reliance on AI outputs, depending on learners’ self-regulatory capacity (Sweller, 2011; Panadero, 2017). Likewise, evaluative stakes may influence metacognitive regulation by affecting students’ motivation (Fan et al., 2025; Xu et al., 2025). Consequently, understanding how learners self-regulate their use of GenAI during assessments, balancing the efficiency of cognitive support against the necessity of active knowledge construction, remains a critical yet underexplored area of research.

2.4. Assessment, Authenticity, and Academic Integrity in the Age of GenAI

In light of the developments reviewed above, recent scholarship suggests that the integration of GenAI calls for a fundamental rethinking of assessment in higher education, extending beyond concerns of misconduct toward broader questions of pedagogical purpose and institutional readiness (Kasneci et al., 2023; Kurtz et al., 2024; Zawacki-Richter et al., 2024). In a comprehensive scoping review, Xia et al. (2024) further argue that GenAI challenges conventional assessment paradigms by foregrounding the need to cultivate students’ self-regulated learning capabilities and institutional AI literacy, as well as to develop holistic pedagogical approaches and updated policy frameworks.
Traditionally, assessment practices in higher education have been grounded in assumptions of individual authorship, closed-resource conditions, and the direct measurement of student knowledge. The introduction of GenAI disrupts these assumptions, raising concerns regarding academic integrity, fairness, and the validity of assessment outcomes. In keeping with such concerns, early institutional responses largely framed GenAI as a threat to academic honesty and reacted by increasing reliance on surveillance mechanisms, detection tools, and restrictive policies (Cotton et al., 2024; Eke, 2023).
More recent scholarship, however, advocates a shift from integrity-focused discourse toward concepts of authenticity and realism in assessment design. González-Calatayud et al. (2021) argue that assessments should reflect the tools and practices students are likely to encounter beyond academia, including AI-supported problem solving. Similarly, Yusuf et al. (2024) suggest that the question is no longer whether GenAI should be permitted in assessments, but how its use can be aligned with learning objectives and transparent evaluation criteria.
Following these conceptual advances, gaining insight into students’ lived experiences is essential for informing assessment practices that balance integrity, authenticity, and meaningful learning. Yet much of the existing literature remains focused on assessment design and institutional policy at an abstract level. Empirical research examining how students themselves experience the legitimacy, fairness, and educational value of GenAI use during tests and examinations remains limited.

2.5. Research Gap and Contribution

Taken together, prior research highlights both the transformative potential and the pedagogical risks of generative AI in higher education. While a growing body of work has examined GenAI in learning and assessment contexts, major gaps remain.
From the perspective of the current research, most existing studies that address the use of GenAI in assessment do so indirectly, focusing on academic integrity, detection mechanisms, or the redesign of assessment formats, rather than on students’ real-world experiences. Very few empirical studies investigate how students actually employ GenAI during examinations, and how this experience reshapes their perceptions, strategies, and metacognitive engagement.
The present study addresses these gaps by examining students’ perceptions before and after a GenAI-supported exam, using a mixed-methods approach that combines quantitative measures with qualitative thematic analysis. In particular, it employs a pre–post design to capture differences in students’ understanding before and after direct participation in GenAI-supported exams, as well as qualitative insights into how students conceptualize learning, fairness, and self-regulation in AI-mediated assessment settings. It also integrates GenAI literacy with established theories of self-regulated learning when examining assessment contexts, something largely neglected by existing work. Overall, by conceptualizing GenAI as an integral cognitive scaffold rather than an external disruption, this study contributes novel empirical evidence on how GenAI reshapes student cognition, metacognitive regulation, and attitudes toward academic evaluation in assessment environments.

3. Research Objectives

This study examines differences between students’ initial expectations and their post-experience perceptions regarding the integration of generative AI into academic assessments. By investigating the transition from abstract anticipation to practical engagement, the research seeks to uncover how direct interaction with GenAI, as a “cognitive scaffold,” influences self-regulated learning strategies, metacognitive monitoring, and students’ evolving views on assessment authenticity. A central objective is to characterize the dynamics through which students navigate the use of these tools.
Following the above, we formulated two research questions:
RQ1: 
How do undergraduate students retrospectively describe their perceptions and expectations regarding the integration of generative AI into academic assessments prior to direct experience?
RQ2: 
How do students’ perceptions and attitudes change following direct practical experience with GenAI during an academic assessment?

4. Methodology and Research Characteristics

Our methodological framework was grounded in an interpretive qualitative paradigm, examining student responses to open-ended questions. This methodology facilitates in-depth exploration of how participants construct personal meaning from their experiences, accommodating varied interpretations, complexity, and potential inconsistencies within respondents’ perspectives. Given that our research design prioritized meaning construction over hypothesis validation within an interpretive qualitative lens, inferential statistical analyses were not central to the study design. Descriptive statistics were used in a complementary function to characterize participants’ demographic profiles and to analyze student preferences regarding the two assessment formats, providing a contextual baseline for the qualitative findings.

4.1. Participants

The study was conducted in two undergraduate courses within a university Engineering faculty. The two courses were taught by the same instructor and had similar pedagogical structures, assessment formats, and cognitive demands.
The first cohort consisted of students taking a required Data Structures course in the university’s Computer Science (CS) program. The second cohort comprised students taking a Data Structures and Algorithms course required for Engineering and Management (E&M) students specializing in Information Systems and Data Science. The former course was taught in the second term of the CS program (out of six), and the latter in the sixth term of the E&M program (out of eight).
The final sample for the current study comprises 90 participants who submitted full responses. This final sample includes 45 E&M students and 45 CS students, 39 females (43%) and 51 males (57%), aged 22 to 30. Participant demographics are summarized in Table 1.

4.2. Experimental Procedure

The research procedure had two phases (both identical for the two cohorts):
Phase 1: During week 11 of the academic term, participants sat for a brief classroom-based assessment (hereafter: quiz). Students were notified about the quiz and its content areas (AVL tree algorithms for CS students and BFS/DFS algorithms for E&M students) approximately 14 days in advance. Students were authorized to bring a single page of notes for reference during the quiz. The allotted time for completing the quiz was roughly 30 min.
Before starting the quiz, all participants provided written informed consent for use of their data in the study. Participants also provided basic demographic information, including age and gender.
Phase 2: In week 12 of the term, a second quiz was administered. As before, students were given notice of the quiz and its contents (non-linear algorithms for CS students and the PRIM algorithm for E&M students) about two weeks in advance. However, this time, participants were notified that the quiz would be conducted in a computing lab and that they would have access to GenAI resources throughout the assessment. The quiz was delivered electronically through the Moodle platform (https://moodle.com/). Again, participants were given roughly 30 min to complete the quiz. As in the prior session, informed consent was obtained from all participants in advance. To preserve naturalistic conditions, participants received no special guidance on implementing GenAI and were free to use their preferred tools and approaches.
Following the week-12 quiz, students were asked to complete a brief post-quiz survey, which comprised the main research instrument. See Section 4.3.

4.3. Research Instrument: Post-Quiz Questionnaire

As noted, immediately after completing the week-12 assessment, students completed an online post-quiz questionnaire designed to capture their experiences and perceptions of using AI tools during the assessment. The questionnaire included three closed-ended items and two open-ended questions. Questionnaires for the two cohorts were identical.
The two open-ended questions examined students’ perceptions before and after experiencing the week-12 quiz, which allowed internet access and the use of GenAI tools. The first asked students to describe their thoughts and expectations about the GenAI-enabled assessment format before taking the quiz (‘What were your thoughts before the quiz regarding this assessment method, which includes the use of open resources, internet access, and the possibility of using AI tools?’), while the second invited them to articulate their views after experiencing this assessment approach (‘What were your thoughts after the quiz regarding this assessment method, which includes the use of open resources, internet access, and the possibility of using AI tools?’). This retrospective pre–post design was intended to capture potential shifts in student attitudes and to generate insights into how direct experience with GenAI-assisted testing influences learner perspectives. However, it is important to note that because the ‘pre-quiz’ expectations were collected retrospectively, these reported prior attitudes may have been subconsciously influenced by the students’ subsequent direct experience. Therefore, the identified shifts are interpreted with appropriate caution, representing students’ reconstructed narratives of their changing perspectives rather than independent pre- and post-measurements.
After completing the open-ended questions, participants responded to the three closed-ended items. The first two of these concerned which GenAI tools were used during the quiz and the extent to which the student checked GenAI-provided responses. Although the detailed evaluation of these behavioral metrics is reserved for other work, it was found that most students used prompts focused on the specific exam topics (i.e., AVL tree algorithms for CS students and BFS/DFS algorithms for E&M students), and typically exhibited behaviors ranging from direct copy-pasting of AI outputs to adopting a skeptical stance, actively debating with the AI, and rigorously verifying the provided responses. This snapshot conceptually grounds the qualitative statements regarding metacognitive monitoring discussed later in this paper.
The third item asked participants to indicate their preferred assessment format: examinations with access to the internet and GenAI tools versus traditional closed-book examinations. This item was integral to our research design. Together, the two open-ended items and the closed-ended preference measure constitute the core of the present study.

4.4. Data Analysis

Data analysis followed a qualitative-dominant, mixed-methods approach, integrating descriptive quantitative analysis of student preferences with a rigorous qualitative thematic inquiry of participants’ open-ended reflections. This combined approach allows for a comprehensive understanding of both the overall distribution of student attitudes and the nuanced meanings they ascribe to their experiences.
Quantitative data derived from the closed-ended preference item were analyzed using descriptive statistics. Frequencies and percentages were computed to determine the overall distribution of student preferences for the two assessment formats (GenAI-enabled versus traditional).
The qualitative component, which formed the primary focus of our analysis, was grounded in the thematic analysis methodology articulated by Braun and Clarke (2006). To facilitate a systematic and rigorous analysis, the qualitative data were organized and managed using ATLAS.ti (version 25) software. To maintain analytical precision and reliability, we implemented a structured, recursive coding procedure. Initially, coding proceeded inductively without predefined thematic categories. In the initial phase, two independent coders examined a subset of 15 student responses to each of the two open-ended questions, sampled randomly from both cohorts (the CS and E&M classes). This exploratory stage enabled us to establish preliminary codes and detect emergent thematic patterns. For instance, when analyzing responses to the question about pre-examination expectations for the GenAI-supported assessment format, one participant noted: “I thought it would make everything easier and faster, but I was concerned about whether I could trust the AI answers.” From this excerpt, we identified three preliminary themes: anticipated efficiency benefits, time-related expectations, and trust concerns. Another student’s pre-examination response stated: “I expected it would level the playing field for everyone and reduce memorization pressure, though I worried about technical issues.” This response was also analyzed as including three themes: equity perceptions, reduced cognitive load, and technical apprehensions. Based on this process of collaborative deliberation on the initial codes and emergent themes, we constructed a provisional codebook with explicitly defined categories and precise inclusion and exclusion criteria for each theme. As coding progressed, related preliminary codes were iteratively reviewed, refined, and consolidated into broader themes. The main themes and subthemes reported in the findings thus represent the outcome of this consolidation process, rather than the initial open coding stage.
The frequency percentages reported in the findings reflect the proportion of coded thematic references to the total number of coded references for each question. Because individual responses could contain multiple thematic elements, a single participant’s comment could be counted under more than one theme. In the next phase, two coders independently analyzed an additional sample of 25 responses for each open-ended question (50 responses total, again sampled randomly from the CS and E&M cohorts). They then reconciled differences in their coding through structured dialogue to create a definitive codebook. By the end of this process, substantial intercoder concordance was attained, with 47 of 50 coded responses (94%) demonstrating complete alignment. An additional evaluation of intercoder reliability using Cohen’s Kappa yielded κ = 0.788 (p < 0.001), indicating near-perfect agreement. The absence of novel themes during this validation phase confirmed the comprehensiveness and stability of our coding architecture.
In the third phase, using this definitive codebook, the lead researcher systematically coded a calibration dataset taken from the responses used to develop the codebook (30 responses per open-ended question, 60 total). A second researcher then conducted an audit trail review by independently examining a randomized subset of 20 responses per question, for a total of 40 responses. Complete concordance was observed between the two coders, further substantiating the internal consistency and credibility of the analysis. Following this validation procedure, the principal investigator proceeded to code the entire corpus, completing analysis of all participant responses to both open-ended questions.
In total, the coding process yielded 167 thematic segments from the pre-quiz reflections and 140 thematic segments from the post-quiz reflections. Figure 1 provides a visual overview of this rigorous three-phase data analysis process, outlining the inductive coding, validation steps, and intercoder reliability metrics.

5. Findings

This section presents the results of our analyses. We first report the results for the two open-ended questionnaire items, structured around the two research questions. This is followed by a brief exploration of how individual student attitudes evolved across the assessment phases. We then report the results for the closed-ended preferred assessment format question. As the present study did not examine cohort-specific differences, data from the two cohorts (CS and E&M) are reported together.

5.1. RQ1: Perceptions Prior to the Intervention

Our first research question asked how undergraduate students retrospectively describe their expectations and perceptions regarding the use of GenAI in academic assessments prior to direct experience. Analysis of students’ responses revealed six main themes: Emotional and Experiential Aspects (25%), Learning and Understanding (22%), Metacognitive Skills and Monitoring (17%), GenAI Literacy (14%), Time Factor (11%), and Systemic and Social Perceptions (11%). These themes consist of several subthemes, as detailed next.
Emotional and Experiential Aspects (25%). This theme was among the most prominent in students’ prior expectations, focusing on the anticipated affective impact of integrating GenAI into the assessment process. This theme is organized into two analytically related subthemes: positive expectations of reassurance and reduced anxiety (19%), and negative feelings, such as stress or discomfort (6%).
Many students reported that, prior to the intervention, they anticipated that access to GenAI would alleviate the pressure typically associated with academic assessments. One student shared the hope that the method would provide emotional support: “I thought it would be a great method, it would make me feel more secure and calm” (Student 9). In contrast, other students expressed initial resistance or unease about the change to the class’s familiar format. For instance, one participant described the inclusion of GenAI as “strange and unnecessary” (Student 15). Another student worried that the complexity of the new format might be counterproductive: “I thought it would be confusing and stressful rather than helpful” (Student 48).
Learning and Understanding (25%). This theme was equally prominent, accounting—like the Emotional and Experiential Aspects theme—for 25% of the responses. Within this theme, three primary subthemes were identified: potential contribution to understanding (13%), concerns about superficial understanding, cognitive atrophy, or over-reliance (9%), and a preference for quick answers over a full cognitive process (3%).
Many students anticipated that support from GenAI would allow for easy access to details, letting the student focus more on engaging with the overall task rather than remembering every piece of information learned: “I thought it would be easier with the chatbot’s help, because even if I am not sure about some details, I will have access to them” (Student 54). In addition, students expected that the tool would serve as a scaffold for their understanding in real-time and strengthen their self-confidence: “It is an amazing method because if I forget something or want to double-check, it really helps me” (Student 2).
Other students reported concerns that having GenAI as a supportive resource might undermine learning by providing rapid answers that bypass deep thinking: “Before the quiz, I thought that using GenAI might interfere with deep understanding, because it provides answers too quickly” (Student 17); “It seems nice, but it reduces the quality of learning for the quiz because you feel like you have a safety net” (Student 30). This concern extended to a fear of developing dependency, where the ease of access might erode independent problem-solving habits: “I’m afraid it will make me not even try to solve it on my own, but run straight to the chat” (Student 72).
Finally, a smaller group of students anticipated that the mere availability of the tool might undermine their ability to engage in independent thought, leading to a preference for quick answers: “Because I have the tool, I am less likely to think for myself and go directly to asking the chatbot” (Student 76). Relatedly, another student worried that adopting a GenAI-supported assessment method could harm the quality of quiz responses: “Before the quiz, I thought that assessment with a GenAI tool could become problematic, leading to answers that are more hasty and superficial” (Student 80).
Metacognitive Skills and Monitoring (17%). This theme reflects students’ critical and reflective approach to GenAI prior to the intervention. Their responses highlighted three primary subthemes: the need for verification and validation of GenAI-provided responses (10%), use as a tool for self-monitoring (4%), and GenAI fallibility as a prompt for deeper examination (3%).
Regarding the need for verification and validation of GenAI-provided responses, students reported feeling a high degree of skepticism regarding the reliability of GenAI, repeatedly emphasizing that its outputs should not be accepted at face value. This critical stance was rooted in an understanding that the technology often lacks precision, produces “nonsense,” or misses subtle nuances, thereby requiring constant external validation: “Using such an open-resource tool does not really test me, except for how much I trust it, and from my experience, I always need to double-check it a lot” (Student 12). This wariness was reflected in a similar comment by another student: “The GenAI can miss the small details, so it cannot be fully trusted” (Student 44). Such concerns even prompted some to question the overall validity of the proposed assessment method (subtheme four): “I thought that assessment with a GenAI tool could become problematic” (Student 80). In addition, a small number of students articulated ethical reservations regarding the appropriate role of AI in assessment contexts (subtheme five): “I believe that artificial intelligence should help us learn the material, not solve the exam. The challenge must remain with the examinee, not the machine” (Student 71). Interestingly, however, this fallibility was also seen as a catalyst for deeper learning (subtheme three); the realization that the tool is not always correct allows for a deeper examination of the student’s understanding of the subject matter: “Yes, because the GenAI is not always correct, and this allows you to examine more deeply the student’s understanding” (Student 61).
Despite these reservations, several students envisioned leveraging GenAI as a strategic partner for cognitive regulation (subtheme two). For these participants, the tool was seen as a valuable asset for clarifying task requirements and verifying their own reasoning: “It is good for self-checking and for helping with questions and understanding them” (Student 24). Similarly, some students framed the interaction with GenAI as a dialectical process, allowing for real-time feedback and the refinement of their ideas through external critique: “I thought this method would give me an opportunity to consult with another source that could critique what I think and explain it to me in real time” (Student 56).
GenAI Literacy (11%). This theme concerns the technical and conceptual competencies required for effective use of GenAI. It includes two primary subthemes: proficiency in prompt engineering (7%) and understanding the technology’s operational principles (4%).
Regarding the former, students emphasized that the utility of the tool in an academic context depends on the user’s ability to operate it in an informed manner, not merely through basic use. They noted that the quality of the output is directly linked to the prompt chosen by the user: “It requires skill in writing the right prompt to get accurate information” (Student 35). In terms of the latter, participants recognized that effective interaction with GenAI requires an understanding of the principles of the technology’s operation: “It’s good if you know how to use it correctly… you need to understand how the system works” (Student 18).
The Time Factor (11%). This theme addresses the complexity of time in students’ perceptions of GenAI use prior to the intervention, as retrospectively reported. The theme examines time from two contrasting angles: as a resource for increased efficiency on the one hand, and as a constraint that may induce pressure on the other. Within this theme, two subthemes were identified: time constraints for thinking and monitoring (6%), and reliance on the tool due to time pressure (5%).
Some participants viewed the technology as an efficiency aid that would streamline information retrieval and facilitate a smoother assessment process: “It saves time searching through slides, it helps me meet the deadlines” (Student 14). Conversely, other students warned that the pressure of the clock might incentivize haste and compromise the verification process (subtheme two). This tension suggests that when time is perceived as a scarce resource, the ease of generating rapid responses may lead to less rigorous cognitive engagement: “In a time crunch, it’s easy to just copy what it says without thinking” (Student 4).
Systemic and Social Perceptions (11%). The final theme identified for RQ1 highlights students’ broader considerations regarding the evolving role of GenAI in higher education and the professional world. Within this theme, two primary subthemes were identified: adapting academia to the new era (6%) and preparing for the labor market (5%).
Participants argued that educational institutions must evolve to align with global technological trends rather than remaining isolated from them: “Academia needs to move forward; this is the real world today” (Student 40). This perspective extended to professional readiness, as students perceived proficiency in GenAI not as an optional skill but as a prerequisite for remaining relevant in the modern job market: “It is important to learn how to use this because it is what will be required of us at work” (Student 63).
For a systematic and comprehensive overview of the thematic distribution, Appendix A provides a detailed table summarizing the frequencies of all themes and subthemes across the pre-assessment phases.

5.2. RQ2: Attitudes and Experiences After the Intervention

Our second research question asked how students’ perceptions and attitudes change following direct practical experience with a GenAI-supported assessment. To address this question, we examined students’ reflective responses following completion of the quiz. The analysis revealed seven main themes, including six recurring themes consistent with those identified prior to the intervention: Emotional and Experiential Aspects (18%), Learning and Understanding (24%), Metacognitive Skills and Monitoring (20%), GenAI Literacy (12%), Time Factor (12%), and Systemic and Social Perceptions (8%). There was one newly emerged theme, Authenticity and Fairness of Assessment (6%). As with RQ1, each theme includes several subthemes, as detailed below.
Emotional and Experiential Aspects (18%). The frequency of this theme decreased from 25% in the responses for RQ1 to 18% in those for RQ2. This shift suggests that practical experience may have tempered some of the students’ anticipated affective concerns and expectations toward more practical and cognitive elements. As before, the theme included two subthemes: positive feelings (12%) and negative feelings (6%).
In terms of the former, direct interaction with GenAI during the quiz was associated with decreased stress levels and increased self-confidence among some students. Access to generative AI provided a sense of security during the assessment and allowed students to focus on answering the questions: “It was a great method, I could understand the questions well without feeling stressed” (Student 10); “After the quiz, I felt that the assessment really tested our ability to think… it was less stressful and more enriching” (Student 90).
Conversely, for other participants, the experience elicited frustration and disillusionment. For these students, what they perceived as forced reliance on the tool due to external constraints such as time pressure—perhaps along with failure to adequately prepare—created an emotional sense of letdown regarding their own performance: “I was disappointed because I felt that I could have solved the quiz very well by myself, but the short time made me depend on the chatbot’s answers” (Student 50).
Learning and Understanding (24%). This theme remained a primary focus of students’ reflections following completion of the GenAI-supported quiz. It comprised three subthemes: contribution to understanding during the quiz (11%), real-time verification and reassurance (7%), and the learning effects of anticipated reliance on GenAI in light of time pressure (6%).
Several students emphasized that the integration of GenAI supported their comprehension during the quiz by providing immediate explanations and examples: “After the quiz, I realized that integrating tools like GenAI has advantages. You can receive an immediate explanation… and compare sorting methods… This helps to understand the advantages and disadvantages of each method” (Student 71). GenAI was also perceived as an interpretative aid that clarified the logic of the questions: “The GenAI not only explained the question but also provided examples of solutions and counterexamples when necessary” (Student 38).
Students also highlighted the tool’s value as a mechanism for verifying their answers in real time, which bolstered their confidence and provided them with a “safety net”: “It helped me to make sure my answers were correct” (Student 27); “The GenAI helped me confirm my understanding of the question and the solution, giving me assurance that I was on the right track” (Student 82).
Finally, several students linked the time constraints of the quiz to a compromise in their learning process and their autonomy. For some, the need for speed overrode the opportunity for deep understanding: “I had to provide answers quickly under time pressure, rather than really learn during the quiz” (Student 14). For others, the knowledge that AI would be available led to reduced preparation, which, combined with time pressure, led them to rely on the tool: “I prepared less for the quiz because I had the AI. With more time, I could have solved the questions on my own and only used GenAI for checking, but the limited time forced me to rely on it” (Student 68).
Metacognitive Skills and Monitoring (20%). The frequency of this theme increased from 17% in the responses for RQ1 (students’ reports of their expectations prior to the quiz) to 20% in the responses for RQ2, suggesting that direct experience with GenAI during the quiz made students more aware of both the limitations and the value of GenAI in terms of metacognitive monitoring, verification, and control. The coding produced three subthemes, all of which had also appeared in the responses for RQ1: the need for verification and validation (8%), the use of GenAI for real-time self-monitoring and regulation (7%), and awareness of potential errors (5%). As in RQ1, awareness of potential errors refers to the perceived fallibility of GenAI as a source of information, which may prompt deeper examination, whereas the need for verification captures the regulatory act of checking and validating responses in practice.
Most often, direct interaction with the tool influenced students to adopt a more skeptical and critical stance toward AI-generated content, as they realized that the tool’s potential for errors requires constant control and verification: “There is a great need to verify ChatGPT because I felt it made many mistakes and could easily mislead me if I was not careful” (Student 29). In other cases, the tool enabled students to verify their own reasoning and check their answers in real time during the quiz, providing a way to monitor their problem-solving steps: “It helped me confirm my answers and review the way I solved the question” (Student 27). Notably, references to verifying and controlling GenAI-produced outputs were more frequent than references to using GenAI to validate one’s own reasoning, suggesting that students were more preoccupied with managing the tool’s reliability than leveraging it for self-regulation.
GenAI Literacy (12%). The frequency of this theme decreased from 14% to 12% in responses for RQ1 and RQ2, respectively. The theme included two subthemes: the importance of prompt engineering (7%) and student responsibility for the final answer (5%).
Direct interaction with GenAI during the quiz influenced students to recognize that the quality of the suggested answers depended heavily on the user’s ability to provide precise instructions. They realized that the way they phrased the prompt directly affected the accuracy of the information they received: “I understood that the result I get is only as good as the prompt I write, so I had to be very specific to get a helpful answer” (Student 41).
Furthermore, the experience reinforced the understanding that the student, rather than the AI, is ultimately responsible for selecting the correct answer in the quiz: “Even though the AI suggests an answer, the responsibility for the final selection is mine alone, and I must know how to direct the tool correctly” (Student 6).
The Time Factor (12%). The frequency of this theme slightly increased from 11% in the responses for RQ1 to 12% in the responses to RQ2. While students initially anticipated that GenAI would primarily serve as a time-saving tool, their experience revealed that time was a significant constraint that hindered their ability to verify the tool’s answers.
Direct interaction with GenAI during the quiz highlighted how time pressure influenced students’ use of the tool. For many, the limited time frame transformed GenAI from an efficient tool into a factor that prevented them from performing a thorough review and verification of the suggested answers: “The time was very limited, which made it difficult to check the answers provided by the AI” (Student 45). Furthermore, interaction with generative AI during the quiz showed that the speed of generating answers can be misleading when the user is required to validate information under pressure: “I felt like I was racing against the clock, and even though the AI provided answers quickly, I did not have enough time to ensure that they were correct” (Student 22).
Systemic and Social Perceptions (8%). The frequency of this theme decreased from 11% to 8% in responses for RQ1 and RQ2, respectively. This decline suggests that following practical experience, students’ reflections shifted somewhat away from broader systemic considerations toward more immediate cognitive and experiential aspects of GenAI use during the assessment. Nevertheless, students continued to express views regarding the inevitable integration of GenAI into the academic and professional worlds. The theme included two subthemes: preparation for the future labor market (5%) and the integration of GenAI into higher education (3%).
Direct interaction with GenAI during the quiz influenced students’ views of these tools as integral to their future professional development. They noted that academic institutions must adapt to these technological changes to remain relevant: “In today’s reality, most of the labor market already uses GenAI tools; it is important and valuable that we learn how to work and study alongside GenAI in a thoughtful and responsible way” (Student 31). Likewise, the experience highlighted the systemic need for a shift in how knowledge is assessed in the digital age: “This is a cool method; the world now relies more on GenAI, so I think academia also needs to adapt to the changing world” (Student 15).
Authenticity and Fairness of Assessment (6%). This theme was only marginally present prior to the intervention, where ethical concerns appeared as a minor subtheme under Metacognitive Skills and Monitoring. Following the practical experience, however, it emerged as a distinct theme, suggesting that direct engagement influenced students to consider more explicitly the ethical implications and validity of the assessment process when using GenAI. The theme included two subthemes: integrity of the assessment process (4%) and the challenge of evaluating individual knowledge (2%).
Direct interaction with GenAI during the quiz led students to question the quiz’s authenticity, given that AI provided substantial assistance. Students expressed concerns that the GenAI-supported format might not accurately reflect a student’s true understanding: “If everyone uses the AI, the quiz no longer tests our knowledge but rather our ability to operate the tool, which raises questions about the fairness of the grade” (Student 55).
Furthermore, this experience highlights a new perspective on academic integrity in the age of AI. Some students realized that the traditional definition of an independent assessment is changing, leading to a sense of ambiguity regarding the value of the assessment: “It felt a bit like cheating even though it was allowed, because the tool did most of the logical work for me” (Student 12).
For a systematic and comprehensive overview of the thematic distribution, Appendix B provides a detailed table summarizing the frequencies of all themes and subthemes across the post-assessment phases.
Table 2 summarizes the frequency of citations for each theme before and after the intervention. This comparative overview highlights shifts in thematic prevalence and identifies whether the frequency of each theme in the students’ remarks increased or decreased as a result of practical experience.

5.3. Longitudinal Tracking of Individual Student Attitudes

Beyond general trends across the entire sample, we tracked individual-level changes by comparing each student’s pre- and post-quiz responses. This tracking enabled the identification of cases in which the same student changed their perspective after the practical experience. Notably, the majority of participants (91%) maintained consistent views between the two stages: only 9% shifted between primary themes, and 13% shifted between subthemes. These changes were not radical reversals, but rather subtle changes in emphasis that illustrate how direct experience with GenAI could refine or reframe existing attitudes. We identified five main patterns, each relating to one of the main themes identified: skepticism to appreciation (learning and understanding); trust to frustration (emotional aspects); abstract concerns to practical understanding (metacognitive monitoring); increased focus on the concrete cost of time; and enthusiasm to caution (systemic and social perceptions). The five patterns presented below were identified through an iterative and systematic review of the entire dataset. After reading all student reflections multiple times, we identified the primary dynamics of change that emerged from the data. We then selected representative cases and quotes that best illustrate the variety and spectrum of these experiences across the two cohorts.
From Skepticism to Appreciation: Some students moved from concerns about the effects of GenAI on deep learning to recognizing the tool’s pedagogical value. For instance, Student 17, reporting on perceptions before the quiz, noted: “I thought that using artificial intelligence might interfere with deep understanding because it provides answers too quickly.” After the quiz, however, the same student described a different effect: “I realized that integrating tools like GenAI has advantages… it helps to understand the strengths and weaknesses of each method… Yet the understanding comes from the student and not only from the GenAI tool.”
From Trust to Frustration: Another student experienced an emotional shift. Before the quiz, she expressed trust in the tool and anticipated feeling more confident and secure with its support: “I thought it would be simpler with the chatbot’s help. Even if I was unsure about some details, I would have access to them, and that would make it much easier.” After the quiz, her experience was characterized by disappointment: “I was disappointed because I felt I could have solved the questions very well by myself and only used the GenAI for checking. Instead, my answers reflected time pressure rather than full comprehension” (Student 50).
From Abstract Concerns to Practical Understanding: Shifts were also observed in students’ practical understanding of the metacognitive monitoring required when using GenAI. One student, in reporting his thoughts before the quiz, expressed a general sense of caution about the tool’s reliability: “It is a good method, but the tool is not always stable and might suddenly give wrong answers.” Following the quiz, his understanding became more concrete and practice-based, as reflected in his comment: “There is a great need to verify ChatGPT. I felt it made many mistakes… therefore, I could not trust it blindly” (Student 21).
The Concrete Cost of Time: Regarding time management, Student 77 emphasized a general rule when reporting his perceptions before the quiz: “You always need to check the GenAI before relying on it.” Afterward, he described how the reality of the quiz shaped this process: “The fear that the GenAI might be wrong cost me time during the exam… I kept comparing its answers with the lecture slides, which took time instead of solving independently” (Student 77).
From Enthusiasm to Caution: In some cases, systemic perspectives shifted. Before the quiz, Student 67 expressed enthusiasm toward the GenAI-supported format, perceiving it as an innovative assessment approach. However, after the quiz, the same student expressed caution and doubt regarding its validity: “I think this does not really test knowledge and understanding, because you can just copy answers from the chatbot” (Student 67).

5.4. Preferred Assessment Format

Following their experience with the GenAI-enabled quiz, participants indicated their preferred assessment format in a single closed-ended question. Of the 90 participants in the study, 89 responded to this item. Of those respondents, 68.5% (n = 61) preferred the GenAI-supported format, while 31.5% (n = 28) preferred the traditional method. To investigate whether these preferences were driven by academic success, a point-biserial correlation was computed between final quiz grades and preferred format. The analysis revealed a near-zero, non-significant correlation (r = 0.014, p = 0.89), and an independent samples t-test confirmed no significant difference in grades between students who preferred the GenAI format (M = 70.44, SD = 19.78) and those who preferred the traditional format (M = 71.00, SD = 15.46).

6. Discussion

The present study examined students’ feelings and attitudes toward the use of GenAI in academic assessments before and after direct experience with using GenAI tools during a computer science or information systems course quiz. We explored how this experience influences students’ perceptions of learning, self-monitoring, cognitive regulation, and attitudes about academic assessments with and without access to GenAI.
In summary, our findings demonstrate that at the end of the study, 68.5% of respondents reported preferring GenAI-enabled assessments. Yet while participants initially anticipated that GenAI would provide reassurance and efficiency, the qualitative data revealed a subtle shift to what we might call measured pragmatism. More precisely, the experience heightened students’ understanding of GenAI as a cognitive scaffold that demands high-level metacognitive monitoring, such as the need to continuously verify potentially erroneous outputs (Tankelevitch et al., 2024). Notably, students’ quiz performance showed no significant correlation with a preference for AI-enabled assessments, meaning that these preferences are driven by perceived pedagogical and professional value rather than individual academic success (Yusuf et al., 2024).
The findings indicate that students’ direct experience using GenAI during the quiz did not substantially change their overall perceptions about the effects of integrating such tools into academic assessments. In both their retrospective reports about pre-experience attitudes and their post-experience reflections, students explicitly referred to process-related factors, such as using the tool for verification and checking, managing time constraints, and assuming personal responsibility for the final answer. This pattern aligns with the self-regulated learning literature, which emphasizes that regulatory processes do not change abruptly but are gradually adapted to context, task demands, and situational conditions (Panadero, 2017; Zimmerman, 2002).
One of the central contributions of this study concerns the metacognitive dimension of students’ engagement with GenAI during assessments. Although the increase in references to metacognitive skills was quantitatively modest, qualitative analysis revealed a substantive shift in the nature of students’ monitoring practices. Prior to direct experience, students primarily expressed a general awareness of the need to verify AI-generated outputs and a cautious stance toward the tool’s reliability. Following the exam, however, students described concrete, real-time monitoring practices, including checking AI-generated answers against their own reasoning, comparing outputs with additional sources, and making deliberate decisions about when to rely on the tool and when to refrain from using it. Based on our interpretive analysis of the qualitative material, we conclude that GenAI does not necessarily reduce metacognitive engagement but appears to increase the metacognitive demands placed on learners, particularly when AI outputs are perceived as uncertain and require ongoing critical evaluation (Ng et al., 2021; Tankelevitch et al., 2024). In this sense, AI literacy is reflected not merely in technical proficiency, but in learners’ capacity to exercise judgment, critically evaluate information, and assume responsibility for their learning-related decisions during assessment (Annapureddy et al., 2025; Ng et al., 2021).
Similarly, the findings on time highlight its central regulatory role in students’ engagement with GenAI during assessments. While changes in the salience of this theme were modest, a clear shift emerged in how participants experienced and interpreted time. Prior to the experience, time was primarily framed as an efficiency resource that GenAI could help save. Following the exam, however, time was identified as a constraint that limited students’ ability to verify AI-generated outputs and fully engage in metacognitive monitoring. This pattern is consistent with insights from cognitive load theory, which suggests that time pressure increases extraneous cognitive load and may undermine deep processing and effective self-regulation (Sweller, 2011). In assessment contexts, the combination of rapid AI output and limited time appears to create concrete and immediate tension between the need for efficiency and learners’ desire to improve their understanding.
Notably, alongside an increased focus on the cognitive and metacognitive aspects of GenAI use, students’ reflections showed a diminished emphasis on broader institutional and social considerations, including the role of higher education institutions, assessment policy, and systemic preparedness for the AI era. This pattern does not suggest that such considerations disappeared, but rather that they receded into the background in favor of more immediate and practical engagement with the exam task itself. This shift is consistent with prior literature, indicating that direct experience with GenAI tends to redirect attention away from abstract discussions of policy and institutional adaptation toward personal decision-making, responsibility, and self-regulation in concrete assessment contexts (González-Calatayud et al., 2021; Xia et al., 2024).
Finally, the emergence of concerns about assessment authenticity and fairness only after direct experience underscores the role of practical engagement as a catalyst for ethical awareness in the field. Prior to the GenAI-supported quiz, students primarily articulated general considerations regarding the future of higher education and the labor market; following the quiz, they raised concrete questions about the validity of such assessments, the evaluation of individual knowledge, and the fairness of grading practices. This finding reinforces the argument that issues of integrity and authenticity do not arise solely at an abstract or principled level, but are instead shaped through direct encounters with changing assessment conditions (Xia et al., 2024; Yusuf et al., 2024).
The quantitative findings complement the patterns identified in the qualitative analysis. Although most students preferred an exam format that integrated GenAI, this preference was not associated with actual academic performance. This finding indicates that students’ positive perceptions of GenAI-supported assessment are not driven by instrumental considerations related to performance improvement, but rather by broader pedagogical and professional considerations, as well as concerns about the authenticity of learning. These statistical findings, combined with the qualitative insights, point toward an interpretive conclusion of a measured pragmatism in students’ attitudes-a recognition that with GenAI here to stay, it behooves all of us (students, educators, and the public) to work out its implications for teaching and learning.
The observed emphasis on verification and the shift toward measured pragmatism should be viewed through the lens of the participants’ disciplinary background. As computer science and engineering students, they are professionally socialized to debug, validate, and exercise skepticism toward automated outputs. This technical literacy, which likely includes more sophisticated prompting strategies, acts as a critical catalyst for the high-level metacognitive monitoring and individual responsibility identified in our findings.

Practical Contributions

Our findings indicate that academic assessments that integrate GenAI place greater demands on students’ ability to exercise judgment, verify information, and assume responsibility for final answers. Accordingly, the design of GenAI-supported assessments cannot be limited to binary decisions on whether to permit or prohibit these tools, but rather requires explicit clarification of expectations concerning how GenAI may be used, the limits of reliance on system-generated outputs, and students’ active role in monitoring the quality of responses and making final decisions.
In addition, the findings underscore the role of time pressure in shaping students’ engagement with GenAI during examinations. Under time constraints, students experience tension between using the tool for efficiency and the need to thoroughly verify and evaluate responses provided by the tool. This finding suggests that time allocation is not merely a technical parameter in the design of GenAI-supported assessment, but a variable that directly influences how students balance efficiency, depth of understanding, and self-regulatory control. Consequently, assessment designs that permit the use of GenAI must attend to the pedagogical implications of time constraints and their role in shaping the processes of verification, judgment, and responsibility during examinations.
Finally, the emergence of concerns related to assessment authenticity and fairness only after direct experience indicates that ethical awareness in the context of GenAI-supported assessment develops through practical engagement with assessment conditions rather than solely through abstract principles or general guidelines. This finding highlights the importance of embedding opportunities for reflection on assessment validity, individual knowledge evaluation, and grading fairness as integral components of assessment designs that incorporate GenAI. Crucially, our data indicates that from the students’ perspective, issues of authenticity and academic integrity are not merely about preventing misconduct, but rather about ensuring a fair assessment environment where true cognitive effort can be reliably demonstrated and evaluated.

7. Limitations

While this study provides nuanced insights into students’ evolving perceptions of GenAI in assessment settings, several limitations should be considered.
Disciplinary and Institutional Specificity: The research was conducted within a single institution and focused exclusively on computer science and engineering students. Given the technical nature of these disciplines, students’ perceptions of AI as a problem-solving or coding aid may differ significantly from those in the humanities or social sciences. Therefore, the identified themes, such as the emphasis on algorithmic verification, may not fully reflect the attitudes of the broader student population. Specifically, the participants’ background in programming and algorithmic design likely predisposed them toward a rigorous verification-oriented approach to AI outputs. This technical affinity suggests that their prompt engineering skills might be higher than those of students in non-technical fields, potentially shaping their transition toward measured pragmatism. In accordance with Guba and Lincoln (1994), these findings should be interpreted through the lens of “transferability” rather than statistical generalizability, acknowledging that the specific academic culture of the participants is an inseparable part of the data’s context.
Retrospective Pre-Post Design and Recall Bias: A central limitation is the retrospective nature of the pre-quiz reflection. Participants were asked to describe their initial expectations after having completed the AI-enabled quiz. This direct experience likely influenced their memory of their prior attitudes, potentially leading to a recalibration of their reported expectations based on the challenges or successes they just encountered. From a conceptual standpoint, these ‘before’ data are essentially retrospective reconstructions subject to memory distortion and hindsight bias. Therefore, the reported pre-assessment attitudes should be interpreted as the students’ synthesized understanding of their experience, rather than an absolute baseline measurement. This phenomenon, often termed “recall bias,” is a well-recognized challenge in educational self-reporting (Cohen et al., 2002), where current knowledge or experiences can unconsciously reframe an individual’s interpretation of their past beliefs.
Influence of Task Context and Time Pressure: The qualitative reflections were deeply rooted in a specific experimental context: a 30 min time-limited quiz. The dominance of the “Time Pressure” theme (12% of post-quiz reflections) suggests that students’ attitudes toward GenAI were shaped to a considerable extent by the limited time frame. In a less constrained setting, such as a take-home assignment, students might report different perceptions of GenAI’s role in deep learning versus its role in efficiency.
Self-Report and Social Desirability Bias: Because the data relied on open-ended self-reports, students’ responses may have been influenced by social desirability, leading them to express attitudes they perceived as “appropriate” for an academic setting, particularly regarding academic integrity and verification. While the anonymity of the responses was expected to mitigate this, the possibility of social desirability bias cannot be excluded. As noted by Podsakoff et al. (2003), such common-method biases are inherent in behavioral research, particularly when students may feel a subtle pressure to align their reflections with the values of the academic institution, even in anonymous settings.

8. Future Research Directions

Future research should systematically examine the role of time as a central contextual variable in GenAI-supported assessment. Investigating different time conditions may deepen our understanding of how time allocation shapes metacognitive monitoring, judgment, and responsibility during academic assessments that incorporate GenAI. In addition, comparative studies across assessment formats, such as time-limited exams, take-home assignments, and open-ended projects, may help clarify whether the patterns identified in the present study are specific to the present context or reflect the broader dynamics of GenAI-mediated evaluation. Such research could contribute to a more nuanced understanding of how the contextual features of assessment interact with the use of GenAI to shape students’ regulatory engagement.
Furthermore, a far research gap remains regarding the specific factors influencing these interactions. Future studies should examine students’ competence in prompting—which is likely high in our sample due to the degree program—as well as their competence in reflecting on the quality of the results (e.g., competence in Computational Thinking 2.0) as key influencing factors.

Author Contributions

Conceptualization, M.A., R.Z.-G. and D.B.; Methodology, M.A., R.Z.-G. and D.B.; Software, M.A., R.Z.-G. and D.B.; Validation, M.A., R.Z.-G. and D.B.; Formal analysis, M.A., R.Z.-G. and D.B.; Investigation, M.A., R.Z.-G. and D.B.; Resources, M.A., R.Z.-G. and D.B.; Data curation, M.A., R.Z.-G. and D.B.; Writing—original draft, M.A., R.Z.-G. and D.B.; Writing—review and editing, M.A., R.Z.-G. and D.B.; Visualization, M.A., R.Z.-G. and D.B.; Supervision, M.A., R.Z.-G. and D.B.; Project administration, M.A., R.Z.-G. and D.B.; Funding acquisition, M.A., R.Z.-G. and D.B. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by Ruppin Academic Center, Grant Number 33200.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki and approved by the Institutional Review Board at Ruppin Academic Center on 11 October 2024 (Approval Number 245).

Informed Consent Statement

Informed consent was obtained from all subjects involved in the study.

Data Availability Statement

The data presented in this study are available upon request from the corresponding author, as they are subject to privacy preservation.

Conflicts of Interest

The authors declare that they have no conflicts of interest.

Appendix A

Table A1. Thematic Categorization and Frequency Distribution of Pre-Assessment Expectations (N = 167 segments).
Table A1. Thematic Categorization and Frequency Distribution of Pre-Assessment Expectations (N = 167 segments).
Main Theme and Subthemes% Coded Segments
Emotional and Experiential Aspects (25%)
Positive expectations of reassurance and reduced anxiety19%
Negative feelings, such as stress or discomfort6%
Learning and Understanding (25%)
Potential contribution to understanding concerns about superficial understanding13%
Cognitive atrophy, or over-reliance9%
A preference for quick answers over a full cognitive process3%
Metacognitive Skills and Monitoring (17%)
The need for verification and validation of GenAI-provided responses10%
Use as a tool for self-monitoring4%
GenAI fallibility as a prompt for deeper examination3%
GenAI Literacy (11%)
Proficiency in prompt engineering7%
Understanding technology’s operational principles4%
Time Factor (11%)
Time constraints for thinking and monitoring6%
Reliance on the tool due to time pressure5%
Systemic and Social Perceptions (11%)
Adapting academia to the new era6%
Preparing for the labor market5%

Appendix B

Table A2. Thematic Categorization and Frequency Distribution of Post-Assessment Expectations (N = 140 segments).
Table A2. Thematic Categorization and Frequency Distribution of Post-Assessment Expectations (N = 140 segments).
Main Theme and Subthemes% Coded Segments
Emotional and Experiential Aspects (18%)
Positive feelings12%
Negative feelings6%
Learning and Understanding (24%)
Contribution to understanding during the quiz11%
Real-time verification and reassurance7%
Learning effects of anticipated reliance on GenAI in light of time pressure6%
Metacognitive Skills and Monitoring (20%)
The need for verification and validation8%
The use of GenAI for real-time self-monitoring and regulation7%
Awareness of potential errors5%
GenAI Literacy (12%)
The importance of prompt engineering7%
Student’s responsibility for the final answer5%
Time Factor (12%)
Limited time constrains thinking and monitoring (7%)7%
Reliance on the tool due to time pressure (5%)5%
Systemic and Social Perceptions (8%)
Preparation for the future labor market to the new era5%
The integration of GenAI into higher education3%
Authenticity and Fairness of Assessment (6%).
Integrity of the assessment process4%
The challenge of evaluating individual knowledge2%

References

  1. Amzalag, M., Beimel, D., & Zviel-Girshin, R. (2025). Learning in hybrid times: Comparing student experiences in traditional and GenAI-supported instruction. Computers and Education Open, 9, 100313. [Google Scholar] [CrossRef] [Scilit]
  2. Amzalag, M., & Kurtz, G. (2025). Generative AI in higher education: Uses and ethical dilemmas of students. In S. Kadry (Ed.), Artificial intelligence in education—Creating an equitable, creative, and effective learning environment. IntechOpen. [Google Scholar] [CrossRef] [Scilit]
  3. Annapureddy, R., Fornaroli, A., & Gatica-Perez, D. (2025). Generative AI literacy: Twelve defining competencies. Digital Government: Research and Practice, 6(1), 13. [Google Scholar] [CrossRef] [Scilit]
  4. Bada, S. O., & Olusegun, S. (2015). Constructivism learning theory: A paradigm for teaching and learning. Journal of Research & Method in Education, 5(6), 66–70. [Google Scholar]
  5. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2024). Generative AI can harm learning. The Wharton School Research Paper. Available online: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4895486 (accessed on 29 October 2024).
  6. Bates, T., Cobo, C., Mariño, O., & Wheeler, S. (2020). Can artificial intelligence transform higher education? International Journal of Educational Technology in Higher Education, 17(1), 42. [Google Scholar] [CrossRef] [Scilit]
  7. Bauer, E., Greiff, S., Graesser, A. C., Scheiter, K., & Sailer, M. (2025). Looking beyond the hype: Understanding the effects of AI on learning. Educational Psychology Review, 37(2), 45. [Google Scholar] [CrossRef] [Scilit]
  8. Beimel, D., Amzalag, M., Zviel-Girshin, R., & Voloch, N. (2025). Blending generative AI and instructor-led learning: Empirical insights on student motivation, learning experience, and academic performance in higher education. Education Sciences, 15(11), 1480. [Google Scholar] [CrossRef] [Scilit]
  9. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. [Google Scholar] [CrossRef] [Scilit]
  10. Chen, L., Chen, P., & Lin, Z. (2020). Artificial intelligence in education: A review. IEEE Access, 8, 75264–75278. [Google Scholar] [CrossRef] [Scilit]
  11. Chiu, T. K., Xia, Q., Zhou, X., Chai, C. S., & Cheng, M. (2023). Systematic literature review on opportunities, challenges, and future research recommendations of artificial intelligence in education. Computers and Education: Artificial Intelligence, 4, 100118. [Google Scholar] [CrossRef] [Scilit]
  12. Cohen, L., Manion, L., & Morrison, K. (2002). Research methods in education. Routledge. [Google Scholar]
  13. Cotton, D. R., Cotton, P. A., & Shipway, J. R. (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in Education and Teaching International, 61(2), 228–239. [Google Scholar] [CrossRef] [Scilit]
  14. Deng, R., Jiang, M., Yu, X., Lu, Y., & Liu, S. (2025). Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Computers & Education, 227, 105224. [Google Scholar]
  15. Denny, P., Prather, J., Becker, B. A., Finnie-Ansley, J., Hellas, A., Leinonen, J., Luxton-Reilly, A., Reeves, B. N., Santos, E. A., & Sarsa, S. (2024). Computing education in the era of generative AI. Communications of the ACM, 67(2), 56–67. [Google Scholar] [CrossRef] [Scilit]
  16. Eke, D. O. (2023). ChatGPT and the rise of generative AI: Threat to academic integrity? Journal of Responsible Technology, 13, 100060. [Google Scholar] [CrossRef] [Scilit]
  17. Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. [Google Scholar] [CrossRef] [Scilit]
  18. González-Calatayud, V., Prendes-Espinosa, P., & Roig-Vila, R. (2021). Artificial intelligence for student assessment: A systematic review. Applied Sciences, 11(12), 5467. [Google Scholar] [CrossRef] [Scilit]
  19. Guba, E. G., & Lincoln, Y. S. (1994). Competing paradigms in qualitative research. In Handbook of Qualitative Research (Vol. 2, pp. 163–194). Sage Publications, Inc. [Google Scholar]
  20. Holmes, W., Bialik, M., & Fadel, C. (2019). Artificial intelligence in education promises and implications for teaching and learning. Center for Curriculum Redesign. [Google Scholar]
  21. Hou, C., Zhu, G., & Sudarshan, V. (2025). The role of critical thinking on undergraduates’ reliance behaviours on generative AI in problem-solving. British Journal of Educational Technology, 56(5), 1919–1941. [Google Scholar] [CrossRef] [Scilit]
  22. Kasneci, E., Seßler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeffer, J., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, 102274. [Google Scholar] [CrossRef] [Scilit]
  23. Kurtz, G., Amzalag, M., Shaked, N., Zaguri, Y., Kohen-Vacs, D., Gal, E., Zailer, G., & Barak-Medina, E. (2024). Strategies for integrating generative AI into higher education: Navigating challenges and leveraging opportunities. Education Sciences, 14(5), 503. [Google Scholar] [CrossRef] [Scilit]
  24. Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., & Qiao, M. S. (2021). Conceptualizing AI literacy: An exploratory review. Computers and Education: Artificial Intelligence, 2, 100041. [Google Scholar] [CrossRef] [Scilit]
  25. Panadero, E. (2017). A review of self-regulated learning: Six models and four directions for research. Frontiers in Psychology, 8, 422. [Google Scholar] [CrossRef] [Scilit]
  26. Piaget, J. (1954). The construction of reality in the child. Basic Books. [Google Scholar]
  27. Podsakoff, P. M., MacKenzie, S. B., Lee, J. Y., & Podsakoff, N. P. (2003). Common method biases in behavioral research: A critical review of the literature and recommended remedies. Journal of Applied Psychology, 88(5), 879. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Qian, Y. (2025). Pedagogical applications of generative AI in higher education: A systematic review of the field. TechTrends, 69, 1105–1120. [Google Scholar] [CrossRef] [Scilit]
  29. Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. [Google Scholar] [CrossRef] [PubMed]
  30. Sweller, J. (2011). Cognitive load theory. In J. P. Mestre, & B. H. Ross (Eds.), Psychology of learning and motivation (Vol. 55, pp. 37–76). Academic Press. [Google Scholar] [CrossRef] [Scilit]
  31. Tankelevitch, L., Kewenig, V., Simkute, A., Scott, A. E., Sarkar, A., Sellen, A., & Rintel, S. (2024). The metacognitive demands and opportunities of generative AI. In Proceedings of the 2024 CHI conference on human factors in computing systems (pp. 1–24). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  32. Wang, J., & Fan, W. (2025). The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: Insights from a meta-analysis. Humanities and Social Sciences Communications, 12(1), 621. [Google Scholar] [CrossRef] [Scilit]
  33. Wecks, J. O., Voshaar, J., Plate, B. J., & Zimmermann, J. (2024). Generative AI usage and exam performance. arXiv, arXiv:2404.19699. [Google Scholar]
  34. Xia, Q., Weng, X., Ouyang, F., Lin, T. J., & Chiu, T. K. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education, 21(1), 40. [Google Scholar] [CrossRef] [Scilit]
  35. Xu, X., Qiao, L., Cheng, N., Liu, H., & Zhao, W. (2025). Enhancing self-regulated learning and learning experience in generative AI environments: The critical role of metacognitive support. British Journal of Educational Technology, 56(5), 1842–1863. [Google Scholar] [CrossRef] [Scilit]
  36. Yusuf, A., Pervin, N., & Román-González, M. (2024). Generative AI and the future of higher education: A threat to academic integrity or reformation? Evidence from multicultural perspectives. International Journal of Educational Technology in Higher Education, 21(1), 21. [Google Scholar] [CrossRef] [Scilit]
  37. Zawacki-Richter, O., Bai, J. Y., Lee, K., Slagter van Tryon, P. J., & Prinsloo, P. (2024). New advances in artificial intelligence applications in higher education? International Journal of Educational Technology in Higher Education, 21(1), 32. [Google Scholar] [CrossRef] [Scilit]
  38. Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education—Where are the educators? International Journal of Educational Technology in Higher Education, 16(1), 39. [Google Scholar] [CrossRef] [Scilit]
  39. Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory into Practice, 41(2), 64–70. [Google Scholar] [CrossRef] [Scilit]
  40. Zviel-Girshin, R. (2024). The good and bad of AI tools in novice programming education. Education Sciences, 14(10), 1089. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Qualitative Data Analysis Flowchart.
Figure 1. Qualitative Data Analysis Flowchart.
Education 16 00642 g001
Table 1. Participant Demographics.
Table 1. Participant Demographics.
Program
CS45 (50%)
E&M45 (50%)
Gender
Male51 (57%)
Female39 (43%)
Age
Range22–30 years old
Table 2. Comparative Distribution of Themes Pre- and Post-Intervention (N = 90).
Table 2. Comparative Distribution of Themes Pre- and Post-Intervention (N = 90).
ThemePre-Quiz Mentions:
N (%)
Post-Quiz Mentions:
N (%)
Directional Shift
Emotional and Experiential Aspects42 (25%)25 (18%)Decrease
Learning and Understanding42 (25%)34 (24%)Slight decrease
Metacognitive Skills and Monitoring28 (17%)28 (20%)Increase
GenAI Literacy18 (11%)17 (12%)Decrease
Time Factor18 (11%)17 (12%)Slight increase
Systemic and Social Perceptions19 (11%)11 (8%)Decrease
Authenticity and Fairness of Assessment0 (0%)8 (6%)New Theme
Total167 (100%)140 (100%)
Note: The total number of coded segments exceeds the number of participants (N = 90) because individual student responses often contained multiple thematic elements.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Amzalag, M.; Zviel-Girshin, R.; Beimel, D. From Expectations to Measured Pragmatism: A Pre- and Post-Experience Study of Student Engagement in AI-Supported Academic Exams. Educ. Sci. 2026, 16, 642. https://doi.org/10.3390/educsci16040642

AMA Style

Amzalag M, Zviel-Girshin R, Beimel D. From Expectations to Measured Pragmatism: A Pre- and Post-Experience Study of Student Engagement in AI-Supported Academic Exams. Education Sciences. 2026; 16(4):642. https://doi.org/10.3390/educsci16040642

Chicago/Turabian Style

Amzalag, Meital, Rina Zviel-Girshin, and Dizza Beimel. 2026. "From Expectations to Measured Pragmatism: A Pre- and Post-Experience Study of Student Engagement in AI-Supported Academic Exams" Education Sciences 16, no. 4: 642. https://doi.org/10.3390/educsci16040642

APA Style

Amzalag, M., Zviel-Girshin, R., & Beimel, D. (2026). From Expectations to Measured Pragmatism: A Pre- and Post-Experience Study of Student Engagement in AI-Supported Academic Exams. Education Sciences, 16(4), 642. https://doi.org/10.3390/educsci16040642

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop