1. Introduction
The advent of generative artificial intelligence (GenAI), particularly large language models (LLMs) such as ChatGPT, has fundamentally disrupted traditional practices in higher education, including assessments (
Xia et al., 2024). Since the emergence of this new technology in late 2022, academic institutions have moved rapidly from initial prohibition and uncertainty toward a more cautious and intentional integration of these tools into pedagogical frameworks (
Amzalag & Kurtz, 2025;
Kasneci et al., 2023;
Zawacki-Richter et al., 2024). Yet while GenAI is increasingly recognized as an essential component of contemporary digital literacy, there is as yet no consensus regarding its appropriate role in educational environments (
Bauer et al., 2025;
Kasneci et al., 2023). Some scholars frame GenAI as a robust cognitive scaffold capable of reducing barriers to learning, while others emphasize the risks of cognitive atrophy, over-reliance, and the erosion of foundational problem-solving skills (
Bates et al., 2020;
Kasneci et al., 2023).
Despite the growing body of literature on GenAI in education, most existing studies focus on formative learning, general student attitudes, or institutional policy at a theoretical level (
Deng et al., 2025;
Wang & Fan, 2025). The present study examines how students actually engage with GenAI during classroom-based assessments, an area marked by a notable scarcity of empirical research (
Chiu et al., 2023;
Xia et al., 2024). In assessment environments, conditions such as strict time pressure and evaluative consequences fundamentally shape students’ cognitive strategies and perceptions (
Yusuf et al., 2024).
The present study was conducted in the context of computer science and engineering examinations, where traditional assumptions about individual knowledge production and algorithmic reasoning are being directly challenged (
Cotton et al., 2024;
Xia et al., 2024). We explore undergraduate students’ perceptions through the lens of self-regulated learning (SRL) theory (
Zimmerman, 2002) and cognitive load theory (
Sweller, 1988,
2011). Using a retrospective pre–post design, we examine how direct practical experience with GenAI during a classroom assessment affected students’ attitudes. In this respect, our study is one of the few to employ longitudinal or pre–post designs to examine how students’ conceptualization of AI changes following direct, practical application. The study integrates qualitative thematic analysis with quantitative measures to provide novel empirical evidence on the evolving role of GenAI in academic evaluations.
This study contributes to the field by offering a process-oriented exploration of how direct experience with GenAI-supported assessments reshapes student perceptions. Theoretically, it bridges the gap between GenAI literacy and self-regulated learning (SRL) frameworks; practically, it provides empirical evidence to guide the design of authentic, fair, and cognitively engaging assessment environments in the AI era.
The remainder of this paper is organized as follows.
Section 2 reviews the relevant literature on GenAI in higher education and self-regulated learning.
Section 2 presents research questions.
Section 3 presents research objectives.
Section 4 details the research methodology, including data collection and analysis procedures.
Section 5 presents the study’s findings, while
Section 6 discusses these results in the context of broader educational implications and concluding remarks. Finally,
Section 7 and
Section 8 offer limitations, and suggestions for future research.
4. Methodology and Research Characteristics
Our methodological framework was grounded in an interpretive qualitative paradigm, examining student responses to open-ended questions. This methodology facilitates in-depth exploration of how participants construct personal meaning from their experiences, accommodating varied interpretations, complexity, and potential inconsistencies within respondents’ perspectives. Given that our research design prioritized meaning construction over hypothesis validation within an interpretive qualitative lens, inferential statistical analyses were not central to the study design. Descriptive statistics were used in a complementary function to characterize participants’ demographic profiles and to analyze student preferences regarding the two assessment formats, providing a contextual baseline for the qualitative findings.
4.1. Participants
The study was conducted in two undergraduate courses within a university Engineering faculty. The two courses were taught by the same instructor and had similar pedagogical structures, assessment formats, and cognitive demands.
The first cohort consisted of students taking a required Data Structures course in the university’s Computer Science (CS) program. The second cohort comprised students taking a Data Structures and Algorithms course required for Engineering and Management (E&M) students specializing in Information Systems and Data Science. The former course was taught in the second term of the CS program (out of six), and the latter in the sixth term of the E&M program (out of eight).
The final sample for the current study comprises 90 participants who submitted full responses. This final sample includes 45 E&M students and 45 CS students, 39 females (43%) and 51 males (57%), aged 22 to 30. Participant demographics are summarized in
Table 1.
4.2. Experimental Procedure
The research procedure had two phases (both identical for the two cohorts):
Phase 1: During week 11 of the academic term, participants sat for a brief classroom-based assessment (hereafter: quiz). Students were notified about the quiz and its content areas (AVL tree algorithms for CS students and BFS/DFS algorithms for E&M students) approximately 14 days in advance. Students were authorized to bring a single page of notes for reference during the quiz. The allotted time for completing the quiz was roughly 30 min.
Before starting the quiz, all participants provided written informed consent for use of their data in the study. Participants also provided basic demographic information, including age and gender.
Phase 2: In week 12 of the term, a second quiz was administered. As before, students were given notice of the quiz and its contents (non-linear algorithms for CS students and the PRIM algorithm for E&M students) about two weeks in advance. However, this time, participants were notified that the quiz would be conducted in a computing lab and that they would have access to GenAI resources throughout the assessment. The quiz was delivered electronically through the Moodle platform (
https://moodle.com/). Again, participants were given roughly 30 min to complete the quiz. As in the prior session, informed consent was obtained from all participants in advance. To preserve naturalistic conditions, participants received no special guidance on implementing GenAI and were free to use their preferred tools and approaches.
Following the week-12 quiz, students were asked to complete a brief post-quiz survey, which comprised the main research instrument. See
Section 4.3.
4.3. Research Instrument: Post-Quiz Questionnaire
As noted, immediately after completing the week-12 assessment, students completed an online post-quiz questionnaire designed to capture their experiences and perceptions of using AI tools during the assessment. The questionnaire included three closed-ended items and two open-ended questions. Questionnaires for the two cohorts were identical.
The two open-ended questions examined students’ perceptions before and after experiencing the week-12 quiz, which allowed internet access and the use of GenAI tools. The first asked students to describe their thoughts and expectations about the GenAI-enabled assessment format before taking the quiz (‘What were your thoughts before the quiz regarding this assessment method, which includes the use of open resources, internet access, and the possibility of using AI tools?’), while the second invited them to articulate their views after experiencing this assessment approach (‘What were your thoughts after the quiz regarding this assessment method, which includes the use of open resources, internet access, and the possibility of using AI tools?’). This retrospective pre–post design was intended to capture potential shifts in student attitudes and to generate insights into how direct experience with GenAI-assisted testing influences learner perspectives. However, it is important to note that because the ‘pre-quiz’ expectations were collected retrospectively, these reported prior attitudes may have been subconsciously influenced by the students’ subsequent direct experience. Therefore, the identified shifts are interpreted with appropriate caution, representing students’ reconstructed narratives of their changing perspectives rather than independent pre- and post-measurements.
After completing the open-ended questions, participants responded to the three closed-ended items. The first two of these concerned which GenAI tools were used during the quiz and the extent to which the student checked GenAI-provided responses. Although the detailed evaluation of these behavioral metrics is reserved for other work, it was found that most students used prompts focused on the specific exam topics (i.e., AVL tree algorithms for CS students and BFS/DFS algorithms for E&M students), and typically exhibited behaviors ranging from direct copy-pasting of AI outputs to adopting a skeptical stance, actively debating with the AI, and rigorously verifying the provided responses. This snapshot conceptually grounds the qualitative statements regarding metacognitive monitoring discussed later in this paper.
The third item asked participants to indicate their preferred assessment format: examinations with access to the internet and GenAI tools versus traditional closed-book examinations. This item was integral to our research design. Together, the two open-ended items and the closed-ended preference measure constitute the core of the present study.
4.4. Data Analysis
Data analysis followed a qualitative-dominant, mixed-methods approach, integrating descriptive quantitative analysis of student preferences with a rigorous qualitative thematic inquiry of participants’ open-ended reflections. This combined approach allows for a comprehensive understanding of both the overall distribution of student attitudes and the nuanced meanings they ascribe to their experiences.
Quantitative data derived from the closed-ended preference item were analyzed using descriptive statistics. Frequencies and percentages were computed to determine the overall distribution of student preferences for the two assessment formats (GenAI-enabled versus traditional).
The qualitative component, which formed the primary focus of our analysis, was grounded in the thematic analysis methodology articulated by
Braun and Clarke (
2006). To facilitate a systematic and rigorous analysis, the qualitative data were organized and managed using ATLAS.ti (version 25) software. To maintain analytical precision and reliability, we implemented a structured, recursive coding procedure. Initially, coding proceeded inductively without predefined thematic categories. In the initial phase, two independent coders examined a subset of 15 student responses to each of the two open-ended questions, sampled randomly from both cohorts (the CS and E&M classes). This exploratory stage enabled us to establish preliminary codes and detect emergent thematic patterns. For instance, when analyzing responses to the question about pre-examination expectations for the GenAI-supported assessment format, one participant noted: “I thought it would make everything easier and faster, but I was concerned about whether I could trust the AI answers.” From this excerpt, we identified three preliminary themes:
anticipated efficiency benefits,
time-related expectations, and
trust concerns. Another student’s pre-examination response stated: “I expected it would level the playing field for everyone and reduce memorization pressure, though I worried about technical issues.” This response was also analyzed as including three themes:
equity perceptions,
reduced cognitive load, and
technical apprehensions. Based on this process of collaborative deliberation on the initial codes and emergent themes, we constructed a provisional codebook with explicitly defined categories and precise inclusion and exclusion criteria for each theme. As coding progressed, related preliminary codes were iteratively reviewed, refined, and consolidated into broader themes. The main themes and subthemes reported in the findings thus represent the outcome of this consolidation process, rather than the initial open coding stage.
The frequency percentages reported in the findings reflect the proportion of coded thematic references to the total number of coded references for each question. Because individual responses could contain multiple thematic elements, a single participant’s comment could be counted under more than one theme. In the next phase, two coders independently analyzed an additional sample of 25 responses for each open-ended question (50 responses total, again sampled randomly from the CS and E&M cohorts). They then reconciled differences in their coding through structured dialogue to create a definitive codebook. By the end of this process, substantial intercoder concordance was attained, with 47 of 50 coded responses (94%) demonstrating complete alignment. An additional evaluation of intercoder reliability using Cohen’s Kappa yielded κ = 0.788 (p < 0.001), indicating near-perfect agreement. The absence of novel themes during this validation phase confirmed the comprehensiveness and stability of our coding architecture.
In the third phase, using this definitive codebook, the lead researcher systematically coded a calibration dataset taken from the responses used to develop the codebook (30 responses per open-ended question, 60 total). A second researcher then conducted an audit trail review by independently examining a randomized subset of 20 responses per question, for a total of 40 responses. Complete concordance was observed between the two coders, further substantiating the internal consistency and credibility of the analysis. Following this validation procedure, the principal investigator proceeded to code the entire corpus, completing analysis of all participant responses to both open-ended questions.
In total, the coding process yielded 167 thematic segments from the pre-quiz reflections and 140 thematic segments from the post-quiz reflections.
Figure 1 provides a visual overview of this rigorous three-phase data analysis process, outlining the inductive coding, validation steps, and intercoder reliability metrics.
5. Findings
This section presents the results of our analyses. We first report the results for the two open-ended questionnaire items, structured around the two research questions. This is followed by a brief exploration of how individual student attitudes evolved across the assessment phases. We then report the results for the closed-ended preferred assessment format question. As the present study did not examine cohort-specific differences, data from the two cohorts (CS and E&M) are reported together.
5.1. RQ1: Perceptions Prior to the Intervention
Our first research question asked how undergraduate students retrospectively describe their expectations and perceptions regarding the use of GenAI in academic assessments prior to direct experience. Analysis of students’ responses revealed six main themes: Emotional and Experiential Aspects (25%), Learning and Understanding (22%), Metacognitive Skills and Monitoring (17%), GenAI Literacy (14%), Time Factor (11%), and Systemic and Social Perceptions (11%). These themes consist of several subthemes, as detailed next.
Emotional and Experiential Aspects (25%). This theme was among the most prominent in students’ prior expectations, focusing on the anticipated affective impact of integrating GenAI into the assessment process. This theme is organized into two analytically related subthemes: positive expectations of reassurance and reduced anxiety (19%), and negative feelings, such as stress or discomfort (6%).
Many students reported that, prior to the intervention, they anticipated that access to GenAI would alleviate the pressure typically associated with academic assessments. One student shared the hope that the method would provide emotional support: “I thought it would be a great method, it would make me feel more secure and calm” (Student 9). In contrast, other students expressed initial resistance or unease about the change to the class’s familiar format. For instance, one participant described the inclusion of GenAI as “strange and unnecessary” (Student 15). Another student worried that the complexity of the new format might be counterproductive: “I thought it would be confusing and stressful rather than helpful” (Student 48).
Learning and Understanding (25%). This theme was equally prominent, accounting—like the Emotional and Experiential Aspects theme—for 25% of the responses. Within this theme, three primary subthemes were identified: potential contribution to understanding (13%), concerns about superficial understanding, cognitive atrophy, or over-reliance (9%), and a preference for quick answers over a full cognitive process (3%).
Many students anticipated that support from GenAI would allow for easy access to details, letting the student focus more on engaging with the overall task rather than remembering every piece of information learned: “I thought it would be easier with the chatbot’s help, because even if I am not sure about some details, I will have access to them” (Student 54). In addition, students expected that the tool would serve as a scaffold for their understanding in real-time and strengthen their self-confidence: “It is an amazing method because if I forget something or want to double-check, it really helps me” (Student 2).
Other students reported concerns that having GenAI as a supportive resource might undermine learning by providing rapid answers that bypass deep thinking: “Before the quiz, I thought that using GenAI might interfere with deep understanding, because it provides answers too quickly” (Student 17); “It seems nice, but it reduces the quality of learning for the quiz because you feel like you have a safety net” (Student 30). This concern extended to a fear of developing dependency, where the ease of access might erode independent problem-solving habits: “I’m afraid it will make me not even try to solve it on my own, but run straight to the chat” (Student 72).
Finally, a smaller group of students anticipated that the mere availability of the tool might undermine their ability to engage in independent thought, leading to a preference for quick answers: “Because I have the tool, I am less likely to think for myself and go directly to asking the chatbot” (Student 76). Relatedly, another student worried that adopting a GenAI-supported assessment method could harm the quality of quiz responses: “Before the quiz, I thought that assessment with a GenAI tool could become problematic, leading to answers that are more hasty and superficial” (Student 80).
Metacognitive Skills and Monitoring (17%). This theme reflects students’ critical and reflective approach to GenAI prior to the intervention. Their responses highlighted three primary subthemes: the need for verification and validation of GenAI-provided responses (10%), use as a tool for self-monitoring (4%), and GenAI fallibility as a prompt for deeper examination (3%).
Regarding the need for verification and validation of GenAI-provided responses, students reported feeling a high degree of skepticism regarding the reliability of GenAI, repeatedly emphasizing that its outputs should not be accepted at face value. This critical stance was rooted in an understanding that the technology often lacks precision, produces “nonsense,” or misses subtle nuances, thereby requiring constant external validation: “Using such an open-resource tool does not really test me, except for how much I trust it, and from my experience, I always need to double-check it a lot” (Student 12). This wariness was reflected in a similar comment by another student: “The GenAI can miss the small details, so it cannot be fully trusted” (Student 44). Such concerns even prompted some to question the overall validity of the proposed assessment method (subtheme four): “I thought that assessment with a GenAI tool could become problematic” (Student 80). In addition, a small number of students articulated ethical reservations regarding the appropriate role of AI in assessment contexts (subtheme five): “I believe that artificial intelligence should help us learn the material, not solve the exam. The challenge must remain with the examinee, not the machine” (Student 71). Interestingly, however, this fallibility was also seen as a catalyst for deeper learning (subtheme three); the realization that the tool is not always correct allows for a deeper examination of the student’s understanding of the subject matter: “Yes, because the GenAI is not always correct, and this allows you to examine more deeply the student’s understanding” (Student 61).
Despite these reservations, several students envisioned leveraging GenAI as a strategic partner for cognitive regulation (subtheme two). For these participants, the tool was seen as a valuable asset for clarifying task requirements and verifying their own reasoning: “It is good for self-checking and for helping with questions and understanding them” (Student 24). Similarly, some students framed the interaction with GenAI as a dialectical process, allowing for real-time feedback and the refinement of their ideas through external critique: “I thought this method would give me an opportunity to consult with another source that could critique what I think and explain it to me in real time” (Student 56).
GenAI Literacy (11%). This theme concerns the technical and conceptual competencies required for effective use of GenAI. It includes two primary subthemes: proficiency in prompt engineering (7%) and understanding the technology’s operational principles (4%).
Regarding the former, students emphasized that the utility of the tool in an academic context depends on the user’s ability to operate it in an informed manner, not merely through basic use. They noted that the quality of the output is directly linked to the prompt chosen by the user: “It requires skill in writing the right prompt to get accurate information” (Student 35). In terms of the latter, participants recognized that effective interaction with GenAI requires an understanding of the principles of the technology’s operation: “It’s good if you know how to use it correctly… you need to understand how the system works” (Student 18).
The Time Factor (11%). This theme addresses the complexity of time in students’ perceptions of GenAI use prior to the intervention, as retrospectively reported. The theme examines time from two contrasting angles: as a resource for increased efficiency on the one hand, and as a constraint that may induce pressure on the other. Within this theme, two subthemes were identified: time constraints for thinking and monitoring (6%), and reliance on the tool due to time pressure (5%).
Some participants viewed the technology as an efficiency aid that would streamline information retrieval and facilitate a smoother assessment process: “It saves time searching through slides, it helps me meet the deadlines” (Student 14). Conversely, other students warned that the pressure of the clock might incentivize haste and compromise the verification process (subtheme two). This tension suggests that when time is perceived as a scarce resource, the ease of generating rapid responses may lead to less rigorous cognitive engagement: “In a time crunch, it’s easy to just copy what it says without thinking” (Student 4).
Systemic and Social Perceptions (11%). The final theme identified for RQ1 highlights students’ broader considerations regarding the evolving role of GenAI in higher education and the professional world. Within this theme, two primary subthemes were identified: adapting academia to the new era (6%) and preparing for the labor market (5%).
Participants argued that educational institutions must evolve to align with global technological trends rather than remaining isolated from them: “Academia needs to move forward; this is the real world today” (Student 40). This perspective extended to professional readiness, as students perceived proficiency in GenAI not as an optional skill but as a prerequisite for remaining relevant in the modern job market: “It is important to learn how to use this because it is what will be required of us at work” (Student 63).
For a systematic and comprehensive overview of the thematic distribution,
Appendix A provides a detailed table summarizing the frequencies of all themes and subthemes across the pre-assessment phases.
5.2. RQ2: Attitudes and Experiences After the Intervention
Our second research question asked how students’ perceptions and attitudes change following direct practical experience with a GenAI-supported assessment. To address this question, we examined students’ reflective responses following completion of the quiz. The analysis revealed seven main themes, including six recurring themes consistent with those identified prior to the intervention: Emotional and Experiential Aspects (18%), Learning and Understanding (24%), Metacognitive Skills and Monitoring (20%), GenAI Literacy (12%), Time Factor (12%), and Systemic and Social Perceptions (8%). There was one newly emerged theme, Authenticity and Fairness of Assessment (6%). As with RQ1, each theme includes several subthemes, as detailed below.
Emotional and Experiential Aspects (18%). The frequency of this theme decreased from 25% in the responses for RQ1 to 18% in those for RQ2. This shift suggests that practical experience may have tempered some of the students’ anticipated affective concerns and expectations toward more practical and cognitive elements. As before, the theme included two subthemes: positive feelings (12%) and negative feelings (6%).
In terms of the former, direct interaction with GenAI during the quiz was associated with decreased stress levels and increased self-confidence among some students. Access to generative AI provided a sense of security during the assessment and allowed students to focus on answering the questions: “It was a great method, I could understand the questions well without feeling stressed” (Student 10); “After the quiz, I felt that the assessment really tested our ability to think… it was less stressful and more enriching” (Student 90).
Conversely, for other participants, the experience elicited frustration and disillusionment. For these students, what they perceived as forced reliance on the tool due to external constraints such as time pressure—perhaps along with failure to adequately prepare—created an emotional sense of letdown regarding their own performance: “I was disappointed because I felt that I could have solved the quiz very well by myself, but the short time made me depend on the chatbot’s answers” (Student 50).
Learning and Understanding (24%). This theme remained a primary focus of students’ reflections following completion of the GenAI-supported quiz. It comprised three subthemes: contribution to understanding during the quiz (11%), real-time verification and reassurance (7%), and the learning effects of anticipated reliance on GenAI in light of time pressure (6%).
Several students emphasized that the integration of GenAI supported their comprehension during the quiz by providing immediate explanations and examples: “After the quiz, I realized that integrating tools like GenAI has advantages. You can receive an immediate explanation… and compare sorting methods… This helps to understand the advantages and disadvantages of each method” (Student 71). GenAI was also perceived as an interpretative aid that clarified the logic of the questions: “The GenAI not only explained the question but also provided examples of solutions and counterexamples when necessary” (Student 38).
Students also highlighted the tool’s value as a mechanism for verifying their answers in real time, which bolstered their confidence and provided them with a “safety net”: “It helped me to make sure my answers were correct” (Student 27); “The GenAI helped me confirm my understanding of the question and the solution, giving me assurance that I was on the right track” (Student 82).
Finally, several students linked the time constraints of the quiz to a compromise in their learning process and their autonomy. For some, the need for speed overrode the opportunity for deep understanding: “I had to provide answers quickly under time pressure, rather than really learn during the quiz” (Student 14). For others, the knowledge that AI would be available led to reduced preparation, which, combined with time pressure, led them to rely on the tool: “I prepared less for the quiz because I had the AI. With more time, I could have solved the questions on my own and only used GenAI for checking, but the limited time forced me to rely on it” (Student 68).
Metacognitive Skills and Monitoring (20%). The frequency of this theme increased from 17% in the responses for RQ1 (students’ reports of their expectations prior to the quiz) to 20% in the responses for RQ2, suggesting that direct experience with GenAI during the quiz made students more aware of both the limitations and the value of GenAI in terms of metacognitive monitoring, verification, and control. The coding produced three subthemes, all of which had also appeared in the responses for RQ1: the need for verification and validation (8%), the use of GenAI for real-time self-monitoring and regulation (7%), and awareness of potential errors (5%). As in RQ1, awareness of potential errors refers to the perceived fallibility of GenAI as a source of information, which may prompt deeper examination, whereas the need for verification captures the regulatory act of checking and validating responses in practice.
Most often, direct interaction with the tool influenced students to adopt a more skeptical and critical stance toward AI-generated content, as they realized that the tool’s potential for errors requires constant control and verification: “There is a great need to verify ChatGPT because I felt it made many mistakes and could easily mislead me if I was not careful” (Student 29). In other cases, the tool enabled students to verify their own reasoning and check their answers in real time during the quiz, providing a way to monitor their problem-solving steps: “It helped me confirm my answers and review the way I solved the question” (Student 27). Notably, references to verifying and controlling GenAI-produced outputs were more frequent than references to using GenAI to validate one’s own reasoning, suggesting that students were more preoccupied with managing the tool’s reliability than leveraging it for self-regulation.
GenAI Literacy (12%). The frequency of this theme decreased from 14% to 12% in responses for RQ1 and RQ2, respectively. The theme included two subthemes: the importance of prompt engineering (7%) and student responsibility for the final answer (5%).
Direct interaction with GenAI during the quiz influenced students to recognize that the quality of the suggested answers depended heavily on the user’s ability to provide precise instructions. They realized that the way they phrased the prompt directly affected the accuracy of the information they received: “I understood that the result I get is only as good as the prompt I write, so I had to be very specific to get a helpful answer” (Student 41).
Furthermore, the experience reinforced the understanding that the student, rather than the AI, is ultimately responsible for selecting the correct answer in the quiz: “Even though the AI suggests an answer, the responsibility for the final selection is mine alone, and I must know how to direct the tool correctly” (Student 6).
The Time Factor (12%). The frequency of this theme slightly increased from 11% in the responses for RQ1 to 12% in the responses to RQ2. While students initially anticipated that GenAI would primarily serve as a time-saving tool, their experience revealed that time was a significant constraint that hindered their ability to verify the tool’s answers.
Direct interaction with GenAI during the quiz highlighted how time pressure influenced students’ use of the tool. For many, the limited time frame transformed GenAI from an efficient tool into a factor that prevented them from performing a thorough review and verification of the suggested answers: “The time was very limited, which made it difficult to check the answers provided by the AI” (Student 45). Furthermore, interaction with generative AI during the quiz showed that the speed of generating answers can be misleading when the user is required to validate information under pressure: “I felt like I was racing against the clock, and even though the AI provided answers quickly, I did not have enough time to ensure that they were correct” (Student 22).
Systemic and Social Perceptions (8%). The frequency of this theme decreased from 11% to 8% in responses for RQ1 and RQ2, respectively. This decline suggests that following practical experience, students’ reflections shifted somewhat away from broader systemic considerations toward more immediate cognitive and experiential aspects of GenAI use during the assessment. Nevertheless, students continued to express views regarding the inevitable integration of GenAI into the academic and professional worlds. The theme included two subthemes: preparation for the future labor market (5%) and the integration of GenAI into higher education (3%).
Direct interaction with GenAI during the quiz influenced students’ views of these tools as integral to their future professional development. They noted that academic institutions must adapt to these technological changes to remain relevant: “In today’s reality, most of the labor market already uses GenAI tools; it is important and valuable that we learn how to work and study alongside GenAI in a thoughtful and responsible way” (Student 31). Likewise, the experience highlighted the systemic need for a shift in how knowledge is assessed in the digital age: “This is a cool method; the world now relies more on GenAI, so I think academia also needs to adapt to the changing world” (Student 15).
Authenticity and Fairness of Assessment (6%). This theme was only marginally present prior to the intervention, where ethical concerns appeared as a minor subtheme under Metacognitive Skills and Monitoring. Following the practical experience, however, it emerged as a distinct theme, suggesting that direct engagement influenced students to consider more explicitly the ethical implications and validity of the assessment process when using GenAI. The theme included two subthemes: integrity of the assessment process (4%) and the challenge of evaluating individual knowledge (2%).
Direct interaction with GenAI during the quiz led students to question the quiz’s authenticity, given that AI provided substantial assistance. Students expressed concerns that the GenAI-supported format might not accurately reflect a student’s true understanding: “If everyone uses the AI, the quiz no longer tests our knowledge but rather our ability to operate the tool, which raises questions about the fairness of the grade” (Student 55).
Furthermore, this experience highlights a new perspective on academic integrity in the age of AI. Some students realized that the traditional definition of an independent assessment is changing, leading to a sense of ambiguity regarding the value of the assessment: “It felt a bit like cheating even though it was allowed, because the tool did most of the logical work for me” (Student 12).
For a systematic and comprehensive overview of the thematic distribution,
Appendix B provides a detailed table summarizing the frequencies of all themes and subthemes across the post-assessment phases.
Table 2 summarizes the frequency of citations for each theme before and after the intervention. This comparative overview highlights shifts in thematic prevalence and identifies whether the frequency of each theme in the students’ remarks increased or decreased as a result of practical experience.
5.3. Longitudinal Tracking of Individual Student Attitudes
Beyond general trends across the entire sample, we tracked individual-level changes by comparing each student’s pre- and post-quiz responses. This tracking enabled the identification of cases in which the same student changed their perspective after the practical experience. Notably, the majority of participants (91%) maintained consistent views between the two stages: only 9% shifted between primary themes, and 13% shifted between subthemes. These changes were not radical reversals, but rather subtle changes in emphasis that illustrate how direct experience with GenAI could refine or reframe existing attitudes. We identified five main patterns, each relating to one of the main themes identified: skepticism to appreciation (learning and understanding); trust to frustration (emotional aspects); abstract concerns to practical understanding (metacognitive monitoring); increased focus on the concrete cost of time; and enthusiasm to caution (systemic and social perceptions). The five patterns presented below were identified through an iterative and systematic review of the entire dataset. After reading all student reflections multiple times, we identified the primary dynamics of change that emerged from the data. We then selected representative cases and quotes that best illustrate the variety and spectrum of these experiences across the two cohorts.
From Skepticism to Appreciation: Some students moved from concerns about the effects of GenAI on deep learning to recognizing the tool’s pedagogical value. For instance, Student 17, reporting on perceptions before the quiz, noted: “I thought that using artificial intelligence might interfere with deep understanding because it provides answers too quickly.” After the quiz, however, the same student described a different effect: “I realized that integrating tools like GenAI has advantages… it helps to understand the strengths and weaknesses of each method… Yet the understanding comes from the student and not only from the GenAI tool.”
From Trust to Frustration: Another student experienced an emotional shift. Before the quiz, she expressed trust in the tool and anticipated feeling more confident and secure with its support: “I thought it would be simpler with the chatbot’s help. Even if I was unsure about some details, I would have access to them, and that would make it much easier.” After the quiz, her experience was characterized by disappointment: “I was disappointed because I felt I could have solved the questions very well by myself and only used the GenAI for checking. Instead, my answers reflected time pressure rather than full comprehension” (Student 50).
From Abstract Concerns to Practical Understanding: Shifts were also observed in students’ practical understanding of the metacognitive monitoring required when using GenAI. One student, in reporting his thoughts before the quiz, expressed a general sense of caution about the tool’s reliability: “It is a good method, but the tool is not always stable and might suddenly give wrong answers.” Following the quiz, his understanding became more concrete and practice-based, as reflected in his comment: “There is a great need to verify ChatGPT. I felt it made many mistakes… therefore, I could not trust it blindly” (Student 21).
The Concrete Cost of Time: Regarding time management, Student 77 emphasized a general rule when reporting his perceptions before the quiz: “You always need to check the GenAI before relying on it.” Afterward, he described how the reality of the quiz shaped this process: “The fear that the GenAI might be wrong cost me time during the exam… I kept comparing its answers with the lecture slides, which took time instead of solving independently” (Student 77).
From Enthusiasm to Caution: In some cases, systemic perspectives shifted. Before the quiz, Student 67 expressed enthusiasm toward the GenAI-supported format, perceiving it as an innovative assessment approach. However, after the quiz, the same student expressed caution and doubt regarding its validity: “I think this does not really test knowledge and understanding, because you can just copy answers from the chatbot” (Student 67).
5.4. Preferred Assessment Format
Following their experience with the GenAI-enabled quiz, participants indicated their preferred assessment format in a single closed-ended question. Of the 90 participants in the study, 89 responded to this item. Of those respondents, 68.5% (n = 61) preferred the GenAI-supported format, while 31.5% (n = 28) preferred the traditional method. To investigate whether these preferences were driven by academic success, a point-biserial correlation was computed between final quiz grades and preferred format. The analysis revealed a near-zero, non-significant correlation (r = 0.014, p = 0.89), and an independent samples t-test confirmed no significant difference in grades between students who preferred the GenAI format (M = 70.44, SD = 19.78) and those who preferred the traditional format (M = 71.00, SD = 15.46).
6. Discussion
The present study examined students’ feelings and attitudes toward the use of GenAI in academic assessments before and after direct experience with using GenAI tools during a computer science or information systems course quiz. We explored how this experience influences students’ perceptions of learning, self-monitoring, cognitive regulation, and attitudes about academic assessments with and without access to GenAI.
In summary, our findings demonstrate that at the end of the study, 68.5% of respondents reported preferring GenAI-enabled assessments. Yet while participants initially anticipated that GenAI would provide reassurance and efficiency, the qualitative data revealed a subtle shift to what we might call measured pragmatism. More precisely, the experience heightened students’ understanding of GenAI as a cognitive scaffold that demands high-level metacognitive monitoring, such as the need to continuously verify potentially erroneous outputs (
Tankelevitch et al., 2024). Notably, students’ quiz performance showed no significant correlation with a preference for AI-enabled assessments, meaning that these preferences are driven by perceived pedagogical and professional value rather than individual academic success (
Yusuf et al., 2024).
The findings indicate that students’ direct experience using GenAI during the quiz did not substantially change their overall perceptions about the effects of integrating such tools into academic assessments. In both their retrospective reports about pre-experience attitudes and their post-experience reflections, students explicitly referred to process-related factors, such as using the tool for verification and checking, managing time constraints, and assuming personal responsibility for the final answer. This pattern aligns with the self-regulated learning literature, which emphasizes that regulatory processes do not change abruptly but are gradually adapted to context, task demands, and situational conditions (
Panadero, 2017;
Zimmerman, 2002).
One of the central contributions of this study concerns the metacognitive dimension of students’ engagement with GenAI during assessments. Although the increase in references to metacognitive skills was quantitatively modest, qualitative analysis revealed a substantive shift in the nature of students’ monitoring practices. Prior to direct experience, students primarily expressed a general awareness of the need to verify AI-generated outputs and a cautious stance toward the tool’s reliability. Following the exam, however, students described concrete, real-time monitoring practices, including checking AI-generated answers against their own reasoning, comparing outputs with additional sources, and making deliberate decisions about when to rely on the tool and when to refrain from using it. Based on our interpretive analysis of the qualitative material, we conclude that GenAI does not necessarily reduce metacognitive engagement but appears to increase the metacognitive demands placed on learners, particularly when AI outputs are perceived as uncertain and require ongoing critical evaluation (
Ng et al., 2021;
Tankelevitch et al., 2024). In this sense, AI literacy is reflected not merely in technical proficiency, but in learners’ capacity to exercise judgment, critically evaluate information, and assume responsibility for their learning-related decisions during assessment (
Annapureddy et al., 2025;
Ng et al., 2021).
Similarly, the findings on time highlight its central regulatory role in students’ engagement with GenAI during assessments. While changes in the salience of this theme were modest, a clear shift emerged in how participants experienced and interpreted time. Prior to the experience, time was primarily framed as an efficiency resource that GenAI could help save. Following the exam, however, time was identified as a constraint that limited students’ ability to verify AI-generated outputs and fully engage in metacognitive monitoring. This pattern is consistent with insights from cognitive load theory, which suggests that time pressure increases extraneous cognitive load and may undermine deep processing and effective self-regulation (
Sweller, 2011). In assessment contexts, the combination of rapid AI output and limited time appears to create concrete and immediate tension between the need for efficiency and learners’ desire to improve their understanding.
Notably, alongside an increased focus on the cognitive and metacognitive aspects of GenAI use, students’ reflections showed a diminished emphasis on broader institutional and social considerations, including the role of higher education institutions, assessment policy, and systemic preparedness for the AI era. This pattern does not suggest that such considerations disappeared, but rather that they receded into the background in favor of more immediate and practical engagement with the exam task itself. This shift is consistent with prior literature, indicating that direct experience with GenAI tends to redirect attention away from abstract discussions of policy and institutional adaptation toward personal decision-making, responsibility, and self-regulation in concrete assessment contexts (
González-Calatayud et al., 2021;
Xia et al., 2024).
Finally, the emergence of concerns about assessment authenticity and fairness only after direct experience underscores the role of practical engagement as a catalyst for ethical awareness in the field. Prior to the GenAI-supported quiz, students primarily articulated general considerations regarding the future of higher education and the labor market; following the quiz, they raised concrete questions about the validity of such assessments, the evaluation of individual knowledge, and the fairness of grading practices. This finding reinforces the argument that issues of integrity and authenticity do not arise solely at an abstract or principled level, but are instead shaped through direct encounters with changing assessment conditions (
Xia et al., 2024;
Yusuf et al., 2024).
The quantitative findings complement the patterns identified in the qualitative analysis. Although most students preferred an exam format that integrated GenAI, this preference was not associated with actual academic performance. This finding indicates that students’ positive perceptions of GenAI-supported assessment are not driven by instrumental considerations related to performance improvement, but rather by broader pedagogical and professional considerations, as well as concerns about the authenticity of learning. These statistical findings, combined with the qualitative insights, point toward an interpretive conclusion of a measured pragmatism in students’ attitudes-a recognition that with GenAI here to stay, it behooves all of us (students, educators, and the public) to work out its implications for teaching and learning.
The observed emphasis on verification and the shift toward measured pragmatism should be viewed through the lens of the participants’ disciplinary background. As computer science and engineering students, they are professionally socialized to debug, validate, and exercise skepticism toward automated outputs. This technical literacy, which likely includes more sophisticated prompting strategies, acts as a critical catalyst for the high-level metacognitive monitoring and individual responsibility identified in our findings.
Practical Contributions
Our findings indicate that academic assessments that integrate GenAI place greater demands on students’ ability to exercise judgment, verify information, and assume responsibility for final answers. Accordingly, the design of GenAI-supported assessments cannot be limited to binary decisions on whether to permit or prohibit these tools, but rather requires explicit clarification of expectations concerning how GenAI may be used, the limits of reliance on system-generated outputs, and students’ active role in monitoring the quality of responses and making final decisions.
In addition, the findings underscore the role of time pressure in shaping students’ engagement with GenAI during examinations. Under time constraints, students experience tension between using the tool for efficiency and the need to thoroughly verify and evaluate responses provided by the tool. This finding suggests that time allocation is not merely a technical parameter in the design of GenAI-supported assessment, but a variable that directly influences how students balance efficiency, depth of understanding, and self-regulatory control. Consequently, assessment designs that permit the use of GenAI must attend to the pedagogical implications of time constraints and their role in shaping the processes of verification, judgment, and responsibility during examinations.
Finally, the emergence of concerns related to assessment authenticity and fairness only after direct experience indicates that ethical awareness in the context of GenAI-supported assessment develops through practical engagement with assessment conditions rather than solely through abstract principles or general guidelines. This finding highlights the importance of embedding opportunities for reflection on assessment validity, individual knowledge evaluation, and grading fairness as integral components of assessment designs that incorporate GenAI. Crucially, our data indicates that from the students’ perspective, issues of authenticity and academic integrity are not merely about preventing misconduct, but rather about ensuring a fair assessment environment where true cognitive effort can be reliably demonstrated and evaluated.