Next Article in Journal
Preservice Secondary School Teachers’ Knowledge and Competencies When Reflecting on the Incorporation of Gamification in the Teaching of Mathematics
Previous Article in Journal
Teachers’ Handlingsrom Under Cross-Pressure: Developing the CP-Well Model of Well-Being in Gifted Education
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

ChatGPT in Programming Education: An Empirical Study on Its Impact on Student Performance, Creativity, and Teamwork

Faculty of Physics and Technology, Plovdiv University “Paisii Hilendarski”, 4000 Plovdiv, Bulgaria
*
Author to whom correspondence should be addressed.
Educ. Sci. 2026, 16(1), 19; https://doi.org/10.3390/educsci16010019
Submission received: 25 November 2025 / Revised: 18 December 2025 / Accepted: 21 December 2025 / Published: 23 December 2025
(This article belongs to the Section Higher Education)

Abstract

This study employs a two-part research design to explore the impact of ChatGPT on programming education for engineering undergraduates. Study 1 involved 56 third-year students who completed a questionnaire examining the frequency and purposes of ChatGPT use in text-based programming. While no statistically significant association was found between ChatGPT usage frequency and final grades, high-achieving students tended to use the tool less frequently, whereas lower-performing students relied on it more for support. Study 2 employed a counterbalanced repeated-measures design with nine first-year students divided into two groups, who developed desktop applications with and without ChatGPT. Project assessments and focus-group interviews were used to examine the effects of ChatGPT on creativity, confidence, effectiveness, and teamwork in visual programming. The results indicate that ChatGPT use was associated with reduced task completion time and increased coding efficiency; however, it was also linked to decreased creativity, greater reliance on ready-made solutions, and diminished code readability and collaborative engagement. These results highlight the need for teaching strategies that balance AI integration in programming education. The study recommends incorporating tasks based on the Reverse Bloom’s Taxonomy and activities that allow students to work with and without ChatGPT to encourage critical reflection and responsible AI use.

1. Introduction

Since the moment of releasing ChatGPT for free use in November 2022, it has become clear that one of its strongest points is the ability to generate programming code. The chatbot is capable of creating source code in almost all programming languages, converting code from one language to another, detecting errors, and suggesting optimal solutions to a given problem (Biswas, 2023). The chatbot understands and provides answers to natural-language questions, making it effective even for novice programmers (McCulloh et al., 2025).
Over about two years, ChatGPT has quickly gained recognition as a powerful programming tool, used by both software developers and students in programming-related disciplines. Our experience, along with studies by other researchers, shows that students use it for various purposes—from searching for information and detecting errors, to generating complete programming solutions.
This situation poses significant challenges for educators, especially in terms of the reliability of evaluating students’ performance, as students can now easily present chatbot-generated software as their own. In this regard, more and more teachers recommend conducting oral exams, which provide greater certainty that the demonstrated skills and knowledge belong precisely to the student.
Since communication with ChatGPT is via a text interface, its capabilities to generate solutions are more pronounced in text-based programming. Our experience shows that, in such a case, the chatbot is capable of creating entirely functional programs, adapted to the specific requirements of the assigned task. It can provide a ready-made working solution that the student can copy into the corresponding development environment and execute, even without understanding the logic behind it.
According to Stoyanova et al. (2025), as project complexity increases—particularly in cases involving visual programming—the ability of ChatGPT to generate fully functional code becomes significantly more limited. In this case, generating “ready-made solutions” is more difficult, as it requires skills and knowledge to manage visual controls, change their properties, and design the user interface. This makes it more difficult for the chatbot to create a complete, fully functional interface without additional user intervention.
This difference in the specifics of the two types of programming justified our decision to structure our study in two separate parts:
  • The first part focuses on text-based programming and examines whether there is a relationship between ChatGPT use frequency and students’ academic performance. A greater danger here is that students can over-trust the chatbot, which could affect the acquisition of knowledge and skills.
  • The second part examines the effect of using ChatGPT on students’ creativity, confidence, and performance in developing graphical user interfaces. In this context, the role of the chatbot is expected to be more supportive than entirely substitutive, which provides an opportunity to study its impact on teamwork and the decision-making process in more complex software projects.
The study was conducted among Bachelor’s degree students, majoring in engineering at the Faculty of Physics and Technology of Plovdiv University “Paisii Hilendarski”. Its results will help identify potential benefits and risks related to the uncontrolled use of ChatGPT in programming education.
Here are the research questions, guiding us in the study:
  • RQ1: Within the context of Study 1, is there a relationship between the frequency of ChatGPT use and academic performance in text-based programming?
  • RQ2: Within the context of Study 2, how does the use of ChatGPT versus traditional sources of information and support affect students’ effectiveness, confidence, and creativity during software project development?
  • RQ3: Within the context of Study 2, how does the use of ChatGPT versus traditional sources of information and support affect team communication and collaboration during software project development?
The research is relevant and of considerable scientific and practical significance, as it explores the relationship between ChatGPT usage frequency and academic performance, and its impact on students’ creativity, confidence, effectiveness, and teamwork during the development of group software projects. By analyzing both individual achievement and collaborative dynamics, the study offers valuable insights for educators seeking to adapt assessment and teaching strategies to the realities of AI-supported learning. Combined with conclusions from previous studies, these results will help formulate recommendations for adapting educational approaches and enhancing the effectiveness of generative AI in programming education.

2. Literature Review

In recent years, ChatGPT, as a representative of generative large language models (LLMs), has strengthened its position as a universal, multifunctional tool with broad potential for application in teaching and learning programming. The accumulated empirical data show that it can be effectively integrated into all phases of the learning process—from mastering the syntax and basic programming principles, through development of algorithmic thinking, to optimization and support of complex software solutions (Clarke & Konak, 2025; Ouh et al., 2023; Sun et al., 2024a; Yilmaz & Karaoglan Yilmaz, 2023).

2.1. The Most Important Applications of ChatGPT in Programming Education

ChatGPT can generate programming solutions across a wide range of programming languages (Python, Java, C++, JavaScript, etc.), adhering to the preset requirements and specifications. The generated code is usually characterized by correctness, clarity, and good structure, which helps students understand the syntactic and semantic peculiarities of different programming languages (Almanasra & Suwais, 2025; Groothuijsen et al., 2024; Rahman & Watanobe, 2023).
The chatbot can be used to create solution templates, develop alternative implementations, and generate missing code. In addition to creating new solutions, ChatGPT is capable of modifying and adapting already written code, adding new functionalities to it, or changing the logic and adapting a solution to other programming languages or frameworks. When solving programming tasks, ChatGPT not only generates code but also provides detailed explanations of the logic, the used programming structures, and the relationships between the individual components of the code (Chen et al., 2023; Sun et al., 2024b). The explanations, generated by the chatbot, based on code analysis, are a valuable tool, assisting students in developing their programming skills.
ChatGPT is capable of performing optimization of programming code by applying more efficient algorithms and data structures, in order to improve performance, decrease resource consumption, and increase the overall efficiency without changing the functionality of the software (Biswas, 2023; Rahman & Watanobe, 2023). Moreover, the chatbot can perform code refactoring in accordance with established principles and standards of good programming practices, offering more elegant, better structured, and more efficient solutions, aimed at increasing software quality, readability, and maintainability (Haque & Li, 2023; Sun et al., 2024a).
ChatGPT can serve as a useful tool for identification and explanation of syntactic and logical errors by means of programming code analysis, as well as for suggesting specific methods and strategies for their elimination (Biswas, 2023; Groothuijsen et al., 2024; Haque & Li, 2023; Zviel-Girshin, 2024). This functionality is especially valuable when students work independently to prepare their assignments, as it reduces the time and effort required to detect and correct errors (Sun et al., 2024a). According to Biswas (2023) and Surameery and Shakor (2023), the chatbot can also be used to detect erroneous code.
ChatGPT has the potential to generate detailed and context-oriented explanations and examples of the fundamental concepts in programming, encompassing various paradigms, data structures, and algorithms (Haindl & Weinberger, 2024b). Rahman and Watanobe’s (2023) research shows that ChatGPT is capable of generating clear and easily understandable explanations, which help learners for the better acquisition of the basics of programming and the theory of algorithms. Thanks to these features, it gains recognition as an effective complement to traditional methods of teaching, providing constant and interactive support for understanding and applying code.
ChatGPT possesses a wide set of functionalities, which can significantly assist teachers of programming in their pedagogical activities. It can be used for making detailed plans of lessons, for preparing thematically relevant learning content, and delivering presentations (Bringula, 2024; Humble et al., 2023; Husain, 2024; Rahman & Watanobe, 2023). Along with learning content creation, the chatbot is also capable of generating programming tasks, automated tests, and assessment criteria, as well as performing automated reviews and evaluation of code developed by students (Bringula, 2024; Rahman & Watanobe, 2023). As Jukiewicz (2024) and Zviel-Girshin (2024) note, the automation frees up valuable time resources, which could be directed by the teachers toward more complex pedagogical activities, individual consultations, and targeted support for the learners. However, integrating chatbots does not diminish the teacher’s central role in the learning process. As emphasized by Ilieva et al. (2023), instructional planning, assessment design, and course management remain the instructor’s responsibility, with chatbots serving a supportive role rather than replacing pedagogical decision-making and professional expertise.

2.2. Main Advantages of Using ChatGPT in Programming Education

ChatGPT is accessible at any time and from any device with an Internet connection, without requiring the installation of specialized software. It provides constant support for learners by offering immediate and mostly correct answers, easy access, and the opportunity to use it anytime, anywhere (Lin et al., 2023; Sun et al., 2024b; Yilmaz & Karaoglan Yilmaz, 2023).
ChatGPT uses natural-language processing to communicate effectively with users. Its ability to interpret natural-language instructions and convert them into syntactically correct, functional code makes it especially valuable for beginners in programming (Surameery & Shakor, 2023). Its ability to work in a multiplicity of languages removes language barriers and creates conditions for more accessible and efficient training in a multicultural educational environment (Yilmaz & Karaoglan Yilmaz, 2023).
Personalization of the learning process, addressed towards satisfying the individual needs of each learner, has been identified as a key factor for increasing students’ commitment (Pesovski et al., 2024a). In this context, ChatGPT can play an important role when adapting explanations, resources, and feedback to the level of preparation, style of learning, and specific interests of each trainee. Adaptability is particularly important, as it enables effective support for both novice programmers who need basic and more detailed explanations and advanced learners who require more complex, in-depth conceptual discussions (Hartley et al., 2024; Penney et al., 2025). By adapting resources and academic activities to the specific difficulties faced by the learners, ChatGPT can give assistance to the process of purposeful development of their programming skills (Rahman & Watanobe, 2023).
The feedback provided by ChatGPT is distinguished by its high precision, especially at the stage of introducing programming tasks, where it can identify errors, explain their causes, suggest corrections, and provide various options to solve the problem (Kiesler et al., 2023). Such a functionality helps in developing self-reflection skills while reducing the time and effort required to acquire new concepts. By providing different alternative solutions to the same problem, the chatbot stimulates students to compare approaches and evaluate their effectiveness, thus developing their critical thinking (Clarke & Konak, 2025).
In addition, by providing structured guidance for planning, distributing academic tasks, assigning exercises of gradually increasing complexity, recommending learning materials, assessing, and giving immediate feedback, ChatGPT supports self-regulated learning (Hartley et al., 2024). The results of the research, carried out by Wu et al. (2024) and Silva et al. (2024), show that training, supported by ChatGPT, effectively supports the development of self-regulation skills and stimulates knowledge acquisition.

2.3. Limitations and Disadvantages of Using ChatGPT in Programming Education

Despite its significant potential, the use of ChatGPT in programming education is related to a number of limitations and challenges, which can negatively impact the quality of the acquired knowledge and skills. One of the main concerns, highlighted in the literature, is the risk of learners becoming excessively dependent on AI chatbots, which could result in a decline in their higher-order thinking skills, reduced ability of solving problems independently, and perfunctory understanding of key programming concepts (Groothuijsen et al., 2024; Husain, 2024; Sánchez-Ruiz et al., 2023; Zviel-Girshin, 2024). Such dependence can form a passive attitude towards learning. The easy accessibility and the trouble-free integration of ChatGPT can lead to weakening students’ motivation for independent work and critical engagement with programming tasks, as well as to increasing their reliance on ready-made code solutions, offered by the chatbot, the logic and theoretical basis of which they do not fully understand (Humble et al., 2023; Xue et al., 2024; Zviel-Girshin, 2024). Moreover, excessive automation—for example, assigning entirely to the artificial intelligence to generate code comments, to structure or to name variables—can limit the development of practical skills and hinder the formation of a deeper understanding of the principles of programming (Zviel-Girshin, 2024). All of this limits the development of critical thinking and reduces students’ ability to withstand the challenges related to solving complex programming tasks (Sánchez-Ruiz et al., 2023; Silva et al., 2024; Rahman & Watanobe, 2023).
A series of studies have shown that the quality of the programming solutions, generated by ChatGPT, is not always guaranteed, as the chatbot often suggests inaccurate or partially correct code, as well as explanations, which, though sounding convincingly, can mislead beginners (Bucaioni et al., 2024; Haindl & Weinberger, 2024a; Humble et al., 2023; Yilmaz & Karaoglan Yilmaz, 2023). Such cases usually arise from inaccurate interpretation of the task by the chatbot or from unclear and incorrectly formulated requests, leading to wrong or suboptimal solutions (Husain, 2024; Zviel-Girshin, 2024). The following characteristics stand out among the additional limitations of using ChatGPT: limited ability for logical reasoning, poor (inadequate) processing of visual information, as well as lack of integrated environments and tools for programming (Rahman & Watanobe, 2023; Yilmaz & Karaoglan Yilmaz, 2023). Experimental results in the studies of Bucaioni et al. (2024) show that despite the high efficiency of ChatGPT in solving programming tasks of low or medium difficulty, its success rate decreases with more complex logical and algorithmic problems. Analysis of the results reveals that the accuracy of the generated code significantly decreases as the complexity of the tasks increases.

2.4. Ethical Considerations and Challenges in Using ChatGPT in Education

Implementing ChatGPT in the educational process raises a number of complex and multifaceted ethical questions, ranging from bias and discrimination, through plagiarism, confidentiality and data safety, information accuracy and reliability, to transparency and social impact (López-Fernández & Vergaz, 2025; Rahman & Watanobe, 2023). According to Ray (2023), the use of ChatGPT for generating written content raises important questions related to the ownership of intellectual property rights and the proper recognition and attribution of authorship.
ChatGPT is trained on large-scale data sets, which very often contain or reproduce biases, prejudices, and social inequalities existing in society (Adel et al., 2024). An essential challenge is the risk that ChatGPT could reproduce or even deepen existing prejudices, embedded in its training data, including gender, racial, cultural, and linguistic stereotypes (Rahman & Watanobe, 2023; Ray, 2023). This can have an adverse impact on the quality of the learning materials, comprising content, created by ChatGPT (Pesovski et al., 2024a). Anagnostopoulos (2023) notes that biases in the training data can be transferred to the generated by ChatGPT programming code or model, leading to prejudiced or discriminatory suggestions. Such deviations are clearly manifested when the created program processes data generated by the chatbot itself (Huang et al., 2023).
A significant ethical risk is posed by issues related to data confidentiality and security. Storage and processing of users’ information raises serious concerns, especially when it comes to young users with limited knowledge of personal data protection (Adel et al., 2024; Silva et al., 2024). The insufficient transparency about the way in which ChatGPT processes, uses, and potentially reuses the information provided for retraining underlines the need for clear regulations and ensuring informed consent (Husain, 2024; Rahman & Watanobe, 2023; Ray, 2023).
Plagiarism, including programming code, is among the most frequently discussed issues (Rahman & Watanobe, 2023; Silva et al., 2024). The ability of ChatGPT to generate ready-made solutions, which students could present as their own work, undermines the validity of assessment and creates conditions for academic dishonesty, especially in online exams and assignments (Adel et al., 2024; Humble et al., 2023; Husain, 2024). The existing plagiarism detection tools have limited effectiveness in recognizing AI-generated content, which requires the development of new verification and control methods (Adel et al., 2024; Gill et al., 2024; Rahman & Watanobe, 2023).
The accuracy and reliability of the answers generated by ChatGPT also have an important ethical dimension. The chatbot may produce grammatically correct but factually inaccurate or misleading statements (Ray, 2023). This poses a risk of spreading misinformation in the learning materials and forming erroneous knowledge, especially if the content is not subjected to critical scrutiny by learners and educators (Chang et al., 2024; Gill et al., 2024; Silva et al., 2024).
The analysis of the reviewed literature sources shows that ChatGPT possesses significant potential to support programming education by means of a wide spectrum of functionalities—from code generation and optimization, through providing immediate and personalized support, to automating routine activities and facilitating the work of the teachers. However, its effective application requires conscious pedagogical integration, clearly articulated ethical guidelines, and targeted development of AI literacy among both students and lecturers.
The accelerated evolution of AI technologies, including ChatGPT, necessitates additional scientific inquiry into the pedagogical and ethical implications of their integration into programming education. The expected results of such research would contribute to a deeper understanding of the influence of ChatGPT on the learning process in a wide range of educational programs and contexts, as well as to the formulation of relevant and practically applicable recommendations for its responsible and effective use.

3. Materials and Methods

3.1. Context

The study was conducted at the Faculty of Physics and Technology of Plovdiv University “Paisii Hilendarski”, which is a state university, the second largest in Bulgaria. The Faculty of Physics and Technology educates students in several Bachelor’s engineering major programs, in whose curricula software disciplines occupy a significant share. Students study programming basics, object-oriented programming, visual programming, web-based programming, etc. During their programming training, they complete practical assignments with the support of their lecturers and using printed textbooks and various online resources.
Over the past two years, the frequency with which students use ChatGPT during their programming training has increased significantly. This raises the need to study its impact on students’ academic performance, creativity, confidence, and work efficiency when developing team projects in visual programming. Such projects closely resemble real-world software development, where effective communication and coordination are essential. Therefore, analyzing the role of ChatGPT in this educational context is important for examining whether its use supports or hinders the development of students’ teamwork and idea-sharing skills that are essential for their future professional practice.
The present study was conducted in the second semester of the academic year 2024/2025 within the training courses in two disciplines that utilize different programming approaches:
  • “Software Development Practicum 2”, focused on object-oriented C++ programming in a text-based programming environment;
  • “Developing a Graphical User Interface—C#”, aimed at developing a graphical user interface in Visual Studio.

3.2. Design and Participants

The study involved students from two Bachelor’s degree majors. Their participation was voluntary, and we used convenience sampling. The study was structured in two separate parts:
  • Study 1: This research investigates whether the frequency of using ChatGPT affects the academic performance of the students in the course “Software Development Practicum 2”;
  • Study 2: This research examines how the use of ChatGPT affects the creativity, confidence, and work efficiency of the learners when developing team projects in visual programming.
At the beginning of the experiment, students were instructed on how to use ChatGPT effectively (clear prompts, context, and follow-up questions) and responsibly (personal responsibility, critical evaluation of content, and academic honesty).

3.2.1. Study 1

Fifty-six (56) Bachelor’s degree full-time third-year students, majoring in “Information and Computer Engineering” were involved in the study. The course “Software Development Practicum 2” aims to consolidate and enrich the knowledge acquired in “Object-oriented programming”, a course studied by second-year students. During the Practicum, the students developed various assignments, primarily based on text-based programming. The tasks were selected to develop skills across the different levels of Bloom’s taxonomy, thus encouraging both basic understanding and higher levels of thinking.
The training in the discipline concludes with the defense of an individual project developed during the semester. The project is defended through a practical examination, during which students answer questions and complete practical tasks related to the software they developed. The tasks include both modifying the student’s code and creating new methods or functionalities. The main goal of this examination format is to objectively assess the skills acquired by the student in programming, even if the student has used external assistance, including ChatGPT, when preparing the project.
Throughout the course, students were free to use ChatGPT during the completion of assignments and the development of their course projects, without any restrictions or requirements imposed by the instructors. As students used their personal devices and accounts, it was not possible to determine whether the free or paid version of ChatGPT was used, nor to identify the specific language model.
After defending their projects, the students voluntarily complete a questionnaire designed to assess the frequency and purposes for which they have used ChatGPT when completing the assignments and the final programming project. They are informed in advance that the questionnaire is intended only for those who have used ChatGPT in their learning in the respective subject. Out of a total of sixty (60) students, three (3) chose not to participate, and one questionnaire was declared invalid, because the respondent indicated that he had never used ChatGPT in a programming context.

3.2.2. Study 2

Bachelor’s degree, full-time first-year students majoring in “Hardware and software systems” were involved in Study 2. In their second semester, these students study the discipline “Developing a Graphical User Interface—C#”, which follows the completion of the “Basics of Programming” course. The experiment was conducted within two consecutive days in June 2025, after the students had passed their examination in the discipline. All students in the major who had passed the exam were invited to participate in the experiment, without offering any incentives or linking participation to grades or mandatory academic requirements. Nine (9) of the twenty-seven (27) invited students applied to participate. All experimental activities were conducted face-to-face in a classroom setting, and all students used university-provided computers. All students were provided with equal access to the freely available version of ChatGPT on university computers. No paid versions or additional AI tools were available, ensuring uniform access conditions for all participants. Compliance with the experimental conditions was ensured through direct supervision by the instructors.
Students were randomly divided into two groups:
  • Group 1: 5 students;
  • Group 2: 4 students.
A counterbalanced repeated-measures design was applied during the experiment—each group developed two WinForms applications in the Visual Studio environment, using two different approaches:
  • With ChatGPT;
  • Without ChatGPT (using traditional resources such as Internet search engines, electronic learning materials, paper textbooks, etc.).
For the study, two tasks for creating educational games were prepared, intended for first-grade pupils. The tasks were comparable in difficulty, volume, and requirements to eliminate the influence of the task itself on the results.
  • Task 1 (Mathematics game): to create an educational game for 1st class pupils, intended to support them in learning mathematics. The game should help children develop skills in comparing, adding, and subtracting numbers from 1 to 10.
  • Task 2 (Language game): to create an educational game, helping 1st class pupils in their Bulgarian language and literature training. The game should help children develop the ability to recognize the printed letters of the Bulgarian alphabet.
The two tasks were in line with the course material, and all the necessary knowledge for their implementation was covered during the lectures and exercises. The students worked in teams, and the time to develop each assignment was 4 h and 30 min. Each group applied both approaches—with and without ChatGPT.
The functional requirements for the games included:
  • The following components must be used: Label, Button, Picture Box, Timer, Message Box;
  • Only components discussed during the training sessions must be used;
  • Feedback to the player must be provided on correct or incorrect answers;
  • The number of correct answers must be counted;
  • When the preset time expires, the player’s score should be displayed;
  • The interface should be tailored to the age of the target group—8-year-old children.
The sequence of the experimental phases is presented in Figure 1. The initial Task–Method pairing (with or without ChatGPT) was randomly assigned for each group. The exchange of methods between the two groups occurred on the second day. Although a short break was provided between experimental stages, potential learning transfer between stages cannot be fully excluded.

3.3. Data Collection Techniques

3.3.1. Study 1

The questionnaire in the first part of the study was intended to assess the frequency and purposes for which the students had used ChatGPT when completing their assignments and the final programming project. It contained seven (7) questions, which the students had to assess using a 5-point Likert scale (from 1—Never to 5—Very often). Each question was structured to serve a specific purpose for using ChatGPT—from acquiring theoretical knowledge and understanding other people’s code to generating new code, detecting errors, optimizing, and searching for ideas. The questionnaires were originally developed and administered in Bulgarian. For publication purposes, the items were translated into English and reviewed by a faculty member to ensure conceptual equivalence.
Data processing was performed by means of the Software Package for Statistical Processing (SPSS 13.0). Since the questions are considered individual indicators, no Cronbach’s α was calculated to assess internal consistency.

3.3.2. Study 2

Data for this study were collected through scoring the developed projects and through conducting a semi-structured interview with the students at the end of the experiment.
Each completed project was scored jointly by the two lecturers, who had led the lectures and conducted the exercises in the subject “Developing a Graphical User Interface—C#”. The assessment criteria were announced in advance and included: code completeness and efficiency, code organization and readability, and interface quality and creativity. The assessment followed a 5-point scale (1–5), with the teachers justifying the scores they assigned to the above-mentioned parameters of the students’ work.
After the experiment was completed, a semi-structured focus-group interview was conducted with each group. The focus-group format facilitated open communication and stimulated discussion. The interview aimed at studying the participants’ perceptions of the two approaches to work—with and without using ChatGPT—in the context of developing software projects in visual programming. Due to the small number of participants in Study 2, its findings should be interpreted as exploratory rather than statistically generalizable.

3.4. Ethical Considerations

The students were informed in advance about the aims of the study, that participation was entirely voluntary, and how their data would be used. Participation in neither Study 1 nor Study 2 contributed to course grades in any way. In Study 1, students completed the questionnaire after defending their projects and after the corresponding grades had already been awarded. This order was chosen to ensure that participation and responses could not influence the assessment of students’ academic performance. In Study 2, the experimental activities were conducted after the course examination had been completed and final grades had been officially announced. Students were informed of their right to withdraw from participation at any stage without any adverse consequences. Although participation was non-anonymous, all data were treated as confidential, analyzed at the group level, and reported only in aggregated form. Both the questionnaire and the semi-structured interview questions were approved by the ethics committee of the Faculty of Physics and Technology at Plovdiv University “Paisii Hilendarski”.

4. Results

4.1. Study 1

The first question in the questionnaire used in Study 1 aims to determine the frequency with which students use ChatGPT as an aid when completing their programming assignments, including when developing the final project. The distribution of the given answers is presented in Table 1. The largest number of students (50%) indicate that they have used the chatbot “Sometimes”, while the smallest number (5.4%) indicate that they have used it “Very often”. Seven students declare using it “Rarely”.
Correlation analysis was applied to test whether there is a statistically significant relationship between students’ grades in the studied discipline and the frequency of ChatGPT use. The calculated Spearman coefficient is −0.1, which indicates that there is a weak negative relationship—lower assessment grades may be observed with more frequent use of ChatGPT. However, the calculated p-value is 0.471, i.e., greater than the accepted level of significance α = 0.05. Therefore, it cannot be assumed that there is a statistically significant relationship between the assessment grade and the frequency of using ChatGPT.
K-means cluster analysis with two clusters was applied to identify the existence of similar patterns of using ChatGPT among the students. The analysis included six (6) questions (Q2–Q7) from the questionnaire, aiming to determine the main purposes and ways of using ChatGPT by the students in completing their course assignments and the programming project:
Q2. How often do you use ChatGPT as a resource (e.g., for explaining methods, operators or concepts in programming)?
Q3. How often do you use ChatGPT to help you understand the logic of a ready-made program code?
Q4. How often do you use ChatGPT to help you in writing code on a given condition or task?
Q5. How often do you use ChatGPT to detect and debug errors in a piece of code, written by you?
Q6. How often do you use ChatGPT to optimize the code you have already written?
Q7. How often do you use ChatGPT to get ideas or guidance on how to solve a complex software problem, without asking for a direct solution?
The results of the cluster analysis revealed two clearly distinguishable profiles (Table 2):
  • Cluster 1 (n = 32). This is the group of students actively using ChatGPT. It is characterized by higher mean values of the answers to all questions Q2–Q7, varying from 2.97 to 4.22.
  • Cluster 2 (n = 24)—the group of students, rarely using ChatGPT, characterized by lower mean values of the answers to all questions, ranging from 1.71 to 3.00. The highest mean value in this cluster (3.00) is reported for the question “How often do you use ChatGPT as a resource?”
Figure 2 presents a boxplot diagram, visualizing the distribution of the assessment grades in the studied discipline for each of the two clusters, including the mean values. For Cluster 1, the median is 3.0, which means that half of the students in this group have an assessment grade less than or equal to 3. This is a relatively low value and can be regarded as an indicator of a potential relationship between the frequent use of ChatGPT and lower academic achievement. For Cluster 2, the median is 4.5, which suggests that the students in this group have higher academic results.
The calculated mean values confirm the trend—the students in Cluster 2 have higher grades (4.40) compared to those in Cluster 1 (3.72). Although the first quartile (Q1) is the same in both clusters (3.0), the third quartile (Q3) is significantly higher in Cluster 2 (6.0 versus 4.63), which again shows that the highest grades are more common among the students who rarely use ChatGPT in programming.
A single-factor dispersion analysis was applied to test whether there are statistically significant differences in student grades across the different clusters. The calculated p-value is 0.093, which is greater than the accepted significance level of α = 0.05. This means that there are no statistically significant differences between the average grades of the students from the two clusters, formed based on the purposes for using ChatGPT.

4.2. Study 2

4.2.1. Results of the Evaluation of the Developed Projects

Both groups completed their projects within the planned time of 4 h and 30 min. Table 3 presents the mean scores for the completed projects by categories for the two groups of students when using the two approaches—with and without ChatGPT.
The results show that the projects, developed with ChatGPT for both teams, received higher scores on the completeness and code efficiency criteria. The assessment score increases from 3.50 to 4.50 in Group 1, and from 4 to 5 in Group 2. These findings suggest that the use of ChatGPT may support students in producing more optimized code, potentially by providing ready-made code segments and suggestions for improvement.
However, a drop in scores is observed for the “Organization and readability of the code” indicator when using ChatGPT. The lecturer notes that the automatically generated code is often more complex and not tailored to the level of the students. In addition, students often add new functionalities without removing similar or duplicate ones already generated by the chatbot, which disrupts the logical structure of the resulting code. Moreover, they change the variable names, making them more descriptive but longer, which further worsens the code’s readability.
Regarding the “Interface quality” criterion, both groups received lower scores when working with ChatGPT. The reason, according to the lecturer, is the use of standard solutions, not modified or adapted to the project’s specifics. Building a user interface is a creative activity that requires experimentation with the design, but the students obviously trusted the first solution suggested by the chatbot. Even though both approaches were positively evaluated for offering help to the users and for being resilient to errors, the projects developed without using ChatGPT demonstrated more originality and attention toward the target group on the side of the students. They selected interesting images, wording, and color schemes, suitable for young children, and showed more imagination and attention to detail.
This is also reflected in the “Creativity” criterion, where the most significant drop in the score is observed when using ChatGPT—by 2 points in Group 1 (from 5.0 to 3.0) and by 1 point in Group 2 (from 5.0 to 4.0). The offered ready-made solutions, though functional, are more uniform compared to the projects developed without the help of the chatbot. The lecturer notes that the projects developed without ChatGPT are more interesting and contain more original details.
Finally, in terms of execution time, ChatGPT in both groups considerably shortened the process by approximately 55 min in Group 1 and 45 min in Group 2. This may be attributed to several factors. Using the chatbot as a single source of information can reduce development time. At the same time, higher creativity scores were observed in projects developed without ChatGPT, suggesting that students may have devoted more time to conceptualizing application design, thereby extending the overall development process.

4.2.2. Analysis of the Semi-Structured Interview with the Students

During the interview, the students had to answer questions aimed at studying their perceptions of the advantages and limitations of the two approaches (with or without ChatGPT), as well as the impact of these approaches on creativity, effectiveness, confidence, and team interaction between the students when they worked on developing their software projects. The two approaches were compared by several main indicators, through performing qualitative analysis of the students’ answers.
Effectiveness. The participants agree that the use of ChatGPT leads to significantly faster completion of their assignments: “the work with ChatGPT was easier”, “with ChatGPT we worked faster”, “we didn’t have to search the Internet for a long time or to ask for any help from the teacher”. The solutions suggested by the chatbot were accepted in most cases without rationalization and full understanding: “we just accepted the suggestion without thinking much”, “we trusted the suggestions, even if we didn’t fully understand what the code was doing”. When working without being helped by the chatbot, the students wrote and analyzed the code themselves, which “was difficult but very useful”. They realize that this process, though slower and requiring more effort, has significantly higher educational value.
Creativity. According to the participants, the use of ChatGPT limits creativity, as they rely on it for ideas and ready-made suggestions, accepted by them without discussion and critical analysis: “rather, we trusted it, sometimes without checking.” Without ChatGPT, they “improvise more”, show “greater imagination”, and create “more original scenarios”. This time they rely on their own skills and experience: “we had to figure out how to make the game interesting and understandable for children”, “we combined different ideas in our own”.
Teamwork. All participants share that the use of ChatGPT reduced the level of communication and joint action for solving the problems: “everyone worked independently”, “in isolation”, “tried to complete their part faster”, and discussions were only held when the suggested answers were not satisfactory: “only when the answers didn’t suit us, did we discuss them”. When working without using ChatGPT, the students report considerably more active team interaction: “we discussed every decision and the different points of view”, “we talked and looked for a solution together”, “we commented and consulted with each other”.
Confidence. According to the participants, ChatGPT boosted their confidence by suggesting “quick answers, serving as a starting point” and helping “when difficulties occurred”. When working without using the chatbot, they “made more mistakes”, and “asked the teacher for help”. However, sometimes the lack of full understanding of the proposed solutions reduced the sense of competence: “Sometimes ChatGPT provided solutions, which we didn’t fully understand”, “we couldn’t always understand what the code was doing”, and this “made us insecure”.
The students gave the following answers to the question “When using the ChatGPT-supported approach, what were the primary purposes for which you used the chatbot?”: “for generating ready-made code”, “for explanations”, “for debugging”, “for ideas on how to get started”, “for a description of how the components operate”, “for design suggestions”. These answers show that the students relied on ChatGPT in all stages of project development. This likely reduced the need to independently search for solutions, which may explain, in turn, the lower creativity scores with this approach.
The last question of the interview aimed at establishing which of the two approaches—with or without ChatGPT—would be preferred by the participants in their future software developments. All respondents categorically state that they would continue using ChatGPT in the future, with the main motives being related to efficiency and time-saving: “ChatGPT saves time”; “provides ready-made solutions”; “if the task is more complex, ChatGPT is irreplaceable”; “it is ideal for getting ideas”.
However, the students underline that an approach combining ChatGPT with the use of other sources would be more useful. Most participants prefer to use the chatbot as a means of generating ideas and then to continue their project development by individual efforts or teamwork. One participant emphasized that ChatGPT saves time, but creativity improves when “it is combined with other sources of information and team discussion.” Another student highlighted its role “as a starting point for ideas, followed by individual work and teamwork.” Such a model of work, according to them, would contribute to better knowledge acquisition (“I will learn better”), deeper understanding of the content (“I will understand better”), and higher creativity and quality of performance (“we shall work more creatively and qualitatively”). Some respondents differentiate the usefulness of the chatbot according to the complexity of the project. According to some, ChatGPT is “irreplaceable” for more difficult tasks, but otherwise, they think they will learn better without using it.

5. Discussion

RQ1: Within the context of Study 1, is there a relationship between the frequency of ChatGPT use and academic performance in text-based programming?
The results of the study show that students can be divided into two clearly distinguishable groups based on the frequency and purposes of using ChatGPT in programming education. The first group is characterized by active use of the chatbot and its active integration into a wide range of activities, related to training in programming, from reference purposes to code optimization and refactoring, as well as to idea generation. The second group has used ChatGPT significantly less frequently for all tasks examined in the study. Although students who use ChatGPT more frequently and for a wider range of purposes show a tendency toward lower mean assessment grades in their programming course, this difference is not statistically significant at α = 0.05. This suggests that the manner and frequency of using ChatGPT for various purposes in programming education do not significantly affect the final assessment grade in programming.
Comparable results are reported by Sun et al. (2024b), who applied a quasi-experimental design to investigate the impact of using ChatGPT in programming education on student achievement. Their study has been conducted as part of a Python programming course. The results show that there is no statistically significant difference between the academic achievements of students in an experimental group, using ChatGPT, and those in the control group, working traditionally. An identical result is described by Kosar et al. (2024), who have experimented with two groups of freshmen (one group using ChatGPT and the other not using it) working on practical assignments in object-oriented programming. Again, the results show that the use of ChatGPT does not have a statistically significant impact on students’ academic performance.
According to Huesca et al. (2024), however, combining ChatGPT with pedagogical strategies that encourage active learning can significantly improve learning outcomes. In an experiment conducted with programming students, the authors found that integrating ChatGPT into a flipped learning strategy led to a significant increase in normalized learning outcomes and a deeper understanding of the studied concepts compared to the traditional video-based approach.
Although the differences in mean assessment grades between the two groups are not statistically significant, the data suggest that a relatively larger proportion of students with high programming assessment grades belong to the group of infrequent ChatGPT users. The possible reasons for this are that students with a high level of academic preparation often rely on their own knowledge and problem-solving skills, reducing the need for external help. Students with lower academic performance more often resort to ChatGPT to overcome preparation deficits. A similar relationship has been documented in other studies, showing that lower-achieving students are more likely to resort to technological aids (Hazelhurst et al., 2011). Lehmann et al. (2025) state that the impact of LLMs on the learning outcomes depends on how they are used. When students rely on chatbots for automatic solution generation, they acquire knowledge more superficially. Conversely, using LLM to provide step-by-step explanations and clarifications may support students’ conceptual understanding of course material (Raihan et al., 2025).
It is noteworthy that, in both groups, some of the highest mean values are observed for the statement “I use ChatGPT as a resource”. This indicates that students in both groups perceive the chatbot primarily as a rapid and easily accessible source of information. Such a tendency raises important questions about the need to train students to critically evaluate and verify information generated by AI systems.
The lower success rate among the students who use ChatGPT more often, reported in the study, raises the question of whether the increased and prolonged use of artificial intelligence will have in the future an adverse effect on the cognitive development of learners. Rahe and Maalej (2025) have established that students who frequently use ChatGPT are more likely to ask the chatbot to generate direct solutions for them, rather than to engage themselves in understanding and analyzing the specific task. This behavior supports concerns that overreliance on generative artificial intelligence may limit the development of critical thinking and algorithmic logic in future programmers. Some authors, such as León-Domínguez (2024), hypothesize that working with ChatGPT and similar LLMs can reduce the active, problem-solving type of thinking. This could lead to changes in the cognitive development of future generations.
RQ2: Within the context of Study 2, how does the use of ChatGPT versus traditional sources of information and support affect students’ effectiveness, confidence, and creativity during software project development?
The results of the study are consistent with previous findings that ChatGPT may support students in developing effective programming code (Jamil et al., 2025; Uandykova et al., 2024). According to our results, projects developed with ChatGPT were completed faster and received higher assessment scores for code completeness and efficiency. This is confirmed in the present study by the higher assessment scores for completeness and efficiency of the code for the projects, developed with ChatGPT, as well as by the reduced completion time. A similar conclusion has been drawn by Q. Yang (2024), who compares codes generated by artificial intelligence applications (GitHub Copilot, Microsoft Copilot, Tabnine, and ChatGPT) in terms of correctness, complexity, efficiency, and size. The results show that ChatGPT generates efficient and detailed code, which is, however, also the most complex among the codes generated by the studied AI assistants.
On the other hand, the observed decrease in scores for code organization and readability in student projects developed with ChatGPT, based on the lecturers’ assessment, is consistent with concerns reported in the literature that code generated by AI-based assistants is often not optimally structured and that students may experience difficulties when editing and debugging such code (Vaithilingam et al., 2022). Q. Yang (2024) also emphasizes that despite the high quality of the code, the solutions generated by artificial intelligence often differ significantly from those written by software developers. This can create challenges in integrating AI-generated and handwritten code, necessitating further adaptation.
In terms of creativity, the results of the present study show a decline in performance when working with ChatGPT, consistent with the trend established by Rajan and Niranjan (2025) that reliance on ready-made solutions limits students’ original thinking and ideas. The analysis of the interviews shows that the participants relied on the chatbot at all stages of their work—from project structuring and code generation, to debugging and design suggestions. Thus, the need for independent search for solutions was reduced and improvisation limited, especially in the more creative aspects of the task. This is consistent with the findings of Liu et al. (2024), who reported that although ChatGPT initially enhances creative productivity, this effect does not persist over time, and creative performance tends to return to baseline levels once ChatGPT is no longer available. Similar conclusions have been drawn by Farah et al. (2025), whose study among students of software engineering shows that the use of LLM-based chatbots in teamwork for idea generation leads to fewer proposed solutions.
Regarding confidence, the results of the interview with the students show that ChatGPT has had a positive effect on their confidence, as it provides quick answers and guidance for overcoming problems. This is consistent with the results reported by Yilmaz and Karaoglan Yilmaz (2023), who found that, according to their students, the main advantage of ChatGPT is providing quick and mostly correct answers to questions and thus increasing confidence. However, our research shows that in some cases, the lack of a full understanding of the code generated by the chatbot leads to a feeling of uncertainty among the students. T.-C. Yang et al. (2025) found that the use of ChatGPT in programming education is associated with lower levels of student confidence and academic achievement compared to conventional methods, as the chatbot’s support often falls short of students’ initial expectations. They suggest that although ChatGPT can provide extensive knowledge and guidance, its inability to promptly assess learners’ understanding and adapt the level of challenge accordingly may disrupt the learning process. Based on these findings, T.-C. Yang et al. (2025) recommend that the integration of AI technologies be aligned with task complexity and learners’ cognitive levels.
RQ3: Within the context of Study 2, how does the use of ChatGPT versus traditional sources of information and support affect team communication and collaboration during software project development?
The data collected from the interview shows that when working without ChatGPT, the students collectively discuss and jointly solve the problems that arise. Conversely, when working with ChatGPT, they interact with each other less frequently, work more in isolation, and discuss solutions only when problems arise. These findings are consistent with concerns raised by Campillo-Ferrer et al. (2025), who report that teachers perceive the misuse of generative AI as a potential risk to students’ social and communication skills, leading to decreased interpersonal interaction. A similar inference is made by Groothuijsen et al. (2024), who report that the use of ChatGPT in programming education has a negative impact on pair programming and student collaboration. In line with these studies, the present research reports a reduction in peer communication during group work, suggesting that unstructured or unmoderated use of such tools may influence collaborative interaction.
Despite the reported shortcomings, all students have expressed a desire to use ChatGPT in the future, mainly because of the complete and efficient code it generates and its time-saving ability. According to the participants in our study, a combined approach that uses ChatGPT for idea generation and guidance, followed by individual or team-based development, supports better knowledge acquisition, deeper understanding of the content, and higher-quality final project outcomes.

6. Limitations of the Study

Several limitations should be considered when interpreting the results of the present study:
  • The study was conducted at a specific point in time and reflects the students’ opinions, identified benefits, and potential risks associated with the use of ChatGPT in programming education at that specific time.
  • The sample consisted of students from a single faculty at one university.
  • Due to the small number of participants, especially in the second study, it cannot be claimed that the conclusions are statistically significant or representative.
  • Some of the data were collected through self-assessment questions, which poses a risk of bias and inaccuracies in self-reporting.
  • In Study 1, ChatGPT was used in an uncontrolled environment, making it impossible to verify whether participants used the free or paid version of the tool or to identify the specific underlying language model; this lack of control may have affected the consistency of AI support among participants.
  • In Study 2, the experiment was conducted in a controlled university environment with equal access to the free version of ChatGPT; therefore, the findings are limited to this specific access level.
  • As this study focuses specifically on ChatGPT, the findings may not be directly generalizable to other large language models or AI-based programming assistants operating under different access conditions.
  • The experiment in the second study was limited to solving two specific programming tasks, which provided a more focused rather than comprehensive view of ChatGPT’s capabilities, especially when solving complex software problems.
Despite the outlined limitations, the results of the study provide valuable information and expand the existing recommendations regarding the implementation of appropriate teaching and learning methods for more effective and targeted integration of chatbots into programming education.

7. Conclusions

ChatGPT is entering all spheres of public life, including education. It offers significant advantages—quick access to information, personalized support for learners, and opportunities to automate routine tasks. At the same time, however, it also hides risks associated with superficial acquisition of knowledge, reduced creativity, and excessive dependence on technology. The teachers are aware of these aspects, but students should also recognize the benefits and drawbacks of implementing chatbots in education. This understanding cannot be achieved solely through verbal instruction; rather, it requires providing students with opportunities to use ChatGPT, followed by reflection and critical evaluation of its impact on their learning. In the context of programming education, where the potential of ChatGPT is particularly significant, such a reflective approach is of prime importance. A complete ban on the use of ChatGPT could limit the formation of skills for its effective use in future professional practice, while its uncontrolled application carries the risk of significant deterioration of programming knowledge and skills.
Therefore, educators should employ a variety of methods, demonstrating both the strengths and weaknesses of chatbots, including:
  • Providing students with the opportunity to work on projects, using both approaches—with and without ChatGPT—to make an informed comparison and determine for themselves in which situations the tool is useful and in which it is not.
  • Including tasks, structured according to the Reverse Bloom’s Taxonomy model (Pesovski et al., 2024b), starting from higher cognitive levels—for example, modifying, optimizing, and extending existing code. Tasks of this type stimulate active thinking and limit passive copying of ChatGPT-generated solutions, being at the same time consistent with the reverse engineering practices in the software industry. In this context, the chatbot can act as a consultant, assisting in clarifying concepts, detecting errors, and analyzing alternative approaches, and this method of using it can be effective even for students with limited prior experience.
We believe that by using carefully selected methods and learning tasks, programming teachers can help students develop critical, purposeful, and effective skills for working with ChatGPT. These skills can be applied not only in their education but also in their future professional careers.

Author Contributions

The authors contributed to the manuscript equally and are all equally accountable for the work. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the European Union-NextGenerationEU, through the National Recovery and Resilience Plan of the Republic of Bulgaria, project “Digital Sustainable Ecosystems–Technological Solutions and Social Models for Ecosystem Sustainability–DUECOS”, BG-RRP-2.004-0001-C01.

Institutional Review Board Statement

The study was conducted in accordance with the Declaration of Helsinki, and both the questionnaire and the questions for the semi-structured interview were approved by the ethics committee at the Faculty of Physics and Technology of Plovdiv University “Paisii Hilendarski” (protocol code 1, date of approval 8 May 2025 and protocol code 3, date of approval 9 June 2025).

Informed Consent Statement

All students were informed in advance of the study’s purpose, the use of collected data, the protection of personal data, and the voluntary nature of participation. Written informed consent was not obtained for the questionnaire; instead, completing the questionnaire was considered as providing informed consent. For the semi-structured interviews, written informed consent was considered unnecessary because no personally identifiable information was collected. Participants received clear verbal information regarding the study objectives, its voluntary nature, and data confidentiality. They were informed that their responses would be used solely for scientific and statistical purposes and that the data could not be linked to them as individuals.

Data Availability Statement

Data available on request due to restrictions (privacy or ethical reasons).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Adel, A., Ahsan, A., & Davison, C. (2024). ChatGPT promises and challenges in education: Computational and ethical perspectives. Education Sciences, 14(8), 814. [Google Scholar] [CrossRef] [Scilit]
  2. Almanasra, S., & Suwais, K. (2025). Analysis of ChatGPT-generated codes across multiple programming languages. IEEE Access, 13, 23580–23596. [Google Scholar] [CrossRef] [Scilit]
  3. Anagnostopoulos, C.-N. (2023). ChatGPT impacts in programming education: A recent literature overview that debates ChatGPT responses. arXiv, arXiv:2309.12348. [Google Scholar] [CrossRef] [Scilit]
  4. Biswas, S. (2023). Role of ChatGPT in computer programming. Mesopotamian Journal of Computer Science, 2023, 9–15. [Google Scholar] [CrossRef] [Scilit]
  5. Bringula, R. (2024). ChatGPT in a programming course: Benefits and limitations. Frontiers in Education, 9, 1248705. [Google Scholar] [CrossRef] [Scilit]
  6. Bucaioni, A., Ekedahl, H., Helander, V., & Nguyen, P. T. (2024). Programming with ChatGPT: How far can we go? Machine Learning with Applications, 15, 100526. [Google Scholar] [CrossRef] [Scilit]
  7. Campillo-Ferrer, J. M., López-García, A., & Miralles-Sánchez, P. (2025). Student perceptions of the use of gen-AI in a higher education program in Spain. Digital, 5(3), 29. [Google Scholar] [CrossRef] [Scilit]
  8. Chang, C. I., Choi, W. C., & Choi, I. C. (2024). Challenges and limitations of using artificial intelligence generated content (AIGC) with ChatGPT in programming curriculum: A systematic literature review. In Proceedings of the 2024 7th artificial intelligence and cloud computing conference (pp. 372–378). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  9. Chen, E., Huang, R., Chen, H.-S., Tseng, Y.-H., & Li, L.-Y. (2023). GPTutor: A ChatGPT-powered programming tool for code explanation. In N. Wang, G. Rebolledo-Mendez, V. Dimitrova, N. Matsuda, & O. C. Santos (Eds.), Artificial intelligence in education. Posters and late breaking results, workshops and tutorials, industry and innovation tracks, practitioners, doctoral consortium and blue sky (Vol. 1831, pp. 321–327). Springer Nature Switzerland. [Google Scholar] [CrossRef] [Scilit]
  10. Clarke, C. J. S. F., & Konak, A. (2025). The impact of AI use in programming courses on critical thinking skills. Journal of Cybersecurity Education, Research and Practice, 2025(1), 5. [Google Scholar] [CrossRef] [Scilit]
  11. Farah, J. C., La Scala, J., Ingram, S., & Gillet, D. (2025, April 27). Supporting brainstorming activities with bots in software engineering education. 2025 IEEE/ACM International Workshop on Bots in Software Engineering (BotSE) (pp. 23–27), Ottawa, ON, Canada. [Google Scholar] [CrossRef] [Scilit]
  12. Gill, S. S., Xu, M., Patros, P., Wu, H., Kaur, R., Kaur, K., Fuller, S., Singh, M., Arora, P., Parlikad, A. K., Stankovski, V., Abraham, A., Ghosh, S. K., Lutfiyya, H., Kanhere, S. S., Bahsoon, R., Rana, O., Dustdar, S., Sakellariou, R., … Buyya, R. (2024). Transformative effects of ChatGPT on modern education: Emerging era of AI chatbots. Internet of Things and Cyber-Physical Systems, 4, 19–23. [Google Scholar] [CrossRef] [Scilit]
  13. Groothuijsen, S., Van Den Beemt, A., Remmers, J. C., & Van Meeuwen, L. W. (2024). AI chatbots in programming education: Students’ use in a scientific computing course and consequences for learning. Computers and Education: Artificial Intelligence, 7, 100290. [Google Scholar] [CrossRef] [Scilit]
  14. Haindl, P., & Weinberger, G. (2024a). Does ChatGPT help novice programmers write better code? Results from static code analysis. IEEE Access, 12, 114146–114156. [Google Scholar] [CrossRef] [Scilit]
  15. Haindl, P., & Weinberger, G. (2024b). Students’ experiences of using ChatGPT in an undergraduate programming course. IEEE Access, 12, 43519–43529. [Google Scholar] [CrossRef] [Scilit]
  16. Haque, M. A., & Li, S. (2023). The potential use of ChatGPT for debugging and bug fixing. EAI Endorsed Transactions on AI and Robotics, 2(1), e4. [Google Scholar] [CrossRef] [Scilit]
  17. Hartley, K., Hayak, M., & Ko, U. H. (2024). Artificial intelligence supporting independent student learning: An evaluative case study of ChatGPT and learning to code. Education Sciences, 14(2), 120. [Google Scholar] [CrossRef] [Scilit]
  18. Hazelhurst, S., Johnson, Y., & Sanders, I. (2011). An empirical analysis of the relationship between web usage and academic performance in undergraduate students. arXiv, arXiv:1110.6267. [Google Scholar] [CrossRef] [Scilit]
  19. Huang, D., Zhang, J. M., Bu, Q., Xie, X., Chen, J., & Cui, H. (2023). Bias testing and mitigation in LLM-based code generation. arXiv, arXiv:2309.14345. [Google Scholar] [CrossRef] [Scilit]
  20. Huesca, G., Martínez-Treviño, Y., Molina-Espinosa, J. M., Sanromán-Calleros, A. R., Martínez-Román, R., Cendejas-Castro, E. A., & Bustos, R. (2024). Effectiveness of using ChatGPT as a tool to strengthen benefits of the flipped learning strategy. Education Sciences, 14(6), 660. [Google Scholar] [CrossRef] [Scilit]
  21. Humble, N., Boustedt, J., Holmgren, H., Milutinovic, G., Seipel, S., & Östberg, A.-S. (2023). Cheaters or AI-enhanced learners: Consequences of ChatGPT for programming education. Electronic Journal of e-Learning, 22, 16–29. [Google Scholar] [CrossRef] [Scilit]
  22. Husain, A. (2024). Potentials of ChatGPT in computer programming: Insights from programming instructors. Journal of Information Technology Education: Research, 23, 002. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Ilieva, G., Yankova, T., Klisarova-Belcheva, S., Dimitrov, A., Bratkov, M., & Angelov, D. (2023). Effects of Generative chatbots in higher education. Information, 14(9), 492. [Google Scholar] [CrossRef] [Scilit]
  24. Jamil, M. T., Abid, S., & Shamail, S. (2025, April 28–29). Can LLMs generate higher quality code than humans? An empirical study. 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR) (pp. 478–489), Ottawa, ON, Canada. [Google Scholar] [CrossRef] [Scilit]
  25. Jukiewicz, M. (2024). The future of grading programming assignments in education: The role of ChatGPT in automating the assessment and feedback process. Thinking Skills and Creativity, 52, 101522. [Google Scholar] [CrossRef] [Scilit]
  26. Kiesler, N., Lohr, D., & Keuning, H. (2023, October 18–21). Exploring the potential of large language models to generate formative programming feedback. 2023 IEEE Frontiers in Education Conference (FIE) (pp. 1–5), College Station, TX, USA. [Google Scholar] [CrossRef] [Scilit]
  27. Kosar, T., Ostojić, D., Liu, Y. D., & Mernik, M. (2024). Computer science education in ChatGPT era: Experiences from an experiment in a programming course for novice programmers. Mathematics, 12(5), 629. [Google Scholar] [CrossRef] [Scilit]
  28. Lehmann, M., Cornelius, P. B., & Sting, F. J. (2025). AI meets the classroom: When do large language models harm learning? arXiv, arXiv:2409.09047. [Google Scholar] [CrossRef] [Scilit]
  29. León-Domínguez, U. (2024). Potential cognitive risks of generative transformer-based AI chatbots on higher order executive functions. Neuropsychology, 38(4), 293–308. [Google Scholar] [CrossRef] [Scilit]
  30. Lin, C.-C., Huang, A. Y. Q., & Yang, S. J. H. (2023). A review of AI-driven conversational chatbots implementation methodologies and challenges (1999–2022). Sustainability, 15(5), 4012. [Google Scholar] [CrossRef] [Scilit]
  31. Liu, Q., Zhou, Y., Huang, J., & Li, G. (2024). When ChatGPT is gone: Creativity reverts and homogeneity persists. arXiv, arXiv:2401.06816. [Google Scholar] [CrossRef] [Scilit]
  32. López-Fernández, D., & Vergaz, R. (2025). ChatGPT in computer science education: A case study on a database administration course. Applied Sciences, 15(2), 985. [Google Scholar] [CrossRef] [Scilit]
  33. McCulloh, I., Rodriguez, P., Kumar, S., Gupta, M., Sharma, V. R., Johnson, B., & Johnson, A. N. (2025). Generative AI in computer science education: Accelerating python learning with ChatGPT. arXiv, arXiv:2505.20329. [Google Scholar] [CrossRef] [Scilit]
  34. Ouh, E. L., Gan, B. K. S., Shim, K. J., & Wlodkowski, S. (2023). ChatGPT, can you generate solutions for my coding exercises? An evaluation on its effectiveness in an undergraduate java programming course. arXiv, arXiv:2305.13680. [Google Scholar] [CrossRef] [Scilit]
  35. Penney, J., Acharya, P., Hilbert, P., Parekh, P., Sarma, A., Steinmacher, I., & Gerosa, M. A. (2025). Outcomes, perceptions, and interaction strategies of novice programmers studying with ChatGPT. In Proceedings of the 7th ACM conference on conversational user interfaces (pp. 1–15). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  36. Pesovski, I., Santos, R., Henriques, R., & Trajkovik, V. (2024a). Generative AI for customizable learning experiences. Sustainability, 16(7), 3034. [Google Scholar] [CrossRef] [Scilit]
  37. Pesovski, I., Vorkel, D., & Trajkovik, V. (2024b, November 6–8). AI-powered education: Rethinking the way programming is taught using AI tools and reversed bloom’s taxonomy. 2024 21st International Conference on Information Technology Based Higher Education and Training (ITHET) (pp. 1–6), Paris, France. [Google Scholar] [CrossRef] [Scilit]
  38. Rahe, C., & Maalej, W. (2025). How do programming students use generative AI? Proceedings of the ACM on Software Engineering, 2, 978–1000. [Google Scholar] [CrossRef] [Scilit]
  39. Rahman, M. M., & Watanobe, Y. (2023). ChatGPT for education and research: Opportunities, threats, and strategies. Applied Sciences, 13(9), 5783. [Google Scholar] [CrossRef] [Scilit]
  40. Raihan, N., Siddiq, M. L., Santos, J. C. S., & Zampieri, M. (2025). Large language models in computer science education: A systematic literature review. In Proceedings of the 56th ACM technical symposium on computer science education V. 1 (pp. 938–944). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  41. Rajan, S., & Niranjan, L. R. (2025). The double-edged sword of ChatGPT: Fostering and hindering creativity in postgraduate academics in Bengaluru. International Journal of Educational Management, 39(2), 317–337. [Google Scholar] [CrossRef] [Scilit]
  42. Ray, P. P. (2023). ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and Cyber-Physical Systems, 3, 121–154. [Google Scholar] [CrossRef] [Scilit]
  43. Sánchez-Ruiz, L. M., Moll-López, S., Nuñez-Pérez, A., Moraño-Fernández, J. A., & Vega-Fleitas, E. (2023). ChatGPT challenges blended learning methodologies in engineering education: A case study in mathematics. Applied Sciences, 13(10), 6039. [Google Scholar] [CrossRef] [Scilit]
  44. Silva, C. A. G. D., Ramos, F. N., De Moraes, R. V., & Santos, E. L. D. (2024). ChatGPT: Challenges and benefits in software programming for higher education. Sustainability, 16(3), 1245. [Google Scholar] [CrossRef] [Scilit]
  45. Stoyanova, D., Stoyanova-Petrova, S., & Mileva, N. (2025). Exploring students’ and teachers’ perceptions about using ChatGPT in programming education. International Journal of Engineering Pedagogy (iJEP), 15(2), 15–41. [Google Scholar] [CrossRef] [Scilit]
  46. Sun, D., Boudouaia, A., Yang, J., & Xu, J. (2024a). Investigating students’ programming behaviors, interaction qualities and perceptions through prompt-based learning in ChatGPT. Humanities and Social Sciences Communications, 11(1), 1447. [Google Scholar] [CrossRef] [Scilit]
  47. Sun, D., Boudouaia, A., Zhu, C., & Li, Y. (2024b). Would ChatGPT-facilitated programming mode impact college students’ programming behaviors, performances, and perceptions? An empirical study. International Journal of Educational Technology in Higher Education, 21(1), 14. [Google Scholar] [CrossRef] [Scilit]
  48. Surameery, N. M. S., & Shakor, M. Y. (2023). Use chat GPT to solve programming bugs. International Journal of Information Technology and Computer Engineering, 31, 17–22. [Google Scholar] [CrossRef] [Scilit]
  49. Uandykova, M., Baitenova, L., Mukhamejanova, G., Yeleukulova, A., & Mirkassimova, T. (2024). Java coding using artificial intelligence. Frontiers in Computer Science, 6, 1473870. [Google Scholar] [CrossRef] [Scilit]
  50. Vaithilingam, P., Zhang, T., & Glassman, E. L. (2022). Expectation vs. experience: Evaluating the usability of code generation tools powered by large language models. In CHI conference on human factors in computing systems extended abstracts (pp. 1–7). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  51. Wu, T.-T., Lee, H.-Y., Li, P.-H., Huang, C.-N., & Huang, Y.-M. (2024). Promoting self-regulation progress and knowledge construction in blended learning via ChatGPT-based learning aid. Journal of Educational Computing Research, 61(8), 1539–1567. [Google Scholar] [CrossRef] [Scilit]
  52. Xue, Y., Chen, H., Bai, G. R., Tairas, R., & Huang, Y. (2024). Does ChatGPT help with introductory programming? An experiment of students using ChatGPT in CS1. In Proceedings of the 46th international conference on software engineering: Software engineering education and training (pp. 331–341). Association for Computing Machinery. [Google Scholar] [CrossRef] [Scilit]
  53. Yang, Q. (2024). Systematic evaluation of AI-generated python code: A comparative study across progressive programming tasks. Research Square. [Google Scholar] [CrossRef] [Scilit]
  54. Yang, T.-C., Hsu, Y.-C., & Wu, J.-Y. (2025). The effectiveness of ChatGPT in assisting high school students in programming learning: Evidence from a quasi-experimental research. Interactive Learning Environments, 33(6), 3726–3743. [Google Scholar] [CrossRef] [Scilit]
  55. Yilmaz, R., & Karaoglan Yilmaz, F. G. (2023). Augmented intelligence in programming learning: Examining student views on the use of ChatGPT for programming learning. Computers in Human Behavior: Artificial Humans, 1(2), 100005. [Google Scholar] [CrossRef] [Scilit]
  56. Zviel-Girshin, R. (2024). The good and bad of AI tools in novice programming education. Education Sciences, 14(10), 1089. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Experimental design of Study 2.
Figure 1. Experimental design of Study 2.
Education 16 00019 g001
Figure 2. Boxplot of the students’ assessment grades in the studied discipline by clusters.
Figure 2. Boxplot of the students’ assessment grades in the studied discipline by clusters.
Education 16 00019 g002
Table 1. Distribution of the given answers to the question “Q1.Determine the frequency of using ChatGPT as an aid in completing the course assignments”.
Table 1. Distribution of the given answers to the question “Q1.Determine the frequency of using ChatGPT as an aid in completing the course assignments”.
FrequencyPercent
Never00
Rarely712.5
Sometimes2850.0
Often1832.1
Very Often35.4
Total56100.0
Table 2. Mean values for questions Q2 to Q7 for each cluster.
Table 2. Mean values for questions Q2 to Q7 for each cluster.
Cluster 1
Actively Using ChatGPT
Cluster 2
Rarely Using ChatGPT
Q24.003.00
Q33.562.38
Q42.972.17
Q54.222.08
Q63.311.71
Q73.692.25
Table 3. Mean values of the scores for the projects, developed by the two groups, following the two approaches—with and without ChatGPT.
Table 3. Mean values of the scores for the projects, developed by the two groups, following the two approaches—with and without ChatGPT.
CategoryScores Group 1Scores Group 2
With ChatGPTWithout ChatGPTWith ChatGPTWithout ChatGPT
Completeness and effectiveness of the code4.53.554
Organization and readability of the code 44.544.5
Interface quality3.7543.754.25
Creativity3545
Time for completion3 h 15 min4 h 10 min3 h 30 min4 h 15 min
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Stoyanova, D.; Stoyanova-Petrova, S.; Shotarova, S.; Lyubomirov, S.; Mileva, N. ChatGPT in Programming Education: An Empirical Study on Its Impact on Student Performance, Creativity, and Teamwork. Educ. Sci. 2026, 16, 19. https://doi.org/10.3390/educsci16010019

AMA Style

Stoyanova D, Stoyanova-Petrova S, Shotarova S, Lyubomirov S, Mileva N. ChatGPT in Programming Education: An Empirical Study on Its Impact on Student Performance, Creativity, and Teamwork. Education Sciences. 2026; 16(1):19. https://doi.org/10.3390/educsci16010019

Chicago/Turabian Style

Stoyanova, Diana, Silviya Stoyanova-Petrova, Snezha Shotarova, Slavi Lyubomirov, and Nevena Mileva. 2026. "ChatGPT in Programming Education: An Empirical Study on Its Impact on Student Performance, Creativity, and Teamwork" Education Sciences 16, no. 1: 19. https://doi.org/10.3390/educsci16010019

APA Style

Stoyanova, D., Stoyanova-Petrova, S., Shotarova, S., Lyubomirov, S., & Mileva, N. (2026). ChatGPT in Programming Education: An Empirical Study on Its Impact on Student Performance, Creativity, and Teamwork. Education Sciences, 16(1), 19. https://doi.org/10.3390/educsci16010019

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop