Next Article in Journal
A Machine Learning-Driven CRM Approach for Identifying Member Churn in a Brazilian Agro-Industrial Cooperative: A Practical Case Study
Next Article in Special Issue
QGKM: A Quantum Fidelity-Based Graph Clustering Framework for Robust Data Pattern Recognition in Education Social Networks
Previous Article in Journal
AI-Driven IFC Processing for Automated IBS Scoring
Previous Article in Special Issue
Gamification in Education and Its Impact on Student Academic Performance: A Conceptual Model Based on Systematic Literature Review and PLS-SEM Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

A Systematic Literature Review on the Pedagogical Implications and Impact of GenAI on Students’ Critical Thinking

1
Department of Multidisciplinary Engineering, Texas A&M University, College Station, TX 77840, USA
2
School of Teacher Education and Leadership, Utah State University, Logan, UT 84322, USA
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(3), 179; https://doi.org/10.3390/a19030179
Submission received: 21 January 2026 / Revised: 20 February 2026 / Accepted: 24 February 2026 / Published: 27 February 2026
(This article belongs to the Special Issue Artificial Intelligence in Education: Innovations and Implications)

Abstract

Critical Thinking (CT) is recognized as a foundational competency for professional readiness, innovation, and ethical reasoning in higher education, enabling students to analyze information, evaluate evidence, and make reasoned decisions in complex environments. The rapid integration of Generative Artificial Intelligence (GenAI) tools, such as large language models, presents new opportunities and risks for CT development. This study conducts a systematic literature review to synthesize empirical evidence on the pedagogical implications and cognitive impact of GenAI on students’ CT. Following PRISMA guidelines, and search terms around GenAI Tools, Critical Thinking And Higher Education, on five major education research databases—Web of Science; Scopus; EBSCOhost (Education Source, ERIC, and APA PsycInfo); and Compendex and Inspec (Elsevier)—63 empirical studies published between January 2023 and April 2025 were analyzed across higher education contexts, disciplines, and intervention designs. Results indicate that GenAI offers notable cognitive affordances, including scaffolding reflective reasoning, promoting self-regulation, and facilitating iterative dialogue and argument evaluation. Pedagogical strategies clustered into four primary integration typologies: AI-based feedback prompts, dialogue simulation and reflection, AI-supported peer review, and critical engagement with AI-generated content. Nearly half of the studies reported statistically significant CT improvements, particularly when GenAI use was guided by structured prompts, reflective activities, and performance-based assessment. However, multiple risks persist, including cognitive offloading, uncritical acceptance of AI outputs, and diminished intellectual autonomy, especially in unguided or surface-level usage. This review highlights the need for intentional pedagogical design, validated CT assessment tools, and longitudinal studies to ensure GenAI acts as a catalyst rather than a substitute for human reasoning. By identifying effective integration strategies and outlining potential pitfalls, this study provides evidence-informed guidance for educators and institutions aiming to responsibly leverage GenAI to strengthen students’ CT skills.

1. Introduction

Critical Thinking (CT) is widely recognized as an essential skill for professional readiness, innovation, and ethical decision-making in higher education [1,2]. CT enables learners to navigate complex and uncertain environments by interpreting information accurately, analyzing arguments rigorously, evaluating evidence objectively, inferring sound conclusions, explaining reasoning clearly, and self-regulating their cognitive processes [2,3]. Despite its importance and emphasis in curricular frameworks, a persistent gap remains between the goal of cultivating CT and the pedagogical strategies typically used in higher education. Furthermore, current instructional designs continue to emphasize content delivery and rote memorization, rather than the deep, metacognitive engagement that drives genuine CT development [4,5].
Since the release of ChatGPT 3.5 in late 2022, Generative Artificial Intelligence (GenAI) tools have proliferated across academic settings, offering students scaffolds for reflective reasoning, iterative dialogue, and real-time feedback [6] raising questions about how these tools influence the cultivation of cognitive skills and students’ cognitive engagement and reasoning strategies [7]. Indeed, while GenAI (e.g., ChatGPT Copilot) offers real-time feedback, ideation support, and adaptive scaffolding [8,9], it also introduces risks of cognitive offloading, uncritical acceptance of AI outputs, and over-reliance on machine-generated responses [10,11,12].
Recent studies have identified the potential of GenAI tools to activate CT through cognitive scaffolding and metacognitive prompting [11,13,14]. GenAI’s technical affordances can map onto specific core CT processes. For instance, real-time, dialogic interaction can enable iterative questioning, clarification, and “why/how” probing that can scaffold analysis and explanation [8]. Zhou et al. [15] found that GenAI supports metacognitive self-regulation by encouraging students to reflect on their reasoning process or explore multiple perspectives in problem-solving, providing support in evaluation and inference by making it easier for learners to compare competing claims, examine trade-offs, and evidence-based argumentation tasks. Also, GenAI’s ability to provide adaptive feedback and on-demand hints can support self-regulation by helping students plan next steps, monitor progress, and reflect on errors, especially when the interaction is designed to surface reasoning rather than optimize for speed [8,9].
Despite the promising affordances of GenAI for enhancing CT, its integration into higher education is not without challenges, introducing a range of pedagogical and cognitive risks to the academic environment. As these tools become more common, concerns have emerged regarding students’ overreliance on AI-generated outputs, potential cognitive offloading, and the erosion of independent reasoning [11]. Automated generation can create an illusion of understanding when students mistake coherence for correctness, while the convenience of instant answers can reduce verification behaviors unless tasks explicitly require source evaluation, triangulation, and reasoning audits [10]. These risks are particularly pressing in higher education, where CT is not merely a desirable attribute but a foundational skill for professional practice and innovation [1,16]. Without thoughtful integration and guided use, GenAI tools can inadvertently undermine the very skills they are intended to support, such as self-regulation, deep reasoning, and independent judgment [17,18]. In addition, variability in output quality, incomplete attribution, and occasional fabrication amplify the importance of instructional guardrails that foreground evaluation standards (e.g., credibility checks, requirement tracing, and error testing) and make reasoning visible through prompts that demand justification, comparison, and reflection.
Despite the rapid adoption of GenAI tools in educational contexts, empirical research remains limited and conflicted regarding their impact on the development and demonstration of critical thinking skills [19]. The existing literature tends to focus on perceptions of AI usefulness, general attitudes, or performance outcomes, often neglecting how students’ thinking processes evolve when working alongside AI systems [11,20]. Furthermore, little is known about how instructional design mediates GenAI’s impact, especially when comparing guided and unguided use. To address this gap, the present study conducts a SLR guided by the PRISMA framework [21], synthesizing peer-reviewed research published between January 2023 and April 2025. This time frame captures the critical post-ChatGPT period, in which GenAI adoption in academic contexts has rapidly accelerated [6].
This review is guided by the following research questions:
  • What patterns and findings emerge from the recent empirical literature regarding the impact of GenAI tools on the development of students’ critical thinking in higher education?
  • What are the primary cognitive affordances and pedagogical risks associated with GenAI use in CT development, and what methodological or theoretical gaps persist in current research?
  • What implications can be drawn from this synthesis to inform responsible and evidence-based redesign of learning environments, ensuring that GenAI complements, rather than substitutes, human critical thinking?
By systematically analyzing how GenAI intersects with CT-related outcomes, such as self-regulation, reflection, metacognition, ethical reasoning, analysis, evaluation and inference, this review contributes to ongoing efforts to reimagine the role of AI in higher education. It aims to provide theoretical grounding and practical guidance for educators, curriculum designers, and institutional leaders seeking to harness the promise of GenAI while safeguarding the intellectual agency and cognitive rigor essential to CT.

2. Conceptual Frameworks for Critical Thinking

CT is generally understood as a deliberate and reflective process that involves the evaluation of evidence and information, logical reasoning, and informed decision-making [3]. Facione [2] defines critical thinking as a “purposeful, self-regulatory judgment which results in interpretation, analysis, evaluation, and inference, as well as an explanation of the evidential, conceptual, methodological, criteriological, or contextual considerations upon which judgment is based”. This widely adopted conceptualization provides a comprehensive framework for understanding the cognitive dimensions of CT and serves as a foundation for its application in engineering contexts [1].
Facione’s Delphi study [2] brought together 46 experts from various fields to reach a consensus on the essential cognitive skills associated with CT. The study identified six core skills: (1) Interpretation, (2) Analysis, (3) Evaluation, (4) Inference, (5) Explanation, and (6) Self-Regulation, each encompassing a set of subskills [3]. As shown in Figure 1, these subskills provide a clear and structured model for assessing and developing CT skills in students. This comprehensive framework has been extensively applied in higher education and continues to inform curriculum development and assessment practices [3,22].
In Facione’s landmark Delphi study [2], the six core CT skills together form the backbone of disciplined reasoning in any domain. Interpretation involves accurately comprehending the meaning and significance of data or arguments, while analysis requires breaking down complex ideas into their constituent parts and examining underlying assumptions. Evaluation then assesses the credibility of sources and the validity of reasoning against explicit criteria, and inference entails drawing well-supported conclusions or hypotheses from available evidence. Explanation focuses on articulating one’s reasoning clearly, citing appropriate evidence and justifications, and self-regulation demands ongoing reflection and adjustment of one’s cognitive strategies [2,23].
Complementing Facione’s skill-based model, Paul and Elder [24] offer a metacognitive framework that emphasizes the standards and thought structures guiding critical reasoning. Their model identifies eight “elements of thought”, including purpose, question at issue, information, inferences, concepts, assumptions, implications, and point of view, and nine “intellectual standards”, such as clarity, accuracy, depth, and fairness, which serve as quality benchmarks for reasoning. By applying these standards to each element of thought, learners learn not only what they should think but how they should evaluate the quality of their thinking. This dual focus on process and quality transforms CT from a set of isolated skills into a coherent, self-regulated activity whereby students consciously monitor and refine their reasoning at every step.
In addition to the previously mentioned skills, CT also depends on the dispositions that motivate sustained engagement with complex problems and ethical dilemmas. Ennis [25] highlighted the importance of traits such as truth-seeking, open-mindedness, intellectual humility, and perseverance, arguing that without these dispositions, learners may fail to apply their cognitive skills when faced with ambiguity or challenge. Also, cultivating a willingness to reconsider viewpoints in light of new evidence [2,24]. In practice, fostering these dispositions entails creating learning environments in which students are encouraged to question their assumptions, tolerate uncertainty, and persist through iterative cycles of inquiry and revision. In the context of GenAI-mediated learning, such dispositions become even more crucial; as students interact with AI-generated suggestions, they must adopt a healthy skepticism, remain willing to challenge both their own ideas and those offered by the system, and persist in seeking deeper understanding rather than accepting surface-level responses [26].
An equally foundational framework for conceptualizing CT is Bloom’s Taxonomy, first articulated by Benjamin Bloom in 1956, which classifies cognitive processes into hierarchical levels: Knowledge, Comprehension, Application, Analysis, Synthesis, and Evaluation. Anderson and Krathwohl’s 2001 revision reconceptualized these levels as Remember, Understand, Apply, Analyze, Evaluate, and Create, reframing “Synthesis” as “Create” to reflect 21st-century educational demands [27]. The Revised Bloom’s Taxonomy explicitly links lower-order thinking skills (e.g., Remember, Understand) with higher-order skills (Analyze, Evaluate, Create), providing a practical scaffold for designing assessments and learning activities that sequentially develop CT.
Foundational frameworks remain essential to understanding how GenAI may support, or hinder, CT development. When applied to GenAI-enhanced learning environments, these frameworks provide a lens to examine how students interact with, question, and critique AI-generated outputs. Also, as Deng et al. [10] argue, the inclusion of a specific theoretical lens in systematic reviews is essential to make sense of the diverse ways GenAI interacts with learning-related outcomes.

3. Background

3.1. Impact of GenAI on Critical Thinking

GenAI systems, particularly large language models (LLMs) such as ChatGPT, offer a variety of cognitive affordances that can be leveraged to enhance CT in higher education [8]. These tools enable students to engage in iterative dialogue, receive real-time feedback, and explore alternative viewpoints, features aligned with key CT processes such as analysis, synthesis, inference, and evaluation [28]. When intentionally integrated, GenAI can actively scaffold reasoning and metacognitive engagement [29].
A notable affordance is GenAI’s ability to reformulate student ideas and present them in different forms, allowing learners to refine their arguments and clarify understanding [8,30]. These interactions promote intellectual flexibility and encourage iterative refinement of thought. GenAI also supports reflective thinking by prompting students to question their reasoning and justify decisions [31]. Scenario-based simulations and classroom activities using ChatGPT have been shown to foster critical, creative, and meditative thinking [29,32].
While GenAI tools offer significant potential to support CT, their integration into higher education also introduces important pedagogical and cognitive risks. Cognitive offloading may occur when students rely on GenAI to bypass complex reasoning tasks [11,18]. This can hinder CT development and reinforce dependency on external cognitive agents [33]. Additionally, the illusion of understanding arises when confident AI responses are misinterpreted as correct despite containing errors [10]. Uncritical acceptance of AI outputs may inhibit self-regulation and weaken higher-order reasoning [17,18,34]. Studies have shown that when students are not explicitly encouraged to reflect on their own reasoning or compare their answers with AI-generated suggestions, they are less likely to revise their thinking or explore alternate solutions [13].

3.2. Previous Systematic Literature Reviews

A growing number of systematic and scoping reviews have begun to map the different applications of GenAI in higher education [10,35,36,37,38,39], yet its impact on CT remains underexplored [11]. A few reviews have focused on CT as a primary outcome [19,40]. Melisa et al. [19] conducted a focused analysis of ChatGPT’s effects on CT skills, drawing from Scopus and ERIC databases for studies published between January 2023 and 2024. Their review, structured around a four-stage PRISMA protocol and culminating in nineteen empirical higher-education studies, highlighted ChatGPT’s potential to enhance evaluative reasoning and argumentation but underscored a lack of validated, GenAI-specific instruments for measuring CT and a shortage of longitudinal evidence on sustained cognitive gains. However, their search strategy appeared limited both in the number of databases and in the specificity of search terms, potentially omitting relevant studies.
Similarly, Sardi et al. [40] explored both self-regulated learning (SRL) and CT across thirty-eight studies sourced from Scopus, Web of Science, and ScienceDirect spanning 2022–2024. While their synthesis detailed GenAI’s role in scaffolding metacognitive strategies and providing adaptive feedback, it primarily treated CT as a secondary outcome linked to SRL metrics. They noted persistent risks of overreliance on AI, cultural and digital-literacy variability, and the absence of concrete models for integrating human and AI regulation but stopped short of embedding these insights within an established CT theory. A methodological limitation of this review is its reliance on only three databases, which increases the risk of omitting relevant studies from broader interdisciplinary venues such as ERIC, PsycINFO, or IEEE Xplore. In addition, their search strategy, although transparent, relied on generic terms such as “AI AND Critical Thinking AND Education”, which might have missed studies using alternative terminology in each of the 3 dimensions (GenAI Tools, Critical Thinking, and Educational Context), and thereby underrepresents the CT literature. Both reviews offer valuable insights into GenAI’s pedagogical affordances and challenges; however, neither systematically applies a robust CT theoretical lens to drive analysis.
Other reviews have explored GenAI’s disciplinary impacts without centering CT outcomes. Ilic et al. [36] focused on STEM education, observing that AI tools enhance fluency and writing efficiency but may inadvertently encourage cognitive offloading when students rely on AI-generated solutions rather than engaging in deep reasoning. Celik et al. [8] conducted a systematic review of AI-based tools for 21st-century skills in higher education, highlighting that while GenAI platforms can scaffold collaborative problem-solving and iterative reasoning, the majority of studies reported only on engagement and usability metrics rather than direct measures of CT (e.g., analysis, evaluation, inference). Similarly, Deng et al. [10] synthesized experimental investigations of ChatGPT in learning contexts, finding small to moderate effect sizes on knowledge acquisition but noting that few studies incorporated validated CT assessments or disentangled metacognitive gains from surface-level performance improvements.
Finally, a subset of reviews has interrogated the risks of GenAI use, particularly overreliance and the illusion of understanding. Zhai, Wibowo, and Li [18] performed a systematic review of AI dialogue systems, showing that students often accept confidently worded AI outputs uncritically and that few studies measure whether such interactions translate into improved or diminished CT capabilities. Gerlich [11] likewise emphasized the threat of cognitive offloading, warning that unguided use of GenAI can erode independent reasoning and undercut the development of metacognitive regulation.

4. Methodology

This SLR followed the PRISMA guidelines [21], including the steps of abstract screening, full-text screening, and data extraction. This review was registered on Open Science Framework (OSF) under the number: en3p5.

4.1. Search Strategy

Building on the terms identified in previous literature reviews, we searched for terms around the intersection of three dimensions: Generative Artificial Intelligence, Educational Context, and Critical Thinking. For example, in the Generative Artificial Intelligence dimension, we used the same terms used by [10]. For the Educational Context dimension, we used terms recommended in [41]. Table 1 presents the search terms used.
These terms were applied to the search engines of five major education research databases: Web of Science; Scopus; EBSCOhost (Education Source, ERIC, and APA PsycInfo); and Compendex & Inspec (Elsevier). This multi-database search ensured appropriate coverage and scope for the screening process.
Two automatic filters were applied. First, previous research has identified a rapid increase in GenAI-related studies in education since 2023, coinciding with the release of ChatGPT in late 2022 [6]. Thus, we limited the search to studies published from January 2023 to April 2025. Second, due to the language limitations of our research team, we restricted the search to English-language publications.
As a result, we identified 1788 studies across the five databases, with Scopus (570) and EBSCO (521) providing the largest number of studies, shown in Table 2. To facilitate data management, we uploaded all of the studies into Covidence, a tool that assists with management, screening, and data extraction for literature reviews [42]. During this process, Covidence automatically identified 800 duplicate studies, leaving 988 studies for the abstract screening phase.

4.2. Screening

The first and second authors met to develop the criteria for inclusion and exclusion. We agreed on four main criteria: (a) higher education settings (e.g., undergraduate or graduate students); (b) empirical studies; (c) articles written in English; and (d) studies with measurable outcomes or observations related to critical thinking skills, including dimensions such as analysis, evaluation, inference, metacognition, ethical reasoning, or self-regulation.
Although the first author led the screening process, we tested inter-rater reliability to improve the reliability and validity of the process. A trained educational technology researcher (second author) acted as the second coder. After drafting the initial criteria, we met to jointly apply them to a set of publications. We discussed several papers, shared our thoughts on inclusion, and refined the criteria by adding detailed explanations.
Once confident with the data and rubric, we selected a random sample of 20 articles and independently screened them. We obtained 92% agreement between coders. Following this, the first author screened the remaining studies. Table 3 provides more detailed descriptions of the final inclusion and exclusion criteria.
Applying the abstract screening criteria, 170 studies met the requirements and moved to the full-text screening phase. We replicated the same strategy, beginning with a joint review of several full papers. We then independently coded a sample of 20 papers, achieving an initial 86% agreement. Although this is within the acceptable range reported in the literature (REF), we decided to refine the rubric further to improve consistency.
Subsequently, the first and second authors independently coded half of the papers, achieving a final inter-rater agreement of 91%, with 100% agreement on the final 30 papers. The first author then proceeded to screen the remaining papers independently. Disagreements were discussed and resolved collaboratively.
A total of 170 studies were included. The most common reasons for exclusion were: no measurement or analysis of CT (52 papers) and not an empirical study (18 papers). As a result, 63 studies were moved to the data extraction phase. Figure 2 shows the PRISMA diagram.

4.3. Analysis Strategy

To interpret and synthesize findings, we employed a structured extraction and categorization process designed to address the research questions and capture both descriptive and interpretive elements of the included studies. Data extraction was conducted in Covidence using a structured form aligned with predefined inclusion and exclusion criteria. The extracted variables included bibliographic details, study context, participant characteristics, methodological design, GenAI tools used, integration strategies, CT measurement instruments, and reported CT outcomes.
The analysis followed a hybrid inductive-deductive approach. An inductive lens was applied to identify recurring patterns and trends emerging from the extracted data, particularly those related to GenAI integration strategies, pedagogical roles, and students’ patterns of interaction with AI systems. This was essential given the diversity of educational contexts, disciplines, and tool implementations represented in the sample. In parallel, a deductive lens was used to interpret and consolidate these patterns, drawing on established CT models and the prior literature on GenAI integration in higher education. This dual approach allowed us to capture both the emergent nature of the findings and their alignment with existing theoretical and empirical work.
To support the data extraction prior to analysis, we employed a semi-automated protocol that leveraged LLM-assisted extraction under strict human oversight. Structured prompts were used to extract relevant variables (e.g., intervention design, study design, CT constructs, CT instruments). For each included study, the principal reviewer conducted an initial manual and independent extraction. The extracted fields were then compared item-by-item against the LLM output; when discrepancies emerged, the reviewer returned to the full text to corroborate, correct, or complement the extracted information before entering the final record in Covidence. Prompts were constrained to request only information explicitly stated in the source (with “not reported” permitted as a valid output), and all final categorization and interpretive decisions remained with the research team. This process followed recent recommendations for AI-assisted synthesis workflows [43,44], ensuring transparency and consistency while maintaining full human control over final categorization and interpretation.
All studies were analyzed for their methodological rigor, theoretical framing, type and duration of intervention, CT constructs targeted, and outcomes reported. The resulting categories, CT assessment strategies, GenAI pedagogical integration typologies, and reported impacts on CT, directly informed the organization and interpretation of the findings presented in Section 5.

5. Results

A total of 63 empirical studies were included in this systematic review, spanning publications from January 2023 to April 2025, with five, 38 and 20 papers each year respectively. The largest share of studies (40%) were conducted in Asian countries, particularly Taiwan and China, with only a few contributions from the UK, the United States, and other regions. The included studies were published in both peer-reviewed journals and international conference proceedings.
Most studies focused on undergraduate populations (75%). Approximately 20% included only graduate-level learners with 5% addressing both graduate and undergraduate students. Disciplinary backgrounds varied with common fields including technology and computing (20%), health and medical sciences (18%), education (16%), and others such as engineering or English as a foreign language. Most of the study focuses on one discipline with only 17% targeting multiple disciplines.

5.1. Methodological Approaches and Study Designs

In terms of research paradigms, as shown in Figure 3, the majority of studies (55.6%) employed mixed-methods approaches, combining quantitative and qualitative data to evaluate the impact of GenAI on students’ CT. About one-third of the studies (31.7%) used purely quantitative methods, often relying on pre/post-tests or Likert-type scales. In contrast, only 12.7% of the studies adopted qualitative designs, such as thematic analysis of student reflections or interview transcripts. More details of this can be found in Table 4.
Regarding the research designs across the empirical studies included in this review, pre-post quasi-experimental designs with comparison groups were the most common approach, employed by 19 studies. Randomized controlled trials accounted for 11 studies. Pre-experimental (single-group pre-post) designs were used in 10 studies.

5.2. Critical Thinking Constructs and Theoretical Frameworks

Notably, just under half of the studies in this review (27; 42.9%) explicitly situated their work within a recognized CT framework, while the majority (36; 57.1%) did not report any theoretical anchor, instead inferring definitions of CT from their chosen instruments or pedagogical activities. Among those that did specify a framework, Facione’s Delphi Report [2] guided 11% of the investigations (e.g., [53,72]) through its six core skills. Bloom’s Revised Taxonomy [27] underpinned 10 studies (15.9%), which framed GenAI-supported tasks in terms of hierarchical cognitive processes, particularly at the Analyze, Evaluate, and Create levels, and employed corresponding rubrics or assessments to gauge student performance (e.g., [66,67]). Paul and Elder’s Elements of Thought and Intellectual Standards [24] informed four studies, with an emphasis on metacognitive criteria such as clarity, accuracy, depth, and fairness in AI-mediated reasoning. Similarly, Higher-Order Thinking Skills (HOTS) frameworks also appeared in four studies, focusing on CT, problem-solving and creativity.
Frequent CT indicators extracted from the data included analysis, evaluation, creative and reflective thinking with five studies focusing on dispositions, such as truth-seeking, open-mindedness, analyticity, systematicity and inquisitiveness. The variety of indicators reflects broad interpretations of CT across disciplines; however, core skills such as analysis and evaluation remain central to most studies.

5.3. Assessment of Critical Thinking

The studies included in this review employed a diverse range of instruments to evaluate students’ CT, reflecting both innovation and inconsistency in assessment practices. These instruments varied in format, validity, and alignment with established CT frameworks, revealing key trends and gaps in the field.
A limited subset of studies utilized standardized and validated instruments. The Cornell Critical Thinking Test Level Z was one of the few standardized assessments employed, serving as a benchmark to quantify CT skills [59,70]. Similarly, a few studies used versions of the California Critical Thinking Disposition Inventory (CCTDI), including its Chinese adaptation (CTDI-CV), to evaluate learners’ dispositions toward critical thinking [49,61,83]. The Watson-Glaser Critical Thinking Appraisal also appeared in one study as a pre/post assessment [31]. Additional instruments like Sosu’s Critical Thinking Scale (CTS), the Reflective Thinking Scale (RTS), and Kember’s Reflective Thinking Questionnaire were used to assess learners’ capacity for reflective and analytical thought, all reporting high internal reliability (e.g., Cronbach’s α ≥ 0.84) [84]. The Critical Reasoning Assessment (CRA) further offered a multidimensional standardized tool (α = 0.87) [14,69]. While these instruments are theoretically grounded and validated, their use was not present in the majority of studies, highlighting a gap in the consistent application of rigorous CT assessment tools across the studies.
In contrast, many studies relied on custom-developed or adapted Likert-scale questionnaires, often drawn from the existing literature [85]. Internal reliability was reported in some cases with Cronbach’s alpha ranging from 0.78 to 0.89 [86]. However, formal validation procedures, such as construct validity analysis, expert review, or confirmatory factor analysis, were frequently omitted. Although these tools were tailored to specific contexts and populations, the lack of standardization limits their comparability and generalizability. Consequently, reported gains should be interpreted cautiously, as differences may reflect instrument sensitivity and reporting quality as much as intervention effects.
A significant number of studies employed rubrics and performance-based assessments to evaluate CT in students’ outputs, including essays, chatbot interactions, and project submissions [75,76,78,87,88]. These rubrics were often based on Bloom’s Revised Taxonomy, scoring across cognitive levels from remembering to creating [26,66,68]. Others included custom rubrics targeting ethical reasoning, argument structure, or quality of engagement with AI tools. A few studies used chatbot-specific evaluation rubrics, scoring on dimensions such as information accuracy, syntactic coherence, and semantic relevance [89]. Although these approaches offered contextual sensitivity and task alignment, they often lacked inter-rater reliability reporting or connection to validated constructs, which weakens their methodological rigor.
Another common strategy involved qualitative and thematic analysis, particularly in studies focusing on reflective journals, blog posts, and interview transcripts. These studies used custom thematic coding schemes informed by CT indicators, metacognition, or existing frameworks like the Four Resources Model or Facione’s CT skills [65]. Such qualitative methods allowed for deeper insights into students’ reasoning processes, especially in GenAI-supported learning environments. However, the subjective nature of coding and the infrequent use of reliability checks (e.g., inter-coder agreement), introduces risks of interpretive bias and limits reproducibility.

5.4. GenAI Integration Strategies and Pedagogical Designs

Across the 63 empirical studies in our review, GenAI interventions were implemented with a diverse range of methodologies, instructional designs, and learning objectives. Most interventions embedded GenAI tools into existing courses rather than as stand-alone applications. ChatGPT dominated (41 studies, 65%), reflecting its early adoption in higher education (e.g., [60,68,78,86]). Other tools such as Gemini, Ernie Bot, and multi-tool workflows were less common [61,70], and a few studies developed custom GPT-based systems ([50,59,84,89,90]). Most interventions occurred in face-to-face, synchronous settings (66%), while 16 studies (25%) were conducted outside class; fully asynchronous online deployments were rare.
GenAI tools were not introduced as mere technical novelties; instead, they were aligned with pedagogical functions intended to scaffold higher-order cognitive processes. Analysis of the interventions revealed six core pedagogical roles, which frame how GenAI supports learning:
  • Formative feedback engine: Provides adaptive hints, corrective suggestions, or automated scoring to guide reasoning.
  • Metacognitive scaffold: Prompts self-reflection, self-regulation, and monitoring of thought processes.
  • Socratic tutor or coach: Engages learners in structured dialogue, questioning assumptions and prompting deeper analysis.
  • Problem-solving assistant: Supports brainstorming, solution generation, and iterative problem analysis.
  • Simulator/role-play partner: Acts as an expert, stakeholder, or devil’s advocate to stimulate perspective-taking and ethical reasoning.
  • Idea generator and research assistant: Assists with evidence synthesis, source evaluation, and ideation tasks.
These functional roles converge into four main typologies of GenAI integration, which map directly to CT skills and dispositions. The diversity of roles naturally converges into four main typologies of integration, which are shown in Table 5. Each typology is described below with its pedagogical function, activated CT dimensions, representative studies, and observed outcomes.

5.4.1. Feedback Prompts/AI-Based Metacognitive Scaffold

These interventions embedded GenAI within learning platforms to generate real-time hints, guiding questions, automated scoring, or corrective feedback in response to student inputs [32,47,49,59]. Across these studies, AI tools were used to prompt students to analyze assumptions, evaluate evidence, refine arguments, or document reasoning steps. Reported target outcomes most frequently aligned with CT dimensions of analysis, interference, explanation, and self-regulation, as designed by established CT frameworks [2,25].
Methodologically, these interventions were typically short in duration (1-2 sessions or several weeks) and frequently evaluated using quantitative pre/post-test designs. Representative studies illustrate the effectiveness of this approach. For example, Suriano et al. [14] reported that depth of interaction with a chatbot was associated with measured gains in CT performance. Wang [61] implemented the GPT-based Feedback Aid System (GPTFAS) in a virtual reality (VR) learning environment and observed improvements in self-regulation and problem-solving performance relative to comparison conditions. Chen [55] integrated AI tools into argument-mapping tasks, requiring students to construct and rebut claims using AI-generated feedback.
Across this typology, reported outcomes included increases in analytic engagement, structured reasoning processes, and documented self-monitoring behaviors. Both standardized CT measures and custom assessment instruments were used to evaluate impact.

5.4.2. Dialogue Simulation and Reflection

These studies positioned ChatGPT as a dialogic partner to prompt questioning, counterargument generation, perspective-taking, or structured reflection. In these designs, AI agents assumed roles such as peer reviewer, expert stake holder, regulator, or devil’s advocate to challenge students’ reasoning and surface underlying assumptions [56,76,93].
These interventions were predominately synchronous and dialogue-oriented, often embedded within case-based learning, structured debates, or multi-stage problem-solving tasks. Reported outcome measures included both quantitative CT assessments and qualitative analyses of reflective journals and interviews. Gunawan et al. [31] reported strengthened argument evaluation and evidence integration in problem-solving tasks. Zhou et al. [15] observed increased self-reflection and critical interrogation of AI outputs after structured AI-mediated debates. Additional studies documented perspective-taking behaviors and explicit evaluation of AI-generated responses through post-intervention journals [62].
Quantitative gains were reported in some studies [77]; however, several investigations relied on qualitative analyses to document changes in reflective reasoning, argument evaluation, and ethical consideration [62]. Across this typology, GenAI functioned primarily as a conversational scaffold rather than as a content generator. This strategy effectively leveraged dialogue and reflection to deepen learners’ CT [68,94].

5.4.3. AI-Based Peer Review

These studies integrated GenAI into peer review, project-based learning (PBL), programming tasks, or structured critique activities. In these interventions, AI tools supplied exemplar feedback, evaluation criteria, alternative solutions, brainstorming prompts, or code suggestions that students were required to analyze, critique, or revise [13,50,71,93].
In project-based contexts, GenAI was embedded across multiple stages of project development, including idea generation, planning, analysis, and drafting [50,59]. Drushlyak et al. [75] reported positive changes in CT performance following structured critique of AI-generated solutions. Styve [13] implemented an “AI with CT” programming assignment in which students evaluated code quality and robustness using AI-generated suggestions.
Other studies incorporated GenAI into information literacy and research tasks. Lin et al. [84] integrated a GPT-assisted summarization tool into an experiential learning course, requiring students to generate questions and document reasoning steps. Pellas [70] examined associations between attitudes toward machine learning and performance outcomes, reporting relationships with problem-solving and CT measures. Across these studies, CT dimensions most frequently activated included analysis, evaluation, and explanation [2,25]. Assessments ranged from rubric-scored project artifacts to standardized CT instruments.

5.4.4. Critical Engagement with AI-Generated Content

In these studies, students were required to explicitly evaluate, revise, audit, or justify modifications to AI-generated outputs. In contrast to content-generation use cases, these interventions framed AI outputs as provisional artifacts subject to critique, correction, and improvement [65,68,74,75,76,89]. In this typology, GenAI was consistently positioned as an object of analysis rather than an unquestioned source of information. Reported outcomes included changes in evaluation, explanation, interpretation, and reflective monitoring behaviors.
Common tasks included identifying logical gaps, unsupported claims, hallucinations, biases, or weak evidence within AI-generated drafts. Farinetti and Canale [89] implemented an activity in which students analyzed AI-generated ethical case essays and revised them while documenting rationale for changes. They found enhanced evaluation, explanation skills and increased students’ awareness of AI limitations and hallucinations. Chen et al. [68] examined structured prompting strategies and observed differences in reasoning depth between high- and low-performing students. Several studies reported increased awareness of AI limitations and improved evaluative reasoning following structured revision tasks [62,64,75]. These interventions were frequently embedded in performance-based assessments supported by CT-aligned rubrics (e.g., Bloom’s Revised Taxonomy, Paul & Elder’s Elements of Thought) [15,47,95].

5.5. Reported Impacts on Critical Thinking

Among the studies included in this review, 30 (47%) reported statistically significant quantitative gains in students’ CT following GenAI-mediated interventions. As we presented in the previous section, these improvements were observed across a range of standardized tests, custom pre-/post-test instruments, and performance metrics aligned with established CT frameworks. For example, Shalong et al. [59] employed the Cornell Critical Thinking Test Level Z to measure changes over six, 12, and 14 weeks, reporting significant improvements on both SDL and CT scores of approximately 13%, which was sustained at follow-up. Similarly, Essel et al. [96] found a statistically discernible effect on the posttest score of the Critical Thinking Disposition Scale (CTDS) with a partial eta squared (η2p) of 0.229 and a significant increase in the dimensions of critical openness and reflective skepticism.
Liu and Wang [49] utilized the Chinese version of the California Critical Thinking Disposition Inventory (CTDI-CV), finding an improvement of CT scores by an average of 14.57 points (20%, 70-items scale), in dimensions such as truth-seeking, open-mindedness, and analyticity for experimental groups (using AI for mind maps, generating and answering text-related questions, interactive quizzes, and AI-assisted debates) versus non significant changes in control groups (traditional teaching). Similarly, Fabio et al. [69] reported similar trends using the Critical Reasoning Assessment (CRA) with heavy-use groups scoring significantly higher than moderate and non-use groups across most CT dimensions. They reported large multivariate effects (MANOVA: F(2, 123) = 16.43, p < 0.01, η2p = 0.29) and strong paired comparisons (d ≈ 0.82–0.89).
Several studies provided standardized effect sizes, allowing clearer interpretation of the magnitude of gains. Lin [58] reported a large effect (Cohen’s d = 1.20) on the Reflective Thinking Questionnaire, while Wang [91] found partial η2 = 0.575 for the same instrument, also indicative of a substantial impact. Huang [57] observed moderate gains on the higher-order thinking skills (HOTS) scale (partial η2 = 0.22), and Zhang [54] reported a statistically significant increase (Diff-in-Diff = 1.323, p < 0.01) in the HOTS 5-point Likert CT scale. In another intervention, Chang [86] reported large effect sizes for CT (η2 = 0.414) when integrating AI into collaborative and problem-solving tasks, alongside gains in problem-solving (0.343).
Lee [28] conducted a randomized controlled trial in a foundational chemistry course to assess the effects of a guidance-based ChatGPT-assisted learning aid (GCLA) on self-regulated learning (SRL), HOTS, and knowledge construction. The findings revealed large effect sizes across multiple dimensions. Specifically, students who used the GCLA showed significantly higher gains in critical thinking (F = 41.73, η2 = 0.418), problem-solving (F = 38.70, η2 = 0.400), and creativity (F = 8.98, η2 = 0.134) compared to those using traditional ChatGPT. Additionally, the GCLA group demonstrated substantial improvements in cognitive (F = 48.60, η2 = 0.369) and behavioral engagement (F = 41.40, η2 = 0.333), as well as self-efficacy (F = 26.46, η2 = 0.313). Similarly, using the same scale Wang [91] found that students who received real-time, voice/text-based feedback from ChatGPT-integrated feedback aids (GPTFAS) significantly outperformed the control group in problem-solving (F = 178.65, η2 = 0.696) and CT (F = 11.54, η2 = 0.128). However, no significant difference was observed in creativity outcomes. Additionally, GPTFAS had a strong impact on self-regulated learning abilities (F = 113.47, η2 = 0.593), highlighting the potential of GenAI feedback to support metacognitive and strategic learning processes in immersive learning contexts.

6. Limitations

Despite the fact that we follow the PRISMA guidelines and hybrid thematic analysis, several limitations constrain the scope and generalizability of this review. First, our search was restricted to studies published in English between January 2023 and April 2025 and indexed in five major databases (Compendex, Inspec, Web of Science, EBSCO, and Scopus). This restriction likely underrepresents research published in languages other than English and may bias the evidence base toward contexts and institutions that disseminate findings in English, even when studies are conducted in non-English-speaking countries (e.g., China, Taiwan).
Second, the heterogeneity of study designs, pedagogical interventions, and measurement instruments precluded quantitative synthesis of GenAI’s effects on CT. Included studies varied widely in their operational definitions of CT, the validity and reliability of their assessment tools, intervention durations (ranging from a single session to semester-long projects), and disciplinary settings. In particular, many studies relied on custom or adapted instruments with limited psychometric reporting beyond internal consistency, and rubric- or coding-based assessments often omitted inter-rater reliability, reducing reproducibility and comparability across studies. While our analysis captures emerging patterns, the lack of standardized measures and comparable outcome metrics limits our ability to draw definitive conclusions about effect sizes across studies or to identify best-practice designs with statistical rigor.
Third, the emergent nature of GenAI integration in higher education means that many studies are exploratory, descriptive, or quasi-experimental, often with small sample sizes, short follow-up periods, and minimal control for confounding variables. Few investigations employed randomized controlled designs or longitudinal tracking to assess the durability of CT gains over time. Given that CT represents a long-term cognitive skill set and disposition, short interventions and immediate post-tests may capture proximal performance effects rather than durable cognitive development. Moreover, reporting of inter-rater reliability for coding rubrics and thematic analyses was inconsistent, introducing potential interpretive bias. Finally, publication bias may have favored studies reporting positive affordances of GenAI, while null or negative findings, especially around cognitive offloading and diminished autonomy, may remain unpublished.
Lastly, generalizability is further constrained by the geographic and disciplinary concentration of the evidence base. Educational cultures differ in teacher-student interaction norms, assessment practices, and expectations about independent reasoning versus guided scaffolding, which may moderate how GenAI is used and how CT outcomes manifest. Similarly, disciplinary epistemologies (e.g., computing vs. health sciences vs. education) shape what counts as evidence, argument quality, and acceptable verification practices. Future research should test GenAI-integrated CT interventions across more diverse regions, languages, and disciplines, and use comparative designs to examine whether the same pedagogical strategies produce similar CT effects in different educational traditions.

7. Discussion

7.1. Summary of Key Findings

7.1.1. Primary Cognitive Affordances

GenAI tools demonstrate substantial potential to scaffold reflective and iterative reasoning processes, essential components for fostering deep CT [49]. Students benefit from adaptive feedback mechanisms inherent in GenAI tools like ChatGPT, which prompt them to reflect on their assumptions and refine their reasoning iteratively [28,32]. This aligns with studies by Zhou et al. [15] that emphasize the importance of designing GenAI tools that support self-regulated learning. That meta-reflection after AI-mediated debates enhanced self-awareness and critical interrogation of AI outputs, positioning the tool as a scaffold rather than a substitute for human judgment [62]. Similarly, Celik et al. [8] emphasize how GenAI can enable students to clarify and link multiple ideas through recommendation and automated feedback, encouraging more critical engagement. Furthermore, the ability of GenAI tools to provide immediate feedback and model complex argumentative structures can also facilitate the construction and evaluation of opposing perspectives [55].
These cognitive affordances are amplified when paired with pedagogical intentionality [47]. Highlighting the importance of structured integration of AI tools with CT to promote responsible use, deeper reflection, increased awareness of code quality and robustness [13,50,71,93]. Educational contexts where AI-generated prompts explicitly encourage reflective questioning, ethical deliberation, and iterative dialogue tend to yield richer cognitive engagement [94]. By positioning GenAI as a reflective interlocutor rather than a content generator, effectively leveraging dialogue and reflection can deepen learners’ CT [68,94]. Tools providing explicit reflective questions or ethical simulations were particularly effective in engaging students deeply, promoting sustained metacognitive awareness [76]. Drushlyak et al. [75] found positive changes in students CT after analyzing and critiquing AI-generated solutions, commenting that the inevitability of AI errors becomes a learning opportunity, prompting active interrogation and deeper reasoning.
Gonsalves [80] found that students moved fluidly across cognitive, affective, and metacognitive domains and identified 12 AI-specific competencies (e.g., melioration, ethical reasoning) underpinning a revised Bloom’s taxonomy. This revised framework emphasizes the importance of ethical reasoning and reflective thinking in AI-assisted learning environments, ensuring that students engage critically with AI-generated content [80].

7.1.2. Moderating Factors

Key moderating factors were identified, including proper instructional design, where integration of AI in education necessitates a revised framework that incorporates AI-specific competencies to effectively nurture CT [80]. Also, educators must adapt their strategies to encourage students to critically evaluate AI-generated information, fostering independence and discernment in their learning processes [67,73,74,76]. Additionally, structured assessment types and clearly articulated task requirements significantly shaped students’ cognitive engagement with GenAI tools [17]. Thus, students’ self-regulatory abilities appear critical in determining the depth and quality of cognitive engagement facilitated by GenAI [15]. Zhou et al. [15] also discussed the importance of designing GenAI tools with high ease of use to foster self-regulation and deepen CT and integrate self-regulatory strategies (goal setting, self-monitoring, reflective practice) into AI-enhanced curricula. This highlights the importance of emphasizing critical evaluation of AI outputs rather than overreliance on perceived usefulness or learning value.
Chen et al. [68] showed that high achievers followed a structured path, whereas low achievers jumped straight to application with minimal deeper inquiry, showing that proactive, scaffolded questioning of ChatGPT (building understanding before application) supports deeper reasoning. Lastly, the importance of carefully scaffolded tasks and modeling of critical questioning helped students see AI as a thinking partner rather than an answer provider [72], positioning AI more than as a tool as a pedagogical partner [97].
Several studies report that supervised use of GenAI improves students’ ability to plan their learning, monitor their understanding, and regulate their emotions when faced with complex tasks [98]. These skills are reinforced when instructors teach students to formulate precise prompts, examine the reliability of sources and critically evaluate AI outputs to prevent superficial use [65,68]. However, studies also warn that without explicit metacognitive training, students may skip crucial stages of application, analysis, and evaluation [98], making metacognitive support essential for sustaining gains.

7.1.3. Principal Pedagogical Risks

Despite these affordances, significant pedagogical risks remain, notably cognitive offloading and a consequent reduction in independent reasoning and self-regulation due to dependency on AI systems [11]. Studies consistently highlight concerns regarding students’ uncritical acceptance of AI-generated content, exacerbated by the fluent yet potentially inaccurate outputs of GenAI systems [10,18]. Students can bypass rigorous analysis and synthesis, instead relying heavily on AI outputs as final answers, which in turn reinforces passive learning habits and superficial cognitive engagement [30]. As shown by [99], excessive reliance on AI may lead to superficial cognitive thinking engagement, diminished teacher-student interaction, and cognitive dissonance, which can hinder learners’ active construction of meaning. Overreliance on GenAI tools can lead to cognitive dependency, diminishing students’ ability to formulate original arguments, critically assess alternative perspectives and decision-making processes [33,80,100]. This phenomenon calls for an urgent revision of traditional CT instructional models to explicitly account for collaborative interactions with GenAI. Educators must carefully balance GenAI use, providing intentional scaffolding and clear instructions to ensure these tools complement rather than substitute students’ CT development [88].
Lastly, the fluency and coherence of AI outputs also can create the illusion of understanding, encouraging students to accept responses without verification [10]. Such uncritical trust undermines the iterative processes of evaluation, reflection, and justification that are central to CT development. As several studies suggest [81,88], this trend highlights the urgent need to revise CT teaching models to explicitly incorporate GenAI collaboration. Gonsalves [80] urges pedagogical redesign to guide critical engagement, mitigate dependency, and foster ethical, metacognitive, and collaborative skills in AI-assisted learning.

7.2. Methodological Strengths and Gaps

Regarding the methodological strengths, one encouraging trend is the growing adoption of validated CT assessments. The use of mixed-methods and, in some cases, randomized controlled trials adds robustness to evidence on GenAI’s impact. These hybrid designs capture both quantitative changes and qualitative depth, acknowledging the multidimensional nature of CT in AI-mediated environments. This methodological choice aligns with calls in the literature for comprehensive, process-sensitive designs that capture not only outcome measures but also the cognitive and metacognitive processes underlying CT development in GenAI-mediated environments [101,102].
On the other hand, limited theoretical grounding constrains the explanatory power of many studies. Over 57% of the reviewed papers did not reference a formal CT framework, instead treating CT as an implicit outcome of instructional activities or as a composite of general higher-order thinking skills. Embedding research within robust theoretical models is essential to ensure that interventions are explicitly aligned with cognitive targets and that their outcomes are interpretable across contexts [10]. The review also revealed significant inconsistencies in how CT is measured in GenAI-related education research. Validated instruments remain underutilized, while custom tools dominate, often without theoretical or psychometric grounding. Many studies failed to clearly articulate how CT was defined or connected to the instruments they employed. This fragmentation points to an urgent need for the development of context-sensitive yet validated CT assessment tools, particularly suited to AI-enhanced learning. Moreover, researchers should aim for greater transparency and rigor in reporting instrument design, validation processes, and alignment with conceptual models. Without such improvements, it will remain difficult to meaningfully compare outcomes across studies or synthesize evidence about the impact of GenAI on students’ CT.
Another major and widely reported limitation is the lack of longitudinal research [101,103]. Most interventions were short-term, ranging from single-session activities to a few weeks in duration, and very few studies examined whether observed CT gains persisted over time or translated into durable cognitive habits. Without longitudinal tracking, it remains unclear whether initial improvements in analysis, evaluation, or reflective thinking are sustainable or whether reliance on GenAI produces delayed negative effects, such as overreliance or cognitive offloading.

7.3. Pedagogical Implications and Recommendations

To mitigate the risk of GenAI undermining critical thinking skills in academic settings, educators can adopt several strategic approaches. These strategies focus on integrating AI tools in ways that enhance student engagement and promote independent thinking. First, guided and reflective integration is essential. Studies consistently show that students achieve deeper cognitive engagement when GenAI use is paired with structured prompts, metacognitive scaffolds, and reflective tasks that require justification of reasoning [14,93]. For example, interventions that combined AI-generated suggestions with explicit reflection journals or critical evaluation tasks fostered analysis, evaluation, and self-regulation, while reducing risks of uncritical acceptance [77,86]. These models must train students to engage with AI critically, integrating explicit verification of outputs and reflective questioning to transform GenAI from a crutch into a catalyst for deep learning [104]. AI-based dialogues can foster greater inquiry, hypothesis generation, and analytical questioning, which in turn encourage students to evaluate evidence more rigorously and refine their reasoning processes [92]. Also automating low-level tasks frees students to engage in deeper reasoning [26].
Second, curriculum design must integrate GenAI literacy and ethical engagement. There is a need for curricula to combine AI literacy with critical thinking pedagogy in order to harness the benefits without fostering over-reliance [26,49]. Embedding AI-specific competencies, such as reflective skepticism, ethical reasoning, and source verification, into courses can transform potential risks into learning opportunities [80,104]. Also, it is important to teach students to use these tools critically and creatively [104,105]. Activities that require students to critique, revise, and justify AI-generated content [89] not only reinforces core CT skills but also cultivates truth-seeking and open-mindedness dispositions [53].
Third, active and performance-based assessment strategies should be prioritized. Performance tasks, such as AI-augmented debates, problem-based projects, and peer review of AI-generated work, proved particularly effective for promoting evaluation and explanation skills [13,75]. Instructors are advised to combine formative assessments with process-oriented evaluation, monitoring not only final outputs but also students’ engagement with reasoning steps, reflective logs, and revisions of AI-assisted work. This shift toward process-based assessment aligns with the broader movement toward durable skill development in higher education [30,106,107].
Finally, balance and intentionality remain paramount. While GenAI can accelerate ideation, provide feedback, and simulate Socratic dialogue, excessive dependence risks undermining independent reasoning and metacognitive self-regulation [15,18,30]. Effective integration requires structured exposure, explicit reflection cycles, and gradual release of responsibility to ensure that students internalize CT processes rather than externalize them to AI systems. Faculty development programs, institutional policies, and ethical guidelines should support this balanced approach, enabling educators to leverage GenAI to support human skills development. Educators should revise existing frameworks like Bloom’s Taxonomy to include AI-specific competencies, emphasizing skills such as ethical reasoning and reflective thinking [80].
Training for instructors on effective pedagogical integration, should focus on prompting, questioning strategies, metacognitive scaffolding. Implementing active teaching methods encourages students to engage critically with AI tools, promoting skills like information verification and comprehensive analysis. Development of educational guidelines emphasizing balanced use of GenAI, promotes student agency and autonomy to prevent overreliance on AI [90]. Also, assessment strategies should evolve to incorporate AI-generated responses, focusing on evaluation and synthesis rather than rote memorization.

8. Conclusions and Future Directions

This SLR synthesized 63 empirical studies published between January 2023 and April 2025 that examined the impact of GenAI tools on students’ CT in higher education. The findings reveal that while GenAI offers promising cognitive affordances, such as scaffolding, reflective reasoning, fostering iterative dialogue, and supporting metacognitive awareness, its pedagogical impact is highly dependent on intentional integration into instructional design. When used strategically, GenAI can facilitate both CT skills and dispositions, yet risks such as cognitive offloading, uncritical acceptance of AI-generated content, and overreliance remain significant concerns.
Currently, the empirical landscape is heavily concentrated in Asian higher education and in fields such as computing, health sciences, and education, with limited attention to interdisciplinary, Western, or Latin American contexts. Expanding research to include varied disciplines and cultural settings will strengthen the generalizability of findings and capture how sociocultural factors mediate GenAI adoption and CT development. Additionally, future studies should examine institutional and systemic dimensions, including the effectiveness of faculty development programs, the role of institutional policies in shaping responsible GenAI use, and the impact of co-agency models where human and AI reasoning are intentionally integrated.
For educators and policymakers, the results underscore the need to align GenAI adoption with evidence-based CT pedagogies, embedding structured reflection, comparative reasoning, and explicit evaluation of AI outputs into learning activities. Future research should prioritize robust measurement approaches, longitudinal evaluations, and cross-disciplinary investigations that examine not only whether GenAI can support CT but under what conditions it is most effective. Given the rapid evolution of GenAI tools and the accelerating pace of related publications, periodic updates of this review, or a ‘living’ systematic review approach, may be necessary to keep the synthesis current as the evidence base expands and stabilizes. This synthesis reflects the evidence available through April 2025; ongoing updates will be needed as new GenAI capabilities and empirical studies emerge.
Ultimately, GenAI’s role in higher education should be framed not as a replacement for human reasoning but as a catalyst for more purposeful, self-regulated, and reflective thinking. Unlocking its potential for CT will require sustained collaboration among educators, researchers, technologists, and students to design learning environments that balance innovation and prioritize human skills development.

Author Contributions

Conceptualization, T.B., B.D., and K.S.; methodology, T.B., and B.D.; screening, T.B., and B.D.; formal analysis, T.B.; investigation, T.B.; data curation, T.B.; writing, original draft preparation, T.B.; writing, review and editing, T.B., B.D., and K.S.; visualization, T.B.; supervision, B.D., and K.S. All authors have read and agreed to the published version of the manuscript.

Funding

We gratefully acknowledge the support received from ANID/PIA/Basal Funds for Centers of Excellence FB0003.

Data Availability Statement

The original contributions presented in this study are included in the article. The data supporting the findings of this review are derived from published studies cited in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Ahern, A.; Dominguez, C.; McNally, C.; O’Sullivan, J.J.; Pedrosa, D. A Literature Review of Critical Thinking in Engineering Education. Stud. High. Educ. 2019, 44, 816–828. [Google Scholar] [CrossRef]
  2. Facione, P. Critical Thinking: A Statement of Expert Consensus for Purposes of Educational Assessment and Instruction (The Delphi Report). In Educational Resources Information Center (ERIC); California Academic Press: Millbrae, CA, USA, 1990; pp. 1–112. [Google Scholar]
  3. Deo, S.; Hölttä-Otto, K. Critical Thinking Assessment in Engineering Education: A Scopus-Based Literature Review. J. Mech. Des. 2024, 146, 072301. [Google Scholar] [CrossRef]
  4. Bhuttah, T.M.; Xusheng, Q.; Abid, M.N.; Sharma, S. Enhancing Student Critical Thinking and Learning Outcomes through Innovative Pedagogical Approaches in Higher Education: The Mediating Role of Inclusive Leadership. Sci. Rep. 2024, 14, 24362. [Google Scholar] [CrossRef] [PubMed]
  5. Van Damme, D.; Zahner, D.; Cortellini, O.; Dawber, T.; Rotholz, K. Assessing and Developing Critical-thinking Skills in Higher Education. Eur. J. Educ. 2023, 58, 369–386. [Google Scholar] [CrossRef]
  6. Belkina, M.; Daniel, S.; Nikolic, S.; Haque, R.; Lyden, S.; Neal, P.; Grundy, S.; Hassan, G.M. Implementing Generative AI (GenAI) in Higher Education: A Systematic Review of Case Studies. Comput. Educ. Artif. Intell. 2025, 8, 100407. [Google Scholar] [CrossRef]
  7. Shanto, S.S.; Ahmed, Z.; Jony, A.I. Enriching Learning Process with Generative AI: A Proposed Framework to Cultivate Critical Thinking in Higher Education Using Chat GPT. Tuijin Jishu J. Propuls. Technol. 2024, 45, 3019–3029. [Google Scholar]
  8. Celik, I.; Gedrimiene, E.; Siklander, S.; Muukkonen, H. The Affordances of Artificial Intelligence-Based Tools for Supporting 21st-Century Skills: A Systematic Review of Empirical Research in Higher Education. Australas. J. Educ. Technol. 2024, 40, 19–38. [Google Scholar] [CrossRef]
  9. Husain, N. Educational Renaissance: Unleashing the Potential of Interdisciplinary AI Integration. J. Environ. Sci. Technol. 2023, 2, 159–166. [Google Scholar]
  10. Deng, R.; Jiang, M.; Yu, X.; Lu, Y.; Liu, S. Does ChatGPT Enhance Student Learning? A Systematic Review and Meta-Analysis of Experimental Studies. Comput. Educ. 2025, 227, 105224. [Google Scholar] [CrossRef]
  11. Gerlich, M. AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies 2025, 15, 6. [Google Scholar] [CrossRef]
  12. Zhai, X.; Nyaaba, M.; Ma, W. Can Generative AI and ChatGPT Outperform Humans on Cognitive-Demanding Problem-Solving Tasks in Science? Sci. Educ. 2025, 34, 649–670. [Google Scholar] [CrossRef]
  13. Styve, A.; Virkki, O.T.; Naeem, U. Developing Critical Thinking Practices Interwoven with Generative AI Usage in an Introductory Programming Course. In Proceedings of the 2024 IEEE Global Engineering Education Conference (EDUCON), Kos Island, Greece, 8–11 May 2024; IEEE: Kos Island, Greece, 2024; pp. 01–08. [Google Scholar]
  14. Suriano, R.; Plebe, A.; Acciai, A.; Fabio, R.A. Student Interaction with ChatGPT Can Promote Complex Critical Thinking Skills. Learn. Instr. 2025, 95, 102011. [Google Scholar] [CrossRef]
  15. Zhou, X.; Teng, D.; Al-Samarraie, H. The Mediating Role of Generative AI Self-Regulation on Students’ Critical Thinking and Problem-Solving. Educ. Sci. 2024, 14, 1302. [Google Scholar] [CrossRef]
  16. Kalita, A.; Singh, P.; Saikia, A.P. An Outcome-Based Learning Framework for Promoting Critical Thinking in Undergraduate Engineering Students. In Advances in Higher Education and Professional Development; Barua, K., Radwan, N., Singh, V., Figueiredo, R., Eds.; IGI Global: Hershey, PA, USA, 2023; pp. 271–296. ISBN 978-1-66849-472-1. [Google Scholar]
  17. Heimdal, A.; Lande, I.; Almlie, G.S.; Roaldsøy, E.Ø. Students’ Experiences from Using AI in Engineering Education. In Proceedings of the International Conference on Engineering and Product Design Education, EPDE 2024, Birmingham, UK, 5–6 September 2024; The Design Society: Glasgow, UK, 2024; pp. 497–502. [Google Scholar]
  18. Zhai, C.; Wibowo, S.; Li, L.D. The Effects of Over-Reliance on AI Dialogue Systems on Students’ Cognitive Abilities: A Systematic Review. Smart Learn. Environ. 2024, 11, 28. [Google Scholar] [CrossRef]
  19. Melisa, R.; Ashadi, A.; Triastuti, A.; Hidayati, S.; Salido, A.; Ero, P.E.L.; Marlini, C.; Zefrin, Z.; Al Fuad, Z. Critical Thinking in the Age of AI: A Systematic Review of AI’s Effects on Higher Education. Educ. Process Int. J. 2025, 14, e2025031. [Google Scholar] [CrossRef]
  20. Benvenuti, M.; Cangelosi, A.; Weinberger, A.; Mazzoni, E.; Benassi, M.; Barbaresi, M.; Orsoni, M. Artificial Intelligence and Human Behavioral Development: A Perspective on New Skills and Competences Acquisition for the Educational Context. Comput. Hum. Behav. 2023, 148, 107903. [Google Scholar] [CrossRef]
  21. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, 71. [Google Scholar] [CrossRef] [PubMed]
  22. Garcia Castro, R.A.; Mayta Cachicatari, N.A.; Bartesaghi Aste, W.M.; Llapa Medina, M.P. Exploration of ChatGPT in Basic Education: Advantages, Disadvantages, and Its Impact on School Tasks. Contemp. Educ. Technol. 2024, 16, ep484. [Google Scholar] [CrossRef]
  23. Facione, P.A. Critical Thinking: What It Is and Why It Counts. Insight Assess. 2011, 1, 1–23. [Google Scholar]
  24. Paul, R.; Elder, L. Critical Thinking: Tools for Taking Charge of Your Professional and Personal Life; Pearson Education: Upper Saddle River, NJ, USA, 2013; ISBN 0-13-311528-3. [Google Scholar]
  25. Ennis, R.H. A Logical Basis for Measuring Critical Thinking Skills. Educ. Leadersh. 1985, 43, 44–48. [Google Scholar]
  26. Essien, A.; Bukoye, O.T.; O’Dea, X.; Kremantzis, M. The Influence of AI Text Generators on Critical Thinking Skills in UK Business Schools. Stud. High. Educ. 2024, 49, 865–882. [Google Scholar] [CrossRef]
  27. Krathwohl, D.R. A Revision of Bloom’s Taxonomy: An Overview. Theory Pract. 2002, 41, 212–218. [Google Scholar] [CrossRef]
  28. Lee, H.-Y.; Chen, P.-H.; Wang, W.-S.; Huang, Y.-M.; Wu, T.-T. Empowering ChatGPT with Guidance Mechanism in Blended Learning: Effect of Self-Regulated Learning, Higher-Order Thinking Skills, and Knowledge Construction. Int. J. Educ. Technol. High. Educ. 2024, 21, 16. [Google Scholar] [CrossRef]
  29. ElSayary, A. Integrating Generative AI in Active Learning Environments: Enhancing Metacognition and Technological Skills. J. Syst. Cybern. Inf. 2024, 22, 34–37. [Google Scholar] [CrossRef]
  30. Hutson, J.; Plate, D. Human-AI Collaboration for Smart Education: Reframing Applied Learning to Support Metacognition. In Advanced Virtual Assistants-A Window to the Virtual Future; IntechOpen: London, UK, 2023. [Google Scholar]
  31. Gunawan, A.; Wiputra, R. Exploring the Impact of Generative AI on Personalized Learning in Higher Education; IEEE: Piscataway, NJ, USA, 2024; pp. 280–285. [Google Scholar]
  32. Youssef, E.; Medhat, M.; Abdellatif, S.; Al Malek, M. Examining the Effect of ChatGPT Usage on Students’ Academic Learning and Achievement: A Survey-Based Study in Ajman, UAE. Comput. Educ. Artif. Intell. 2024, 7, 100316. [Google Scholar] [CrossRef]
  33. Sabharwal, D.; Kabha, R.; Srivastava, K. Artificial Intelligence (AI)-Powered Virtual Assistants and Their Effect on Human Productivity and Laziness: Study on Students of Delhi-NCR (India) & Fujairah (UAE). J. Content Community Commun. 2023, 17, 162–174. [Google Scholar] [CrossRef]
  34. Modran, H.A.; Chamunorwa, T.; Ursuțiu, D.; Samoilă, C. Integrating Artificial Intelligence and ChatGPT into Higher Engineering Education. In Towards a Hybrid, Flexible and Socially Engaged Higher Education; Auer, M.E., Cukierman, U.R., Vendrell Vidal, E., Tovar Caro, E., Eds.; Lecture Notes in Networks and Systems; Springer Nature Switzerland: Cham, Switzerland, 2024; Volume 899, pp. 499–510. ISBN 978-3-031-51978-9. [Google Scholar]
  35. Cai, L.Y.; Msafiri, M.M.; Kangwa, D. Exploring the Impact of Integrating AI Tools in Higher Education Using the Zone of Proximal Development. Educ. Inf. Technol. 2025, 30, 7191–7264. [Google Scholar] [CrossRef]
  36. Ilic, J.; Ivanovic, M.; Klasnja-Milicevic, A. The Impact of ChatGPT on Student Learning Experience in Higher STEM Education: A Systematic Literature Review. In Proceedings of the 2024 IEEE 11th International Conference on e-Learning e-Education and Online Education (ICILE), Belgrade, Serbia, 23–25 May 2024; pp. 1–9. [Google Scholar]
  37. Nguyen, K.V. The Use of Generative AI Tools in Higher Education: Ethical and Pedagogical Principles. J. Acad. Ethics 2025, 23, 1435–1455. [Google Scholar] [CrossRef]
  38. Tillmanns, T.; Salomão Filho, A.; Rudra, S.; Weber, P.; Dawitz, J.; Wiersma, E.; Dudenaite, D.; Reynolds, S. Mapping Tomorrow’s Teaching and Learning Spaces: A Systematic Review on GenAI in Higher Education. Trends High. Educ. 2025, 4, 2. [Google Scholar] [CrossRef]
  39. Vargas-Murillo, A.R.; de la Asuncion Pari-Bedoya, I.N.M.; de Jesús Guevara-Soto, F. Challenges and Opportunities of AI-Assisted Learning: A Systematic Literature Review on the Impact of ChatGPT Usage in Higher Education. Int. J. Learn. Teach. Educ. Res. 2023, 22, 122–135. [Google Scholar] [CrossRef]
  40. Sardi, J.; Candra, O.; Yuliana, D.F.; Yanto, D.T.P.; Eliza, F. How Generative AI Influences Students’ Self-Regulated Learning and Critical Thinking Skills? A Systematic Review. Int. J. Eng. Pedagog. 2025, 15, 94–108. [Google Scholar] [CrossRef]
  41. Hooshyar, D.; Pedaste, M.; Saks, K.; Leijen, Ä.; Bardone, E.; Wang, M. Open Learner Models in Supporting Self-Regulated Learning in Higher Education: A Systematic Literature Review. Comput. Educ. 2020, 154, 103878. [Google Scholar] [CrossRef]
  42. Kellermeyer, L.; Harnke, B.; Knight, S. Covidence and Rayyan. J. Med. Libr. Assoc. 2018, 106, 580–583. [Google Scholar] [CrossRef]
  43. Galli, C.; Gavrilova, A.V.; Calciolari, E. Large Language Models in Systematic Review Screening: Opportunities, Challenges, and Methodological Considerations. Information 2025, 16, 378. [Google Scholar] [CrossRef]
  44. Van Dinter, R.; Tekinerdogan, B.; Catal, C. Automation of Systematic Literature Reviews: A Systematic Literature Review. Inf. Softw. Technol. 2021, 136, 106589. [Google Scholar] [CrossRef]
  45. Darmawansah, D.; Rachman, D.; Febiyani, F.; Hwang, G.-J. ChatGPT-Supported Collaborative Argumentation: Integrating Collaboration Script and Argument Mapping to Enhance EFL Students’ Argumentation Skills. Educ. Inf. Technol. 2025, 30, 3803–3827. [Google Scholar] [CrossRef]
  46. Dindorf, C.; Weisenburger, F.; Bartaguiz, E.; Dully, J.; Klappenberger, L.; Lang, V.; Zimmermann, L.; Fröhlich, M.; Seibert, J.-N. Exploring Decision-Making Competence in Sugar-Substitute Choices: A Cross-Disciplinary Investigation among Chemistry and Sports and Health Students. Educ. Sci. 2024, 14, 531. [Google Scholar] [CrossRef]
  47. Lu, J.; Zheng, R.; Gong, Z.; Xu, H. Supporting Teachers’ Professional Development with Generative AI: The Effects on Higher Order Thinking and Self-Efficacy. IEEE Trans. Learn. Technol. 2024, 17, 1279–1289. [Google Scholar] [CrossRef]
  48. Li, T.; Ji, Y.; Zhan, Z. Expert or Machine? Comparing the Effect of Pairing Student Teacher with in-Service Teacher and ChatGPT on Their Critical Thinking, Learning Performance, and Cognitive Load in an Integrated-STEM Course. Asia Pac. J. Educ. 2024, 44, 45–60. [Google Scholar] [CrossRef]
  49. Liu, W.; Wang, Y. The Effects of Using AI Tools on Critical Thinking in English Literature Classes Among EFL Learners: An Intervention Study. Eur. J. Educ. 2024, 59, e12804. [Google Scholar] [CrossRef]
  50. Naatonis, R.N.; Rusijono, R.; Jannah, M.; Malahina, E.A.U. Evaluation of Problem Based Gamification Learning (PBGL) Model on Critical Thinking Ability with Artificial Intelligence Approach Integrated with ChatGPT API: An Experimental Study. Qubahan Acad. J. 2024, 4, 485–520. [Google Scholar] [CrossRef]
  51. Saritepeci, M.; Durak, H. Effectiveness of Artificial Intelligence Integration in Design-Based Learning on Design Thinking Mindset, Creative and Reflective Thinking Skills: An Experimental Study. Educ. Inf. Technol. 2024, 29, 25175–25209. [Google Scholar] [CrossRef]
  52. Shin, H.; De Gagne, J.C.; Kim, S.S.; Hong, M. The Impact of Artificial Intelligence-Assisted Learning on Nursing Students’ Ethical Decision-Making and Clinical Reasoning in Pediatric Care: A Quasi-Experimental Study. CIN-Comput. Inform. Nurs. 2024, 42, 704–711. [Google Scholar] [CrossRef] [PubMed]
  53. Wang, P.; Yin, K.; Zhang, M.; Zheng, Y.; Zhang, T.; Kang, Y.; Feng, X. The Effect of Incorporating Large Language Models into the Teaching on Critical Thinking Disposition: An “AI + Constructivism Learning Theory” Attempt. Educ. Inf. Technol. 2025, 30, 11625–11647. [Google Scholar] [CrossRef]
  54. Zhang, Y.; Lai, X.; Yi, S.; Lu, Y. Does ChatGPT-Based Reading Platform Impact Foreign Language Paper Reading? Evidence from a Quasi-Experimental Study on Chinese Undergraduate Students. Educ. Inf. Technol. 2025, 30, 9737–9754. [Google Scholar] [CrossRef]
  55. Chen, X.; Jia, B.; Peng, X.; Zhao, H.; Yao, J.; Wang, Z.; Zhu, S. Effects of ChatGPT and Argument Map(AM)-Supported Online Argumentation on College Students’ Critical Thinking Skills and Perceptions. Educ. Inf. Technol. 2025, 30, 17623–17658. [Google Scholar] [CrossRef]
  56. De La Puente, M.; Torres, J.; Troncoso, A.L.B.; Meza, Y.Y.H.; Carrascal, J.X.M. Investigating the Use of chatGPT as a Tool for Enhancing Critical Thinking and Argumentation Skills in International Relations Debates among Undergraduate Students. Smart Learn. Environ. 2024, 11, 55. [Google Scholar] [CrossRef]
  57. Huang, Y.-M.; Chen, P.-H.; Lee, H.-Y.; Sandnes, F.E.; Wu, T.-T. ChatGPT-Enhanced Mobile Instant Messaging in Online Learning: Effects on Student Outcomes and Perceptions. Comput. Hum. Behav. 2025, 168, 108659. [Google Scholar] [CrossRef]
  58. Lin, C.-J.; Lee, H.-Y.; Wang, W.-S.; Huang, Y.-M.; Wu, T.-T. Enhancing Reflective Thinking in STEM Education through Experiential Learning: The Role of Generative AI as a Learning Aid. Educ. Inf. Technol. 2025, 30, 6315–6337. [Google Scholar] [CrossRef]
  59. Shalong, W.; Yi, Z.; Bin, Z.; Ganglei, L.; Jinyu, Z.; Yanwen, Z.; Zequn, Z.; Lianwen, Y.; Feng, R. Enhancing Self-Directed Learning with Custom GPT AI Facilitation among Medical Students: A Randomized Controlled Trial. Med. Teach. 2024, 47, 1126–1133. [Google Scholar] [CrossRef] [PubMed]
  60. Singh, A.; Brooks, C.; Wang, X.; Li, W.; Kim, J.; Wilson, D. Bridging Learnersourcing and AI: Exploring the Dynamics of Student-AI Collaborative Feedback Generation. In Proceedings of the 14th Learning Analytics and Knowledge Conference, Kyoto, Japan, 18–22 March 2024; ACM: Kyoto Japan, 2024; pp. 742–748. [Google Scholar]
  61. Wang, W.-S.; Lin, C.-J.; Lee, H.-Y.; Huang, Y.-M.; Wu, T.-T. Integrating Feedback Mechanisms and ChatGPT for VR-Based Experiential Learning: Impacts on Reflective Thinking and AIoT Physical Hands-on Tasks. Interact. Learn. Environ. 2025, 33, 1770–1787. [Google Scholar] [CrossRef]
  62. Aure, P.A.; Cuenca, O. Fostering Social-Emotional Learning through Human-Centered Use of Generative AI in Business Research Education: An Insider Case Study. J. Res. Innov. Teach. Learn. JRIT 2024, 17, 168–181. [Google Scholar] [CrossRef]
  63. Aydın Yıldız, T. Exploring the Impact of ChatGPT on Improving 21st-Century Skills for Future English Teachers During Lesson Planning. Comput. Sch. 2024, 42, 119–142. [Google Scholar] [CrossRef]
  64. Datskiv, O.; Zadorozhna, I.; Shon, O. Developing Future Teachers’ Academic Writing and Critical Thinking Skills Using ChatGPT. Int. J. Emerg. Technol. Learn. 2024, 19, 126–136. [Google Scholar] [CrossRef]
  65. Shen, Y.; Chen, L. ‘Critical Chatting’ or ‘Casual Cheating’: How Graduate Efl Students Utilize Chatgpt for Academic Writing. Comput. Assist. Lang. Learn. 2025, 1–29. [Google Scholar] [CrossRef]
  66. Jiang, X.; Li, J.; Chen, C.H. Enhancing Critical Thinking Skills with ChatGPT-Powered Activities in Chinese Language Classrooms. J. Chin. Lang. Teach. 2024, 5, 47–73. [Google Scholar]
  67. Dilekli, Y.; Boyraz, S. From “Can AI Think?” to “Can AI Help Thinking Deeper?”: Is Use of ChatGPT in Higher Education a Tool of Transformation or Fraud? Int. J. Mod. Educ. Stud. 2024, 8, 49–71. [Google Scholar] [CrossRef]
  68. Chen, M.-S.; Hsu, T.-P.; Hsu, T.-C. GAI-Assisted Personal Discussion Process Analysis. In Innovative Technologies and Learning; Lecture Notes in Computer Science; Springer Nature Switzerland: Cham, Switzerland, 2024; pp. 194–204. ISBN 978-3-031-65883-9. [Google Scholar]
  69. Fabio, R.A.; Plebe, A.; Suriano, R. Ai-Based Chatbot Interactions and Critical Thinking Skills: An Exploratory Study. Curr. Psychol. J. Divers. Perspect. Divers. Psychol. Issues 2024, 44, 8082–8095. [Google Scholar] [CrossRef]
  70. Pellas, N. The Role of Students’ Higher-Order Thinking Skills in the Relationship between Academic Achievements and Machine Learning Using Generative AI Chatbots. Res. Pract. Technol. Enhanc. Learn. 2024, 20, 36. [Google Scholar] [CrossRef]
  71. Wu, C.-H.; Weng, T.-S.; Liu, C.-H. Exploring ChatGPT’s Potential to Enhance Problem-Solving and Critical Thinking in Education. Educ. Technol. Soc. 2025, 28, 310–326. [Google Scholar]
  72. Avsheniuk, N.; Lutsenko, O.; Svyrydiuk, T.; Seminikhyna, N. Empowering Language Learners’ Critical Thinking: Evaluating ChatGPT’s Role in English Course Implementation. Arab World Engl. J. 2024, 1, 210–224. [Google Scholar] [CrossRef]
  73. Barana, A.; Marchisio, M.; Roman, F. Fostering Problem Solving and Critical Thinking in Mathematics through Generative Artificial Intelligence. In Proceedings of the 20th International Conference on Cognition and Exploratory Learning in Digital Age (CELDA 2023), Funchal, Portugal, 21–23 October 2023. [Google Scholar]
  74. Sharma, D. Critical Thinking and Problem-Solving in the Age of ChatGPT: Experiential-Bibliotherapy-Blogging Project. Bus. Prof. Commun. Q. 2024, 87, 630–653. [Google Scholar] [CrossRef]
  75. Drushlyak, M.; Lukashova, T.; Shamonia, V.; Semenikhina, O. ChatGPT-Based Simulation Helps to Develop the Pre-Service Mathematics Teachers’ Critical Thinking. Int. J. Instr. 2025, 18, 153–172. [Google Scholar] [CrossRef]
  76. Exintaris, B.; Karunaratne, N.; Yuriev, E. Metacognition and Critical Thinking: Using ChatGPT-Generated Responses as Prompts for Critique in a Problem-Solving Workshop (SMARTCHEMPer). J. Chem. Educ. 2023, 100, 2972–2980. [Google Scholar] [CrossRef]
  77. Chu, H.C.; Lu, Y.C.; Tu, Y.F. How GenAI-Supported Multi-Modal Presentations Benefit Students with Different Motivation Levels: Evidence from Digital Storytelling Performance, Critical Thinking Awareness, and Learning Attitude. Educ. Technol. Soc. 2025, 28, 250–269. [Google Scholar]
  78. Kim, D.; Majdara, A.; Olson, W. A Pilot Study Inquiring into the Impact of ChatGPT on Lab Report Writing in Introductory Engineering Labs. Int. J. Technol. Educ. IJTE 2024, 7, 259–289. [Google Scholar] [CrossRef]
  79. Michalon, B.; Camacho-Zuñiga, C. ChatGPT, a Brand-New Tool to Strengthen Timeless Competencies. Front. Educ. 2023, 8, 1251163. [Google Scholar] [CrossRef]
  80. Gonsalves, C. Generative AI’s Impact on Critical Thinking: Revisiting Bloom’s Taxonomy. J. Mark. Educ. 2024, 48, 02734753241305980. [Google Scholar] [CrossRef]
  81. Huang, K.; Liu, Y.; Dong, M. Incorporating AIGC into Design Ideation: A Study on Self-Efficacy and Learning Experience Acceptance under Higher-Order Thinking. Think. Ski. Creat. 2024, 52, 101508. [Google Scholar] [CrossRef]
  82. Kartal, G. The Influence of ChatGPT on Thinking Skills and Creativity of EFL Student Teachers: A Narrative Inquiry. J. Educ. Teach. 2024, 50, 627–642. [Google Scholar] [CrossRef]
  83. Li, Y.; Keung, J.; Ma, X. Integrating Generative AI in Software Engineering Education: Practical Strategies. In Proceedings of the 2024 International Symposium on Educational Technology (ISET), Macau, Macao, 29 July–1 August 2024; IEEE: Macau, Macao; pp. 49–53.
  84. Lin, C.-H.; Zhou, K.; Li, L.; Sun, L. Integrating Generative AI into Digital Multimodal Composition: A Study of Multicultural Second-Language Classrooms. Comput. Compos. 2025, 75, 102895. [Google Scholar] [CrossRef]
  85. Sanchez-Guerrero, J.; Miranda-Lopez, X.; Buenano, H.; Lozada-Miranda, D. ChatGPT to Motivate Critical Thinking in the Teaching–Learning Process of First Semester Students at the Technical University of Ambato; Springer Science and Business Media Deutschland GmbH: Bogota, Colombia, 2024; Volume 380, pp. 169–179. [Google Scholar]
  86. Chang, C.-Y.; Yang, C.-L.; Jen, H.-J.; Ogata, H.; Hwang, G.-H. Facilitating Nursing and Health Education by Incorporating ChatGPT into Learning Designs. Educ. Technol. Soc. 2024, 27, 215–230. [Google Scholar]
  87. Hwang, G.-J.; Cheng, P.-Y.; Chang, C.-Y. Facilitating Students’ Critical Thinking, Metacognition and Problem-Solving Tendencies in Geriatric Nursing Class: A Mixed-Method Study. Nurse Educ. Pract. 2025, 83, 104112. [Google Scholar] [CrossRef]
  88. Liu, W.; Cui, X. Generative AI-Assisted Collaborative Argumentation: Implications for the Argumentation Process and Outcome; IEEE: Piscataway, NJ, USA, 2024; pp. 220–225. [Google Scholar]
  89. Farinetti, L.; Canale, L. Chatbot Development Using LangChain: A Case Study to Foster Critical Thinking and Creativity; Association for Computing Machinery: Milan, Italy, 2024; Volume 1, pp. 401–407. [Google Scholar]
  90. Zhang, J. Generative AI in Higher Education: Challenges and Opportunities for Course Learning. Adv. Soc. Sci. Res. J. 2025, 12, 11–18. [Google Scholar] [CrossRef]
  91. Wang, W.; Lin, C.; Lee, H.; Huang, Y.; Wu, T. Enhancing Self-Regulated Learning and Higher-Order Thinking Skills in Virtual Reality: The Impact of ChatGPT-Integrated Feedback Aids. Educ. Inf. Technol. 2025, 30, 19419–19445. [Google Scholar] [CrossRef]
  92. Palupi, S.; Hardi, R.; Pribadi, A.S.; Zulkarnain, R.; Setiawan, M.N.; Sari, N.W.W.; Utomo, D.T. Enhancing Computational Strategies to Decode ChatGPT’s Influence on the Critical Thinking Abilities of University Students. In Proceedings of the 2024 4th International Conference of Science and Information Technology in Smart Administration (ICSINTESA), Balikpapan, Indonesia, 12 July 2024; IEEE: Balikpapan, Indonesia; pp. 27–32.
  93. Baláž, N.; Porubän, J.; Horváth, M.; Kormaník, T. Using ChatGPT During Implementation of Programs in Education. In OpenAccess Series in Informatics (OASIcs); Schloss Dagstuhl–Leibniz-Zentrum für Informatik: Wadern, Germany, 2024; Volume 122, pp. 18:1–18:9. [Google Scholar] [CrossRef]
  94. Santamaría-Velasco, J.; Núñez-Naranjo, A.; Morales-Urrutia, X. Critical Thinking and AI: Enhancing History Teaching through ChatGPT Simulations. Int. J. Innov. Res. Sci. Stud. 2025, 8, 564–575. [Google Scholar] [CrossRef]
  95. Yusuf, A.; Bello, S.; Pervin, N.; Tukur, A.K. Implementing a Proposed Framework for Enhancing Critical Thinking Skills in Synthesizing AI-Generated Texts. Think. Ski. Creat. 2024, 53, 101625. [Google Scholar] [CrossRef]
  96. Essel, H.B.; Vlachopoulos, D.; Essuman, A.B.; Amankwa, J.O. ChatGPT Effects on Cognitive Skills of Undergraduate Students: Receiving Instant Responses from AI-Based Conversational Large Language Models (LLMs). Comput. Educ. Artif. Intell. 2024, 6, 100198. [Google Scholar] [CrossRef]
  97. Moustaghfir, S.; Brigui, H. Navigating Critical Thinking in the Digital Era: An Informative Exploration. Int. J. Linguist. Lit. Transl. 2024, 7, 137–143. [Google Scholar] [CrossRef]
  98. Díaz, B.; Delgado, C. Artificial Intelligence: Tool or Teammate? J. Res. Sci. Teach. 2024, 61, 2575–2584. [Google Scholar] [CrossRef]
  99. Singh, A.; Taneja, K.; Guan, Z.; Ghosh, A. Protecting Human Cognition in the Age of AI. arXiv 2025, arXiv:2502.12447. [Google Scholar] [CrossRef]
  100. Ji, Y.; Zhan, Z.; Li, T.; Zou, X.; Lyu, S. Human-Machine Co-Creation: The Effects of ChatGPT on Students’ Learning Performance, AI Awareness, Critical Thinking, and Cognitive Load in a STEM Course towards Entrepreneurship. IEEE Trans. Learn. Technol. 2025, 18, 402–415. [Google Scholar] [CrossRef]
  101. Huang, C.-W.; Coleman, M.; Gachago, D.; Van Belle, J.-P. Using ChatGPT to Encourage Critical AI Literacy Skills and for Assessment in Higher Education; Springer Science and Business Media Deutschland GmbH: Gauteng, South Africa, 2024; Volume 1862, pp. 105–118. [Google Scholar]
  102. Bond, M.; Khosravi, H.; De Laat, M.; Bergdahl, N.; Negrea, V.; Oxley, E.; Pham, P.; Chong, S.W.; Siemens, G. A Meta Systematic Review of Artificial Intelligence in Higher Education: A Call for Increased Ethics, Collaboration, and Rigour. Int. J. Educ. Technol. High. Educ. 2024, 21, 4. [Google Scholar] [CrossRef]
  103. Abu Saa, A.; Al-Emran, M.; Shaalan, K. Factors Affecting Students’ Performance in Higher Education: A Systematic Review of Predictive Data Mining Techniques. Technol. Knowl. Learn. 2019, 24, 567–598. [Google Scholar] [CrossRef]
  104. Premkumar, P.P.; Yatigammana, M.R.K.N.; Kannangara, S. Impact of Generative AI on Critical Thinking Skills in Undergraduates: A Systematic Review. J. Desk Res. Rev. Anal. 2024, 2, 199–215. [Google Scholar] [CrossRef]
  105. Perifanou, M.; Economides, A.A. Collaborative Uses of GenAI Tools in Project-Based Learning. Educ. Sci. 2025, 15, 354. [Google Scholar] [CrossRef]
  106. Jin, Y.; Yan, L.; Echeverria, V.; Gašević, D.; Martinez-Maldonado, R. Generative AI in Higher Education: A Global Perspective of Institutional Adoption Policies and Guidelines. Comput. Educ. Artif. Intell. 2025, 8, 100348. [Google Scholar] [CrossRef]
  107. Balart, T.S.; Shryock, K.J. A Framework for Integrating Artificial General Intelligence into Engineering Education. Int. J. Eng. Educ. 2025, 41, 171–194. [Google Scholar]
Figure 1. CT skills identified by Facione [2,3].
Figure 1. CT skills identified by Facione [2,3].
Algorithms 19 00179 g001
Figure 2. Article search and selection process.
Figure 2. Article search and selection process.
Algorithms 19 00179 g002
Figure 3. Distribution of research paradigms.
Figure 3. Distribution of research paradigms.
Algorithms 19 00179 g003
Table 1. Search terms used.
Table 1. Search terms used.
CategoryTerms Used
GenAI Tools“Generative Artificial Intelligence” OR “GenAI” OR “Generative AI” OR “GAI” OR “ChatGPT” OR “chat generative pre-trained transformer” OR “GPT-3.5” OR “GPT-4” OR “GPT-4o” OR “generative model” OR “artificial intelligence generated content” OR “AIGC” OR “AI-generated”
Critical Thinking“critical thinking” OR “CT” OR “thinking skills” OR “critical literacy”
Educational Context“higher education” OR “student” OR “universit*” OR “postgrad*” OR “undergrad*” OR “sophomore” OR “college” OR “course” OR “freshman” OR “tertiary” OR “post-secondary education” OR “pupil” OR “teacher” OR “lecturer” OR “professor” OR “faculty” OR “Instructor”
LanguageEnglish
Date RangeJanuary 2023–April 2025
* The asterisk denotes a truncation/wildcard operator used to capture word variants (e.g., “universit*” retrieves university/universities; “undergrad*” retrieves undergraduate/undergraduates; “postgrad*” retrieves postgraduate/postgraduates).
Table 2. Results of database searches.
Table 2. Results of database searches.
CategoryTerms Used
Compendex & Inspec (Elsevier)386
Web of Science311
Education Source + ERIC + APA PsycInfo (EBSCO)521
Scopus570
Total1788
(800 duplicates removed)
988
(After duplicates extraction)
Table 3. Screening criteria.
Table 3. Screening criteria.
Inclusion CriteriaExclusion Criteria
- Studies published between January 2023 and April 2025, to capture recent developments in GenAI applications.
- Articles written in English.
- Research conducted in higher education settings (e.g., undergraduate or graduate students).
- Studies addressing GenAI tools.
- Empirical studies that include measurable outcomes or observations related to critical thinking skills, including dimensions such as analysis, evaluation, inference, metacognition, ethical reasoning, or self-regulation.
Theoretical papers, conceptual frameworks, opinion pieces, or editorials without empirical data.
- Studies analyzing perspectives on GenAI.
- Studies focused solely on K-12 or informal education contexts.
- Studies examining student interaction with GenAI tools such as ChatGPT, large language models, or other GenAI systems used in educational contexts that do not explicitly examine CT or related constructs as an outcome or objective.
- Studies examining AI in education not related to generative technologies (e.g., chatbots, personalized learning, predictive analytics, automated grading).
- Research not involving student participants (e.g., faculty-only perspectives, professional development, system design papers).
Table 4. Studies designs distribution.
Table 4. Studies designs distribution.
Study DesignNumber of StudiesExamples
Quasi-experimental (pre-post with comparison group)19[26,45,46,47,48,49,50,51,52,53,54]
Randomized controlled trial11[28,55,56,57,58,59,60,61]
Case study6[62,63,64,65,66,67]
Observational survey6[14,15,68,69,70,71]
Pre-experimental
(single-group pre-post)
10[72,73,74,75,76,77,78]
Action research3[13,66,79]
Others (e.g., narrative inquiry)8[32,80,81,82]
Table 5. Main typologies of integration.
Table 5. Main typologies of integration.
Four Main Typologies of IntegrationExplanationExamples
Feedback prompts (AI-based metacognitive scaffold)AI-generated formative quizzes, automated scoring, or real-time feedback prompts (e.g., GPTFAS in VR, test-style assessments, auto-scored reflections).[26,32,45,55,68,86,91]
Dialogue simulation and reflectionConversational agents taking on roles (e.g., regulator, peer, expert) to challenge student reasoning, support perspective-taking, and surface hidden biases.
Structured activities in which students and instructors jointly interpret AI outputs, fostering dialogic meaning-making and shared metacognition.
[49,62,77,88,92]
AI-Based peer reviewMulti-week PBL interventions where ChatGPT guides students through project phases (idea generation, analysis, report drafting).
GenAI tools that supply exemplar feedback, suggested criteria, or consistency checks to help students critique one another’s work more rigorously. Help with coding tasks.
[13,26,56,65,71,80]
Critical Engagement with AI-Generated ContentTasks requiring students to audit, correct, or improve AI-generated outputs, spotting “hallucinations”, logical flaws, or unsupported claims.
Structured activities in which students and instructors jointly interpret AI outputs, fostering dialogic. meaning-making and shared metacognition.
[65,68,74,75,76,90]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Balart, T.; Díaz, B.; Shryock, K. A Systematic Literature Review on the Pedagogical Implications and Impact of GenAI on Students’ Critical Thinking. Algorithms 2026, 19, 179. https://doi.org/10.3390/a19030179

AMA Style

Balart T, Díaz B, Shryock K. A Systematic Literature Review on the Pedagogical Implications and Impact of GenAI on Students’ Critical Thinking. Algorithms. 2026; 19(3):179. https://doi.org/10.3390/a19030179

Chicago/Turabian Style

Balart, Trini, Brayan Díaz, and Kristi Shryock. 2026. "A Systematic Literature Review on the Pedagogical Implications and Impact of GenAI on Students’ Critical Thinking" Algorithms 19, no. 3: 179. https://doi.org/10.3390/a19030179

APA Style

Balart, T., Díaz, B., & Shryock, K. (2026). A Systematic Literature Review on the Pedagogical Implications and Impact of GenAI on Students’ Critical Thinking. Algorithms, 19(3), 179. https://doi.org/10.3390/a19030179

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop