Previous Article in Journal
Discrepancies Between Self-Reported and Peer-Attributed Frequencies of AI-Assisted Academic Cheating Among Undergraduate Students
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Redesigning STEM Higher Education in the Era of Generative AI: From Curriculum Design to Classroom Practice

by
Christos Papaneophytou
and
Stella A. Nicolaou
*
Department of Life Sciences, School of Life and Health Sciences, University of Nicosia, 2417 Nicosia, Cyprus
*
Author to whom correspondence should be addressed.
Trends High. Educ. 2026, 5(3), 84; https://doi.org/10.3390/higheredu5030084
Submission received: 7 July 2026 / Revised: 11 August 2026 / Accepted: 21 August 2026 / Published: 26 August 2026

Abstract

Generative artificial intelligence (GenAI) has moved from an emerging educational tool to a structural challenge for science, technology, engineering, and mathematics (STEM) higher education. This narrative review argues that the most consequential effect of GenAI is not the automation of existing teaching practices but the need to redesign curricula, learning outcomes, pedagogies, and assessment around disciplinary judgment, critical verification, intellectual independence, and transparent, ethical use of GenAI. Its distinctive contribution is to frame GenAI as a problem of curriculum and assessment validity rather than primarily as a question of tool adoption or academic integrity. Because widely available systems can generate code, solve quantitative problems, summarize literature, draft laboratory reports, and produce fluent scientific prose, conventional submitted artifacts have become weaker indicators of the reasoning and competence they are intended to demonstrate. The review therefore examines the full programme-to-classroom pathway, connecting definitions of graduate competence with course design, classroom and laboratory practice, assessment, feedback, faculty capability, technology adoption, and iterative evaluation. The analysis integrates cognitive load theory, constructive alignment, constructivist perspectives, and frameworks of faculty capability and technology adoption. The biological sciences serve as a recurring disciplinary case because they combine conceptual knowledge, laboratory practice, computational analysis, and ethical decision-making, and are also being transformed by AI-based scientific methods. A worked cell biology example, structured using the Analysis, Design, Development, Implementation, and Evaluation model, operationalizes the review’s conceptual argument and demonstrates how GenAI integration can translate into needs analysis, outcome specification, resource development, blended laboratory implementation, assessment, and iterative redesign. The resulting design logic is generalized into a transferable five-step template for STEM curriculum redesign, with recommendations at programme, course, and institutional levels.

1. Introduction

The integration of artificial intelligence (AI) into higher education has evolved from a gradual technological development into a structural transformation, driven in particular by the widespread availability and rapid adoption of generative AI systems based on large language models [1]. This transformation is especially significant for science, technology, engineering, and mathematics (STEM) education. STEM programmes have traditionally relied on cumulative conceptual knowledge, quantitative and analytical reasoning, experimental and computational competence, and assessment practices that assume a clear boundary between students’ own work and contributions from external tools [2]. Generative AI (GenAI) challenges these assumptions simultaneously. Systems capable of generating code, solving or explaining quantitative problems, synthesizing scientific literature, proposing experimental approaches, and producing fluent scientific prose can alter not only how students complete academic tasks but also how disciplinary competence is defined, developed, and demonstrated [3].
The educational implications of AI are neither uniform across STEM disciplines nor fully understood. Its advantages and limitations vary substantially among proof-intensive mathematics, code-centered computer science, design-oriented engineering, and the experimental and increasingly data-intensive natural sciences [4]. The biological sciences offer a particularly useful case because the discipline occupies a distinctive dual position: AI is at once a tool for teaching and learning biology and a transformative methodological force within biological research [5]. Teaching biology in the era of AI therefore requires students to learn both how to use AI responsibly as an educational and scientific tool and how to understand, evaluate, and question the AI-based methods that are reshaping their discipline [6].
The early literature on generative AI in higher education has often wavered between enthusiasm for its transformative potential and concern about its disruptive effects. Moreover, much of the available evidence comes from small-scale pilot studies, single-course or single-cohort experiences, self-reports, and conceptual commentaries rather than longitudinal or comparative investigations of learning effectiveness [7]. Two concerns are particularly prominent. The first concerns academic integrity and assessment validity. Generative AI systems’ ability to produce plausible, contextually appropriate responses, combined with the limited reliability of AI-detection tools, has exposed the vulnerability of conventional assessment formats [8]. The second concerns institutional and professional capacity, including changes in the role of instructors, faculty-development requirements, workload, unequal access to advanced tools, algorithmic bias, data privacy, intellectual property, and the environmental costs associated with training and operating large-scale AI systems [9]. Together, these tensions suggest that the most durable educational consequences of AI may lie not in accelerating the production or delivery of content, but in redefining what learners are expected to know, what they must be able to do independently or in collaboration with AI, and how institutions can establish that such competence has genuinely been achieved [10].
Existing reviews have begun to map this rapidly developing field, but important gaps remain. Some treat STEM as largely homogeneous, obscuring substantial differences in disciplinary epistemologies, professional practices, and modes of assessment. Others focus on a single component of the educational process, most commonly classroom use or assessment, without examining alignment among programme-level curriculum objectives, course and laboratory design, teaching activities, assessment, feedback, and evaluation. Furthermore, relatively few reviews ground their analyses in established accounts of learning, curriculum alignment, faculty capability, and technology adoption when examining where AI enters the educational process and how changes at one stage affect the others [11,12].
Recent reviews have also established both the rapid expansion of AI applications in education and persistent limitations in the available empirical evidence. Almasri [13] synthesized empirical research on AI-supported science education, highlighting reported learning benefits alongside implementation challenges and methodological limitations. Yusuf et al. [14] mapped the broader GenAI literature across education and research and identified a rapidly developing but methodologically heterogeneous evidence base. These reviews reinforce the need for rigorous, course- and outcome-specific evaluation of AI-supported educational interventions. However, a prior design question remains insufficiently resolved: how broad calls for curriculum change can be translated into a theoretically grounded, constructively aligned, and testable redesign extending from programme-level definitions of competence to course activities, assessment, implementation, and evaluation.
The present review reframes GenAI integration as a problem of curriculum and assessment validity rather than primarily as a question of tool adoption or academic integrity. It connects programme-level definitions of competence with learning outcomes, course design, classroom and laboratory practice, assessment, feedback, faculty capability, institutional adoption, and iterative evaluation. It integrates conceptual lenses operating at different levels, so that design decisions follow from established accounts of learning, alignment, capability, and adoption rather than from technological novelty. It further distinguishes principles shared across STEM from requirements that must remain discipline-specific, translating the synthesis into an evidence-informed worked example and a transferable redesign process. These four contributions form the basis of the organisation of the paper. Specifically, the curriculum and assessment-validity problem is established in Section 4, the programme-to-classroom redesign pathway is developed in Section 5, the integration of conceptual frameworks is applied decision-by-decision in the worked example of Section 6, and the separation of shared from discipline-specific requirements is set out in Section 7 and Section 8.
Accordingly, this narrative review examines GenAI-driven change across the full instructional pipeline, from curriculum and course design to classroom and laboratory practice, assessment, feedback, evaluation, and iterative redesign. The analysis is organized through established conceptual and pedagogical lenses, including cognitive load theory, constructive alignment, constructivist perspectives, and frameworks of faculty capability and technology adoption, particularly Technological Pedagogical Content Knowledge (TPACK) and the Unified Theory of Acceptance and Use of Technology (UTAUT) [15]. The biological sciences are used throughout as a recurring disciplinary case to translate general principles into practice. The review culminates in a fully worked example of cell biology course design, in which the Analysis, Design, Development, Implementation, and Evaluation (ADDIE) model provides the instructional-design scaffold for developing, implementing, evaluating, and iteratively refining a GenAI-integrated course [16].
In this review, artificial intelligence (AI) is used as an umbrella term encompassing generative and non-generative systems, including machine learning, learning analytics, intelligent tutoring systems, automated assessment, and adaptive or virtual-learning technologies. Generative artificial intelligence (GenAI) refers specifically to models that generate new content in response to user input. Within GenAI, this review distinguishes text-generating large language models and conversational agents, code-generating models, and multimodal systems capable of processing or generating combinations of text, images, and other data. These categories have different educational affordances and risks: text-generating systems primarily affect writing, explanation, feedback, and literature synthesis; code-generating systems affect programming and data-analysis tasks; and multimodal systems may influence visual interpretation, scientific representation, and laboratory preparation. The terms AI, GenAI, and LLM are therefore used according to the technological scope of the evidence or application being discussed.

2. Conceptual Framework and Review Approach

This section establishes the conceptual and methodological foundations of the review. It first introduces the theoretical and pedagogical lenses used consistently throughout the analysis, ensuring that the discussion extends beyond a catalog of AI tools and capabilities. It then explains how the ADDIE model is used as the instructional-design scaffold for the biological sciences worked example, in relation to alternative instructional-design approaches. Finally, it defines the scope and methodology of the narrative review to make transparent the basis of the arguments developed in subsequent sections.

2.1. Conceptual and Pedagogical Lenses

A recurrent limitation of the emerging literature on AI in education is that descriptions of technological capabilities are not always grounded in established accounts of learning, curriculum alignment, cognitive development, or technology adoption [17]. To address this limitation, the present review draws principally on cognitive load theory and constructive alignment. Constructivist perspectives, the TPACK framework, and the UTAUT provide additional support for specific dimensions of the analysis, with TPACK and UTAUT used together to address faculty capability and adoption. These lenses operate at deliberately different levels: cognitive load theory at the level of the individual learner’s mental processing, constructive alignment at the level of curriculum coherence, constructivism at the level of pedagogical design, and TPACK and UTAUT at the levels of professional knowledge and adoption behaviour respectively. The revised Bloom’s taxonomy is used as a supporting framework, primarily in the Design phase of the biological sciences worked example, where it helps classify intended learning outcomes and distinguish the cognitive processes students are expected to demonstrate from the apparent complexity of AI-generated outputs.

2.1.1. Cognitive Load Theory

Cognitive load theory distinguishes intrinsic cognitive load, arising from the inherent complexity and element interactivity of the material being learned, from extraneous cognitive load created by the way information, instructions, or tasks are presented [18]. It also considers how limited working-memory resources are directed towards constructing and refining durable knowledge structures. In GenAI-supported STEM education, this distinction is especially important because the same technology may either support learning by reducing avoidable demands or undermine it by performing the reasoning processes that students are expected to develop and demonstrate.
AI may reduce unnecessary cognitive demands by clarifying instructions, generating worked examples, supporting routine calculations, or providing immediate explanations. Such support may allow learners to direct greater attention towards conceptual relationships, problem-solving strategies, and disciplinary interpretation. However, GenAI interaction may also increase cognitive demands. Students may need to formulate and iteratively refine prompts, retain the original problem and evaluation criteria in working memory, compare alternative outputs, verify claims and sources, identify hallucinations or unsupported assumptions, and resolve conflicting responses. These demands may be productive when prompting and verification are themselves intended learning outcomes, but they may constitute extraneous load when they arise from unfamiliar interfaces, unclear instructions, poorly structured outputs, or tool-management requirements unrelated to the disciplinary objective.
The intrinsic cognitive load of a STEM task is similarly dependent on the interaction between task complexity and the learner’s prior knowledge rather than on the topic alone. For a novice, unfamiliar scientific terminology, symbols, and representations may impose substantial demands because the relevant concepts have not yet been organized into established knowledge structures. Reasoning about causal pathways may create greater element interactivity because learners must simultaneously coordinate mechanisms, variables, temporal relationships, alternative explanations, confounding factors, and the strength of the available evidence. For a more experienced learner, established schemas may reduce these demands. Critical thinking should therefore not be treated as a separate and uniformly more demanding curriculum component; its cognitive demands depend on the disciplinary knowledge, number of interacting elements, and reasoning processes required by the particular task.
The relevant question is consequently not simply whether AI makes a task easier or harder, but which forms of cognitive effort it removes, which it introduces, which it preserves, and which it redirects [19,20]. This distinction is especially important in STEM disciplines, where apparent task completion may conceal weak conceptual understanding. A student may use AI to produce executable code, solve a quantitative problem, interpret a graph, or draft an experimental explanation without developing the knowledge required to verify the result. Cognitive load theory therefore provides a basis for distinguishing beneficial scaffolding from premature or excessive cognitive outsourcing [18,21].
In this review, cognitive load theory is used as an analytical and design framework for determining which forms of GenAI support reduce avoidable processing demands, which introduce additional demands, and which risk displacing the cognitive activity necessary for learning. AI-supported activities are considered educationally appropriate when they manage task complexity and redirect effort towards disciplinary interpretation, verification, and justification. Independent performance and carefully sequenced scaffolding are retained where terminology, conceptual relationships, procedural fluency, or causal reasoning must first be developed to support subsequent judgement. This distinction informs the later analysis of learner needs, content sequencing, learning activities, assessment design, and evaluation of the illustrative course.

2.1.2. Constructive Alignment and Constructivism

Constructive alignment provides the principal bridge between the learning-science arguments and curriculum redesign. It requires coherence among intended learning outcomes, teaching and learning activities, and assessment [22]. In an AI-supported curriculum, a fourth element must be made explicit: the role assigned to AI within each activity and assessment [23].
Constructive alignment is used throughout this review to connect program-level graduate attributes with course-level outcomes, learning activities, laboratory practice, feedback, and assessment. It also supports the distinction between activities in which students must demonstrate independent foundational competence and those in which effective human-AI collaboration forms part of authentic disciplinary practice.
Constructivism provides a complementary account of how the aligned activities produce learning. Constructivist perspectives hold that learners actively develop understanding by connecting new information with prior knowledge, testing explanations, and revising their conceptual structures [24]. From this perspective, AI is most educationally valuable when it supports questioning, comparison, explanation, experimentation, and reflection rather than supplying finished answers for passive acceptance. Constructionism extends this principle to the creation of tangible artifacts, such as code, models, visualizations, experimental plans, and scientific explanations [25]. The educational value of AI-supported creation lies not only in the final product but also in the cycles of design, critique, verification, revision, and justification through which the product is developed. Read together, constructive alignment specifies what must cohere, while constructivism and constructionism specify how the aligned learning experiences should be designed so that AI scaffolds active knowledge construction rather than displacing it.
In the subsequent analysis, these perspectives inform activities in which students compare AI-generated explanations with disciplinary evidence, identify errors or unsupported assumptions, revise generated artefacts, and justify the resulting decisions. The emphasis is therefore not simply on producing an AI-assisted product, but on making the process of knowledge construction visible through questioning, verification, reflection, and iterative refinement. These principles inform the later discussion of learning activities, laboratory design, process-oriented assessment, and the Development phase of the illustrative course.

2.1.3. Faculty Capability and Adoption: TPACK and UTAUT

Faculty capability and adoption are analysed through two complementary frameworks. The TPACK framework emphasizes that effective AI integration requires more than technical familiarity [26]. Educators must understand how the technology interacts with pedagogical strategies and with the specific knowledge practices of their discipline. This is particularly important in STEM, where technically fluent AI output may nevertheless be conceptually, quantitatively, or experimentally invalid [27].
TPACK explains what educators must know to integrate AI responsibly, but it does not, by itself, explain whether they will choose to adopt it. The UTAUT addresses this gap by holding that behavioural intention and use are shaped by four core determinants: performance expectancy, effort expectancy, social influence, and facilitating conditions, with the influence of these determinants moderated by factors such as experience [28]. UTAUT is preferred here over earlier acceptance models because its inclusion of social influence and facilitating conditions aligns with the institutional, governance, and equity emphases of this review, in which adoption depends not only on individual perceptions but also on support structures, access to tools, and professional norms.
Used together, TPACK and UTAUT frame faculty development as a dual problem: building the technological, pedagogical, and disciplinary knowledge required for sound integration, and creating the conditions of performance expectancy, effort expectancy, social support, and facilitating infrastructure under which educators are willing and able to adopt AI tools in practice. Across the framework as a whole, each lens performs a distinct function, from the cognitive demands of learning through curriculum coherence and active knowledge construction to professional knowledge and adoption, while the revised Bloom’s taxonomy is reserved for the worked example, where it classifies the intended learning outcomes of the illustrative course.

3. Methods

This review is reported as a narrative review, and its design, conduct, and reporting were guided by the Scale for the Assessment of Narrative Review Articles (SANRA) [29]. The subsections below are organized to address these criteria explicitly (Supplementary Table S1).

3.1. Rationale and Importance of the Review

A narrative approach was chosen because research on AI in higher education is rapidly evolving, conceptually diverse, and distributed across educational research, discipline-based education, computer science, ethics, and institutional policy. Much of the relevant evidence consists of small-scale pilots, single-cohort studies, conceptual commentaries, and authoritative sector guidance that a narrowly defined systematic review would either exclude or fail to integrate. A narrative synthesis is therefore better suited to the present aim, which is to connect empirical studies, conceptual analyses, reviews, and policy guidance into a coherent, theoretically grounded account of curriculum transformation across the programme-to-classroom pipeline.

3.2. Aims and Guiding Questions

The review is structured around five research questions:
  • Does the dissemination of GenAI create a need for curriculum redesign in STEM higher education?
  • If such a need exists, can it be met through incremental, classroom-level adjustment, or is it structural and therefore located at the curriculum level?
  • To what extent can a redesign be shared across STEM disciplines, and where must it be tailored to the distinct epistemic, practical, and assessment traditions of each field?
  • How should curriculum and course design respond in the era of GenAI, at the levels of needs analysis, learning outcomes, constructive alignment, and programme design?
  • How can generative AI tools be applied within the biological sciences, as demonstrated through an ADDIE-structured worked example?

3.3. Literature Search

Five databases and platforms were searched: Scopus, Web of Science, ERIC, and PubMed for peer-reviewed literature, and Google Scholar for additional coverage including grey literature. Searches were conducted between April and June 2026, and a representative search string combined three concept blocks using Boolean logic:
  • AI block: “artificial intelligence” OR “generative AI” OR “large language model” OR “ChatGPT” OR “machine learning”.
  • Education block: “higher education” OR “curriculum” OR “instructional design” OR “assessment” OR “pedagogy”.
  • STEM or discipline block: “STEM” OR “biology” OR “biological sciences” OR “laboratory” OR “bioinformatics”.
The three blocks were combined with the AND operator, and database-specific syntax (field tags, truncation, and controlled vocabulary) was adapted for each platform. Reference lists of key articles and reviews were screened to identify additional relevant sources (backward citation searching).
The combined searches returned several hundred records. Titles and abstracts were screened for relevance against the eligibility criteria in Section 3.4, and the full texts of the eligible records were examined. Sources were retained where they contributed to the conceptual framework, the disciplinary analysis, or the biological sciences worked example. These sources were supplemented by foundational documents identified through backward citation searching. Consistent with the SANRA guidance for narrative reviews, these figures are reported to indicate the breadth of the search rather than as the output of a systematic screening protocol; formal record counts at each stage, dual independent screening, and a PRISMA flow diagram were therefore not produced.

3.4. Eligibility and Selection

Eligibility was determined against five criteria. The review concentrated on literature published from 2018 to 2026, with particular emphasis on the post-2022 GenAI literature, excluding earlier work unless it was foundational to the conceptual or instructional-design framework, such as work on ADDIE, constructive alignment, and UTAUT. Eligible sources addressed higher education, so studies confined to primary or secondary education and professional training were excluded. Peer-reviewed studies and reviews were prioritized for claims concerning learning, assessment, and educational effectiveness. Any foundational theoretical and instructional-design sources predating 2018 were retained where they underpin the conceptual framework; and sources were chosen for their relevance to curriculum and course design, classroom or laboratory teaching, assessment, feedback, faculty development, or institutional implementation.

3.5. Scope, Synthesis, and Reasoning

The review focuses on undergraduate and postgraduate STEM education and treats STEM as a collection of disciplines with distinct epistemic, practical, and assessment traditions rather than as a homogeneous field. Consistent with the terminology defined in Section 1, AI is used as an umbrella term encompassing generative AI, intelligent tutoring systems, adaptive learning, machine learning, learning analytics, and automated assessment, with generative AI treated as the principal recent catalyst for curriculum redesign. The biological sciences serve as a recurring disciplinary case and provide the worked example in Section 6 because they combine conceptual learning, laboratory practice, computational analysis, and ethical decision-making.
Consistent with the SANRA criteria on scientific reasoning and the appropriate presentation of evidence, the synthesis was organized through the conceptual lenses set out in Section 2 (cognitive load theory, constructive alignment, constructivist perspectives, TPACK, and UTAUT), so that claims are interpreted through established accounts of learning, alignment, faculty capability, and adoption rather than according to technological novelty alone. Where evidence is mixed or contested, competing findings are presented together, and the strengths and limitations of the underlying studies are noted in the text. The two following examples illustrate the process followed. First, on AI-assisted grading, we juxtapose evidence that well-prompted models can approach teaching-assistant quality for some formative feedback with evidence that model grading agrees only moderately to poorly with human graders for higher-order work. The recommendation is that summative judgment should remain human-led (Section 6.7 and Section 7). Second, on virtual-laboratory effects, the review reports positive gains in knowledge acquisition and misconception reduction while noting that some effect sizes are inflated by the programmed constraints of virtual environments. The magnitude of benefit is therefore treated as provisional (Section 6.6 and Section 9). Where primary evidence was sparse, claims are presented as reasoned extrapolation rather than as established findings, and are identified as such.

4. Why GenAI Creates a Curriculum and Assessment-Validity Problem

4.1. The Automation of Routine STEM Competence

Much of what STEM curricula currently teach and reward consists of tasks that are now within the reach of widely available AI systems. Generating syntactically correct code, carrying out standard algebraic manipulations, summarizing primary literature, drafting conventional laboratory reports, and performing routine data cleaning and basic statistical analysis can all be produced, at least in draft form, by current large language models and allied tools [30]. The significance of this development is frequently misstated. The issue is not that these skills have become worthless; competent practitioners still rely on them daily. The issue is that teaching and assessing such tasks as terminal outcomes no longer certify the underlying competence the curriculum is intended to guarantee. When a student submits working code or a fluent literature summary, that artifact is now a weak signal of the reasoning, fluency, and understanding it once reliably indicated [20]. A curriculum that continues to define and verify competence through artifacts that AI can generate is, in effect, certifying something it no longer measures.
This is a different and more structural problem than the familiar worry that students might use AI to cheat. Even a student acting in complete good faith, using AI exactly as a future professional would, exposes the same gap: the learning outcome as written, and the assessment as designed, no longer capture what we actually want the graduate to be able to do. The mismatch is built into the curriculum, not into the student’s conduct.

4.2. From Recall and Reproduction Toward Judgment and Verification

If AI can produce fluent output that is sometimes confidently wrong, the scarce and genuinely teachable human capability shifts toward a different cluster of skills: framing a problem well, selecting an appropriate method, critically evaluating AI-generated output, and verifying whether a result is correct and trustworthy [31,32]. This reorientation has a clear basis in established learning theory. A cognitive-load perspective suggests that delegating routine, lower-order operations to AI can free working-memory resources, but only if that freed capacity is deliberately redirected toward higher-order reasoning; otherwise the offloading simply removes the productive struggle through which understanding is built [33,34]. The same pattern can be described as a compression of routine, reproducible tasks and a corresponding rise in the relative importance of judgment, evaluation, and the creation of novel work, a distinction this review makes explicit through the revised Bloom’s taxonomy when it specifies intended learning outcomes in the worked example of Section 6 [35,36].
The curricular implication is direct. Outcomes framed around reproducing knowledge or producing standard artifacts describe precisely the territory AI now occupies. Outcomes framed around judgment, justification, and verification describe what remains distinctively human and therefore what the curriculum must now foreground. This is the concrete content of the review’s central argument: the durable transformation is less about delivering existing material faster and more about redefining what competence means and how it is certified.

4.3. The Disciplinary Mismatch

STEM is not a single field facing a single challenge, and treating it as one obscures the most important dynamics. The affordances and pressures of AI differ sharply across disciplines, and the redesign each requires differs accordingly [37]. The discipline-specific pressures and corresponding curriculum and assessment responses are summarized in Table 1.
  • Computer science feels the most direct pressure, because code generation overlaps the curriculum itself: the artifact that introductory programming courses ask students to produce is exactly what these tools generate most readily, forcing a rethink of how foundational programming competence is taught and evidenced [38].
  • Mathematics confronts symbolic computation and proof assistance, which press on derivation and problem-solving in a different way, raising questions about what should be done by hand and why [39].
  • Engineering encounters AI in design, simulation, and optimization, where the question becomes how to teach sound design judgment when generation of candidate solutions is cheap.
  • The data-intensive and laboratory-based natural sciences face AI across experimental design, data analysis, and interpretation, a particularly broad surface of contact [40].
Because these pressures are not uniform, a single institution-wide policy or a generic “AI literacy” module cannot answer them. Each discipline must redefine, in its own terms, which competencies are now table stakes, which become more valuable, and which can be safely delegated. This is inherently a curriculum-design task carried out at the disciplinary level, not a classroom workaround.
Table 1. Discipline-specific AI pressure points and redesign responses across STEM disciplines.
Table 1. Discipline-specific AI pressure points and redesign responses across STEM disciplines.
DisciplineAI Pressure PointCurriculum ResponseAssessment ResponseRef.
Computer scienceCode generationEmphasize debugging, explanation, architecture, testingLive coding, code review, oral defense[41]
MathematicsSymbolic solving and proof assistancePreserve foundational fluency; teach verificationProof critique, supervised derivation, explanation[42]
EngineeringDesign generation and optimizationTeach constraints, trade-offs, safety, ethicsDesign justification, simulations, design review[43]
BiologyAI-assisted literature, data analysis, experimental interpretationTeach AI literacy, biological plausibility, wet-lab competenceAI-critique tasks, lab portfolios, practical exams[44]

4.4. Why This Is a Curriculum Problem, Not Only a Classroom One

The preceding subsections converge on a single conclusion: the gap between what STEM curricula currently certify and what the AI-altered environment now requires is structural, and it cannot be closed by individual instructors acting alone. An instructor can revise an assignment, add an oral component, or restrict tool use in a particular sitting, and such measures matter for the integrity of specific assessments. But none of them can repair a program-level mismatch in defined learning outcomes and graduate attributes, because the definition of competence is set above the level of the individual course [45]. When the stated outcomes themselves describe automatable performance, every downstream course inherits the mismatch.
Recognizing the problem as curricular has two consequences for the rest of this review. First, it justifies a systematic, program-to-classroom treatment rather than a collection of classroom tips, since the change has to propagate from defined outcomes through course and laboratory design to assessment. Second, it warrants both a principled account of how curriculum and course design should respond and a concrete demonstration of that response in a single discipline. Accordingly, the review first sets out the design task at the programme and course level (Section 5), and then develops a fully worked disciplinary example in the biological sciences (Section 6), where the ADDIE model serves as the instructional-design scaffold for tracing AI-driven change from needs analysis and design through development, implementation, and evaluation.

5. From Programme-Level Competence to Course and Assessment Redesign

Section 4 argued that the gap between what STEM curricula currently certify and what an AI-altered environment requires is structural, and therefore a design problem rather than a classroom workaround. This section takes up the first half of that design task at the programme and course level. Working from the conceptual lenses of Section 2.1, it addresses two linked questions: how AI reshapes the analysis of learner needs and required competencies, and how it reshapes the specification of learning outcomes, the sequencing of content, and the blueprint that links outcomes to assessment. The procedural realization of these principles within a single discipline, organized through the ADDIE model, is presented in the worked example of Section 6.
This curricular framing matters because, unlike a syllabus, a curriculum encompasses not only the content and organization of the subject matter to be taught. Although approaches vary in practice, a curriculum generally establishes an educational framework comprising the learning objectives pursued through teaching the subject, the design of learning experiences intended to achieve those objectives, the organization and support of students’ educational experiences (that is, the underlying pedagogical philosophy), and the evaluation methods through which attainment of the anticipated objectives is judged [46]. AI exerts pressure on all four components at once, which is precisely why the response cannot be confined to a single syllabus or classroom: a redefinition of required competence (Section 5.1) propagates through objectives (Section 5.3), learning experiences and their alignment with assessment (Section 5.4), and the programme-level organization that sequences them (Section 5.5).

5.1. Renegotiating Required Competence Before Outcomes Are Written

In conventional course planning, the analysis of needs establishes learner characteristics, prior knowledge, and the scope of content to be covered [47]. AI alters this task in two distinct ways, one procedural and one substantive.
The procedural change is that AI tools can now assist the analysis itself: surveying prior-knowledge gaps, clustering common misconceptions, and helping map the conceptual territory of a course [48]. This is a genuine efficiency, but the less important of the two changes, and it should be treated with care, since analyses driven by historical data can entrench existing assumptions about what a course should contain.
The substantive change follows directly from Section 4: when a meaningful share of routine tasks becomes automatable, the analysis can no longer take the list of required competencies as given. It must expand to include a deliberate re-examination of which competencies remain essential, which gain value, and which can be delegated to tools [49]. Needs analysis in the AI era is therefore about redefining the meaning of competence before any outcomes are written.

5.2. AI Literacy as a Discipline-Specific Graduate Attribute

One output of this expanded analysis is the recognition that critical, informed use of AI is itself a competency the curriculum must name and develop, rather than an incidental skill students acquire on their own. A growing body of institutional and disciplinary guidance frames AI literacy as a graduate attribute spanning technical understanding, critical evaluation of outputs, and ethical and responsible use [50,51]. For STEM programmes, two cautions keep this from becoming an empty slogan.
First, AI literacy must be discipline-specific, not generic. As Section 4.3 argued, the relevant competencies differ across fields: critically reading AI-generated code is not the same skill as critically interpreting an AI-proposed experimental design or an AI-assisted statistical analysis. A single institution-wide module cannot substitute for AI literacy embedded within each discipline’s own content and standards of judgment [52].
Second, AI literacy should be framed as a higher-order competency in the sense of Section 2.1: the goal is not fluency in operating a tool that quickly dates, but the durable capacity to evaluate, verify, and appropriately distrust AI output. Defined this way, AI literacy is continuous with the broader curricular shift toward judgment and verification rather than a separate add-on.

5.3. Writing Outcomes for Judgment and Verification

Translating the redefined competencies into stated learning outcomes is where the review’s central prescription becomes concrete. As Section 4.1 argued, outcomes phrased around reproducing knowledge or producing standard artifacts now describe territory AI occupies, so certifying them through their usual artifacts certifies less than it appears to [53]. Outcomes phrased around analysis, evaluation, justification, and verification describe what remains distinctively the student’s own to demonstrate.
Examples of how conventional STEM learning outcomes can be reframed for the GenAI era are shown in Figure 1.
These examples do not imply that foundational disciplinary tasks should be abandoned. Rather, they illustrate a shift in the evidence of learning: from the production of an artifact alone toward verification, justification, attribution, and disciplinary judgment. A student may still need to summarize literature, analyse data, or report laboratory work, but the revised outcome requires the student to demonstrate ownership of the reasoning process and accountability for the final product.
Rewriting outcomes accordingly requires more than moving to higher-order cognitive demands, a reorientation the review makes explicit through the revised Bloom’s taxonomy in the worked example of Section 6. It also requires specifying the conditions under which competence is demonstrated, including whether and how AI may be used. An outcome such as “evaluate and correct AI-generated solutions to a class of problems, justifying each change” both presupposes AI’s presence and targets a higher-order capability, showing how outcome design can absorb the technology rather than ignore or merely prohibit it [54]. This also resolves the good-faith-student problem of Section 4.1: when the outcome concerns judgment about AI output, authentic AI use becomes part of demonstrating competence rather than a threat to it.

5.4. Constructive Alignment and the Assessment Blueprint

Outcomes are only as meaningful as their alignment with teaching activities and assessment. The principle of constructive alignment holds that outcomes, learning activities, and assessment tasks must form a coherent triad, each consistent with the others [55]. AI raises the stakes of alignment because it can break the triad silently: an outcome may be rewritten toward higher-order judgment while the assessment that supposedly measures it still rewards a producible artifact, leaving the apparent alignment intact but the actual measurement hollow.
Two implications follow. First, the logic of backward design, beginning with the desired outcome and the evidence that would demonstrate it before planning instruction, is particularly valuable under AI conditions because it forces the question of which evidence remains valid when artifacts are automatable [56]. Second, the assessment blueprint must be designed alongside the outcomes rather than appended afterward, with explicit attention to which tasks retain validity in the presence of AI and which require redesign toward process evidence, oral defense, or authentic application.

5.5. Programme-Level Coherence and Sequencing

Because the mismatch identified in Section 4 is curricular rather than confined to single courses, design must also operate at the programme level. Three sequencing questions arise. First, where AI literacy is introduced and revisited: treating it as a single early module risks the same staleness that afflicts tool-specific teaching, whereas distributing and deepening it across the curriculum aligns with its character as a higher-order, discipline-specific competency [57]. Second, how prerequisites change when foundational tasks are partly automatable: a programme may decide that certain procedural fluencies remain prerequisites precisely because they underpin the judgment to evaluate AI output, a decision that should be made explicitly rather than by inertia [58]. Third, how higher-order competencies are scaffolded across years, since judgment and verification cannot be acquired in a single course and require graduated, repeated practice consistent with the constructivist commitment of Section 2.1.
Designing for programme-level coherence is where the curricular framing of this review pays off. Decisions about AI literacy placement, prerequisite retention, and the scaffolding of higher-order competence cannot be made coherently by individual instructors acting in isolation; they are properties of the curriculum as a whole, and they set the conditions within which the course-level development and implementation choices examined in the worked example of Section 6 are made.

6. Worked Disciplinary Application: An ADDIE-Structured Cell Biology Model

This section applies the principles developed above to a single discipline and a single course. ADDIE provides the procedural sequence (Analysis, Design, Development, Implementation, and Evaluation), while the conceptual lenses of Section 2.1 supply the pedagogical reasoning within each phase.
The course design presented here is illustrative and evidence-informed rather than empirically validated. It has not been implemented, delivered, or tested as a whole, and no single study evaluated the exact configuration described. Each component is grounded in cited evidence drawn from adjacent courses, disciplines, and course levels, and the components are assembled into a coherent worked example to demonstrate how the design principles of Section 4 and Section 5 can be operationalised. The example is therefore a testable design proposal rather than a report of proven practice.

6.1. ADDIE as the Instructional-Design Scaffold for the Worked Example

The conceptual framework outlined in Section 2.1 explains how AI may influence learning, cognition, curriculum alignment, faculty capability, and adoption, but it does not provide a procedural template for building and refining a course. For that purpose, this worked example adopts the ADDIE model [59] as the scaffold for the design of an illustrative biological sciences course. Its five phases map closely onto the practical work of course design, situating AI-related change within the broader relationships among course aims, learning outcomes, instructional activities, assessment, faculty capability, and governance rather than treating AI as an isolated tool. In the Analysis phase, educators identify curriculum gaps, learner needs, disciplinary expectations, and institutional constraints; Design translates these into learning outcomes, pedagogical strategies, assessment plans, and decisions about effective and ethical AI use. Development creates resources, activities, assessments, rubrics, and guidance; Implementation addresses their practical implementation; and Evaluation would assess whether the redesign improved learning, assessment validity, equity, and graduate preparedness, informing subsequent revisions. The application of these phases to AI-enabled course redesign, with examples from biological sciences, is summarized in Figure 2.
Since ADDIE is a process model rather than a theory of learning, it is therefore paired with the conceptual lenses of Section 2.1. Specifically, cognitive load theory distinguishes productive scaffolding from cognitive outsourcing; constructive alignment ensures coherence among outcomes, activities, permitted AI use, and assessment; and TPACK and UTAUT inform the faculty-capability and adoption dimensions of the Implementation and Evaluation phases. The revised Bloom’s taxonomy is introduced at the Design phase, where writing and classifying intended learning outcomes requires an explicit taxonomy of cognitive processes.

6.2. Why the Biological Sciences?

Biological sciences offer a productive case for examining how AI is reshaping teaching and assessment. This is because the discipline couples conceptually dense, data-intensive content with laboratory practice and a longstanding reliance on written assessment. In a recent review by our group, we argue that AI is transforming research methodologies, data analysis, and hypothesis generation across the biological sciences, making the deliberate cultivation of critical thinking, skepticism toward AI outputs, contextual understanding, and ethical judgment a central task for higher education [5]. Bawn and colleagues add that teaching in the biological sciences, perhaps more than in other STEM subjects, has historically depended heavily on written assessment, which makes the discipline distinctively exposed to the rapid, plausible text generation of GenAI [60].
Cell and molecular biology laboratories are especially fitting as a case. A recent multi-country analysis of 138 universities across 28 European countries identified the common biology laboratory curriculum and proposed a competency-based curriculum that explicitly recommends virtual labs and AI to modernise laboratory education [61]. Alongside this curriculum-level evidence, biology and the wider life and health sciences have an established record of technology-enabled innovation, from interactive cloud laboratories [62] and mobile cooperative-learning platforms [63] to recent GenAI applications in course design [60,64], laboratory instruction [34,65,66,67], and assessment [60,68,69]. Taken together, these strands frame the central design question of this illustrative course-design example: how can educators integrate AI so that it augments, rather than replaces, the inquiry, interaction, reflection, and critical thinking that characterise deep learning in the biological sciences [5,60,67]?

6.3. Cell Biology: The Model Course

Cell biology is a foundational course and prerequisite across the health sciences, and it is frequently paired with a laboratory component, making it a useful model course. In addition to the lecture component, an AI-integrated course may adopt a blended laboratory model, using AI-enhanced virtual labs for preparation and physical labs for hands-on practice. Throughout, the design treats AI as a supplemental scaffold under human oversight intended to enhance rather than bypass students’ reasoning [5].

6.4. Analysis Phase: Needs, Learners, and Required Competence

The Analysis phase establishes learner characteristics and the competencies the course must develop. For a mixed health-sciences cohort, three needs are identifiable. First, students arrive with uneven prior disciplinary knowledge and GenAI familiarity, alongside documented differences in AI attitudes and access across disciplines and demographic groups [70]. From a cognitive load perspective, these differences are consequential because the same task may impose substantially different demands depending on students’ prior knowledge and familiarity with GenAI. For novices, unfamiliar terminology and conceptual relationships may contribute to intrinsic load, while prompt formulation, tool navigation, output comparison, and source verification may introduce additional demands. The Analysis phase should therefore establish not only what students know, but also whether GenAI use reduces avoidable demands or introduces new demands that compete with the disciplinary reasoning the course is intended to develop. Second, the discipline’s heavy reliance on written assessment leaves it exposed to GenAI, requiring identification of which competencies remain valid to assess through conventional methods and which require redesign [60]. Third, critical evaluation of AI should be treated as a required competency, consistent with the discipline-specific AI-literacy argument developed in Section 5.2 [5,71].
The shift documented by Ng and colleagues away from programming-heavy instruction towards interdisciplinary, scaffolded designs further supports embedding AI literacy within a mainstream biology course rather than confining it to computational specialists [72].

6.5. Design Phase: Outcomes, Alignment, and Ethical AI Use

The Design phase translates these needs into intended learning outcomes (ILOs) and an aligned assessment plan. Following Section 5.3, outcomes are written across the breadth of the topic using the revised Bloom’s taxonomy, with higher-order outcomes explicitly targeting judgment about AI. The revised Bloom’s taxonomy is useful for writing and classifying intended learning outcomes, as it provides an ordered taxonomy of cognitive processes. Objectives and aligned items would be AI-drafted and instructor-verified, given evidence that chatbots frequently misassign cognitive level [73].
  • ILO1: Describe core cell structures, organelles, and their functions (Understand).
  • ILO2: Perform basic microscopy and staining safely and accurately (Apply).
  • ILO3: Interpret simple cell observations and basic data (Analyse).
  • ILO4: Critically evaluate evidence and AI-generated outputs for accuracy, bias, and ethical implications (Evaluate) [5].
  • ILO5: Communicate findings clearly to a non-specialist audience (Create).
  • ILO6: Demonstrate AI literacy: responsible, disclosed, verified AI use (Create) [71].
ILO4 and ILO6 derive directly from the insistence that students be trained not only to use AI but to question it, checking sources and considering broader implications [5] and from the AI-literacy-as-competency argument [71]. The laboratory outcomes (ILO2, ILO3) follow the competency progression of the standardized European blueprint [61].
Design also fixes permitted AI use against each outcome, as constructive alignment under AI conditions requires (Section 5.4). Where the outcome is independent foundational competence (ILO1, ILO2), AI is restricted or excluded at the point of assessment. Where the outcome is critical evaluation or responsible use of AI (ILO4, ILO6), AI is deliberately present and is itself the object of scrutiny. This distinction follows directly from constructive alignment rather than from a general preference either for or against GenAI use. If an intended outcome requires independent performance, allowing GenAI to perform the target reasoning would weaken the validity of the assessment evidence. Conversely, where critical evaluation or responsible human–AI collaboration is itself the intended outcome, excluding GenAI would make the assessment less authentic. Permitted AI use is therefore determined by the competence being assessed rather than by a uniform course-level rule.
The course-design literature converges on this active, outcome-aligned orientation: active-learning pedagogies (inquiry-based learning, problem- and project-based learning, flipped classrooms, cooperative learning, and gamification) can be strategically combined with AI tools such as chatbots and intelligent tutoring systems, with the educator remaining essential for interpreting AI outputs, contextualising knowledge, and embedding ethical considerations [5]; institutional models frame AI literacy as a competency every student should develop through curriculum-wide infusion rather than isolated technical courses [71]; the GRATL framework embeds AI into experiential learning with explicit research competencies and mandatory source verification [64]; and frontier chatbots can draft objectives and aligned assessment items but satisfy all quality criteria less than half the time, requiring instructor oversight [73].
Table 2 combines all elements of this worked example together, and operationalises the central methodological claim of the review. Table 2 shows how each ADDIE phase is realised in practice while the “Guiding conceptual lens(es)” column shows why each phase asks the question it asks. In the Analysis phase, cognitive load theory frames the decision of which competencies to retain and which to delegate. In Design, constructive alignment and the revised Bloom’s taxonomy govern how outcomes are written and how permitted AI use is fixed against each outcome. In Development, constructivist and constructionist principles shape resources and activities that ask students to build, critique, and revise rather than passively receive. In Implementation, TPACK and UTAUT govern faculty capability and adoption, while constructive alignment continues to regulate permitted AI use. In Evaluation, cognitive load theory distinguishes genuine reasoning from apparent competence, and TPACK and UTAUT frame feasibility, workload, and equity. The final phase, iterative redesign, re-enters the cycle and therefore draws on all lenses at once. The remaining columns then specify, for each phase, the relevant AI tool category, the support available to teachers and to students, and a concrete biological sciences instantiation, with supporting references. Presented this way, the table is not a catalogue of tools but a demonstration that each design decision is anchored in an explicit account of learning, alignment, or adoption, which is precisely what distinguishes intentional redesign from ad hoc tool adoption.

6.6. Development Phase: Building Resources, Activities, and the Course Map

The Development phase creates the resources, activities, and assessments specified in design. The course map below integrates GenAI and critical human oversight. This is summarized in Table 3 and it is intentionally generic and thus applicable to STEM courses.
The choice of these activities follows from the constructivist and constructionist perspectives outlined in Section 2.1.2. GenAI-generated material is therefore not treated as content for passive consumption but as an artefact that students must interrogate, compare with disciplinary evidence, modify, and justify. Activities such as correcting a GenAI-generated protocol, evaluating generated code, revising concept maps, or challenging a GenAI-produced explanation make the processes of knowledge construction and revision observable rather than rewarding the generated product alone.
The following paragraphs address these components in more detail, beginning with lecture resources and tutoring. In our example, the instructor would use GenAI to draft quizzes, concept summaries, and lay-language explainers (useful for translating cell biology into clinically relevant terms for health-sciences students), then edit them for accuracy before release [60]. Students would be permitted to use an intelligent tutoring system or chatbot for self-paced review, but would be explicitly taught that large language models can hallucinate (for example, fabricating citations or misidentifying protein functions) and must therefore cross-check outputs and reflect on detected errors as a critical-thinking exercise [5,74]. The move toward scaffolded, interdisciplinary AI use rather than programming-heavy instruction follows the trajectory documented by Ng and colleagues [72].
Regarding laboratory and inquiry resources, the selected laboratory topics are among the most frequently integrated biology laboratories in health-related undergraduate programmes, and align with a standardized, competency-based curriculum mapped to European Tuning and Vision and Change that explicitly recommends virtual labs and AI [61]. Each wet-lab would be preceded by an AI-enhanced, inquiry-aligned virtual-laboratory module with intelligent tutoring, automated formative checks, and real-time feedback, justified by evidence that AI-enhanced virtual labs improved knowledge acquisition, skill development, and misconception reduction while reducing cost [66], by the cloud-lab model that scales authentic inquiry [62], and by the interactive-science-laboratory mechanisms that develop critical thinking [5]. A dry-lab analogue from computational genetics is instructive: students with little programming background generated Python scripts to analyse real RNA-sequencing datasets, with the decisive safeguard that disciplinary understanding was required to validate outputs, since identical prompts often produced different and sometimes non-functional code [65]. A reflective bioethics strand would integrate large language models to develop ethical competence, building on Wang’s finding of deeper ethical reflection while addressing the same study’s concerns about inaccuracy, bias, and erosion of independent thinking [34], and reinforced by GRATL’s integrity safeguards [64]. Cooperative observation tasks would adapt the visible, real-time, rubric-graded cooperative-work model in which small groups build labelled structure-to-function diagrams, optionally correcting and extending GenAI-generated first-draft concept maps [5,63].
From a cognitive load perspective, the proposed virtual pre-laboratory activity is intended to reduce avoidable extraneous demands associated with unfamiliar equipment, procedural sequence, laboratory terminology, and safety expectations before students enter the physical laboratory. This preparatory role should not be interpreted as a replacement for practical laboratory work. Microscopy, staining, observation, troubleshooting, and biological interpretation remain central learning targets and should therefore be demonstrated in the wet laboratory, where students must coordinate procedural actions with disciplinary judgment. The proposed design uses virtual preparation to support entry into the task while preserving the cognitive activity required for practical competence and biological interpretation.
The assessment blueprint would be developed alongside the learning outcomes, to ensure alignment between teaching and assessment. A key feature is that assessment is AI resilient (Figure 3 and Supplementary Table S3b). The proposed assessment blueprint combines complementary assessment methods designed to promote authentic learning while strengthening resilience to inappropriate reliance on generative AI. The lab portfolio would emphasize process over product by requiring students to document observations and experimental work through annotated images and logs, providing evidence of ongoing engagement and learning. The practical examination would assess laboratory skills through a live demonstration under supervised conditions, ensuring that competence is demonstrated directly by the student. The AI-critique task would develop students’ ability to critique AI outputs by requiring them to evaluate the accuracy, reliability, and limitations of GenAI-generated explanations, thereby fostering critical digital literacy. The bioethics reflection would promote evidence-based thinking by using credible sources to support reasoned perspectives on complex scientific and societal issues. The oral group presentation would strengthen attributed contribution by making individual participation visible and verifiable through presentation delivery and questioning. Finally, the supervised concept test would provide assessment under controlled conditions, allowing students to demonstrate their understanding of key biological concepts independently. Collectively, these assessment approaches would balance knowledge, practical skills, critical thinking, reflection, and accountability, creating an assessment framework intended to uphold academic integrity and to prepare students to engage thoughtfully and responsibly with AI technologies.
Explicit consideration should also be given to the feasibility of these assessment modes, since the formats most resilient to GenAI are generally more demanding to develop and to mark [76]. Designers may weigh four parameters: development effort, marking workload and how it scales with cohort size, the class or group size for which each format is appropriate, and the degree to which the grade reflects genuine student ability rather than tool use or presentation (reliability and proportionality). These parameters apply across the blueprint’s formats, from scalable supervised concept tests and portfolios to lower-volume, high-validity practical examinations, oral defences, and AI-critique tasks, as discussed below.
Development effort and marking workload scale differently. Some formats are inexpensive to build but costly to mark at scale, such as oral presentations and practical examinations [77], whereas others carry a higher one-off design cost but then mark efficiently, such as supervised concept tests [76]. AI-critique tasks occupy a middle position: they draw on a reusable prompt bank but still require human marking, because they target the higher-order evaluation that AI grades least reliably [68,69].
The assessment blueprint is therefore intended to be balanced rather than uniformly labour-intensive. Consistent with programmatic approaches to assessment, no single format is expected to carry the evidential load, and high-validity formats can be sampled rather than applied exhaustively [78]. The blueprint combines scalable formats such as concept tests and portfolios, which remain workable in large cohorts, with lower-volume, high-validity formats such as oral defence and practical examinations, which are best used in smaller classes, tutorial groups, or laboratory sections, or applied on a sampling basis in large cohorts [77,78].
Finally, resilience and proportionality trade-off against automation needs to be addressed. AI-assisted marking can reduce workload for routine components and for first-pass formative feedback, but because model grading concords only moderately to poorly with human graders for higher-order work [68,69], human markers must retain responsibility for the attributable, supervised judgments on which proportionality depends; the clarity of rubrics and the attribution of authorship remain central to that proportionality [79]. The design thus preserves validity at a manageable, though non-trivial, staffing cost. The workload and reliability estimates offered here are indicative and are themselves identified as objects for empirical evaluation (Section 9 and Section 10).

6.7. Implementation Phase

The Implementation phase introduces the redesigned course in practice, with attention to orientation, equitable access, faculty support, transparency, and data protection. Five elements structure this phase: AI-literacy onboarding, blended pre-laboratory preparation, in-laboratory AI support, cooperative and reflective work, and the boundaries governing feedback and grading.
In the proposed design, implementation would begin with AI-literacy onboarding. The course would open with a session covering permitted tools, disclosure norms, the duty to verify outputs, and the known risks of inaccuracy and bias. This operationalises the principle that AI literacy is a core competency for all students regardless of discipline [71], attends to documented differences in AI attitudes by gender and discipline [70], and reflects the requirement that AI be treated as a supplemental, verifiable aid [64]. Implementation also depends on instructor capability, and the capability this requires is specific rather than generic. An instructor must be able to recognise that a fluent GenAI account of, for example, organelle function, subcellular localisation, or the interpretation of a staining result can be disciplinarily wrong in ways that are invisible to a reader without cell biology expertise, and that generic prompting or verification training will not surface. This is the technology-content intersection TPACK identifies, and it explains why professional development for this design cannot be delivered as generic institutional training detached from the discipline, just as Section 5.2 argued for students.
Each wet lab would be preceded by an inquiry-aligned virtual laboratory module in which students rehearse microscope operation, staining, and cell structure identification. This blended design, virtual pre-lab followed by physical wet-lab, is directly supported by educator-adoption evidence: in a four-country European study, educators valued virtual labs specifically for pre-lab preparation, safe repetition without material waste, and flexible asynchronous access, while insisting that hands-on manipulation and learning from physical mistakes remain essential and cannot be replaced [67]. Read through the UTAUT lens of Section 2.1; this pattern reflects performance expectancy and facilitating conditions outweighing mere ease of use: educators adopt the toolkit because it aligns with inquiry-based teaching and is institutionally supported, not because it is technically novel. The pre-labs are accordingly built around inquiry-based learning (questioning, hypothesis testing, and evidence-based reasoning), since pedagogical alignment with inquiry was the factor most strongly associated with educator intention to adopt [67]. The design preserves hands-on practice in the physical laboratory, consistent with the caution that virtual environments lack haptic feedback and that programmed constraints can inflate apparent safety gains [66].
In staffed laboratory sessions, a GenAI virtual teaching assistant could provide first-line support for routine procedural questions, freeing human teaching assistants to focus on higher-value interactions. This is implemented with a safety net: because human teaching assistants are more accurate and more pedagogically effective than GenAI, and students can identify and prefer human responses, GenAI answers serve as first-line support for basic queries only, with a human teaching assistant confirming guidance before students act, especially where safety is involved [75].
Cooperative and reflective work runs alongside the practical strand. Small groups build labelled organelle diagrams under live instructor monitoring [5,63], while a reflective strand requires students to interrogate GenAI outputs using explicit guiding questions: what assumptions does the model make, whose data might be missing, and what are the ethical implications [5,34,64].
Finally, feedback and grading boundaries are made explicit. GenAI-assisted formative feedback could be provided on drafts throughout, supported by evidence that well-prompted large language models can approach human teaching-assistant quality for feedback [68]. Corrective, explanatory feedback is prioritised, since its absence was the most consistent educator critique of virtual-laboratory tools across four countries [67]. All summative grading remains human-led, because LLM grading of inquiry-style work showed only moderate-to-poor concordance with human graders and was weakest for higher-order domains [69]. Every assessment requires an AI-use disclosure statement and, where relevant, evidence of original work (annotated images, logs, prompts, and drafts), citation of credible sources, and balanced viewpoints, mirroring GRATL integrity mechanisms [64] and the emphasis on human criticality and ethical oversight [5].

6.8. Evaluation Phase: Did the Redesign Preserve Learning and Assessment Validity?

The Evaluation phase asks whether the redesign improved learning and preserved the validity of assessment, and it feeds the next design cycle. Four evaluation questions follow directly from the outcomes and the central argument of the review.
First, can students independently demonstrate foundational competence? In the proposed design, the supervised practical examination and concept test would provide evidence for ILO1 to ILO3 under controlled conditions, isolating the student’s own capability from GenAI assistance [61,69].
Second, can students critically evaluate AI rather than merely operate it? In the proposed design, the AI-critique task and bioethics reflection would be intended to provide direct evidence for ILO4 and ILO6 by assessing whether students can identify inaccuracies and bias in GenAI-generated output and justify their reasoning, which represents the higher-order capability foregrounded by the curriculum [5].
Third, did AI support rather than displace cognition? Cognitive load theory requires evaluation of both possibilities. GenAI may provide productive scaffolding when it reduces avoidable demands while leaving students responsible for disciplinary interpretation, verification, and justification. It may instead produce cognitive outsourcing when it performs the reasoning processes that students are expected to develop, or it may increase extraneous demands when students must devote substantial effort to prompt refinement, tool management, or resolving unreliable outputs. Evaluation should therefore examine not simply whether students performed better with AI, but which cognitive activities were transferred to the tool and whether students can subsequently demonstrate the underlying competence independently. Process evidence, including portfolios, logs, and oral vivas, provides one means of distinguishing genuine reasoning from fluent but unowned output [60,63].
Fourth, was the redesign equitable and feasible? Evaluation attends to whether onboarding narrowed rather than widened gaps in AI access and confidence [70], and to faculty workload and capability. Interpreted through the combined TPACK and UTAUT lens, feasibility depends both on whether educators hold the technological, pedagogical, and disciplinary knowledge to integrate AI soundly and on whether facilitating conditions and pedagogical alignment make adoption attractive, since educator adoption is driven primarily by these factors rather than by technological novelty [67,71].

6.9. Iterative Redesign

Consistent with the iterative use of ADDIE described in Section 6.1, evaluation is the basis for the next cycle rather than an endpoint. Likely revision targets would include any activity in which AI found to displace rather than support reasoning, any assessment shown to remain vulnerable to undisclosed AI use, and any virtual pre-lab whose programmed constraints inflated apparent competence relative to bench performance [66]. New AI capabilities, emerging misconceptions, and student and faculty feedback all feed the renewed analysis of needs, closing the loop back to Section 6.4.

6.10. Synthesis: A Generalisable Design Principle

Across the ADDIE phases, a common design principle emerges: GenAI should be integrated only where it supports, rather than displaces, the disciplinary reasoning and competence that students are expected to develop. Course and laboratory design should therefore align GenAI use with intended learning outcomes, preserve independent performance where foundational competence must be demonstrated, and use process-oriented assessment to make reasoning, verification, and attribution visible [5,60,61,64,68,69,71,72,73]. From this perspective, the value of GenAI lies less in automating tasks than in supporting active learning, formative feedback, critical evaluation, and appropriately scaffolded preparation under human oversight.
The biological sciences example illustrates how this principle can be operationalised within a blended, competency-based model that combines virtual preparation with physical laboratory practice and explicit AI-literacy support [61,66,67,70,71]. Its broader relevance lies not in the specific cell biology content, but in the design logic linking learner needs, outcomes, permitted AI use, assessment evidence, implementation conditions, and iterative evaluation. The extent to which this logic transfers to other health-sciences and STEM contexts remains an empirical question addressed in Section 7, Section 8, Section 9 and Section 10.

7. Discussion: From Biological Sciences to STEM Curriculum Redesign

The current review has done more than restate that curricula must change. It has traced how that change propagates from programme-level definitions of competence through outcomes, permitted AI use, and aligned assessment to implementation and evaluation. The underlying problem is one of curriculum and assessment validity. Because widely available GenAI systems can now produce many of the conventional artifacts through which STEM programmes have historically inferred competence, the completion of an assignment, laboratory report, code, calculation, or scientific summary no longer reliably evidences the underlying disciplinary reasoning, and the resulting mismatch concerns what programmes define, develop, and certify rather than academic integrity alone. This is the specific respect in which the review differs from recent syntheses of GenAI in science and higher education: reviews such as those of Almasri [13] and Yusuf et al. [14] establish that AI is expanding and that curricula must respond, but characterise the field and its evidence base rather than specifying how a response should be constructed. The present review begins where those reviews conclude, supplying a theoretically grounded and transferable design procedure and demonstrating it in a single discipline.
The review’s five guiding questions converge on a single interpretation, summarised against their implications for STEM curriculum redesign in Table 4.
The biological sciences example also clarifies the relationship between the conceptual lenses and the ADDIE process model. The conceptual lenses carry the analytical weight of the review, whereas ADDIE provides the procedural structure for translating that analysis into course design. Cognitive load theory helps distinguish productive scaffolding from cognitive outsourcing, particularly in virtual or AI-supported environments, where apparent competence may not be confirmed by observable bench performance [66]. Constructive alignment requires that permitted AI use be specified on an outcome-by-outcome basis rather than through a blanket rule, ensuring that AI is restricted where independent competence is assessed and deliberately included where critical evaluation of AI is itself an intended outcome [5,64]. The combined TPACK and UTAUT perspective explains why a technically feasible design will be implemented only when educators possess the required technological, pedagogical, and disciplinary knowledge and when institutional conditions make adoption worthwhile and manageable [67,71]. The revised Bloom’s taxonomy, operationalized during the Design phase, further distinguishes the apparent sophistication of an AI-generated output from the cognitive processes that students themselves must demonstrate [5,73].
The cell biology example should therefore be understood as an evidence-informed operationalization of the review’s conceptual argument rather than as a validated intervention or a model that can be transferred unchanged. What transfers across STEM is the redesign process: identifying which routine competencies are vulnerable to automation; deciding which competencies must remain independently demonstrable, which may be AI-supported, and which require critical human–AI collaboration; rewriting outcomes around disciplinary judgment and verification; aligning permitted AI use with valid assessment evidence; preserving supervised demonstration of foundational competence where necessary; and evaluating whether AI supported learning or displaced it. What remains discipline-specific is the content and evidence of competence, set out for each cluster in Section 8.4 and Table 5.
Faculty capability and institutional conditions are central to whether this redesign is feasible. The worked example depends on educators who can determine whether AI-supported activities are technically accurate, pedagogically purposeful, and disciplinarily valid, corresponding to the capability dimension captured by TPACK. Implementation also depends on whether educators have sufficient time, infrastructure, institutional support, confidence, and professional incentives, corresponding to the adoption conditions captured by UTAUT. Evidence that educators’ intention to adopt virtual-laboratory tools is associated more strongly with pedagogical alignment and institutional support than with technological usability suggests that successful implementation requires professional development, disciplinary exemplars, workload recognition, and reliable infrastructure rather than tool procurement alone [61,71]. Equity is an equally important institutional condition. Differences in AI access, confidence, and attitudes across disciplines and demographic groups mean that onboarding, tool selection, and access arrangements are questions of fairness as well as pedagogy [70].
Several unresolved tensions temper these conclusions. The boundary between productive scaffolding and cognitive outsourcing is conceptually persuasive but is not yet supported by widely validated measures, leaving educators dependent on professional judgment. GenAI-generated feedback may be useful for selected formative purposes, whereas GenAI-assisted grading remains less reliable for higher-order outcomes and therefore requires human oversight [68,69]. Detection of undisclosed AI use remains an unstable basis for assessing integrity, which is why the proposed approach prioritizes attributable process evidence, supervised performance, transparent disclosure, and oral or practical verification over detection alone [60,63]. Finally, the rapid evolution of AI capabilities means that tool-specific conclusions should be regarded as provisional. The review’s durable contribution lies in identifying curriculum design principles that should remain relevant as specific technologies change: constructive alignment, discipline-specific competence, critical verification, process evidence, human oversight, equity, and ethical accountability.

8. Recommendations

The analysis supports a set of actionable recommendations addressed to the three levels at which the redesign must occur: the programme and curriculum, the course and classroom, and the institution. They are offered as evidence-informed guidance rather than as validated prescriptions. Each recommendation follows from one of the conceptual lenses rather than from the tools themselves. To illustrate, cognitive load theory implies that AI should be permitted where it reduces avoidable processing demands and withdrawn where it would displace the cognitive activity through which durable disciplinary knowledge is built. Constructive alignment implies that permitted GenAI use must be fixed outcome by outcome rather than by a uniform course-level rule. Constructivist and constructionist perspectives imply that activities should require students to interrogate, verify, and revise generated material rather than consume it. The combined TPACK and UTAUT lens implies that adoption depends on educator capability and facilitating conditions rather than on procurement. The recommendations below apply these principles at the programme, course, and institutional levels.

8.1. For Programme and Curriculum Designers

The renegotiation of required competence should be treated as the first design task, not an afterthought. Programme designers should re-examine which competencies remain essential, which gain value, and which may be delegated to AI before learning outcomes are written (cognitive load theory and constructive alignment) [49]. AI literacy should be embedded as a discipline-specific graduate attribute that is distributed and deepened across the curriculum rather than confined to a single early module [57,71]. Prerequisite retention should be made an explicit decision, with procedural fluencies retained precisely where they underpin the judgment to evaluate AI output [58].

8.2. For Course and Classroom Practitioners

Practitioners should write outcomes for judgment and verification, and should specify permitted AI use outcome-by-outcome, restricting GenAI where independent foundational competence is assessed and requiring it where critical evaluation of GenAI is the outcome [5,64]. The assessment blueprint should be designed alongside the outcomes, favouring attributable, process-oriented formats (portfolios, supervised practical work, oral presentations, and AI-critique tasks) that remain valid under ubiquitous AI (constructive alignment) [60,63,69]. LLM-based tools should be used for first-pass formative feedback while summative judgment remains human-led, particularly for higher-order work [68,69]. A blended laboratory model should be adopted, in which AI-enhanced virtual pre-labs prepare students for, but do not replace, hands-on bench practice [66,67].

8.3. For Institutions and Academic Leaders

Institutions and academic leaders should invest in faculty capability and adoption together, pairing TPACK-oriented professional development with the UTAUT-identified facilitating conditions (reliable infrastructure, recognised workload, and pedagogically aligned exemplars) that actually drive adoption (TPACK and UTAUT) [71]. Equity should be treated as a design requirement, with documented differences in AI access and attitudes addressed through structured onboarding and considered tool selection [67,70]. Transparent disclosure norms and integrity safeguards should be established at the institutional level, so that individual courses are not left to improvise them [64].

8.4. From Biology to a STEM Template

Although the worked example is grounded in cell biology, its underlying logic is discipline-agnostic, and the ADDIE sequence used in Section 6 functions as a portable template that any STEM programme can adapt. What transfers is not the microscopy content but the design moves. Renegotiating required competence before writing outcomes, phrasing outcomes around judgment and verification, fixing permitted AI use outcome by outcome, building attributable and process-oriented assessment, and treating faculty capability and adoption (through TPACK and UTAUT) as conditions of feasibility rather than afterthoughts. Applied in this way, the five ADDIE phases become a common procedure whose inputs vary by discipline while its structure remains constant.
The template can be stated as five transferable steps:
  • Analysis. Identify which routine competencies in the discipline are now automatable, and re-examine which remain essential, which gain value, and which may be delegated to AI [49].
  • Design. Rewrite outcomes toward the discipline’s characteristic higher-order work (proof, design judgment, experimental interpretation, code architecture), and specify permitted AI use against each outcome [5,64]
  • Development. Build resources and, critically, AI-resilient assessment formats that require attributable process evidence rather than a producible artifact [60,63,69].
  • Implementation. Introduce discipline-specific AI-literacy onboarding, keep human oversight where safety or higher-order judgment is at stake, and use AI for formative rather than summative judgment [68,69,71].
  • Evaluation and iterative redesign. Verify that students can demonstrate competence independently and can critically evaluate AI, then feed the findings back into the next needs analysis [70].
The discipline-specific content that populates these steps mirrors the disciplinary pressures set out in Table 1. Table 5 shows how the same template is instantiated across the four STEM clusters, so that a mathematics, computer science, or engineering programme can follow the identical procedure while substituting its own competencies, tasks, and assessment evidence.
These recommendations operate at two levels. Every STEM discipline may face the same core AI challenge, by applying the five shared steps above. These steps are discipline-neutral and can be adopted as common institutional policy. At the level of discipline-specific challenge, however, the redesign must diverge because the competence differs. Computer science must protect the reasoning around code where the artifact itself is now automatable, mathematics must decide which derivational and proof fluencies remain prerequisite because they underpin verification, engineering must safeguard design judgment and safety-critical responsibility where candidate solutions are cheap to generate, and the biological and natural sciences must preserve experimental interpretation and hands-on laboratory competence across an unusually broad surface of AI contact. Table 5 sets out, for each cluster, the competence to renegotiate, the higher-order outcome focus, the AI-resilient assessment evidence, and the oversight and AI-literacy emphasis that follow from these distinct challenges. Because the depth of treatment in this review is necessarily greatest for biology, the recommendations for the other clusters are offered as principled starting points for discipline-led design rather than as validated blueprints, and the discipline-specific studies called for in Section 10 are the means by which they should be tested and refined.

9. Study Limitations

As a narrative review reported against the SANRA criteria, this synthesis involves authorial judgment and does not claim exhaustive coverage. Although the literature search was systematic, the review did not employ dual independent screening, formal risk-of-bias appraisal, or quantitative pooling. It is therefore subject to selection and interpretation bias, which SANRA-guided transparency can mitigate but cannot eliminate.
Much of the underlying evidence derives from small-scale, single-cohort, cross-sectional, or exploratory studies. Moreover, the rapidly changing capabilities of AI systems mean that findings relating to particular tools, detection methods, or institutional policies should be regarded as provisional. The choice of biology as the recurring disciplinary case, although justified by its combination of conceptual, laboratory, computational, and ethical dimensions, also means that discipline-specific issues in mathematics, computer science, and engineering are treated less extensively.
These limitations extend to the illustrative cell biology course design. The worked example is an evidence-informed and theoretically grounded design blueprint rather than an empirically validated intervention. No single study has evaluated the complete configuration proposed here, and much of the supporting evidence is drawn from adjacent disciplines and different course levels. Its application to an introductory health-sciences cell biology course therefore represents a reasoned extrapolation [66,69,73], although one supported by multi-country European evidence on common laboratory curricula and educator adoption of virtual-laboratory toolkits [67].
Accordingly, the manuscript does not establish which individual redesign components, combinations of components, or implementation conditions are most effective for particular courses, student populations, or learning outcomes. It also cannot determine whether the proposed configuration improves disciplinary knowledge, critical evaluation of AI, assessment validity, independent competence, or equity relative to conventional or alternative course designs. These questions require prospective implementation, component-level evaluation, and comparative studies using clearly specified learning outcomes and appropriate control or comparison conditions.
Several additional caveats remain. Reported effect sizes for AI-enhanced virtual laboratories may be partly inflated by the programmed constraints of virtual environments [66]. The educator-adoption evidence is exploratory and cross-sectional and is based on a modest, unevenly distributed sample; its associations should therefore not be interpreted as causal or predictive [67]. Similarly, the principle that GenAI should augment rather than replace human judgment is an evidence-informed recommendation rather than a demonstrated outcome of the proposed course design [5].
For these reasons, the review prioritizes durable design principles over particular technologies or claims of demonstrated effectiveness [5,67,71,72]. The proposed design should therefore be interpreted as a testable framework for subsequent empirical evaluation, not as a validated prescription.

10. Future Research

The field now requires empirical evidence capable of determining which GenAI-related curriculum redesign strategies are effective, for whom, under which conditions, and for which learning outcomes. Six priorities follow.
  • Prospective and comparative evaluation of complete redesign models. Studies should compare redesigned, judgement-centred curricula with conventional or alternative course designs and determine whether they improve disciplinary knowledge, independent competence, critical evaluation of AI-generated output, assessment validity, and ethical AI use.
  • Component-level evaluation of redesign strategies. Research should isolate the effects of specific elements, including AI-literacy onboarding, virtual pre-laboratories, AI-critique tasks, process-oriented portfolios, oral defence, supervised practical assessment, and AI-assisted formative feedback. Such studies are needed to identify which components, alone or in combination, are most useful for particular courses and intended learning outcomes.
  • Longitudinal and transfer studies. Research should move beyond single-course and single-cohort evaluations to examine whether students retain and transfer critical verification, disciplinary judgement, and responsible human–AI collaboration across subsequent courses, laboratory settings, and professional contexts.
  • Validated measures of productive scaffolding and cognitive outsourcing. New measures are required to determine whether AI support reduces unnecessary cognitive load while preserving productive reasoning, or instead displaces the cognitive processes that students are expected to develop.
  • Assessment-validity research. Attributable and process-oriented assessment formats should be evaluated for reliability, scalability, authenticity, equity, and resistance to undisclosed AI use across disciplines. Comparative studies should also determine when oral, practical, portfolio-based, and AI-critique assessments provide stronger evidence of competence than conventional submitted artifacts.
  • Discipline-specific implementation, faculty capability, and equity research. Studies in mathematics, computer science, engineering, and other STEM fields should test how far the redesign process derived from the biological sciences generalises and where discipline-specific adaptation is required. Research framed by TPACK and UTAUT should also examine which forms of faculty development, institutional support, workload recognition, infrastructure, and student onboarding enable sustainable adoption while narrowing rather than widening inequalities in access and confidence [67,70,71].
Direct empirical evaluation of the illustrative cell biology design is a natural first step within this agenda. Such evaluation should use prospectively defined outcomes, appropriate comparison conditions, and both student- and implementation-level measures. It should also examine the complete redesign and its individual components to determine which features contribute most strongly to critical-thinking gains, disciplinary competence, assessment validity, and responsible GenAI use across health-sciences cohorts.

11. Conclusions

GenAI has shifted from an incremental classroom aid to a force that destabilises the relationship between what STEM curricula certify and what graduates must be able to do. The durable response is not faster delivery of existing content but the deliberate redesign of learning outcomes, pedagogies, and assessment around disciplinary judgement, verification, intellectual independence, and the transparent and ethical use of GenAI.
The biological sciences make this challenge particularly visible: AI is simultaneously a tool for learning biology, a method transforming biological research, and an object of disciplinary critique. A biology graduate who cannot use, interpret, verify, and question AI-generated outputs may therefore be insufficiently prepared for contemporary scientific and professional practice. Organised through established learning-science, curriculum, capability, and adoption frameworks, and operationalised through an ADDIE-structured cell biology example, this review argues that intentional design—not technological novelty—determines whether GenAI strengthens or weakens disciplinary competence.
The proposed ADDIE-based course model should, however, be interpreted as an evidence-informed and testable design framework rather than an empirically validated intervention. Future studies should implement and compare the model across courses and student populations, evaluate its effects on disciplinary learning, critical verification, independent competence, assessment validity, responsible GenAI use, equity, and feasibility, and determine which individual design components contribute most strongly to these outcomes. Longitudinal research is also needed to establish whether such competencies are retained and transferred across subsequent courses, laboratory settings, and professional contexts.
Although specific tools will continue to change, the principles of constructive alignment, competency-based progression, inquiry, critical thinking, human oversight, equity, and academic integrity should remain central to STEM curriculum redesign.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/higheredu5030084/s1, Table S1: SANRA reporting checklist and manuscript location; Table S2: Eligibility criteria for the narrative review; Table S3a: Cell Biology Course Parameters; Table S3b. Assessment blueprint with AI-resilience features. References [5,34,60,61,62,64,69] are cited in Supplementary Materials.

Author Contributions

Conceptualization, C.P. and S.A.N.; methodology, C.P. and S.A.N.; investigation, C.P. and S.A.N.; writing, original draft preparation, C.P. and S.A.N.; writing, review and editing, C.P. and S.A.N.; supervision, S.A.N. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

During the preparation of this manuscript, the authors used Claude Opus 4.8 for text refinement and language editing. The authors reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ADDIEAnalysis, Design, Development, Implementation, and Evaluation
AIArtificial intelligence
GenAIGenerative artificial intelligence
ILOIntended learning outcome
LLMLarge language model
SANRAScale for the Assessment of Narrative Review Articles
STEMScience, technology, engineering, and mathematics
TPACKTechnological Pedagogical Content Knowledge
UTAUTUnified Theory of Acceptance and Use of Technology

References

  1. Schmidt, D.A.; Alboloushi, B.; Thomas, A.; Magalhaes, R. Integrating artificial intelligence in higher education: Perceptions, challenges, and strategies for academic innovation. Comput. Educ. Open 2025, 9, 100274. [Google Scholar] [CrossRef] [Scilit]
  2. Amar, S.; Benchouk, K. The Power Generative AI for Innovative Assessment Design in STEM Education Programs. In Generative AI in Higher Education Assessment: Theory, Practice, and Ethical Implications; Lahby, M., Schaeffer, S.E., Maleh, Y., Paliktzoglou, V., Eds.; Springer Nature: Cham, Switzerland, 2026; pp. 215–228. [Google Scholar] [CrossRef] [Scilit]
  3. Adejumo, A.A.; Oyelere, S.S.; Sanusi, I.T.; Suhonen, J. A systematic review of the impact of GenAI on learning performance, AI hallucinations, and problem-solving in computer science education. Comput. Educ. Artif. Intell. 2026, 10, 100570. [Google Scholar] [CrossRef] [Scilit]
  4. Qu, Y.; Tan, M.X.Y.; Wang, J. Disciplinary differences in undergraduate students’ engagement with generative artificial intelligence. Smart Learn. Environ. 2024, 11, 51. [Google Scholar] [CrossRef] [Scilit]
  5. Papaneophytou, C.; Nicolaou, S.A. Promoting Critical Thinking in Biological Sciences in the Era of Artificial Intelligence: The Role of Higher Education. Trends High. Educ. 2025, 4, 24. [Google Scholar] [CrossRef] [Scilit]
  6. Cotton, P.A.; Cotton, D.R.E. Beyond ChatGPT: A review of the use of AI tools in biological education. J. Biol. Educ. 2026, 1–25. [Google Scholar] [CrossRef] [Scilit]
  7. Francis, N.J.; Jones, S.; Smith, D.P. Generative AI in Higher Education: Balancing Innovation and Integrity. Br. J. Biomed. Sci. 2024, 81, 14048. [Google Scholar] [CrossRef] [Scilit]
  8. Hadra, M.; Cambridge, K.; Mesbah, M. Evaluating the accuracy and reliability of AI content detectors in academic contexts. Int. J. Educ. Integr. 2026, 22, 4. [Google Scholar] [CrossRef] [Scilit]
  9. Cotilla Conceição, J.M.; van der Stappen, E. The Impact of AI on Inclusivity in Higher Education: A Rapid Review. Educ. Sci. 2025, 15, 1255. [Google Scholar] [CrossRef] [Scilit]
  10. Biström, E.; Mollwing, J. AI in education and the future of teachers’ meaningful work. Front. Educ. 2026, 11, 1844085. [Google Scholar] [CrossRef] [Scilit]
  11. Tarlochan, F.; Alduais, A.; Chaaban, Y.; Du, X. Integrating sustainability into STEM education and career development: A scientometric and narrative review. Int. J. STEM Ed. 2025, 12, 62. [Google Scholar] [CrossRef] [Scilit]
  12. Vedrenne-Gutiérrez, F.; López-Suero, C.D.C.; De Hoyos-Bermea, A.; Mora-Flores, L.P.; Monroy-Fraustro, D.; Orozco-Castillo, M.F.; Martínez-Velasco, J.F.; Altamirano-Bustamante, M.M. The axiological foundations of innovation in STEM education—A systematic review and ethical meta-analysis. Heliyon 2024, 10, e32381. [Google Scholar] [CrossRef] [Scilit]
  13. Almasri, F. Exploring the Impact of Artificial Intelligence in Teaching and Learning of Science: A Systematic Review of Empirical Research. Res. Sci. Educ. 2024, 54, 977–997. [Google Scholar] [CrossRef] [Scilit]
  14. Yusuf, A.; Pervin, N.; Román-González, M.; Noor, N.M. Generative AI in education and research: A systematic mapping review. Rev. Educ. 2024, 12, e3489. [Google Scholar] [CrossRef] [Scilit]
  15. Sun, D.; Ba, S.; Cha, Y.; Yu, J.; Chiang, F.-K.; Dai, H.M.; Lim, C.-P. Empowering university teachers in higher education: A generative AI-responsive competency framework. Comput. Educ. Artif. Intell. 2026, 10, 100542. [Google Scholar] [CrossRef] [Scilit]
  16. Kibar, P.N.; Ilgaz, H. The intersection of artificial intelligence and instructional design practice: A systematic review. Educ. Technol. Res. Dev. 2026, 1–27. [Google Scholar] [CrossRef] [Scilit]
  17. Wang, S.; Wang, F.; Zhu, Z.; Wang, J.; Tran, T.; Du, Z. Artificial intelligence in education: A systematic literature review. Expert Syst. Appl. 2024, 252, 124167. [Google Scholar] [CrossRef] [Scilit]
  18. Baxter, K.A.; Sachdeva, N.; Baker, S. The Application of Cognitive Load Theory to the Design of Health and Behavior Change Programs: Principles and Recommendations. Health Educ. Behav. 2025, 52, 469–477. [Google Scholar] [CrossRef] [Scilit]
  19. Alam, M.N.; Islam, M.A.; Babiker, M.O.A.; Siddiqui, M.S.; Amin, M.B.; Oláh, J. AI-assisted learning tools and student learning outcomes: A cognitive load theory perspective. Comput. Hum. Behav. 2026, 21, 100986. [Google Scholar] [CrossRef] [Scilit]
  20. Jose, B.; Cherian, J.; Verghis, A.M.; Varghise, S.M.; S, M.; Joseph, S. The cognitive paradox of AI in education: Between enhancement and erosion. Front. Psychol. 2025, 16, 1550621. [Google Scholar] [CrossRef] [Scilit]
  21. Sun, Y.; Liao, Y.; Ma, X. Trusting AI to detect AI? A systematic evaluation of the reliability and robustness of current AIGC detection tools for student academic work. Comput. Educ. 2026, 249, 105616. [Google Scholar] [CrossRef] [Scilit]
  22. Stamov Roßnagel, C.; Lo Baido, K.; Fitzallen, N. Revisiting the relationship between constructive alignment and learning approaches: A perceived alignment perspective. PLoS ONE 2021, 16, e0253949. [Google Scholar] [CrossRef] [Scilit]
  23. Funa, A.A.; Gabay, R.A.E. Policy guidelines and recommendations on AI use in teaching and learning: A meta-synthesis study. Soc. Sci. Humanit. Open 2025, 11, 101221. [Google Scholar] [CrossRef] [Scilit]
  24. Taber, K.S. Educational Constructivism. Encyclopedia 2024, 4, 1534–1552. [Google Scholar] [CrossRef] [Scilit]
  25. Downey, R.J.; Youmans, K.; Villanueva Alarcón, I.; Nadelson, L.; Bouwma-Gearhart, J. Building Knowledge Structures in Context: An Exploration of How Constructionism Principles Influence Engineering Student Learning Experiences in Academic Making Spaces. Educ. Sci. 2022, 12, 733. [Google Scholar] [CrossRef] [Scilit]
  26. Ait Ali, D.; El Meniari, A.; El Filali, S.; Morabite, O.; Senhaji, F.; Khabbache, H. Empirical Research on Technological Pedagogical Content Knowledge (TPACK) Framework in Health Professions Education: A Literature Review. Med. Sci. Educ. 2023, 33, 791–803. [Google Scholar] [CrossRef] [Scilit]
  27. Willermark, S. The subject is the subject: Why TPACK matters in the era of GenAI. Soc. Sci. Humanit. Open. 2025, 12, 101893. [Google Scholar] [CrossRef] [Scilit]
  28. Venkatesh, V.; Morris, M.G.; Davis, G.B.; Davis, F.D. User Acceptance of Information Technology: Toward A Unified View. Manag. Inf. Syst. 2003, 27, 425–478. [Google Scholar] [CrossRef] [Scilit]
  29. Baethge, C.; Goldbeck-Wood, S.; Mertens, S. SANRA—A scale for the quality assessment of narrative review articles. Res. Integr. Peer Rev. 2019, 4, 5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Chamola, V.; Dave, D.; Goyal, I.; Sharma, S. Generative Artificial Intelligence in STEM Education: A Review of Applications, Benefits and Challenges. Expert Syst. 2026, 43, e70201. [Google Scholar] [CrossRef] [Scilit]
  31. Mei, P.; Brewis, D.N.; Nwaiwu, F.; Sumanathilaka, D.; Alva-Manchego, F.; Demaree-Cotton, J. If ChatGPT can do it, where is my creativity? Generative AI boosts performance but diminishes experience in creative writing. Comput. Hum. Behav. Artif. Hum. 2025, 4, 100140. [Google Scholar] [CrossRef] [Scilit]
  32. Lee, C.; Kim, J.; Lim, J.S.; Shin, D. Generative AI risks and resilience: How users adapt to hallucination and privacy challenges. Telemat. Inform. Rep. 2025, 19, 100221. [Google Scholar] [CrossRef] [Scilit]
  33. Ouwehand, K.; Lespiau, F.; Tricot, A.; Paas, F. Cognitive Load Theory: Emerging Trends and Innovations. Educ. Sci. 2025, 15, 458. [Google Scholar] [CrossRef] [Scilit]
  34. Wang, Y. Integrating large language models into medical undergraduate laboratory course to enhance bioethical competence: A quasi-experimental study. Front. Med. 2026, 12, 1745975. [Google Scholar] [CrossRef] [Scilit]
  35. Adams, N.E. Bloom’s taxonomy of cognitive learning objectives. J. Med. Libr. Assoc. 2015, 103, 152–153. [Google Scholar] [CrossRef] [Scilit]
  36. Bellas, F.; Naya-Varela, M.; Mallo, A.; Paz-Lopez, A. Education in the AI era: A long-term classroom technology based on intelligent robotics. Humanit. Soc. Sci. Commun. 2024, 11, 1425. [Google Scholar] [CrossRef] [Scilit]
  37. Pacheco-Olea, F.E.E.; Villacís-Macías, C. Artificial intelligence in STEM education: Applications, learning outcomes, and methodological gaps (2016–2025). Front. Educ. 2026, 11, 1791361. [Google Scholar] [CrossRef] [Scilit]
  38. Agbo, F.J.; Olivia, C.; Oguibe, G.; Sanusi, I.T.; Sani, G. Computing education using generative artificial intelligence tools: A systematic literature review. Comput. Educ. Open 2025, 9, 100266. [Google Scholar] [CrossRef] [Scilit]
  39. Gabriel, F.; Kennedy, J.; Marrone, R.; Leonard, S. Pragmatic AI in education and its role in mathematics learning and teaching. npj Sci. Learn. 2025, 10, 26. [Google Scholar] [CrossRef] [Scilit]
  40. He, P.; Krajcik, J. Artificial intelligence in science education: Global insights and future directions. Discip. Interdscip. Sci. Educ. Res. 2026, 8, 6. [Google Scholar] [CrossRef] [Scilit]
  41. Kohen-Vacs, D.; Usher, M.; Jansen, M. Integrating Generative AI into Programming Education: Student Perceptions and the Challenge of Correcting AI Errors. Int. J. Artif. Intell. Educ. 2025, 35, 3166–3184. [Google Scholar] [CrossRef] [Scilit]
  42. Yoon, H.; Hwang, J.; Lee, K.; Roh, K.H.; Kwon, O.N. Students’ use of generative artificial intelligence for proving mathematical statements. ZDM—Math. Educ. 2024, 56, 1531–1551. [Google Scholar] [CrossRef] [Scilit]
  43. Tsakalerou, M.; Akhmadi, S.; Balgynbayeva, A.; Kumisbek, Y. AI-assisted design synthesis and human creativity in engineering education. Front. Artif. Intell. 2026, 9, 1714523. [Google Scholar] [CrossRef] [Scilit]
  44. Tugce, D. Use of Artificial Intelligence in Biology Education: A Systematic Review of Literature. J. Educ. Sci. Environ. Health 2025, 11, 314–333. [Google Scholar] [CrossRef] [Scilit]
  45. Alam, M.M.; Farhaz, S.; Haq, S.M.A.; Ferdous, M. Generative Artificial Intelligence integration in higher education: A constructivist learning theory approach. Comput. Educ. Open 2026, 10, 100378. [Google Scholar] [CrossRef] [Scilit]
  46. Yiatas, D.; Jimoyiannis, A. Enacting Computer Science Curriculum Reform: The Case of Model and Experimental Lower Secondary Schools in Greece. Computers 2026, 15, 140. [Google Scholar] [CrossRef] [Scilit]
  47. Spatioti, A.G.; Kazanidis, I.; Pange, J. A Comparative Study of the ADDIE Instructional Design Model in Distance Education. Information 2022, 13, 402. [Google Scholar] [CrossRef] [Scilit]
  48. Rodríguez-Ortiz, M.Á.; Santana-Mancilla, P.C.; Anido-Rifón, L.E. Machine Learning and Generative AI in Learning Analytics for Higher Education: A Systematic Review of Models, Trends, and Challenges. Appl. Sci. 2025, 15, 8679. [Google Scholar] [CrossRef] [Scilit]
  49. Schöttl, F.; Zeller-Lanzl, J.; Gimpel, H. Basics for the Future—Basic Competency Shifts in the AI Era. Bus. Inf. Syst. Eng. 2026, 1–24. [Google Scholar] [CrossRef] [Scilit]
  50. Chiu, T.K.F. AI literacy and competency: Definitions, frameworks, development and future research directions. Interact. Learn. Environ. 2025, 33, 3225–3229. [Google Scholar] [CrossRef] [Scilit]
  51. Kong, S.C.; Hu, W. Unleashing human potential: An artificial intelligence competency framework for K–12 education. Comput. Educ. Artif. Intell. 2026, 10, 100556. [Google Scholar] [CrossRef] [Scilit]
  52. Stolpe, K.; Larsson, A.; Johansson Falck, M. Discipline-Specific AI literacy (DiSAIL): A theoretical framework for situated engagement with generative AI in education. Int. J. Technol. Des. Educ. 2026, 36, 1917–1932. [Google Scholar] [CrossRef] [Scilit]
  53. Brattli, H.; Utne, A.; Lynch, M. Assessment Validity in the Age of Generative AI: A Natural Experiment. Informatics 2026, 13, 56. [Google Scholar] [CrossRef] [Scilit]
  54. Laidlaw, P. Supporting assessment redesign in the age of AI: A critical practice-based framework for equitable change. Teach. High. Educ. 2026, 1–24. [Google Scholar] [CrossRef] [Scilit]
  55. Biggs, J. Enhancing teaching through constructive alignment. High. Educ. 1996, 32, 347–364. [Google Scholar] [CrossRef] [Scilit]
  56. Kiyan Tsunami, C.; Henríquez-Trujillo, A.R.; Ferreira-Meyers, K.; Mwanda, Z.; Rimal, J.; Pozu-Franco, J.; Delvaux, T.; Sibongwere, D.K.; Montalvo Navarrete, H.J.; Dasgupta, A.; et al. Guidelines for Integrating actionable A-SMART Learning Outcomes into the Backward Design Process. MedEdPublish 2024, 14, 242. [Google Scholar] [CrossRef] [Scilit]
  57. Kohout-Diaz, M. Making sense of AI in teacher education: A qualitative study of perceptions, practices and pedagogical tensions. Teach. Teach. Educ. 2026, 171, 105342. [Google Scholar] [CrossRef] [Scilit]
  58. Klein, C.R.R.; Klein, R. The extended hollowed mind: Why foundational knowledge is indispensable in the age of AI. Front. Artif. Intell. 2025, 8, 1719019. [Google Scholar] [CrossRef] [Scilit]
  59. Branch, R.M.; Merril, M.D. Characteristics of Instructional Design Models. In Trends and Issues in Instructional Design and Technology; Pearson: Boston, MA, USA, 2012; Volume 3, pp. 8–16. [Google Scholar]
  60. Bawn, M.; Francis, N.; Alvey, E.; Hassall, C.; Pires-daSilva, A.; Barra, P.; Hough, D.; Campbell, H.; Hardy, M.; Canet-Perez, J. Perspectives from a workshop: Intelligent assessment in the age of artificial intelligence. Adv. Physiol. Educ. 2026, 50, 73–82. [Google Scholar] [CrossRef] [Scilit]
  61. Nicolaou, S.A.; Nicolaou, P.; Dafli, E.; Bamidis, P.D.; Puig, B.; Lazar, G. Standardizing biology laboratory curriculum in health education: A blueprint for European undergraduate programs. Adv. Physiol. Educ. 2026, 50, 57–64. [Google Scholar] [CrossRef] [Scilit]
  62. Hossain, Z.; Bumbacher, E.; Brauneis, A.; Diaz, M.; Saltarelli, A.; Blikstein, P.; Riedel-Kruse, I.H. Design Guidelines and Empirical Case Study for Scaling Authentic Inquiry-based Science Learning via Open Online Courses and Interactive Biology Cloud Labs. Int. J. Artif. Intell. Educ. 2018, 28, 478–507. [Google Scholar] [CrossRef] [Scilit]
  63. Zhou, C.; Lewis, M. A mobile technology-based cooperative learning platform for undergraduate biology courses in common college classrooms. Biochem. Mol. Biol. Educ. 2021, 49, 427–440. [Google Scholar] [CrossRef] [Scilit]
  64. Cheng, Z.; Yang, J.; Shelnutt, K.P. Implementing GRATL and artificial intelligence in experiential learning of obesity physiology and etiology. Adv. Physiol. Educ. 2025, 49, 871–878. [Google Scholar] [CrossRef] [Scilit]
  65. Delcher, H.A.; Alsatari, E.S.; Haastrup, A.I.; Naaz, S.; Hayes-Guastella, L.A.; McDaniel, A.M.; Clark, O.G.; Katerski, D.M.; Prinsloo, F.O.; Roberts, O.R.; et al. Using ChatGPT as a tool for training nonprogrammers to generate genomic sequence analysis code. Biochem. Mol. Biol. Educ. 2025, 53, 433–444. [Google Scholar] [CrossRef] [Scilit]
  66. Ge, Y.; Wang, H.; Wang, D.; Yang, B. AI-enhanced virtual laboratories improve learning outcomes and student engagement in food microbiology education. Sci. Rep. 2026, 16, 22792. [Google Scholar] [CrossRef] [Scilit]
  67. Dafli, E.; Dratsiou, I.; Nisiforou, E.; Mylona, P.; Puig, B.; Lazar, G.; Nicolaou, P.; Bamidis, P.D.; Nicolaou, S.A. Toward Sustainable Digital Education in Biology: Evaluating Educators’ Perceptions and Adoption Intentions for a Virtual Laboratory Toolkit from Four European Contexts. Sustainability 2026, 18, 6445. [Google Scholar] [CrossRef] [Scilit]
  68. Poličar, P.G.; Špendl, M.; Curk, T.; Zupan, B. Automated assignment grading with large language models: Insights from a bioinformatics course. Bioinformatics 2025, 41, i21–i29. [Google Scholar] [CrossRef] [Scilit]
  69. Teckwani, S.H.; Wong, A.H.-P.; Luke, N.V.; Low, I.C.C. Accuracy and reliability of large language models in assessing learning outcomes achievement across cognitive domains. Adv. Physiol. Educ. 2024, 48, 904–914. [Google Scholar] [CrossRef] [Scilit]
  70. Kim, J.; Klopfer, M.; Grohs, J.R.; Eldardiry, H.; Weichert, J.; Cox, L.A.; Pike, D. Examining Faculty and Student Perceptions of Generative AI in University Courses. Innov. High. Educ. 2025, 50, 1281–1313. [Google Scholar] [CrossRef] [Scilit]
  71. Southworth, J.; Migliaccio, K.; Glover, J.; Glover, J.N.; Reed, D.; McCarty, C.; Brendemuhl, J.; Thomas, A. Developing a model for AI Across the curriculum: Transforming the higher education landscape via innovation in AI literacy. Comput. Educ. Artif. Intell. 2023, 4, 100127. [Google Scholar] [CrossRef] [Scilit]
  72. Ng, D.T.K.; Lee, M.; Tan, R.J.Y.; Hu, X.; Downie, J.S.; Chu, S.K.W. A review of AI teaching and learning from 2000 to 2020. Educ. Inf. Technol. 2023, 28, 8445–8501. [Google Scholar] [CrossRef] [Scilit]
  73. Crowther, G.J.; Funk, M.D.; Hennessey, K.M.; Lawrence, M.M. Frontier model chatbots can help instructors create, improve, and use learning objectives. Adv. Physiol. Educ. 2025, 49, 219–229. [Google Scholar] [CrossRef] [Scilit]
  74. Dhanvijay, A.K.D.; Pinjar, M.J.; Dhokane, N.; Sorte, S.R.; Kumari, A.; Mondal, H. Performance of Large Language Models (ChatGPT, Bing Search, and Google Bard) in Solving Case Vignettes in Physiology. Cureus 2023, 15, e42972. [Google Scholar] [CrossRef] [Scilit]
  75. Doğru, M.S.; Faulconer, E.K. ChatGPT as a Virtual Laboratory Teaching Assistant in Undergraduate Biology. Res. Sci. Educ. 2026, 56, 379–399. [Google Scholar] [CrossRef] [Scilit]
  76. Villarroel, V.; Bloxham, S.; Bruna, D.; Bruna, C.; Herrera-Seda, C. Authentic assessment: Creating a blueprint for course design. Assess. Eval. High. Educ. 2018, 43, 840–854. [Google Scholar] [CrossRef] [Scilit]
  77. Di Giusto, F. Digital assessment formats in higher education: An empirical analysis of the process efficiency of e-assessments compared to paper-pen exams. Int. J. Eval. Res. Educ. 2026, 15, 1151. [Google Scholar] [CrossRef] [Scilit]
  78. van der Vleuten, C.P.; Schuwirth, L.W. Assessing professional competence: From methods to programmes. Med. Educ. 2005, 39, 309–317. [Google Scholar] [CrossRef] [Scilit]
  79. Dawson, P. Assessment rubrics: Towards clearer and more replicable design, research and practice. Assess. Eval. High. Educ. 2017, 42, 347–360. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Examples of traditional and GenAI-era revised learning outcomes in STEM education.
Figure 1. Examples of traditional and GenAI-era revised learning outcomes in STEM education.
Higheredu 05 00084 g001
Figure 2. ADDIE model for systematically integrating GenAI into Course Design.
Figure 2. ADDIE model for systematically integrating GenAI into Course Design.
Higheredu 05 00084 g002
Figure 3. GenAI-resilient assessment.
Figure 3. GenAI-resilient assessment.
Higheredu 05 00084 g003
Table 2. The ADDIE model as a scaffold for GenAI-integrated course redesign, mapped to the guiding conceptual lenses.
Table 2. The ADDIE model as a scaffold for GenAI-integrated course redesign, mapped to the guiding conceptual lenses.
ADDIE PhaseGuiding Conceptual Lens(es)Core Design FocusAI Tool CategoryTeacher SupportStudent SupportBiological Sciences ApplicationReferences
AnalysisCognitive load theory (which competencies to retain vs. delegate)Identify learner needs, disciplinary expectations, and curriculum gapsLearning analytics; LLM-supported misconception mappingMap prior knowledge, AI familiarity, misconceptions, access gaps, and assessment vulnerabilitiesComplete diagnostic self-checks on knowledge and AI familiarityDetermine whether students can evaluate AI-generated biological explanations, data analyses, or experimental recommendations[61,70,72]
DesignConstructive alignment; revised Bloom’s taxonomy (outcome level)Specify outcomes, assessment evidence, and permitted AI useLLMs for outcome and rubric drafting; intelligent tutoring systemsDraft and refine ILOs; align outcomes, activities, assessment, and permitted AI useReview plain-language outcomes; self-check progress against course expectationsInclude outcomes requiring students to judge biological plausibility and verify GenAI outputs[5,71,73]
DevelopmentConstructivism and constructionismBuild resources, activities, rubrics, and safeguardsLLMs; AI-enhanced virtual labs; formative feedback toolsCreate verified summaries, quizzes, prompts, pre-labs, rubrics, and disclosure templatesUse virtual pre-labs, GenAI practice questions, and guided critique tasksCompare an AI-generated protocol with an established laboratory method[5,60,61,62,66,68,73]
ImplementationTPACK and UTAUT (capability and adoption); constructive alignment (permitted AI use)Run the course with transparency, equity, and human oversightChatbots; virtual teaching assistants; institutional AI-literacy resourcesProvide AI-literacy onboarding; monitor safety; confirm GenAI guidance before actionUse GenAI for preparation and review; disclose use; verify outputs against course materialsUse guided AI support in pre-labs, ethics tasks, and routine procedural queries[64,66,67,70,71,72,74,75]
EvaluationCognitive load theory (scaffolding vs. outsourcing); TPACK and UTAUT (feasibility)Test learning, assessment validity, equity, and feasibilityLearning analytics; formative feedback toolsEvaluate independent competence, GenAI-critique quality, workload, equity, and GenAI dependenceReceive formative feedback; demonstrate independent and GenAI-assisted reasoningAssess whether students can defend AI-assisted conclusions using biological evidence[5,67,68,69,70,71]
Iterative redesignAll lenses (re-enter the cycle)Revise the course for the next cycleLLM-supported reflection and redesign toolsRevise activities where GenAI displaced reasoning; update safeguards for new AI capabilitiesProvide structured feedback on AI use, access, and perceived learning supportModify pre-labs or assessment tasks where apparent competence exceeded bench performance[5,66,67]
Table 3. Proposed generic STEM course map with GenAI integration and human oversight.
Table 3. Proposed generic STEM course map with GenAI integration and human oversight.
Week(s)ComponentAI Integration PointHuman Oversight/Critical-Thinking SafeguardSupporting Evidence for Component
1OrientationDiscipline-specific AI-literacy briefing; tool selection; disclosureInstructor sets policy; students sign disclosure[64,70,71]
1–12Lecture/core contentGenAI-generated quizzes, summaries, and plain-language explainersInstructor verifies disciplinary accuracy before release[60,73]
1–12Self-paced studyIntelligent tutoring system or chatbot for reviewStudents cross-check outputs; reflect on GenAI errors[5,72,74]
2–10Practical preparation (pre-lab, pre-studio, or pre-problem set)GenAI-enhanced virtual or simulated preparation module (inquiry-aligned)Preparation quiz; in-person technique or method check[61,62,66,67]
2–10Practical session (laboratory, studio, or supervised computation)GenAI virtual teaching assistant for routine procedural questionsSafety net; human instructor confirms before action[66,75]
3–9Reflection and ethicsLLM-supported critical-thinking and disciplinary-ethics tasksBalanced view, source citation, AI-critique required[5,34,64]
4–11Collaborative workCooperative tasks with GenAI-assisted concept mapping or draft generationVisible group work; instructor live feedback[5,63]
OngoingFeedbackGenAI-assisted formative feedback on drafts, code, proofs, or designsHuman grading of summative work[67,68,69]
Table 4. Synthesis of the review questions and distinctive implications for STEM curriculum redesign.
Table 4. Synthesis of the review questions and distinctive implications for STEM curriculum redesign.
Review QuestionSynthesis from the ReviewDistinctive Implication for
STEM Curriculum Redesign
Does GenAI create a need for curriculum redesign in STEM higher education?Yes. GenAI can produce many conventional academic artifacts, including code, summaries, explanations, reports, and routine analyses, weakening their value as evidence of student competence.Curricula must move beyond artifact production toward judgment, verification, attribution, and accountable reasoning.
Is the required response classroom-level or curriculum-level?The problem is structural because it concerns what programmes define, teach, and certify as competence.Redesign must begin with graduate attributes, programme outcomes, and assessment strategy, not only with individual assignments.
Can redesign be shared across STEM disciplines?A shared redesign logic is possible, but disciplinary implementation differs across mathematics, computer science, engineering, and the natural sciences.Institutions can use a common design process, but each discipline must specify its own AI pressure points, competencies, and valid evidence of assessment.
How should curriculum and course design respond?Outcomes should be rewritten around higher-order disciplinary reasoning, AI-use conditions should be specified, and assessment should be constructively aligned with permitted AI use.Course teams should redesign learning outcomes, learning activities, and assessment blueprints together rather than treating AI as an add-on.
What does the biological sciences example demonstrate?The cell biology example shows how AI can support preparation, inquiry, feedback, and critique while preserving wet-lab competence and human oversight.Biology provides one worked model of how AI-enabled redesign can be implemented, evaluated, and adapted for other STEM fields.
Table 5. Applying the ADDIE-based template across STEM disciplines.
Table 5. Applying the ADDIE-based template across STEM disciplines.
DisciplineAnalysis: Competence to RenegotiateDesign: Higher-Order Outcome FocusDevelopment and Assessment: GenAI-Resilient EvidenceImplementation and Evaluation: Oversight and AI-Literacy Emphasis
Computer scienceWhich coding tasks are now generated automatically; where fluency still underpins judgmentDebugging, explanation, architecture, testing, and verification of generated codeLive coding, code review, oral defense of design decisionsAI-literacy in prompt critique and code verification; human review of higher-order design work
MathematicsWhich manipulations and proofs are handled by symbolic tools; which fluencies remain prerequisiteVerification, proof critique, and justification of method choiceProof critique, supervised derivation, oral explanationPreserve foundational fluency; human oversight of reasoning; verify rather than trust AI output
EngineeringWhich design and optimization steps AI can generate cheaplyConstraints, trade-offs, safety, and ethical judgment in designDesign justification, simulation critique, design reviewHuman sign-off on safety-critical judgment; AI-literacy in evaluating candidate solutions
Biology (worked example)AI-assisted literature, data analysis, and experimental interpretationBiological plausibility, critical evaluation of GenAI output, wet-lab competenceAI-critique tasks, lab portfolios, supervised practical examsBlended virtual and physical labs; AI-literacy onboarding; human-led summative grading
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Papaneophytou, C.; Nicolaou, S.A. Redesigning STEM Higher Education in the Era of Generative AI: From Curriculum Design to Classroom Practice. Trends High. Educ. 2026, 5, 84. https://doi.org/10.3390/higheredu5030084

AMA Style

Papaneophytou C, Nicolaou SA. Redesigning STEM Higher Education in the Era of Generative AI: From Curriculum Design to Classroom Practice. Trends in Higher Education. 2026; 5(3):84. https://doi.org/10.3390/higheredu5030084

Chicago/Turabian Style

Papaneophytou, Christos, and Stella A. Nicolaou. 2026. "Redesigning STEM Higher Education in the Era of Generative AI: From Curriculum Design to Classroom Practice" Trends in Higher Education 5, no. 3: 84. https://doi.org/10.3390/higheredu5030084

APA Style

Papaneophytou, C., & Nicolaou, S. A. (2026). Redesigning STEM Higher Education in the Era of Generative AI: From Curriculum Design to Classroom Practice. Trends in Higher Education, 5(3), 84. https://doi.org/10.3390/higheredu5030084

Article Metrics

Back to TopTop