Skip to Content
Education SciencesEducation Sciences
  • Perspective
  • Open Access

20 September 2026

18 Pages

Turning a Rock and a Hard Place into a Stepping Stone: Authentic STEM Assessment in the Age of Generative AI

Department of Curriculum and Pedagogy, Faculty of Education, The University of British Columbia, Vancouver, BC V6T1Z4, Canada

Abstract

The rapid emergence of generative artificial intelligence (AI) has intensified concerns about academic integrity, student cheating, and the future of assessment in STEM education. This paper argues that AI has not created a STEM assessment crisis; rather, it has exposed longstanding weaknesses that reward answer production, procedural fluency, and information reproduction rather than meaningful understanding, transfer, and application of knowledge. To operationalize this approach, the paper proposes a STEM-focused framework that distinguishes evidence of the disciplinary outcome, the learner’s reasoning process, the critical use and evaluation of AI, and individual understanding. Drawing on educational research and examples from STEM contexts, the paper contends that the central question is not how to prevent students from using AI, but what knowledge, skills, and disciplinary practices educators genuinely value and seek to assess. The paper proposes that AI functions as a litmus test for revealing assessment tasks that can be completed successfully without demonstrating disciplinary understanding. In response, it advocates a shift toward authentic assessment that emphasizes reasoning, creativity, communication, judgment, and the meaningful application of knowledge. Five examples are discussed: laboratory reports, problem solving, design challenges from the UBC Physics Olympics, scientific modelling and data interpretation, and STEM teacher education. Together, these examples illustrate how AI can support learning while authentic assessment focuses on students’ abilities to explain, justify, apply, evaluate, and defend their understanding in complex, meaningful contexts.

1. Introduction

Today, STEM educators find themselves between the proverbial rock and a hard place. Generative artificial intelligence (GenAI) has fundamentally changed the landscape of education. AI systems can now generate sophisticated explanations, solve complex scientific and mathematical problems, write computer code, analyze data, and produce reports that closely resemble human work (Kayyali, 2026; Maciejewski, 2025; Milner-Bolotin, 2025). As these tools become increasingly accessible, educators across secondary and post-secondary education are questioning how student learning can be assessed meaningfully while maintaining academic integrity and supporting deep learning (Beale, 2026; Ben-David Kolikant et al., 2020; Cotton et al., 2025; Kestin et al., 2025; Khan, 2023).
The response of many educational institutions has been swift. Universities and schools have introduced AI policies, adopted AI-detection software, increased invigilation, redesigned examinations, and, in some cases, attempted to prohibit AI altogether (Crompton et al., 2026; Pikhart & Al-Obaydi, 2025; Singh & Strzelecki, 2025; Thanh et al., 2025; Tomczyk, 2025; Wang, 2025). Such responses are understandable, but they focus primarily on preventing AI use rather than reconsidering what students should learn and how that learning should be demonstrated (Farrelly & Baker, 2023; Salinas-Navarro et al., 2024). This “AI ostrich” approach (Milner-Bolotin, 2025, 2026a), illustrated in Figure 1, assumes that assessment can be preserved simply by restricting access to new technologies.
Figure 1. Illustration of the ‘AI Ostrich’ approach to educational technology adoption.
The fundamental challenge, however, is not how to prevent students from using AI. AI is rapidly becoming an integral part of professional and everyday life, making its use increasingly inevitable. The more important question is how to design learning environments and assessments that remain meaningful when information generation, routine problem solving, and even scientific writing can increasingly be delegated to intelligent systems. As Mollick (2024) argues, education should prepare students to become co-intelligent with AI rather than encouraging them either to ignore it or rely on it uncritically. Before asking how to assess students, educators must first ask a more fundamental question: What knowledge, skills, and habits of mind should STEM education cultivate in an AI-rich world?
This question is particularly significant in STEM education (Milner-Bolotin, 2026a). Throughout history, STEM disciplines have continually adopted technologies that extended human capabilities: from slide rules and calculators to statistical software, computer simulations, and sophisticated laboratory instrumentation. GenAI represents the latest, and perhaps most transformative, stage of this evolution because it performs tasks traditionally associated with human reasoning rather than computation alone (Figure 2). Rather than viewing AI as an external disruption, STEM educators should recognize it as another powerful cognitive tool that requires a reconsideration of educational goals, teaching practices, and assessment.
Figure 2. The evolution of STEM tools over the last four centuries.
Thus, the emergence of GenAI has exposed a deeper issue. If students can successfully complete assessments using AI without demonstrating meaningful understanding, the problem may lie not only with the technology but also with the assessment itself. AI has become a powerful litmus test, revealing long-standing weaknesses in assessment practices that reward information reproduction and routine procedures more readily than conceptual understanding, disciplinary reasoning, creativity, and knowledge transfer.
The distinctive contribution of this Perspective is not the general claim that GenAI necessitates more authentic assessment, which is already well established. Rather, the paper extends this argument into STEM education by identifying AI-mediated assessment as a dual assessment problem: educators must evaluate both the quality of the disciplinary work and the learner’s capacity to exercise informed judgment while using AI. This paper therefore translates principles of authentic assessment into a STEM-focused framework that distinguishes evidence of disciplinary understanding, reasoning and decision-making, critical engagement with AI, and individual understanding (Table 1). This distinction is particularly important in STEM, where a technically correct answer may conceal faulty reasoning, inappropriate assumptions, fabricated evidence, or uncritical reliance on an AI-generated procedure. The framework is intended to help educators design assessments in which GenAI use is visible and examinable while foundational knowledge, disciplinary standards, and individual accountability remain central.
Table 1. A STEM-focused framework for assessing AI-mediated learning.
Accordingly, this paper examines how GenAI can serve as a catalyst for assessment reform in STEM education. Rather than focusing exclusively on preventing students from using AI, STEM educators should also consider which forms of knowledge, reasoning, collaboration, creativity, and professional judgment remain valuable in an AI-rich world and how assessments can intentionally cultivate and evaluate these capabilities. The paper argues that authentic assessment (Wiggins, 1998) provides a promising foundation for this work, while recognizing the continuing importance of foundational knowledge, structured assessment, assessment security, and individual accountability. Through STEM-focused examples, the paper illustrates how these principles can be translated into assessment practices that prepare learners for meaningful participation in future STEM study, work, and citizenship (Chachashvili-Bolotin et al., 2016).

2. AI Challenges as a Litmus Test for Meaningful STEM Assessment

Much of the current discussion surrounding AI in education frames it as a threat to learning, fueling an educational arms race in which AI tools become more sophisticated while institutions develop increasingly elaborate methods to detect or restrict their use. Although such efforts may provide temporary reassurance, they fail to address a more fundamental issue. If students can successfully complete assessments using AI without demonstrating the intended learning outcomes, the problem lies not only with the technology but also with the assessment itself. Rather than creating an assessment crisis, GenAI has exposed long-standing weaknesses in assessment design. In this sense, AI serves as a litmus test, revealing tasks that can be completed without demonstrating meaningful understanding.
This challenge is particularly significant in STEM education, where foundational knowledge is essential but insufficient. STEM graduates must also be able to apply knowledge in unfamiliar contexts, solve authentic problems, and transfer their understanding across situations. Yet fostering such transferable understanding has long been one of STEM education’s greatest challenges. Richard Feynman famously observed that students often develop “fragile” knowledge by learning through rote memorization rather than genuine understanding (Feynman & Leighton, 1985). Today, GenAI has made this fragility more visible than ever.
This concern is not new. Since as early as the 1990s, physics education researchers have shown that students earning high grades in traditional courses often struggle to apply fundamental concepts beyond familiar textbook problems (Hake, 1998, 2007; Mazur, 1997). Similar findings across STEM education have demonstrated that conventional assessments frequently reward memorization, procedural fluency, and algorithmic problem solving, rather than conceptual understanding, critical reasoning, or knowledge transfer (Gallagher et al., 2012; Greiff et al., 2013; Pellegrino & Hilton, 2012; Schneps et al., 1989; Wieman, 2012; Wieman & Perkins, 2005). These concerns have motivated decades of calls for more authentic assessments that require students to apply knowledge in novel contexts, justify their reasoning, communicate their thinking, and engage in genuine disciplinary practices (Finkelstein et al., 2005; Milner-Bolotin & Milner, 2023).
GenAI has not created these vulnerabilities; it has made them impossible to ignore. Consequently, the central question is not whether AI makes cheating easier, but whether our assessments genuinely measure the knowledge, skills, habits of mind, and disciplinary practices that we value. If they do not, increasingly sophisticated surveillance and detection technologies are unlikely to improve learning. Instead, AI challenges educators to redefine what constitutes convincing evidence of understanding and what learners should be able to accomplish independently, collaboratively, and with intelligent tools in an AI-rich world (Milner-Bolotin, 2025).

3. Beginning with the End in Mind

Assessment should begin with a clear understanding of desired learning outcomes (Wieman, 2012). Yet discussions of AI often focus on assessment methods rather than educational goals. If STEM education is intended to prepare learners for participation in STEM communities, then assessment should reflect the knowledge, skills, and practices valued by those communities. STEM professionals do far more than recall facts or produce correct answers. They formulate questions, investigate problems, analyze evidence, construct and evaluate explanations, communicate their ideas, collaborate with others, and refine their thinking in response to new evidence (Martinovic & Milner-Bolotin, 2022; Milner-Bolotin & Martinovic, 2025).
Accordingly, STEM assessment should prioritize conceptual understanding, critical thinking, problem framing, modelling, design, communication, collaboration, ethical judgment, and lifelong learning. These capacities become even more important in an AI-rich world because they rely on human judgment rather than the production of information alone. Although AI can generate answers, summarize information, and propose solutions, learners themselves must remain responsible for curiosity, judgment, ethical decision-making, and the consequences of how AI-generated information is used.

4. Authentic STEM Assessment as a Response to AI

Calls for more authentic approaches to assessment are not new. For more than three decades, STEM education researchers have argued that assessment should extend beyond the reproduction of information to evaluate students’ ability to apply knowledge, solve meaningful problems, communicate their reasoning, and engage in authentic disciplinary practices (Angelo & Cross, 1993; Hake, 2007; Tobias, 2000; Van Driel et al., 2001). Although authentic assessment, inquiry-based learning, project-based learning, and portfolio assessment have been widely advocated, their implementation in STEM education has remained uneven (Martinovic & Milner-Bolotin, 2020). Designing, implementing, and evaluating authentic assessments requires considerable pedagogical expertise, time, and institutional support: resources that many educators continue to lack despite growing expectations for educational innovation (Ben-David Kolikant et al., 2020; Dawson, 2017; Martinovic & Milner-Bolotin, 2020; Rodrigues et al., 2025).
Authentic assessment has long been recognized as an essential component of effective STEM education because it evaluates students’ ability to apply knowledge in meaningful contexts rather than simply recall information or follow routine procedures (Wiggins, 1998). Instead of asking students to reproduce what they know, authentic assessment requires them to investigate meaningful questions, solve realistic problems, justify their reasoning, communicate their thinking, and reflect on their learning. Such practices closely mirror the work of scientists, engineers, and mathematicians, who construct, evaluate, and communicate knowledge rather than merely produce correct answers.
The emergence of GenAI makes these characteristics even more important. If AI can successfully complete an assessment by retrieving or generating information, then the assessment may reveal relatively little about students’ conceptual understanding or their ability to apply knowledge. Authentic assessment therefore shifts the emphasis from information reproduction to disciplinary thinking, creativity, judgment, collaboration, and knowledge transfer. Rather than asking whether students used AI, educators should ask whether an assessment requires learners to engage in forms of thinking that remain educationally valuable in an AI-rich world.
A widely accepted framework for understanding authentic assessment was proposed by Gulikers et al. (2004); they argue that authenticity extends beyond the assessment task itself. They identify five interrelated dimensions of authentic assessment: task, physical context, social context, assessment product or performance, and assessment criteria. Together, these dimensions encourage educators to design assessments that more closely resemble the ways knowledge is created, applied, and evaluated in professional STEM practice. Subsequent work has further refined these principles. Ashford-Rowe et al. (2014) emphasize informed judgment, meaningful performances, collaboration, and reflection, while Villarroel et al. (2018) argue that authentic assessment should be intentionally embedded within course design rather than treated as an isolated assignment. Collectively, this body of research suggests that authenticity is not defined by a particular assessment format but by the extent to which learners engage in meaningful disciplinary practices while developing transferable knowledge and skills.
The Gulikers et al. (2004) framework provides the conceptual lens for the five examples presented in this paper. Rather than attempting to maximize every dimension of authenticity simultaneously, each example highlights different aspects of authentic assessment depending on its educational purpose, disciplinary context, and learning outcomes. Collectively, however, the examples illustrate how authentic tasks, contexts, collaboration, assessment products, and transparent criteria can be combined to create assessments that remain meaningful in an AI-rich educational environment.
Importantly, authentic assessment should not be interpreted as replacing foundational knowledge with open-ended projects. Deep conceptual understanding, procedural fluency, and disciplinary knowledge remain essential prerequisites for meaningful inquiry and problem solving. Rather, authentic assessment seeks to integrate foundational knowledge with meaningful application, enabling learners to transfer what they know to unfamiliar situations, justify their decisions, collaborate effectively, and exercise informed judgment.
Building on the conceptual framework of Gulikers et al. (2004), we distinguish four complementary components of evidence that instructors can use to assess AI-mediated STEM learning: the disciplinary outcome, the student’s reasoning process, the critical use and evaluation of AI, and evidence of individual understanding (Table 1). These components operationalize authentic assessment by clarifying not only what students produce, but also what instructors should examine when determining whether meaningful disciplinary learning has occurred.
These components need not be weighted equally in every assessment. Their relative importance should reflect the intended learning outcomes, the disciplinary context, and the role permitted for GenAI. For example, an engineering-design task may place greater emphasis on the disciplinary product and iterative decision-making, whereas a teacher-education task may emphasize professional judgment and critical evaluation of AI-generated materials. A brief oral defence, individualized follow-up question, or transfer task can provide evidence of individual understanding without requiring an additional major assessment.
The five examples that follow demonstrate how these principles can be enacted across diverse STEM learning environments. Each example is analyzed through Gulikers et al.’s (2004) five dimensions of authenticity and the four components of assessment evidence summarized in Table 1. Together, the examples illustrate how authentic assessment can position GenAI as a cognitive tool that supports learning while strengthening opportunities for students to provide meaningful evidence of their understanding.

5. Examples of Authentic STEM Assessment in the Age of AI

The examples that follow draw on my experience of more than three decades of teaching science at the secondary and post-secondary levels, conducting STEM education research, and nearly two decades of preparing and mentoring STEM teachers. While grounded in my own educational practice, each example is interpreted using Gulikers et al.’s (2004) framework to illustrate how different dimensions of authentic assessment can be implemented in STEM education in the age of AI.

5.1. Authentic Assessment of STEM Laboratory Activities

Traditional STEM laboratory activities often culminate in students submitting a written laboratory report. Today, however, GenAI can produce substantial portions of the text, generate graphs, assist with statistical analyses, and even create plausible datasets. Faced with this reality, educators may be tempted to focus on detecting AI-generated content. A more productive response is to redesign the assessment so that it evaluates scientific reasoning rather than the production of a polished report.
Instead of assessing only the final written product, students can be asked to justify their experimental design, defend their interpretation of evidence, explain unexpected results, critique methodological limitations, evaluate alternative explanations, and reflect on how their understanding evolved throughout the investigation. They may also discuss how AI was used during the investigation and critically evaluate both its contributions and its limitations.
Viewed through the lens of Gulikers et al.’s (2004) framework, this assessment extends authenticity well beyond the laboratory report itself. The task reflects the practices of scientific inquiry by requiring students to analyze evidence and justify their decisions. The physical context is grounded in an authentic laboratory investigation, while the social context can be strengthened through peer discussions or collaborative investigations. The emphasis on oral or written scientific argumentation broadens the assessment product from a single report to multiple forms of evidence demonstrating understanding. Finally, transparent assessment criteria focus students’ attention on the quality of their reasoning, interpretation of evidence, and reflective thinking rather than on producing a flawless document.
By shifting the focus from generating a report to demonstrating scientific reasoning, this approach positions AI as a cognitive tool that can support learning while ensuring that assessment remains centred on capacities for which learners themselves must remain intellectually and ethically responsible: interpreting evidence, exercising judgment, communicating scientific ideas, and constructing meaningful understanding.

5.2. Authentic Assessment of STEM Problem Solving

Rather than assigning routine problem sets with predetermined solutions, educators can engage students in analyzing authentic STEM problems characterized by incomplete information, multiple constraints, and competing solutions. Students may use GenAI as one resource among many during the problem-solving process, but they remain responsible for evaluating AI-generated suggestions, identifying their limitations, integrating multiple sources of evidence, and justifying their decisions.
Assessment therefore shifts from evaluating whether students arrive at the correct answer to examining how they reason through complex problems. Students may be asked to explain why they selected a particular solution, defend their assumptions, critique alternative approaches, and reflect on how AI influenced their thinking. In this way, AI becomes part of the learning environment rather than an external threat, while assessment focuses on disciplinary reasoning, judgment, and decision-making.
Viewed through the lens of Gulikers et al.’s (2004) framework, this approach reflects several dimensions of authentic assessment. The task mirrors the open-ended nature of real scientific and engineering problem solving, where solutions are rarely straightforward. The social context can be strengthened through collaborative discussions in which students compare and critique alternative solutions, while the assessment product extends beyond a final answer to include evidence of reasoning, reflection, and justification. Transparent assessment criteria emphasize the quality of students’ thinking, their evaluation of evidence, and their responsible use of AI rather than simply the correctness of the solution.
AI can also play an important role as a learning partner throughout this process. Rather than supplying answers, it can provide hints, alternative representations, guiding questions, explanations, and timely feedback that help students overcome conceptual obstacles and continue making progress (Khan, 2023; Mollick, 2024). This form of scaffolding can benefit struggling learners who require additional support while also challenging advanced students to explore more sophisticated extensions.
The potential value of individualized instructional support has a long history in educational research. Bloom’s (1984) influential “2 Sigma Problem” reported exceptionally large learning gains for students receiving one-to-one tutoring compared with conventional classroom instruction. However, the often-cited two-standard-deviation advantage should not be treated as a general estimate of tutoring effectiveness. In a later review, VanLehn (2011) reported substantially more modest average effects, approximately d = 0.79 for human tutoring and d = 0.76 for intelligent tutoring systems relative to non-tutored instruction. These findings also underscore an important distinction: expert human tutoring, conventional intelligent tutoring systems, and contemporary generative-AI tutors represent different forms of instructional support and should not be treated as pedagogically or empirically equivalent.
Recent evidence nevertheless suggests that carefully designed GenAI tutoring can support learning under particular conditions. In a randomized controlled trial in undergraduate physics, Kestin et al. (2025) found greater learning gains for students using a specially designed AI tutor than for students receiving in-class active-learning instruction, with reported effects ranging from approximately 0.73 to 1.3 standard deviations depending on the analysis. However, these findings should be interpreted within the conditions of the study: the tutor used expert-designed prompts and pedagogical scaffolding, focused on specific introductory physics topics and learning objectives, and was evaluated over a relatively short intervention. The study therefore provides promising evidence for purposefully designed AI-supported tutoring, rather than evidence that generic GenAI use reproduces the effects of expert human tutoring.
The broader educational goal of providing timely, responsive support at scale also predates GenAI. Approaches such as Just-in-Time Teaching (Novak et al., 1999), for example, use evidence of students’ emerging understanding to adapt instruction before misconceptions become entrenched. GenAI creates new possibilities for extending such responsiveness, but its educational value depends on the quality of its pedagogical design rather than personalization alone.
Ultimately, the value of AI depends on how learning is assessed. When assessments reward only correct answers, AI can become a substitute for thinking. When they require students to justify decisions, evaluate evidence, transfer knowledge, and reflect on their reasoning, AI becomes a tool for learning rather than a shortcut around it. In this way, authentic assessment keeps the focus on the judgment, responsibility, and disciplinary understanding that learners themselves must demonstrate, while allowing AI to support rather than replace their intellectual development.

5.3. Authentic Assessment of Design Challenges: Lessons from the UBC Physics Olympics

A compelling example of authentic STEM assessment is provided by the University of British Columbia (UBC) Physics Olympics (Milner-Bolotin, 2026b; Milner-Bolotin et al., 2019; UBC Department of Physics and Astronomy, 2026; Zhang et al., 2024). This annual outreach event engages more than one thousand British Columbia secondary school students in a series of physics challenges that require both conceptual understanding and practical application. Unlike traditional examinations, success depends not on recalling formulas or following routine procedures but on applying physics creatively in unfamiliar situations (Table 1).
One example is the pre-build challenge, in which student teams design, construct, test, and refine a device that meets specified performance criteria (Figure 3). Throughout the design process, students make predictions, evaluate evidence, troubleshoot unexpected problems, and iteratively improve their designs. GenAI can support this process by suggesting design ideas, explaining relevant physics concepts, or helping students troubleshoot technical issues. However, successful completion of the challenge still requires students to engage in experimental testing, collaborative decision-making, and iterative refinement of the physics design. Ultimately, success depends on students’ ability to apply disciplinary knowledge, learn from failure, and make informed design decisions based on evidence.
Figure 3. Students demonstrate their pre-built pendulum apparatus during the 2026 UBC Physics Olympics.
Similarly, students may use AI to prepare for Fermi estimation challenges by exploring solution strategies or practising estimation techniques (Ärlebäck & Albarracín, 2019, 2024). During the competition, however, they must identify relevant variables, make reasonable assumptions, judge the plausibility of their estimates, and justify their reasoning under significant time constraints. These activities require scientific judgment and quantitative reasoning that extend well beyond generating numerical answers.
Viewed through the lens of Gulikers et al.’s (2004) framework, the UBC Physics Olympics exemplifies multiple dimensions of authentic assessment. The tasks closely resemble the open-ended challenges encountered by scientists and engineers, while the competition itself provides an authentic physical context in which students test ideas under realistic conditions. The team-based format creates a rich social context that requires collaboration, communication, and shared decision-making. Students demonstrate their learning through tangible performances and products: working prototypes, experimental results, and reasoned solutions, rather than through written responses alone. Finally, clearly defined assessment criteria, communicated in advance, evaluate not only successful outcomes but also the effective application of physics principles within realistic constraints.
The UBC Physics Olympics illustrates a central principle of authentic assessment in the age of AI: foundational knowledge remains essential, but it is no longer sufficient. Students must be able to transfer their knowledge to unfamiliar situations, collaborate effectively, make informed judgments, and adapt their thinking in response to evidence. By aligning assessment with the authentic practices of STEM professionals, activities such as the UBC Physics Olympics position AI as a valuable design and learning partner while ensuring that assessment remains focused on building learners’ capacities for creativity, critical thinking, and evidence-based decision-making.

5.4. Assessing STEM Modelling and Data Interpretation

Scientific modelling and data interpretation are fundamental practices across STEM disciplines (Martinovic & Milner-Bolotin, 2021, 2022). Scientists, engineers, and mathematicians develop models to explain phenomena, test hypotheses, make predictions, and refine their understanding in light of new evidence. Similarly, interpreting data requires identifying patterns, evaluating evidence, recognizing uncertainty, and drawing scientifically justified conclusions (Table 1). These practices have become even more important as Gen.
AI enables students to rapidly generate graphs, fit mathematical models, perform statistical analyses, and propose interpretations.
Rather than assessing the final model or data analysis alone, authentic assessment should focus on the reasoning that underpins these products. Students should justify why a particular model was selected, explain its underlying assumptions, evaluate its limitations, compare alternative models, and defend their conclusions using evidence. For example, when investigating population growth, climate change, motion, or electrical circuits, students may use AI to generate multiple candidate models and visualizations. Their task, however, is to determine which model best represents the phenomenon, identify sources of uncertainty, evaluate the quality of the available evidence, and explain how new evidence might require revising the model.
Viewed through the lens of Gulikers et al.’s (2004) framework, modelling activities naturally embody several dimensions of authentic assessment. The task reflects a core practice of STEM professionals who routinely develop, evaluate, and refine models to understand complex systems. The physical context can be grounded in authentic datasets collected through laboratory investigations, simulations, or real-world observations. Students demonstrate their understanding through meaningful assessment products, including models, evidence-based interpretations, and scientific arguments rather than isolated calculations. When modelling is conducted collaboratively, the social context encourages students to critique alternative explanations, negotiate interpretations, and communicate their reasoning. Finally, transparent assessment criteria emphasize the quality of modelling decisions, evaluation of evidence, and critical reflection rather than simply producing accurate graphs or statistical outputs.
Importantly, authentic assessment recognizes that scientists increasingly use computational and AI-assisted tools in their professional work. The goal is therefore not to prohibit AI, but to ensure that students understand the assumptions, limitations, and implications of the models and analyses it generates. Just as calculators did not replace mathematical reasoning, AI does not replace scientific reasoning.
This approach shifts assessment from evaluating the products generated by AI to evaluating students’ capacity to interpret evidence, construct and revise models, justify conclusions, and communicate scientific arguments. By aligning assessment with authentic STEM practice, modelling activities can position AI as a powerful analytical partner while keeping students responsible for demonstrating critical evaluation, evidence-based judgment, and scientific sense-making.

5.5. Authentic Assessment in STEM Teacher Education

The challenges posed by GenAI are particularly significant in STEM teacher education because today’s teacher candidates will enter classrooms where AI is readily available to both teachers and students (Magic School, 2026). Preparing future teachers as if AI does not exist is therefore neither realistic nor responsible. Instead, STEM teacher education should prepare them to use AI critically, ethically, and pedagogically while maintaining a clear focus on student learning.
This perspective has led me to redesign assessment in my own STEM teacher education courses. Rather than attempting to prevent teacher candidates from using AI, I assume that they will use it. The assessment challenge is therefore not whether AI was used, but whether future teachers can demonstrate the pedagogical knowledge, professional judgment, and reflective decision-making required of effective STEM educators.
For example, submitting a polished lesson plan is no longer sufficient evidence of teaching competence because GenAI can now produce complete lesson plans within seconds (Magic School, 2026). Instead, teacher candidates are asked to justify their instructional decisions, anticipate student misconceptions, adapt lessons for diverse learners, select appropriate technologies, design meaningful assessments, and critically evaluate and revise AI-generated suggestions. In doing so, they demonstrate not only what they have designed but also why particular pedagogical choices are appropriate for a specific learning context.
Viewed through the lens of Gulikers et al.’s (2004) framework, these assessments reflect authentic professional practice. The task mirrors the work of practising teachers who routinely design, adapt, and evaluate instruction. The physical context is grounded in realistic classroom situations, curriculum expectations, and student needs. Opportunities for peer review, collaborative lesson design, and professional discussion strengthen the social context, reflecting the collaborative nature of teaching. Rather than submitting a lesson plan alone, teacher candidates produce meaningful assessment products that may include lesson plans, teaching demonstrations, reflective commentaries, and revisions based on feedback. Finally, transparent assessment criteria emphasize pedagogical reasoning, responsiveness to learners, critical evaluation of AI-generated ideas, and evidence-informed instructional decision-making rather than the production of polished documents.
Applying the four components in Table 1 makes the evidence used to evaluate this work more explicit. The disciplinary and professional outcome is represented by the lesson plan, teaching demonstration, or assessment design and is evaluated for pedagogical appropriateness, content accuracy, and responsiveness to learners. The reasoning process is documented through written or oral justifications, anticipated misconceptions, and revisions made in response to feedback. Critical use of AI is demonstrated through selected AI outputs, verification of their accuracy, identification of limitations, and explanation of which suggestions were accepted, modified, or rejected. Finally, individual understanding can be assessed through a brief oral defence or an individualized classroom scenario requiring the teacher candidate to explain or adapt a pedagogical decision. These sources of evidence allow the instructor to evaluate both the quality of the professional product and the teacher candidate’s independent judgment.
Importantly, this approach recognizes that effective teachers routinely consult colleagues, curriculum resources, research literature, and increasingly AI-powered tools when planning instruction. The goal is therefore not to assess whether teacher candidates can work without assistance, but whether they can evaluate available resources critically and make informed decisions that enhance student learning. AI becomes a professional tool that supports planning, creativity, and reflection, while assessment remains focused on dimensions of professional teaching for which teacher candidates themselves must remain responsible: pedagogical judgment, ethical decision-making, adaptability, empathy, and responsiveness to learners.
Authentic assessment in STEM teacher education therefore mirrors the realities of professional practice while preparing future teachers to use AI responsibly and effectively. By aligning assessment with the authentic work of teachers, it ensures that AI supports professional learning without replacing the reflective judgment, creativity, and educational expertise that define effective STEM teaching.
Taken together, these five examples illustrate how authentic assessment can shift attention from the production of answers and polished artifacts toward richer evidence of disciplinary reasoning and professional practice. Through Gulikers et al.’s (2004) five dimensions (task, physical context, social context, assessment product or performance, and assessment criteria), the examples demonstrate different ways of aligning assessment with the practices STEM learners are ultimately being prepared to undertake. They also illustrate an important boundary to this argument: authenticity does not make an assessment inherently secure from AI assistance, nor does it by itself guarantee validity, reliability, equity, or knowledge transfer. Rather, authentic assessment provides a framework for deciding what meaningful performance should look like; questions about how confidently that performance can be attributed to an individual learner and how fairly and reliably it can be evaluated require additional assessment-design considerations (Dawson, 2020).

6. Limitations and Words of Caution

This paper does not argue that authentic assessment should replace all traditional forms of assessment. Foundational knowledge, conceptual understanding, procedural fluency, and the ability to recall and apply core disciplinary concepts remain essential components of STEM learning. In many contexts, carefully designed examinations and other structured assessments continue to provide valuable evidence of students’ readiness for more advanced learning. The challenge is not to choose between traditional and authentic assessment, but to achieve an appropriate balance between them.
Authenticity, however, should not be conflated with assessment security or with broader claims about assessment quality. An assessment may closely resemble disciplinary or professional practice while still allowing substantial assistance from GenAI or other sources. Laboratory reports, design documentation, lesson plans, modelling tasks, and reflective writing can all be authentic in the sense described by Gulikers et al. (2004), yet provide limited assurance about which aspects of the submitted work represent a student’s independent capability. Dawson (2020) distinguishes assessment security from broader questions of assessment quality, emphasizing both authentication of student work and control over assessment conditions. Authentic assessment should therefore be understood as one component of a broader assessment system rather than as an inherently AI-proof form of assessment. Depending on the purpose and stakes of the assessment, educators may need to combine authentic tasks with complementary evidence such as observation of performance, brief oral questioning, staged work, or carefully designed individual assessment.
Authentic assessment also presents important practical challenges. Designing meaningful tasks, evaluating complex student work, and providing rich formative feedback require considerable time, expertise, and institutional support. Large class sizes, accreditation requirements, limited resources, and competing faculty responsibilities may constrain the extent to which authentic assessment can be implemented. At the same time, AI offers opportunities to reduce routine administrative work, enabling educators to devote more time to designing learning experiences, providing feedback, and supporting student learning.
Authentic assessment also raises questions of reliability, comparability, equity, and feasibility. Complex performances can be more difficult to evaluate consistently across students, assessors, and contexts, particularly in large classes or high-stakes settings. Moreover, strategies intended to strengthen evidence of individual learning, such as oral defence, process documentation, or additional reflection, may create disproportionate demands for multilingual learners, students with disabilities, or students with limited time and access to academic support. These considerations reinforce the need for transparent criteria, multiple forms of evidence, reasonable accommodations, and assessment designs that balance authenticity with fairness and feasibility. Similarly, although authentic tasks can create opportunities for students to apply knowledge across contexts, knowledge transfer should be treated as an intended and empirically examinable outcome rather than an automatic consequence of authenticity.
Equally important, AI itself should be approached with caution. Despite rapid advances, AI systems remain susceptible to inaccuracies, bias, hallucinations, and uneven performance across domains. Overreliance on AI may also limit the development of essential disciplinary knowledge and independent reasoning if used uncritically. Consequently, this paper advocates neither the uncritical adoption of AI nor the wholesale replacement of traditional assessment. Rather, it argues for balanced assessment approaches that preserve the strengths of foundational knowledge assessment while creating opportunities for students to demonstrate disciplinary reasoning, creativity, professional judgment, and the responsible use of AI-supported tools.

7. Moving Forward

Rather than asking how assessment can survive AI, educators should ask what kinds of learning matter most in an AI-rich world. This shift requires rethinking assessment from first principles, beginning with the knowledge, skills, and dispositions that STEM education seeks to cultivate.
Several principles may help guide this transition. First, assessment should prioritize capacities for which learners themselves must remain intellectually and ethically responsible, including critical judgment, ethical reasoning, creativity, collaboration, communication, and the ability to navigate uncertainty. Second, authentic assessment should value learning processes as well as final products by encouraging students to document their thinking, justify decisions, reflect on their learning, and demonstrate knowledge transfer across contexts. Students should own their learning (Milner-Bolotin, 2001) and take pride not only in what they produce but also in how they learn.
Third, transparency should replace secrecy. Rather than attempting to detect or prohibit AI use, educators should create assessment environments in which students openly document how AI contributed to their work, critically evaluate its outputs, and justify when, why, and how it was used. Such transparency shifts attention from AI use itself to students’ capacity to exercise informed professional judgment.
Fourth, assessment should mirror authentic disciplinary practice. Scientists, engineers, mathematicians, and educators increasingly use AI-supported tools in their professional work. Educational assessment should acknowledge this reality while ensuring that students develop the expertise required to evaluate, adapt, and responsibly employ these technologies.
Finally, assessment should remain grounded in educational purpose. Technology should not determine what is valued; educational goals should determine how technology is used. This principle reflects the Deliberate Pedagogical Thinking with Technology framework (Milner-Bolotin, 2020), which emphasizes intentional instructional decision-making grounded in pedagogical goals rather than technological novelty. Achieving this vision will require sustained investment in teacher education and high-quality professional development.
An essential outcome of this transformation is AI literacy (Crompton et al., 2026; Tomczyk, 2025; Walter, 2024). Just as scientific literacy requires learners to evaluate evidence critically and understand the limitations of scientific models, AI literacy requires students to understand the capabilities, limitations, and biases of AI systems. Students should learn to formulate effective prompts, evaluate AI-generated outputs, verify claims using credible sources, recognize hallucinations, and make informed decisions about when AI should, and should not, be used. Authentic assessment provides an ideal context for developing these capabilities because it requires students to explain, justify, critique, and reflect upon their interactions with AI rather than simply producing AI-assisted products.

8. Conclusions

GenAI has not created a crisis in educational assessment; rather, it has made longstanding weaknesses in assessment increasingly difficult to ignore. Like litmus paper revealing the underlying properties of a solution, AI exposes assessment tasks that can be completed successfully through information reproduction, routine procedures, or externally generated responses without necessarily demonstrating meaningful understanding. As AI becomes increasingly capable of performing many tasks traditionally assigned to students, educators are compelled to revisit fundamental questions: What does it mean to learn? What constitutes expertise? What knowledge and skills should students be able to demonstrate independently, and what should they be able to accomplish responsibly with intelligent tools?
This paper contributes to the growing conversation on AI and assessment by framing AI-mediated STEM assessment as a dual assessment problem: educators must evaluate both the quality of disciplinary work and the learner’s capacity to exercise informed judgment while using AI. Building on Gulikers et al.’s (2004) five-dimensional model of authentic assessment, the paper proposes a STEM-focused framework that distinguishes four complementary forms of evidence: the disciplinary outcome, the learner’s reasoning process, the critical use and evaluation of AI, and individual understanding. The five examples illustrate how these forms of evidence can be elicited across laboratory investigation, complex problem solving, engineering design, scientific modelling and data interpretation, and STEM teacher education. Across these contexts, assessment should provide opportunities for learners not merely to produce answers or products, but to explain, justify, apply, evaluate, adapt, and defend their thinking.
At the same time, authentic assessment should not be regarded as an inherently AI-proof solution. Authenticity and assessment security are related but distinct design considerations. An assessment may closely reflect disciplinary or professional practice while still allowing substantial assistance from GenAI or other sources. Nor does authenticity by itself guarantee validity, reliability, comparability, equity, or knowledge transfer. These considerations are especially important in large classes and high-stakes settings, where educators must balance meaningful disciplinary performance with feasible, fair, and sufficiently trustworthy evidence of individual learning (Dawson, 2020). Authentic assessment should therefore be viewed as one component of a broader assessment system that may incorporate multiple forms of evidence, transparent criteria, appropriate accommodations, and, where warranted, opportunities for students to explain or defend their reasoning.
This perspective also changes the educational response to GenAI. Rather than relying primarily on surveillance, AI-detection software, or increasingly restrictive policies, educators can create assessment environments in which AI use is transparent and students remain accountable for the quality of their reasoning and decisions. This does not mean that all AI use should be permitted in all assessments. There will continue to be important contexts in which students must demonstrate foundational knowledge or particular capabilities independently. The challenge is to make these decisions deliberately, based on learning goals rather than fear of technological change.
These limitations do not diminish the value of authentic assessment; rather, they clarify its purpose. Its strength lies not in making assessment impervious to AI, but in aligning what students are asked to demonstrate with the intellectual and professional practices that STEM education seeks to cultivate. Scientists, engineers, mathematicians, and educators increasingly work with sophisticated technological and AI-supported tools. Educational assessment should acknowledge this reality while ensuring that learners themselves develop and demonstrate the disciplinary knowledge, critical judgment, creativity, ethical responsibility, and capacity to evaluate evidence required to use such tools wisely.
Ultimately, the future of STEM assessment depends less on distinguishing human-generated work from AI-generated work than on clarifying what learners themselves need to know, understand, and be able to do with and without AI, and on designing assessments that make those capabilities visible. This requires AI literacy, thoughtful assessment design, and deliberate pedagogical decision-making grounded in educational purpose rather than technological novelty.
STEM educators therefore do find themselves between a rock and a hard place: rapidly advancing AI on one side and long-established assessment practices that no longer fully reflect the goals and realities of contemporary education on the other. Yet this tension also creates an opportunity. By using AI to expose weaknesses in existing assessment, embracing authentic disciplinary practice while recognizing its limitations, and remaining focused on meaningful evidence of student learning, educators can turn this apparent obstacle into a stepping stone toward more thoughtful, rigorous, and future-oriented STEM assessment.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

The illustrations presented in Figure 1 and Figure 2 were generated with the assistance of ChatGPT (GPT-5.6, OpenAI 2026) and subsequently reviewed and adapted by the author. The photograph used in Figure 3 is the property of the author. AI was used to help proofread the final version of the paper. I would also like to thank the anonymous reviewers for their thoughtful feedback, which helped me strengthen the paper’s arguments and improve its readability.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Angelo, T. K., & Cross, K. P. (1993). Classroom assessment techniques: A handbook for college teachers (2nd ed.). Jossey-Bass Publishers. [Google Scholar]
  2. Ashford-Rowe, K., Herrington, J., & Brown, C. (2014). Establishing the critical elements that determine authentic assessment. Assessment & Evaluation in Higher Education, 39(2), 205–222. [Google Scholar] [CrossRef] [Scilit]
  3. Ärlebäck, J. B., & Albarracín, L. (2019). The use and potential of fermi problems in the STEM disciplines to support the development of twenty-first century competencies. ZDM, 51(6), 979–990. [Google Scholar] [CrossRef] [Scilit]
  4. Ärlebäck, J. B., & Albarracín, L. (2024). Fermi problems as a hub for task design in mathematics and STEM education. Teaching Mathematics and its Applications: An International Journal of the International Mathematical Association, 43(1), 25–37. [Google Scholar] [CrossRef] [Scilit]
  5. Beale, R. (2026). Adapting university policies for generative AI: Opportunities, challenges, and policy solutions in higher education. arXiv, arXiv:2506.22231. [Google Scholar]
  6. Ben-David Kolikant, Y., Martinovic, D., & Milner-Bolotin, M. (Eds.). (2020). STEM teachers and teaching in the digital era: Professional expectations and advancement in 21st century schools. Springer. [Google Scholar]
  7. Bloom, B. S. (1984). The 2 Sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4–16. [Google Scholar] [CrossRef]
  8. Chachashvili-Bolotin, S., Milner-Bolotin, M., & Lissitsa, S. (2016). Examination of factors predicting secondary students’ interest in tertiary STEM education. International Journal of Science Education, 38(2), 366–390. [Google Scholar] [CrossRef] [Scilit]
  9. Cotton, D. R. E., Wyness, L., Jane, B., & Cotton, P. A. (2025). Redefining assessments in the age of AI. In J. R. Corbeil, & M. E. Corbeil (Eds.), Teaching and learning in the age of generative AI (pp. 283–308). Routledge. [Google Scholar] [CrossRef] [Scilit]
  10. Crompton, H., Burke, D., Nickel, C., Bozkurt, A., Miao, F., Sharples, M., Greene, J. A., Parsons, D., Gill-Simmen, L., Edmett, A., Pegrum, M., de Waard, I., Bonk, C. J., Garcia, M. B., Curry, J. H., Lindsey, L., Yang, M., Marshall, S., Bali, M., … Yu, S. (2026). Governing generative AI in higher education: A global Delphi study on policy and practice. International Journal of Educational Technology in Higher Education, 23(1), 21. [Google Scholar] [CrossRef] [Scilit]
  11. Dawson, P. (2017). Assessment rubrics: Towards clearer and more replicable design, research and practice. Assessment & Evaluation in Higher Education, 42(3), 347–360. [Google Scholar] [CrossRef] [Scilit]
  12. Dawson, P. (2020). Defending assessment security in a digital world: Preventing E-cheating and supporting academic integrity in higher education (1st ed.). Routledge. [Google Scholar] [CrossRef] [Scilit]
  13. Farrelly, T., & Baker, N. (2023). Generative artificial intelligence: Implications and considerations for higher education practice. Education Sciences, 13(11), 1109. [Google Scholar] [CrossRef] [Scilit]
  14. Feynman, R. P., & Leighton, R. (1985). “Surely you’re joking, Mr. Feynman!”: Adventures of a curious character. W.W. Norton and Company. [Google Scholar]
  15. Finkelstein, N. D., Adams, W. K., Keller, C. J., Kohl, P. B., Perkins, K. K., Podolefsky, N. S., Reid, S., & LeMaster, R. (2005). When learning about the real world is better done virtually: A study of substituting computer simulations for laboratory equipment. Physical Review Special Topics—Physics Education Research, 1(1), 010103. [Google Scholar] [CrossRef] [Scilit]
  16. Gallagher, C., Hipkins, R., & Zohar, A. (2012). Positioning thinking within national curriculum and assessment systems: Perspectives from Israel, New Zealand and Northern Ireland. Thinking Skills and Creativity, 7(2), 134–143. [Google Scholar] [CrossRef] [Scilit]
  17. Greiff, S., Holt, D. V., & Funke, J. (2013). Perspectives on problem solving in educational assessment: Analytical, interactive, and collaborative problem solving. The Journal of Problem Solving, 5(2), 5. [Google Scholar] [CrossRef] [Scilit]
  18. Gulikers, J. T. M., Bastiaens, T. J., & Kirschner, P. A. (2004). A five-dimensional framework for authentic assessment. Educational Technology Research and Development, 52(3), 67–86. [Google Scholar] [CrossRef] [Scilit]
  19. Hake, R. R. (1998). Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses. American Journal of Physics, 66(1), 64–74. [Google Scholar] [CrossRef] [Scilit]
  20. Hake, R. R. (2007). Six lessons from the physics education reform effort. Latin-American Journal of Physics Education, 1(1), 24–31. [Google Scholar]
  21. Kayyali, M. (2026). The future of AI in special education: Tools for individualized support. In Teacher perspectives and responsible practice for integrating AI in the classroom (p. 26). IGI Global Scientific Publishing. [Google Scholar] [CrossRef] [Scilit]
  22. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15(1), 17458. [Google Scholar] [CrossRef] [Scilit]
  23. Khan, S. (2023). Khanmigo: Bringing AI into classrooms responsibly. Khan Academy. Available online: https://www.khanmigo.ai/ (accessed on 1 September 2026).
  24. Maciejewski, W. (2025). A tale of two exams. Canadian Journal of Science, Mathematics and Technology Education, 25(1), 66–78. [Google Scholar] [CrossRef] [Scilit]
  25. Magic School. (2026). Inside a K–12 AI integration that worked for every teacher. Available online: https://www.magicschool.ai/case-studies/west-vancouver-schools (accessed on 4 January 2026).
  26. Martinovic, D., & Milner-Bolotin, M. (2020). Discussion: Teacher professional development in the era of change. In Y. Ben-David Kolikant, D. Martinovic, & M. Milner-Bolotin (Eds.), STEM teachers and teaching in the era of change: Professional expectations and advancement in 21st century schools (pp. 185–197). Springer. [Google Scholar]
  27. Martinovic, D., & Milner-Bolotin, M. (2021). Examination of modelling in K-12 STEM teacher education: Connecting theory with practice. STEM Education, 1(4), 279–298. [Google Scholar] [CrossRef] [Scilit]
  28. Martinovic, D., & Milner-Bolotin, M. (2022). Problematizing STEM: What it is, what it is not, and why it matters. In C. Michelsen, A. Beckmann, V. Freiman, U. Thomas Jankvist, & A. Savard (Eds.), 15 Years of MACAS (mathematics and its connections to the arts and sciences) (pp. 135–162). Springer Nature. [Google Scholar] [CrossRef] [Scilit]
  29. Mazur, E. (1997). Peer instruction: User’s manual. Prentice Hall. [Google Scholar]
  30. Milner-Bolotin, M. (2001). The effects of the topic choice in project-based instruction on undergraduate physical science students’ interest, ownership, and motivation [Doctoral dissertation, The University of TX at Austin]. [Google Scholar]
  31. Milner-Bolotin, M. (2020). Deliberate pedagogical thinking with technology in STEM teacher education. In Y. Ben-David Kolikant, D. Martinovic, & M. Milner-Bolotin (Eds.), STEM teachers and teaching in the era of change: Professional expectations and advancement in 21st century schools (pp. 201–219). Springer. [Google Scholar]
  32. Milner-Bolotin, M. (2025). From TV to AI: Evolving challenges and enduring questions in STEM education. Canadian Journal of Science Mathematics and Technology Education, 25(3), 809–818. [Google Scholar] [CrossRef] [Scilit]
  33. Milner-Bolotin, M. (2026a). The AI ostrich dilemma: Why STEM teacher education cannot look away. In M. Kayyali (Ed.), Encyclopedia of higher education (Vol. 6, p. 20). Apple Academic Press and Routledge/Taylor & Francis. [Google Scholar]
  34. Milner-Bolotin, M. (2026b). UBC physics olympics as a model of educational transformation: STEM outreach bridging schools, universities, and future STEM teachers. In M. Kayyali (Ed.), Encyclopedia of higher education (Vol. 6, p. 25). Apple Academic Press and Routledge/Taylor & Francis. [Google Scholar]
  35. Milner-Bolotin, M., Liao, T., & McKenna, J. (2019). UBC physics olympics: Forty-one years of province-wide physics outreach. International Newsletter on Physics Education: International Commission on Physics Education—International Union of Pure and Applied Physics, 70, 5–6. [Google Scholar]
  36. Milner-Bolotin, M., & Martinovic, D. (2025). Creative approaches for 21st-century science, technology, engineering, and mathematics teacher education: From theory to practice to policy. Future in Education Research, 3(1), 5–12. [Google Scholar] [CrossRef] [Scilit]
  37. Milner-Bolotin, M., & Milner, V. (2023). Breaking the vicious circle of secondary science education with twenty-first-century technology: Smartphone physics labs. In G. P. Thomas, & H. J. Boon (Eds.), Challenges in science education: Global perspectives for the future (pp. 177–199). Springer International Publishing. [Google Scholar] [CrossRef] [Scilit]
  38. Mollick, E. (2024). Co-intelligence: Living and working with AI (1st ed.). Penguin Random House. [Google Scholar]
  39. Novak, G. M., Gavrini, A., Christian, W., & Patterson, E. (1999). Just-in-time teaching: Blending active learning with web technology. Prentice Hall. [Google Scholar]
  40. Pellegrino, J. W., & Hilton, M. L. (Eds.). (2012). Education for life and work: Developing transferable knowledge and skills in the 21st century. The National Academies Press. [Google Scholar] [CrossRef] [Scilit]
  41. Pikhart, M., & Al-Obaydi, L. H. (2025). Reporting the potential risk of using AI in higher education: Subjective perspectives of educators. Computers in Human Behavior Reports, 18, 100693. [Google Scholar] [CrossRef] [Scilit]
  42. Rodrigues, S., Milner-Bolotin, M., & Moosvi, F. (2025, June 27–July 2). Instructor experiences with alternative grading at the university of British Columbia. 30th ACM Conference on Innovation and Technology in Computer Science Education V. 2, Nijmegen, The Netherlands. [Google Scholar] [CrossRef] [Scilit]
  43. Salinas-Navarro, D. E., Vilalta-Perdomo, E., Michel-Villarreal, R., & Montesinos, L. (2024). Using generative artificial intelligence tools to explain and enhance experiential learning for authentic assessment. Education Sciences, 14(1), 83. [Google Scholar] [CrossRef] [Scilit]
  44. Schneps, M. H., Sadler, P. M., Woll, S., & Crouse, L. (1989). A private universe [Grant]. S. Burlington. Available online: http://www.learner.org/teacherslab/pup/ (accessed on 1 September 2026).
  45. Singh, S., & Strzelecki, A. (2025). Academics as adopters of generative AI: An application of diffusion of innovations theory. Education and Information Technologies, 31(2), 621–645. [Google Scholar] [CrossRef] [Scilit]
  46. Thanh, D. T., Dinh, V. P., & Nguen Van, H. (2025). Exploring AI integration in biology self-learning: A TAM-based analysis among high school students in Ho Chi Minh City. Eurasia Journal of Mathematics, Science & Technology Education, 21(10), em 2716. [Google Scholar] [CrossRef] [Scilit]
  47. Tobias, S. (2000). From innovation to change: Forging a physics education reform agenda for the 21st century. American Journal of Physics, 68(2), 103–104. [Google Scholar] [CrossRef] [Scilit]
  48. Tomczyk, Ł. (2025). AI in education—Mapping theoretical frameworks for digital literacy: DigComp 2.2, AI literacy competency framework, AI literacy TPACK, the machine learning education framework, the UNESCO AI competency framework for teachers and other innovative approaches. In New media pedagogy: Research trends, methodological challenges, and successful implementations. Springer. [Google Scholar]
  49. UBC Department of Physics and Astronomy. (2026). UBC physics olympics. UBC. Available online: http://physoly.phas.ubc.ca/ (accessed on 4 March 2026).
  50. Van Driel, J. H., Beijaard, D., & Verloop, N. (2001). Professional development and reform in science education: The role of teachers’ practical knowledge. Journal of Research in Science Teaching, 38(2), 137–158. [Google Scholar] [CrossRef]
  51. VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221. [Google Scholar] [CrossRef] [Scilit]
  52. Villarroel, V., Bloxham, S., Bruna, D., Bruna, C., & Herrera-Seda, C. (2018). Authentic assessment: Creating a blueprint for course design. Assessment & Evaluation in Higher Education, 43(5), 840–854. [Google Scholar] [CrossRef] [Scilit]
  53. Walter, Y. (2024). Embracing the future of artificial intelligence in the classroom: The relevance of AI literacy, prompt engineering, and critical thinking in modern education. International Journal of Educational Technology in Higher Education, 21(1), 15. [Google Scholar] [CrossRef] [Scilit]
  54. Wang, C. (2025). Exploring students’ generative AI-assisted writing processes: Perceptions and experiences from native and nonnative English speakers. Technology, Knowledge and Learning, 30(3), 1825–1846. [Google Scholar] [CrossRef] [Scilit]
  55. Wieman, C. E. (2012). Applying new research to improve science education. Issues in Science and Technology, 29(1), 25–32. [Google Scholar]
  56. Wieman, C. E., & Perkins, K. (2005). Transforming physics education. Physics Today, 58(11), 36–42. [Google Scholar] [CrossRef] [Scilit]
  57. Wiggins, G. (1998). Educative assessment. Designing assessments to inform and improve student performance. Jossey Bass Publishers. [Google Scholar]
  58. Zhang, B., Ren, H., & Milner-Bolotin, M. (2024). 科学拓展活动:实践·体验·合作 ——以UBC物理奥林匹克活动为例 [Science outreach activities: Practice · experience · cooperation—Take UBC physics olympics as an example]. Physics Teacher (China), 43(11), 90–96. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.