Skip to Content
Education SciencesEducation Sciences
  • Article
  • Open Access

23 April 2026

Educator–GenAI Partnership Model for Assessment Design to Foster Higher-Order Thinking

,
,
,
,
and
1
School of IT, National Academy of Professional Studies (NAPS), Melbourne, VIC 3000, Australia
2
School of IT, National Academy of Professional Studies (NAPS), Sydney, NSW 2010, Australia
3
School of IT and Engineering, Melbourne Institute of Technology (MIT), Melbourne, VIC 3000, Australia
*
Authors to whom correspondence should be addressed.

Abstract

The rise of generative artificial intelligence (GenAI) is creating new opportunities for assessment design in universities, particularly in subjects that emphasize analytical and creative skills. This paper introduces the Educator–GenAI Partnership Model, an iterative five-stage model that helps educators create assessments that foster higher-order thinking (HOT). The model is grounded in constructive alignment and Bloom’s taxonomy, with a central emphasis on preserving human oversight to ensure educators retain control over assessment validity, academic integrity, and the ethical use of AI. The model maps out the unique strengths and responsibilities of both educators and GenAI, showing how each plays a distinct role in the assessment design process. It illustrates how GenAI can support the rapid generation of assessment tasks and marking rubrics, while positioning educators as critical decision-makers who only review, adapt, and iteratively refine AI-generated outputs to ensure alignment with higher-order learning outcomes. Overall, this paper presents a structured and practical model for utilizing GenAI responsibly in assessment design, thereby strengthening academic rigor while enhancing efficiency for educators.

1. Introduction

Universities around the world aim to develop higher-order thinking (HOT) skills in their students, including the ability to analyze, evaluate, and create knowledge or solutions (Lubbe et al., 2025). These skills are vital for building strong problem-solving and innovation abilities, making them key learning goals in today’s education system, particularly as generative artificial intelligence (GenAI) continues to shape how students learn and demonstrate their understanding (Weng et al., 2024). Despite the widespread adoption of GenAI tools by students, many universities and higher education providers continue to rely heavily on traditional assessment approaches (Bittle & El-Gayar, 2025). These older methods often focus on lower-level skills, such as memorization and comprehension, and reflect a pre-GenAI mindset that fails to align with the changing learning context (Reyna, 2025; Weng et al., 2024).
Continuing to rely on traditional assessment methods creates a significant mismatch with the demands of today’s workforce and industry for graduates (Bouckaert, 2023; Reyna, 2025). Modern GenAI tools can increasingly be capable of generating accurate and complex responses to tasks requiring only lower-level cognitive skills, which weakens the credibility of traditional assessments (Corbin et al., 2025; Zhao et al., 2024). This gap also increases the risk of compromising academic integrity, as students may turn to GenAI to complete tasks that do not require genuine higher-level thinking or personal engagement. Guidelines on assessment design introduced in Australian universities are reported in (Babai Shishavan, 2024), which highlights common concerns about the challenges GenAI tools pose to traditional assessments, and the assessments in the era of GenAI should emphasize critical thinking, incorporate contextual elements, focus on process-oriented and staged assessments, and use collaborative or in-class assessments.

1.1. Urgent Need Meets Reality

There is a growing need to redesign assessment tasks to focus on those HOT skills, such as analyzing information, synthesizing concepts, and making sound judgments, in the age of GenAI (Clark, 2025; Lubbe et al., 2025). Experts are warning that if we fail to adapt, we risk increased academic dishonesty and a loss of public trust in university credentials (Bittle & El-Gayar, 2025; Whitham et al., 2023). Accordingly, institutional policies should be updated to include explicit guidance on AI literacy and ethical use within the assessment framework (Ilieva et al., 2025; Shailendra et al., 2024; Xia et al., 2024). This shift necessitates moving traditional assessment tasks towards authentic, process-oriented tasks that require creativity, critical reasoning, and contextual judgment—the kind of nuanced, complex skills that even the best AI struggles to replicate effectively (Kadel et al., 2024). While a complete overhaul of assessment practices is evidently necessary, academics face significant constraints such as the lack of time, comprehensive training in new pedagogical skill sets, and robust institutional infrastructure support to tackle such a massive redesign effort (Association of Pacific Rim Universities [APRU], 2025).
The above discussion highlights a persistent gap between the capabilities of GenAI and the ability of current assessment systems to ensure genuine learning. Addressing this gap will require coordinated institutional investment, evidence-based redesign principles, and sustained professional development to support academics in creating assessments that remain rigorous, ethical, and ensure higher-order thinking outcomes.

1.2. Contribution and Outline

This paper proposes a conceptual model which emphasizes the increased role of technology in every step of student assessment design and evaluation. The model reinforces the traditional assessment design constructs; however, in the presence of GenAI being deeply rooted in the assessment lifecycle. It provides a strategy for systematic adoption and usage of GenAI to alleviate challenges introduced by it.
This model contributes to educational theory by positioning GenAI not as a replacement but as a scaffold that amplifies educators’ creativity and supports alignment. This model integrates human pedagogical expertise with the generative capacity of machines to address one of higher education’s enduring challenges: designing assessments that effectively elicit and develop HOT skills. Specifically, the model extends the traditional assessment discourse in two aspects:
  • By placing GenAI in parallel with the educator, this model acknowledges the technology’s substantial influence on assessment design while upholding the educator’s role as the primary authority.
  • By utilizing process-driven checkpoints and feedback loops, the model ensures that HOT-based assessment design remains transparent and measurable, even when GenAI is integrated into the assignment design workflow.
The rest of the paper is structured as follows: Section 2 provides a brief overview of the theoretical background for model development. This section reinforces the need for active involvement of the educator in the entire assessment process to ensure authenticity and pedagogy of uncertainty (Section 2.3). Section 3 presents a review of the recent literature on HOT-based assessment design. Section 4 presents an Educator–GenAI Partnership Model and its applicable phases. While the proposed model is conceptual in nature, its design and applicability have been tailored for higher education environments, followed by implementation aspects and limitations in Section 5. Finally, Section 6 concludes with key insights and outlines future research directions.

2. Theoretical Background

This section presents the critical aspects of the pedagogy that underpins fundamental principles of the assessment design. It further examines how these principles are increasingly important due to the growing integration of GenAI within teaching and learning environments. This section establishes the rationale for a new model of assessment design where human and technology play their roles while ensuring that the use of technology remains subject to appropriate human oversight.

2.1. Higher-Order Thinking (HOT)

In the age of GenAI, HOT has become a central focus in assessment design as educators shift from traditional knowledge recall toward deeper cognitive engagement. Within Bloom’s taxonomy and its later revision (Clark, 2025), HOT encompasses skills such as analysis, evaluation, and creation. These abilities require learners to move beyond memorizing content and instead apply, synthesize, and critique ideas meaningfully. These capacities are increasingly important for preparing learners to navigate complex, uncertain, and interdisciplinary problem spaces. Assessments that target HOT, therefore, aim to capture learners’ ability to interpret evidence, apply concepts to novel situations, generate innovative solutions, and justify their reasoning through coherent argumentation.
Designing assessments that effectively measure HOT requires careful alignment between learning outcomes, instructional approaches, and task design. Unlike traditional tests that can often be answered through rote learning, HOT-focused assessments typically involve open-ended problems, authentic contexts, reflective tasks, case analyses, project-based tasks, or design challenges. These approaches encourage learners to demonstrate not only what they know but how their thinking, reasoning, and decision-making processes work. When integrated thoughtfully, such assessments support deeper learning, foster an awareness of design scaffolding, and promote transferable skills that are essential in industry-ready graduates. As educational contexts evolve, particularly with the integration of GenAI, designing assessments that authentically capture HOT remains both an imperative and a challenge. However, without properly aligned assessments, students tend to focus on procedural or factual tasks that fail to capture critical and creative thinking (Anderson & Krathwohl, 2001; Brookhart, 2010). The challenges in designing HOT assessments are fundamentally rooted in the scientific discourse surrounding construct (e.g., critical thinking) representation and psychometric validity (American Educational Research Association et al., 2014; Messick, 1995).
Firstly, the requirement for high cognitive fidelity significantly affects the ability to design complex assessments. Ensuring that the assessment task triggers the intended construct requires analysis, evaluation, or creation involve extensive planning, alignment to learning outcomes, iterative refinement, and validation. These processes demand far more effort than traditional recall-based tasks. Given the substantial teaching and administrative loads carried by higher education educators, this additional burden can deter the adoption of advanced assessment formats.
The second challenge lies in translating abstract cognitive models into practical assessment tasks. While many educators are familiar with revised versions of Bloom’s taxonomy or other cognitive models, applying these models to develop authentic tasks, prompts, and scenarios requires a high level of design expertise. Designing effective assessments to stimulate deep, complex, recursive thinking requires understanding of how experts in a subject think and the mental steps students must take. It also involves identifying the visible actions that demonstrate mastery. Without clear scaffolding and exemplars, educators may struggle to operationalize HOT principles in assessment design.
The third challenge lies in balancing reliability and validity, i.e., designing reliable rubrics and scoring criteria for HOT assessments (American Educational Research Association et al., 2014). Because tasks aimed at fostering HOT often allow diverse and open-ended responses, creating rubrics that are both rigorous and flexible is challenging. The design process must account for variability in student approaches with multiple valid outcomes, and needs to evaluate reasoning quality as well as correctness. This increases the design workload and requires iterative testing and calibration before implementation.
Finally, the design process for HOT assessments often lacks technological and computational support. Traditional assessment tools are typically built around structured or objective formats, providing limited functionality for designing complex, open-ended tasks. Without dedicated tools or AI-assisted mechanisms, the design process relies heavily on manual effort, creativity, and experience, which can slow innovation and reduce scalability in higher education contexts. Therefore, there is a need for a structured approach that leverages technology to translate complex educational theory into practical assessment reality, paving the way for a collaborative partnership between educators and GenAI.

2.2. Constructive Alignment

Constructive alignment provides a coherent framework for teaching and assessment that directly supports intended learning outcomes (Biggs et al., 2022). The underlying principle is that learning is most effective when students construct meaning through activities and assessments intentionally aligned with the intended learning outcomes. In this view, learning outcomes define the level and type of understanding required; teaching and learning activities scaffold learners toward those outcomes; and assessments measure the extent to which learners have achieved them. Constructive alignment, thus, shifts the role of the evaluation from merely grading performance to guiding and evidencing learning. By centering the design around outcomes rather than content coverage or instructor preference, this framework promotes transparency, coherence, and fairness in curriculum design.
Constructive alignment helps ensure that assessment tasks genuinely reflect and require the depth and nature of learning intended. Accordingly, aligned assessment tasks must move beyond recall-based formats and instead incorporate authentic, performance-based, or inquiry-driven tasks. This alignment also requires criteria and rubrics that clearly articulate performance standards at the appropriate cognitive level. When assessments, learning activities, and intended outcomes are cohesively aligned, they reinforce one another to create a more meaningful and integrative learning experience. However, the increasing presence of GenAI in education further underscores the need for carefully aligned assessments that maintain academic integrity while continuing to capture genuine learner achievement (Biggs et al., 2022).
Assessment design must move beyond task construction to encompass the broader learning environment, ensuring that learning outcomes, supporting teaching strategies (including the integration of GenAI), and assessment methods operate in synergy. Thoughtful, theory-informed assessment design is essential for achieving the higher-level learning goals that contemporary higher education aims to cultivate. Without this alignment, higher-order learning outcomes risk remaining aspirational rather than realized in practice. While the value of constructive alignment in promoting higher-order learning is widely recognized, achieving this level of alignment in practice remains challenging due to the complexity, time demands, and expertise required in designing HOT assessments. GenAI offers new possibilities to support educators in this design process by assisting in task creation, rubric development, and iterative refinement.

2.3. Human Oversight for Pedagogical Authority

Although GenAI can support multiple aspects of assessment development, the role of the educator remains critical in ensuring constructive alignment between outcomes, learning activities, and assessment tasks. GenAI is capable of generating assessment tasks, rubrics, and feedback suggestions, but it cannot independently determine the pedagogical intent, disciplinary standards, or contextual learning priorities of a course. Educators must, therefore, guide GenAI output by interpreting curriculum requirements, determining qualification levels, validating cognitive demand, and refining generated material to ensure that assessments genuinely reflect HOT goals rather than surface-level complexity. Educators must also critically review the GenAI-generated content for accuracy, relevance, and fairness, ensuring it aligns with disciplinary values and institutional expectations (Sharma et al., 2025). In this way, GenAI becomes an assistive co-designer, while educators retain responsibility for making informed pedagogical decisions that uphold constructive alignment and ensure that assessment practices promote meaningful and measurable learning.
GenAI may produce fabricated information, suffer from model drift, or fail to address discipline-specific nuances and cultural contexts (Jackson, 2025). Furthermore, ethical concerns, such as data privacy, academic integrity, and algorithmic bias, necessitate vigilant monitoring by the human educator. Consequently, the integration of GenAI introduces challenges that necessitate strong human oversight despite the considerable efficiency benefits. Human oversight ensures that GenAI consistently serves as a supportive tool rather than a substitute for professional judgment. By maintaining pedagogical authority, educators safeguard the integrity of HOT outcomes and foster meaningful, ethical learning experiences (Bellas et al., 2025; Hau, 2025). This partnership model blends with the pedagogy of uncertainty, where the educator’s expertise lies not in the certainty of the students’ or tools’ output, but in the professional wisdom required to navigate its unpredictability.
The partnership model advocates the human-in-the-loop (HITL) philosophy such that the educator remains the final decision-maker. This ensures all GenAI outputs align with learning outcomes and ethical standards. Their authority helps in validating that GenAI-generated outputs align with curriculum standards, institutional policies, ethical requirements, and the intended learning outcomes. Without such oversight, there is a risk of introducing inaccuracies, bias, or pedagogical misalignment that could compromise academic integrity.

4. Educator–GenAI Partnership Model

The Educator–GenAI Partnership Model is developed based on the foundational principle of Human Oversight for Pedagogical Authority (Bellas et al., 2025; Hau, 2025; Ilieva et al., 2025; Sterz et al., 2024). This principle is guiding the model’s design around assessment validity, Bloom’s higher cognitive levels, balancing the strengths and weaknesses of GenAI, leveraging human judgment, and facilitating continuous improvement. The primary factors that led to the design of the five-phase model are detailed below:
  • AI Literacy and Ethical Transparency: The model emphasizes that educators must develop AI literacy through professional development to critically evaluate and effectively integrate GenAI tools. This includes adhering to institutional and professional ethical guidelines, ensuring ethical AI use, transparency, and proper attribution (Ilieva et al., 2025; Xia et al., 2024). All GenAI-assisted tasks should be fully disclosed to stakeholders (faculty, students, institutions) to maintain trust and accountability (Clark, 2025).
  • Structured Alignment and Oversight: Educators must adopt structured oversight strategies when using GenAI-generated content to maintain pedagogical authority, including systematic review and refinement of tasks, rubrics, and feedback before implementation (Bannister et al., 2025; Nguyen et al., 2025). Additionally, all GenAI-assisted outputs should explicitly align with course learning outcomes and Bloom’s higher-order cognitive levels, following the principles of constructive alignment.
  • Continuous Monitoring and Adaptability: Educators must implement a continuous monitoring process to analyze student performance data and GenAI-generated outputs, identifying gaps or unintended consequences in assessment effectiveness. This process should also include proactive strategies to adapt to evolving GenAI technologies, ensuring that teaching practices, assessment methods, and oversight remain current and effective (Australian Curriculum, Assessment and Reporting Authority [ACARA], 2012; Su & Yang, 2023).
Figure 1 illustrates the proposed Educator–GenAI Partnership Model, which involves five interrelated phases where GenAI is integrated in a supportive, iterative manner while rigorous human academic oversight is maintained. This process supports the assessment lifecycle stages of planning, designing, implementing, evaluating and improving (Holloman et al., 2021). The respective roles of the educator and GenAI tools are categorized as Major, Moderate, and Minor by their level of influence and contribution in the overall assessment lifecycle. Table 2 provides information on the contribution type, with its explanation and example.
Figure 1. Educator–GenAI Partnership Model for HOT assessment design.
Table 2. Contribution type and its explanation.

4.1. Phase 1: Identify Intent and Cognitive Demand

In this phase, educators need to precisely determine the target cognitive level using Bloom’s taxonomy (e.g., analyze or create), identify the specific learning outcomes to be assessed, and define the task’s contextual constraints. This intentional specificity must also include selecting the appropriate assessment format that aligns with HOT skills, such as authentic scenarios (e.g., case studies, client pitches), performance-based tasks (e.g., simulations, product prototyping, design), or inquiry-driven formats (e.g., complex problem sets, critical debate prompts). This high-quality input is crucial, as it provides the necessary guidance to the GenAI tool for effective use in the subsequent generation phase, ensuring the resulting assessment is pedagogically purposeful, not merely plausible (Brown et al., 2020; Wei et al., 2022). A foundational requirement for the phase’s effective use is that educators possess sufficient AI literacy to engage meaningfully with the GenAI tool using a structured prompt. The output from this phase is foundational for the remaining phases. The roles of the educator and GenAI, along with outcomes, are outlined below for this phase:
  • Educator Role (Major): Determine assessment intent, discipline-specific cognitive demands, and precisely define the target HOT skills (e.g., analysis, evaluation, creation) aligned with the specific learning outcomes. This involves selecting the appropriate assessment format (e.g., authentic scenarios, performance-based tasks) to guide subsequent GenAI-supported task generation.
  • GenAI Role (Minor): Suggest and refine appropriate action verbs corresponding to the target Bloom’s level, propose performance indicators, or provide exemplar outcomes aligned with the higher-order cognitive skills identified by the educator.
  • Outcome: A set of refined structured prompts incorporating the precise HOT level, suitable action verbs, assessment format, and learning context and outcome.

Structuring Effective Prompts

Structured prompts are the bridge between this phase and the next one. It is now common knowledge that the output of a large language model (LLM) depends upon the quality of the prompt used to generate output from it (Brown et al., 2020; Wei et al., 2022). The prompts need to cover the aspects and provide clear instructions for the LLM so that it generates the desired output. Moreover, if the instructions are precise and clear, it reduces the risk of tangential or factual outcome called hallucination. There are several approaches and guidelines being provided for structuring prompting. There are popular organizations like HuggingFace (Hugging Face, 2024) that provide the guidelines for writing prompts. There are additional resources available in (A. V. Y. Lee et al., 2024) that provide prompt guidelines and templates for efficient prompting. Harvard Publishing Press (Levy & Pérez Albertos, 2025; Wei et al., 2022) has advocated a task–instruction–context (TIC)-based approach for academic task creation. This ensures that the LLM is generating factual and good-quality output, adhering to the desired requirements. A clear, well-articulated prompt with all required ingredients will ensure the assessment includes the ingredients needed to develop HOT, as identified in this phase. At the same time, an incomplete or bad-quality prompt will lead to substandard output from the LLM. Furthermore, we argue that the correctness and the quality of the output must be evaluated while keeping the human-in-the-loop as the final authority. Therefore, educators need to adopt a suitable strategy for writing a structured prompt based on a well-developed strategy and personal experience, following organizational policy. Real-life examples of effective and ineffective prompts, along with their corresponding outputs generated using the ChatGPT free version (GPT-5.4) using the TIC approach (Levy & Pérez Albertos, 2025), are attached as a supplementary document for the benefit of readers.

4.2. Phase 2: GenAI-Supported Task Generation

In this phase, GenAI is tasked with generating multiple versions of assessments, scenarios, or problems that aim to elicit HOT. This capability directly addresses the institutional time constraint by rapidly generating a diverse pool of complex tasks, shifting the educator’s role from laborious content creation to strategic curation. To manage the variability of GenAI outputs, educators strictly mandate the use of highly constrained and discipline-specific prompts (Levy & Pérez Albertos, 2025; Wei et al., 2022). Educators must supply system prompts that enforce Bloom’s higher-order verbs (e.g., “Justify”, “Critique”, “Develop a novel synthesis”) and specific contextual parameters (Jackson, 2025). This human guidance acts as the primary quality control mechanism, ensuring the AI output is not generic but intentionally designed to elicit HOT skills Sharma et al. (2025). Educators need to transparently acknowledge the use of GenAI tools for assessment creation and responsible usage according to the institutional policy and procedure. The roles of educator and GenAI and outcomes are listed below for this phase:
  • Educator Role (Moderate): Select suitable GenAI tools based upon university recommendation and/or personal experience. The educator’s primary action is structuring prompt, inputting the specific, pre-defined constraints from Phase 1 (target HOT level, learning outcomes, assessment format, and contextual details) into the selected tool to enforce content and structure. Revise the prompt and perform the task generation if required.
  • GenAI Role (Major): Generate multiple drafts of assessment materials rapidly, such as detailed scenario narratives, complex datasets, simulated documents, or nuanced problem descriptions, aligned with the chosen HOT assessment format (e.g., authentic case study or inquiry-based problem set). In this role, GenAI acts as an engine for acceleration and diversification, producing qualitative assessments in significantly less time.
  • Outcome: A set of potential assessment task drafts that are ready for critical human evaluation in the next phase, successfully designed to elicit analysis, synthesis, and evaluation skills.

4.3. Phase 3: Alignment and Validation

In this phase, educators critically evaluate and refine GenAI-generated outputs for academic integrity, contextual relevance, and disciplinary accuracy. This phase is where the human oversight principle comes to life, using the educator’s disciplinary expertise and ethical judgment to validate tasks (Bellas et al., 2025; Hau, 2025; Ilieva et al., 2025; Sterz et al., 2024). The success of this phase relies on the educator utilizing a clear pedagogical alignment checklist. This operationalizes validity by requiring human confirmation that the GenAI-generated task meets four criteria for alignment and validation. The first three criteria are common in the traditional assessment design, while the fourth criterion will try to mitigate the GenAI challenge to academic integrity by evaluating the solvability of the assessment task with GenAI.
  • Mapping to the learning outcomes: Educators map the tasks created with the intended learning outcomes, and expected competencies.
  • Successful elicitation of the target Bloom’s level: Educators evaluate that the task created has effectively prompted learners to demonstrate the specific cognitive skill intended, aligning accurately with the desired level of Bloom’s taxonomy.
  • Demonstration of authenticity (real-world complexity and contextual judgment): Educators evaluate the complexity of the task and contextual settings.
  • Measure AI-solvability using tool: Educators assess the assessment task’s AI-solvability, ensuring that its complexity and contextual nuance require genuine HOT skills from the student, thereby mitigating the risk to academic integrity and AI-misuse.
Various approaches have been utilized to estimate AI-solvability in the literature. An approach for estimating AI-solvability is presented in (Akbar, 2025), where a Python-based system (using a combination of GPT-3.5 Turbo, BERT-based semantic similarity, and TF-IDF metrics) is designed and employed to predict the extent to which AI can solve tasks. Applied to 50 assignments across multiple computer science subdomains, the tool reliably categorized tasks by their vulnerability to AI assistance and generated corresponding solvability scores. Another approach, based on prompting an LLM to generate an answer for a designed assessment with a suitable marking guide, is introduced in (Thanh et al., 2023) for economics; experts determine the AI-solvability of the task according to Bloom’s taxonomy. Furthermore, automatic classification of assessment outcomes according to Bloom’s taxonomy is possible, and various methods (machine learning models, recurrent neural network models, Transformer-based models and LLMS) are presented Kumar et al. (2025); Mazza et al. (2025). In the literature, different approaches have been discussed to categorize assessments according to Bloom’s taxonomy, with accuracy reaching above 90% accuracy using different models with minimal overfitting Kumar et al. (2025); Mazza et al. (2025). These findings demonstrate that an AI-solvability and categorization test can be automated using suitable tools and techniques, but with human academic oversight in the loop.
The AI-solvability test is a validation step in which the educator submits the GenAI-generated assessment task to a separate GenAI tool (Thanh et al., 2023) or to an in-built tool (Akbar, 2025) to evaluate whether the task can be easily solved by the GenAI tool. If the assessment task does not pass the AI-solvability test or does not meet Bloom’s expected level, the task must be revised to increase cognitive rigor. This refinement may include requiring metacognitive justification, where students explain and defend their choice of tools or methods, or incorporating non-public information such as local context, proprietary data, or course-specific resources that are not available to GenAI tools. Refinement may include a demonstration of the process used to obtain the product, ensuring that the assessment remains authentically challenging and continues to target HOT. The roles of educator and GenAI and outcomes are listed below for this phase:
  • Educator Role (Major): Critically evaluate tasks to ensure constructive alignment and remove bias and factual inaccuracies. Check AI-solvability with suitable and available tools to maintain academic integrity.
  • GenAI Role (Minor): Offer revisions, alternative phrasing, or complexity adjustments upon the educator’s request to quickly rectify identified issues (during the four criteria analysis for alignment and validation) to maintain academic integrity, authenticity, and complexity.
  • Outcome: A validated and contextually aligned assessment, confirmed to require genuine higher-order cognitive effort from the student.

4.4. Phase 4: Rubric and Feedback Design

In this phase, GenAI assists with drafting rubrics and feedback statements aligned with cognitive levels in Bloom’s taxonomy. This phase ensures the evaluation criteria precisely reflect the HOT skills embedded in the validated task from the previous phase. GenAI significantly enhances efficiency here by automating the creation of criterion-based rubric skeletons and corresponding descriptors, which are directly mapped to the specific Bloom’s levels defined in Phase 1 (Fernández-Sánchez et al., 2025; Huang et al., 2024; McGowan, 2025). This automation allows educators to concentrate on high-level calibration—refining linguistic clarity, ensuring fairness, and, crucially, designing criteria that measure the process of analysis and creation rather than simple outcomes. Furthermore, GenAI is utilized to draft personalized, formative feedback suggestions tailored to common errors or high-performance indicators, streamlining the post-marking process (Bannister et al., 2025; Ilieva et al., 2025; Pan, 2025). The roles of educator and GenAI and outcomes are listed below for this phase:
  • Educator Role (Moderate): Calibrate rubric criteria and descriptors to ensure clarity, fairness, and validity, focusing on criteria that explicitly capture higher-order cognitive performance indicators.
  • GenAI Role (Moderate): Generate initial rubric templates (e.g., criteria, levels of achievement) and draft specific, targeted formative feedback suggestions aligned with potential student performance outcomes.
  • Outcome: A validated rubric that explicitly captures higher-order cognitive performance indicators and provides a foundation for efficient and high-quality formative feedback.

4.5. Phase 5: Reflection and Continuous Improvement

After assessments are used, educators and students reflect on how effectively tasks fostered HOT. This phase closes the iterative assessment design loop, transforming the model from a one-time tool into a continuous assessment design mechanism. GenAI’s role here is to streamline the analysis of student performance data and qualitative feedback, allowing educators to quickly identify patterns in student engagement, common misconceptions, or areas where the HOT task failed to achieve its intended cognitive level (Mpolomoka, 2025). This efficiency supports the long-term professional development of academics in assessment literacy and also improves the assessment quality. The roles of educator and GenAI and outcomes are listed below for this phase:
  • Educator Role (Major): Collect feedback, analyze student responses, and refine future assessments, focusing on the relationship between the GenAI-assisted task complexity and student HOT performance.
  • GenAI Role (Minor): Summarize large qualitative (e.g., student survey comments) and quantitative (e.g., grade distribution) feedback data or suggest iterative improvements based on educator reflection notes and performance summaries.
  • Outcome: Design of feedback parameters to continuously enhance assessment design for deeper learning engagement, ensuring that the model leads to iterative improvement in the educator’s assessment practice.
Table 3 provides mapping that shows how each phase of the model integrates the assessment design elements (HOT, constructive alignment, and human oversight) outlined in Section 2. The challenges related to design, creation and evaluation to ensure HOT outcome, as discussed in Section 2.1, are being addressed in Phase 1, Phases 2 and 3, and Phase 4, respectively. Additionally, the implementation of this proposed model must adhere strictly to institutional data governance and privacy policies supported by institution provided infrastructure. The model assumes the use of secure, private, or institutionally licensed GenAI instances, rather than public-facing platforms, to prevent the unauthorized transmission of confidential assessment content, proprietary university information, or sensitive course materials to external vendors.
Table 3. Mapping of phases of the model and assessment design elements.

5. Implementation Considerations and Limitations

The proposed model is firmly grounded in the established pedagogy of constructive alignment and Bloom’s Taxonomy. It centers human oversight as the governing principle throughout the assessment lifecycle. While the model offers a structured and practical approach to GenAI-assisted assessment design, its real-life adoption involves a set of implementation considerations and inherent limitations that require careful attention.

5.1. Implementation Considerations

The following considerations highlight the key implementation aspects that educators and institutions should address for effective and responsible implementation.

5.1.1. Educator AI Literacy and Prompt Engineering Competency

A foundational requirement for the model’s effective use is that educators possess sufficient AI literacy to engage meaningfully with GenAI tools across all five phases. In particular, Phases 1, 2 and 4 demand the ability to craft structured, discipline-specific prompts that yield educationally purposeful outputs rather than generic ones. We have included some examples for the readers’ guidance. Institutions must invest in targeted professional development that builds competency in prompt engineering, critical evaluation of AI-generated content, and the ethical use of GenAI in academic contexts. Without this foundational capability, the quality of the assessments generated will be severely constrained.

5.1.2. Defining Clear Human–GenAI Boundaries

The model delineates the roles of the educator and GenAI across phases. In practical application, these boundaries may blur, particularly during the iterative refinement cycles. Educators must remain vigilant that their role is one of critical oversight and pedagogical authority, not passive acceptance of AI-generated outputs. Institutions should develop clear operational guidelines that reinforce these boundaries and prevent over-reliance on GenAI in phases where human judgment is either moderate or major.

5.1.3. Institutional Infrastructure and Data Governance

Effective implementation of this model requires access to secure, institutionally licensed GenAI platforms that comply with data privacy and governance policies. The use of public-facing GenAI tools introduces risks related to the confidentiality of assessment content, sensitive course materials and privacy of the users. Institutions must, therefore, ensure that appropriate technological infrastructure is in place before campus-wide adoption, including suitable tools to support AI-solvability testing, as described in Phase 3.

5.1.4. Evaluation of Model’s Effectiveness

The model’s practical efficacy, including its impact on assessment quality, educator workload, and HOT performance, requires empirical validation before it can be widely recommended for adoption. While Phase 5 provides a built-in mechanism for ongoing evaluation through systematic reflection and analysis of student performance data, this alone may not be sufficient. Measuring effectiveness across diverse disciplines and institutional contexts will necessitate specialized academic evaluation metrics (Shailendra et al., 2024), rubric quality analysis (Morgan, 2025), cognitive alignment audits, and time-efficiency assessments. This broader evaluation process is, therefore, identified as a priority area for future research.

5.2. Limitations

Although the proposed model has robust and well-established theoretical foundations on pedagogical frameworks, it needs to be tested in real instructional settings. Moreover, its effectiveness in practice across different disciplines, qualification levels, and institutional cultures remains to be demonstrated. The model further assumes a baseline level of educator readiness or upskilling through professional development, along with institutional support, which may not be available uniformly. Disparities in AI familiarity, access to licensed GenAI tools, and professional development opportunities across regions and institutions may limit equitable adoption. Furthermore, the evolution of GenAI capabilities at an alarming rate stresses different aspects of the model, particularly the AI-solvability criteria in Phase 3, which will require periodic review and recalibration to remain relevant.
Furthermore, the authors want to highlight that the proposed model has its dependence on a baseline level of digital literacy among educators. In contexts where educators have limited familiarity with digital tools or GenAI systems, effective adoption may be constrained, highlighting the need for targeted competency development and professional support (Caspari-Sadeghi, 2026; Cui et al., 2025; OECD, 2026). Additionally, the model may have reduced applicability in disciplines where assessments are predominantly objective or highly quantitative, as evidence suggests that GenAI effectiveness varies across domains and may offer limited pedagogical enhancement beyond efficiency gains in structured tasks (OECD, 2026; Ogunleye et al., 2024).

6. Conclusions and Future Works

This paper has introduced the Educator–GenAI Partnership Model, a five-phase, iterative conceptual model designed to leverage GenAI tools for the creation of high-quality, HOT assessments. Grounded in the principle of human oversight for pedagogical authority, the model ensures that the educators maintain control over learning outcomes, complexity, cognitive demand (Phase 1), constructive alignment, validation, and AI-solvability (Phase 3), and ethical assessment practices (various phases). By integrating GenAI as a supportive collaborator in task generation (Phase 2) and rubric and feedback design (Phase 4), the model addresses key higher education challenges, enhancing efficiency while maintaining assessment validity and intellectual rigor. The final phase provides a mechanism for reflection and further improvement for the next assessment design.
The proposed model provides a structured approach for the responsible integration of GenAI in assessments. However, its practical efficacy remains to be substantiated through rigorous empirical validation. Future work should focus on testing and validating the model across diverse educational contexts, with particular attention paid to its impact on student learning outcomes, critical thinking, and academic integrity. This may also require adapting the model to different disciplines, assessment modalities and learning environments. In addition, longitudinal studies are needed to examine how iterative use of the model influences assessment literacy, students’ HOT capabilities and educators’ confidence in integrating GenAI into assessment design.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/educsci16050672/s1.

Author Contributions

Conceptualization, R.K. and Z.Z.; literature review, R.K., Z.Z., S.S. and U.R.S.; model development, R.K., Z.Z., S.S., U.R.S., A.S. and I.M.T.; implementation considerations and limitations, R.K., Z.Z., and S.S.; writing—original draft preparation, R.K., Z.Z., and S.S.; writing—review and editing, R.K., Z.Z., S.S., U.R.S., A.S. and I.M.T. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

Supplementary data is available and attached with the manuscript.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Akbar, M. S. (2025). Beyond detection: Designing AI-resilient assessments with automated feedback tool to foster critical thinking. arXiv, arXiv:2503.23622. [Google Scholar] [CrossRef] [Scilit]
  2. American Educational Research Association, American Psychological Association & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. American Educational Research Association. [Google Scholar]
  3. Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of bloom’s taxonomy of educational objectives. Longman. [Google Scholar]
  4. Association of Pacific Rim Universities [APRU]. (2025). Generative AI in higher education: Current practices and ways forward (Whitepaper). Association of Pacific Rim Universities. Available online: https://www.apru.org/wp-content/uploads/2025/01/APRU-Generative-AI-in-Higher-Education-Whitepaper_Jan-2025.pdf (accessed on 15 April 2026).
  5. Australian Curriculum, Assessment and Reporting Authority [ACARA]. (2012). Curriculum development process (version 6). Available online: https://www.acara.edu.au/curriculum/history-of-the-australian-curriculum/development-of-australian-curriculum (accessed on 15 April 2026).
  6. Babai Shishavan, H. (2024). AI in higher education: Guidelines on assessment design from Australian universities. In T. Cochrane, V. Narayan, E. Bone, C. Deneen, M. Saligari, K. Tregloan, & R. Vanderburg (Eds.), Proceedings ascilite 2024 melbourne: Navigating the terrain: Emerging frontiers in learning spaces, pedagogies, and technologies (pp. 118–126). ASCILITE. [Google Scholar] [CrossRef] [Scilit]
  7. Bannister, P., Urbieta, A. S., & Alvira, N. B. (2025). Appraising higher education assessment validity: Development of the PANDORA GenAI susceptibility rubric. Journal of Applied Learning and Teaching, 8(1), 41–55. [Google Scholar] [CrossRef] [Scilit]
  8. Bellas, F., Ooge, J., Roddeck, L., Rashheed, H. A., Skenduli, M. P., Masdoum, F., Zainuddin, N. b., Gori, J. N., Costello, E., Kralj, L., & Dcosta, D. T. (2025). Explainable ai in education: Fostering human oversight and shared responsibility. Available online: https://data.europa.eu/doi/10.2797/6780469 (accessed on 15 April 2026). [CrossRef]
  9. Biggs, J., Tang, C., & Kennedy, G. (2022). Teaching for quality learning at university 5e. McGraw-Hill Education (UK). [Google Scholar]
  10. Bittle, K., & El-Gayar, O. (2025). Generative AI and academic integrity in higher education: A systematic review and research agenda. Information, 16(4), 296. [Google Scholar] [CrossRef] [Scilit]
  11. Borge, M., Smith, B. K., & Aldemir, T. (2024). Using generative AI as a simulation to support higher-order thinking. International Journal of Computer-Supported Collaborative Learning, 19(4), 479–532. [Google Scholar] [CrossRef] [Scilit]
  12. Bouckaert, M. (2023). The assessment of students’ creative and critical thinking skills in higher education across oecd countries: A review of policies and related practices (Tech. Rep. No. 293). Organisation for Economic Co-Operation and Development. Available online: https://one.oecd.org/document/EDU/WKP(2023)8/en/pdf (accessed on 15 April 2026).
  13. Brookhart, S. M. (2010). How to assess higher-order thinking skills in your classroom. ASCD. [Google Scholar]
  14. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., & Agarwal, S. (2020). Language models are few-shot learners. In Advances in neural information processing systems. Curran Associates, Inc. [Google Scholar]
  15. Caspari-Sadeghi, S. (2026). AI literacy for teacher educators: A holistic curriculum for capacity-building in higher education. Frontiers in Education, 11, 1745768. [Google Scholar] [CrossRef] [Scilit]
  16. Clark, L. A. (2025). Reframing bloom’s for the age of AI: A white paper for future-ready educators (White Paper). Anthology. Available online: https://backstage.anthology.com/sites/default/files/2025-09/ReframingBloomsForTheAgeOfAI_WhitePaper_v1.pdf (accessed on 15 April 2026).
  17. Corbin, T., Dawson, P., & Liu, D. (2025). Talk is cheap: Why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 50(7), 1087–1097. [Google Scholar] [CrossRef] [Scilit]
  18. Cui, Y., Meng, Y., & Tang, L. (2025). Reconsidering teacher assessment literacy in GenAI-enhanced environments: A scoping review. Teaching and Teacher Education, 165, 105163. [Google Scholar] [CrossRef] [Scilit]
  19. Fernández-Sánchez, A., Lorenzo-Castiñeiras, J. J., & Sánchez-Bello, A. (2025). Navigating the future of pedagogy: The integration of AI tools in developing educational assessment rubrics. European Journal of Education, 60(1), e12826. [Google Scholar] [CrossRef] [Scilit]
  20. Gonsalves, C. (2025). Contextual assessment design in the age of generative AI. Journal of Learning Development in Higher Education. [Google Scholar] [CrossRef] [Scilit]
  21. Hau, D. (2025). Beyond the loop: Reclaiming pedagogy in an AI age. UNESCO Ideas Lab. Available online: https://www.unesco.org/en/articles/beyond-loop-reclaiming-pedagogy-ai-age (accessed on 15 April 2026).
  22. Holloman, T. K., Lee, W. C., London, J. S., Hawkins Ash, C. D., & Watford, B. A. (2021). The assessment cycle: Insights from a systematic literature review on broadening participation in engineering and computer science. Journal of Engineering Education, 110(4), 1027–1048. [Google Scholar] [CrossRef] [Scilit]
  23. Huang, X., Jen, F.-L., Lian, Y., & Jiao, J. (2024). Can generative Al really empower teachers’ professional practices? A quasi-experiment on human-GenAl collaborative rubric design. In International conference on technology in education (pp. 124–133). Springer. [Google Scholar] [CrossRef] [Scilit]
  24. Hugging Face. (2024). Prompt engineering. Available online: https://huggingface.co/docs/transformers/en/tasks/prompting (accessed on 16 March 2026).
  25. Ilieva, G., Yankova, T., Ruseva, M., & Kabaivanov, S. (2025). A framework for generative AI-driven assessment in higher education. Information, 16(6), 472. [Google Scholar] [CrossRef] [Scilit]
  26. Jackson, J. (2025). Higher order prompting: Applying bloom’s revised taxonomy to the use of large language models in higher education. Studies in Technology Enhanced Learning, 4(1), 1–26. [Google Scholar] [CrossRef] [Scilit]
  27. Kadel, R., Mishra, B. K., Shailendra, S., Abid, S., Rani, M., & Mahato, S. P. (2024). Crafting tomorrow’s evaluations: Assessment design strategies in the era of generative AI. In 2024 international symposium on educational technology (ISET) (pp. 13–17). IEEE. [Google Scholar] [CrossRef] [Scilit]
  28. Kadel, R., Shailendra, S., & Saxena, U. R. (2025). Navigating the new landscape: A conceptual model for project-based assessment (PBA) in the age of GenAI. In 2025 world engineering education forum—Global engineering deans council (WEEF-GEDC) (pp. 1–9). GEDC. [Google Scholar] [CrossRef] [Scilit]
  29. Kumar, R., Gulwani, D., & Singh, S. (2025). Automated analysis of learning outcomes and exam questions based on bloom’s taxonomy. arXiv, arXiv:2511.10903. [Google Scholar] [CrossRef] [Scilit]
  30. Kwan, P., Kadel, R., Memon, T. D., & Hashmi, S. S. (2025). Reimagining flipped learning via bloom’s taxonomy and student–teacher–GenAI interactions. Education Sciences, 15(4), 465. [Google Scholar] [CrossRef] [Scilit]
  31. Lee, A. V. Y., Teo, C. L., & Tan, S. C. (2024). Prompt engineering for knowledge creation: Using Chain-of-thought to support students’ improvable ideas. AI, 5(3), 1446–1461. [Google Scholar] [CrossRef] [Scilit]
  32. Lee, C. C., & Low, M. Y. H. (2024). Using genAI in education: The case for critical thinking. Frontiers in Artificial Intelligence, 7, 1452131. [Google Scholar] [CrossRef] [Scilit]
  33. Lee, J., Hung, J.-T., Soylu, M. Y., Popescu, D., Cui, C. Z., Grigoryan, G., Joyner, D. A., & Harmon, S. W. (2025). Socratic mind: Impact of a novel GenAI-powered assessment tool on student learning and higher-order thinking. arXiv, arXiv:2509.16262. [Google Scholar] [CrossRef] [Scilit]
  34. Levy, D., & Pérez Albertos, A. (2025). Simple AI tips for designing courses. In HBP education’s must reads: Effective course design—Ideas to enrich and evolve your syllabus (Originally published in Inspiring Minds, pp. 31–37). Harvard Business School Publishing. Available online: https://hbsp.harvard.edu/inspiring-minds/simple-ai-tips-designing-courses (accessed on 15 April 2026).
  35. Li, F., Yan, X., Su, H., Shen, R., & Mao, G. (2025). An assessment of human–AI interaction capability in the generative AI era: The influence of critical thinking. Journal of Intelligence, 13(6), 62. [Google Scholar] [CrossRef] [Scilit]
  36. Lodge, J. M., Bearman, M., Dawson, P., Gniel, H., Harper, R., Liu, D., McLean, J., & Ucnik, L. (2025). Enacting assessment reform in a time of artificial intelligence. Tertiary Education Quality and Standards Agency (TEQSA), Australian Government. Available online: https://www.teqsa.gov.au/sites/default/files/2025-09/enacting-assessment-reform-in-a-time-of-artificial-intelligence.pdf (accessed on 15 April 2026).
  37. Lodge, J. M., Howard, S., Bearman, M., & Dawson, P. (2023). Assessment reform for the age of artificial intelligence. Tertiary Education Quality and Standards Agency. Available online: https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/assessment-reform-age-artificial-intelligence (accessed on 15 April 2026).
  38. Lubbe, A., Marais, E., & Kruger, D. (2025). Cultivating independent thinkers: The triad of artificial intelligence, bloom’s taxonomy and critical thinking in assessment pedagogy. Education and Information Technologies, 30, 17589–17622. [Google Scholar] [CrossRef] [Scilit]
  39. Mazza, A., El Makkaoui, K., Ouahbi, I., & Maleh, Y. (2025). Automating educational assessment with AI: Leveraging bloom’s taxonomy and transformer models for question classification. In 2025 international conference on circuit, systems and communication (ICCSC) (pp. 1–5). IEEE. [Google Scholar] [CrossRef] [Scilit]
  40. McGowan, B. (2025). From prompt to practice: Using AI to build better assessment rubrics with SOLO taxonomy. Ulster University. [Google Scholar] [CrossRef]
  41. Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons’ responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741. [Google Scholar] [CrossRef]
  42. Morgan, F. (2025). Gen AI and the essay: Evaluating task specificity and rubric alignment (Tech. Rep.). The University of Melbourne. Available online: https://www.teqsa.gov.au/sites/default/files/2025-06/gen-AI-and-the-essay-evaluating-task-specificity-UoM.pdf (accessed on 15 April 2026).
  43. Mpolomoka, D. L. (2025). Utilizing artificial intelligence for assessment in higher education. Pedagogical Research, 10(3), em0243. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Nguyen, A., Duong, A. T., Nguyen, D. T. B., Lai, V. T. T., & Dang, B. (2025). Guidelines for learning design and assessment for generative artificial intelligence-integrated education: A unified view. Information and Learning Sciences, 126(7/8), 491–512. [Google Scholar] [CrossRef] [Scilit]
  45. OECD. (2026). OECD digital education outlook 2026. Organisation for Economic Co-Operation and Development (OECD) Publishing. [Google Scholar] [CrossRef] [Scilit]
  46. Ogunleye, B., Zakariyyah, K. I., Ajao, O., Olayinka, O., & Sharma, H. (2024). Higher education assessment practice in the era of generative AI tools. Journal of Applied Learning and Teaching, 7(1), 28. [Google Scholar] [CrossRef] [Scilit]
  47. Oiva-Córdova, L. M., Alvarez-Icaza, I., & George-Reyes, C. E. (2025). Evaluation of generative AI use to foster critical thinking in higher education. IEEE Revista Iberoamericana de Tecnologias del Aprendizaje, 20, 237–243. [Google Scholar] [CrossRef] [Scilit]
  48. Pan, Y. (2025). Leveraging generative AI powered rubric-indexed feedback as a formative assessment strategy for enhancing medical English education. Discover Computing, 28(1), 284. [Google Scholar] [CrossRef] [Scilit]
  49. Rana, V., Verhoeven, B., & Sharma, M. (2025). Generative AI in design thinking pedagogy: Enhancing creativity, critical thinking, and ethical reasoning in higher education. Journal of University Teaching and Learning Practice, 22(4), 1–22. [Google Scholar] [CrossRef] [Scilit]
  50. Reyna, J. (2025). Future-proofing academia: Integrating generative AI for innovative learning and assessment in higher education. In T. Bastiaens (Ed.), Proceedings of edmedia 2025 (pp. 418–423). Association for the Advancement of Computing in Education (AACE). Available online: https://www.learntechlib.org/p/226169 (accessed on 15 April 2026).
  51. Salinas-Navarro, D. E., Vilalta-Perdomo, E., Michel-Villarreal, R., & Montesinos, L. (2024). Designing experiential learning activities with generative artificial intelligence tools for authentic assessment. Interactive Technology and Smart Education, 21(4), 708–734. [Google Scholar] [CrossRef] [Scilit]
  52. Shailendra, S., Kadel, R., & Sharma, A. (2024). Framework for adoption of generative artificial intelligence (GenAI) in education. IEEE Transactions on Education, 67(5), 777–785. [Google Scholar] [CrossRef] [Scilit]
  53. Sharma, A., Shailendra, S., & Kadel, R. (2025, February 21–23). Experiences with content development and assessment design in the era of GenAI. 2025 6th International Conference on Computer Science, Engineering, and Education (CSEE) (pp. 1–5), Nanjing, China. [Google Scholar] [CrossRef] [Scilit]
  54. Sterz, S., Baum, K., Biewer, S., Hermanns, H., Lauber-Rönsberg, A., Meinel, P., & Langer, M. (2024). On the quest for effectiveness in human oversight: Interdisciplinary perspectives. In Proceedings of the 2024 ACM conference on fairness, accountability, and transparency (pp. 2495–2507). ACM. [Google Scholar] [CrossRef] [Scilit]
  55. Su, J., & Yang, W. (2023). Unlocking the power of ChatGPT: A framework for applying generative AI in education. ECNU Review of Education, 6(3), 355–366. [Google Scholar] [CrossRef] [Scilit]
  56. Thanh, B. N., Vo, D. T. H., Nhat, M. N., Pham, T. T. T., Trung, H. T., & Xuan, S. H. (2023). Race with the machines: Assessing the capability of generative AI in solving authentic assessments. Australasian Journal of Educational Technology, 39(5), 59–81. [Google Scholar] [CrossRef] [Scilit]
  57. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. [Google Scholar] [CrossRef] [Scilit]
  58. Weng, X., Qi, X., Gu, M., Rajaram, K., & Chiu, T. K. (2024). Assessment and learning outcomes for generative AI in higher education: A scoping review on current research status and trends. Australasian Journal of Educational Technology, 40(6), 37–55. [Google Scholar] [CrossRef] [Scilit]
  59. Whitham, R., Stockton, G., Richards, D., Lindley, J., Jacobs, N., & Coulton, P. (2023). Implications of generative AI on learning and assessment in higher education and design research practice. Available online: https://api.semanticscholar.org/CorpusID:271194224 (accessed on 15 April 2026).
  60. Xia, Q., Weng, X., Ouyang, F., Lin, T. J., & Chiu, T. K. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education, 21(1), 40. [Google Scholar] [CrossRef] [Scilit]
  61. Yusuf, A., Pervin, N., & Román-González, M. (2024). Generative AI and the future of higher education: A threat to academic integrity or reformation? Evidence from multicultural perspectives. International Journal of Educational Technology in Higher Education, 21(1), 21. [Google Scholar] [CrossRef] [Scilit]
  62. Zhao, J., Chapman, E., & Sabet, P. G. (2024). Generative AI and educational assessments: A systematic review. Education Research and Perspectives, 51, 124–155. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.