1. Introduction
Agentic artificial intelligence (AAI) is distinguished by its capacity to autonomously execute multi-step tasks, representing a notable leap forward in educational technologies (
Cetinkaya & Krämer, 2026;
J. Wang et al., 2025). For the purpose of this study, AAI refers specifically to systems capable of autonomous multi-step task execution with limited real-time human intervention, thereby distinguishing them from conventional prompt-response generative AI systems. Unlike traditional generative AI tools that primarily respond to user prompts, AAI systems operate with greater autonomy, fundamentally transforming human–technology interactions in educational contexts. This transition from assistive technologies to delegation-capable systems introduces new considerations for educators, who must assess not only the pedagogical benefits of AI but also the implications of transferring aspects of instructional control to intelligent agents (
Shulner-Tal et al., 2025). Consequently, conventional technology evaluation frameworks centred on functionality, efficiency, or ease of use provide only a partial explanation of how AAI becomes embedded within higher education teaching practices. Although educators are increasingly experimenting with AAI tools, consistent and meaningful pedagogical integration remains uneven (
Baig & Yadegaridehkordi, 2025;
Zheng et al., 2025).
The distinction between initial experimentation and meaningful pedagogical integration is particularly important, as engagement driven by novelty does not necessarily translate into sustained instructional embedding. In an environment where AI capabilities are increasingly integrated into browsers, learning management systems, and productivity platforms, the question of whether educators continue to use AI has become less informative than understanding how deeply, reliably, and pedagogically AAI is integrated into teaching, assessment, and instructional decision-making (
Cukurova, 2026;
Silalahi et al., 2026). This shift redirects scholarly attention from adoption and continuance towards integration depth, emphasising how educators incorporate AAI into their professional practices rather than merely whether they use it (
Singh & Paiva, 2025).
Pedagogical integration entails deeper cognitive, behavioural, and institutional commitments and is especially consequential in educational settings where technology directly influences teaching quality, student outcomes, and professional accountability. Existing research on AI in higher education has largely focused on initial acceptance through constructs such as perceived usefulness, perceived ease of use, and behavioural intention (
Feng et al., 2025;
Khanfar et al., 2025), as well as continuance intention as conceptualised by the Expectation Confirmation Model for IS (
Bhattacherjee, 2001). While these perspectives provide important insights, they generally assume that positive expectation confirmation is sufficient for continued use and often treat enabling and constraining influences as independent factors. In the context of AAI, however, the issues of autonomy and instructional delegation introduce a fundamentally different evaluative challenge. AAI provides enabling affordances, including instructional automation and pedagogical augmentation, while simultaneously raising concerns regarding autonomy, accountability, and data privacy (
Pikhart & Al-Obaydi, 2025;
Park, 2026).
This duality suggests that pedagogical integration emerges not from isolated evaluations of benefits and risks, but from their dynamic interplay (
W. Li, 2025). Nevertheless, prior studies have tended to examine these dimensions separately, offering limited insight into how they jointly shape sustained instructional engagement. To address this gap, the present study develops a mechanism-based framework that explains pedagogical integration through two interconnected pathways: a value pathway and a risk pathway. The value pathway captures educators’ evaluations of pedagogical benefits, whereas the risk pathway reflects concerns associated with delegating instructional responsibilities to autonomous systems. These pathways are further shaped by institutional support and regulatory environments, which represent important contextual conditions influencing educators’ experiences with AAI (
Erdmann & Toro-Dupouy, 2025;
Liu, 2025). By integrating these elements, the study offers a process-oriented perspective on pedagogical integration in agentic AI contexts (
Anomah, 2025).
The study also considers broader institutional and societal influences. Regulatory clarity may reduce uncertainty by providing governance structures, while perceived AI legitimacy reflects the extent to which AAI use is regarded as professionally appropriate and socially endorsed (
Colonna, 2026;
Şen et al., 2026). These contextual conditions shape how educators evaluate the value–risk trade-off and, consequently, their willingness to embed AAI within instructional practice (
Alfiras et al., 2025;
García-López & Trujillo-Liñán, 2025). Pedagogical integration is therefore conceptualised as a dual-evaluation process situated within institutional and normative environments. Building on this discussion, the present study addresses the following research questions:
RQ1: How do institutional capability and perceived complexity jointly shape educators’ evaluations of AAI systems in terms of pedagogical value and autonomy-related risk?
RQ2: How do perceived pedagogical value and perceived autonomy risk interact to shape the depth and consistency of pedagogical integration of AAI into instructional practice in higher education?
RQ3: How do contextual boundary conditions, such as regulatory clarity and AI legitimacy, moderate the relationships within the value–risk mechanism governing pedagogical integration of AAI?
Contributions of the Study
This study makes four principal contributions to the emerging literature on agentic AI in higher education. First, it shifts the analytical focus from continuance intention towards pedagogical integration, emphasising the depth, consistency, and instructional embedding of AAI within teaching practice. Second, it develops a dual-pathway value–risk framework that conceptualises pedagogical integration as the outcome of interconnected evaluations of pedagogical benefits and autonomy-related concerns rather than treating these dimensions as independent predictors. Third, the study integrates affordance theory, institutional theory, and automation risk perspectives into a unified explanatory framework, thereby providing a more comprehensive understanding of educators’ interactions with delegation-capable AI systems. Finally, by drawing on a diverse international sample of higher-education educators and employing multiple robustness checks, including invariance, marker-variable, and endogeneity assessments, the study offers evidence across varied institutional contexts while acknowledging the limitations associated with purposive online recruitment.
The remainder of this paper is structured as follows. The next section presents the theoretical framework and hypothesis development, followed by the research methodology and empirical results. The paper concludes with a discussion of theoretical and practical implications, limitations, and directions for future research.
4. Research Methodology
4.1. Research Design and Data Collection
This study employs a cross-sectional survey design to examine the mechanisms underlying the pedagogical integration of agentic artificial intelligence (AAI) systems in higher education. Given the emergent and complex nature of AAI—characterised by autonomous multi-step task execution, workflow delegation, and limited real-time human intervention—a survey-based approach facilitates the systematic examination of educators’ perceptions, evaluations, and integration practices across diverse institutional and geographic contexts. The design is particularly appropriate for investigating cognitive evaluations, such as perceived pedagogical value and perceived AI complexity, alongside autonomy-related concerns that constitute the proposed affordance–institutional–risk framework.
Consistent with the cross-sectional nature of the data, the findings should be interpreted as associative rather than causal, and the temporal ordering among constructs remains theoretically motivated rather than empirically established. While the proposed framework is grounded in established theoretical perspectives, future longitudinal and experimental studies are necessary to validate the directionality and evolution of these relationships over time.
Data were collected through an online questionnaire administered to higher education educators across multiple countries. Recruitment was conducted through academic mailing lists, professional LinkedIn networks, faculty communities, and scholarly conferences where discussions on educational technology and AI-enabled teaching practices were prevalent. No monetary or non-monetary incentives were offered for participation, thereby reducing the possibility of participation motivated primarily by external rewards.
The study employed purposive sampling to target educators with varying degrees of exposure to AAI-enabled instructional systems. Although this approach ensured the inclusion of respondents familiar with autonomous and semi-autonomous educational technologies, it also introduces the possibility of self-selection bias. Specifically, educators who are digitally confident, interested in AI, or early adopters of emerging technologies may be overrepresented in the sample. Accordingly, the findings should be interpreted as reflecting informed evaluations among educators with at least some exposure to AAI rather than as representative of the entire higher education population.
The questionnaire was pre-tested with a small group of academics to ensure clarity, contextual relevance, and content validity, particularly with respect to concepts involving autonomous task execution and instructional delegation. Minor refinements were incorporated prior to full deployment. A total of 364 responses were initially obtained. Following the implementation of the study’s screening procedures and standard data-quality checks, 26 responses were excluded because respondents did not satisfy the conceptual and behavioural filter requirements designed to verify meaningful exposure to AAI systems. The final analytical sample therefore comprised 338 valid responses.
The final sample demonstrates considerable international diversity. Respondents represented both Global North and Global South contexts, with 60.1% and 39.9% of participants, respectively. Countries represented in the Global North included the United States, Canada, the United Kingdom, Germany, Sweden, the Netherlands, and Australia. Global South participants were drawn from India, Brazil, South Africa, Kenya, Nigeria, Indonesia, Vietnam, the Philippines, Thailand, and the United Arab Emirates. For analytical purposes, these countries were subsequently aggregated into the regional categories reported in
Table 1. This diverse composition provides broader contextual representation while acknowledging the limitations associated with purposive online recruitment.
Although participants varied in their levels of engagement with AI technologies, all retained respondents reported at least minimal validated exposure to AAI systems capable of autonomous or semi-autonomous task execution, ensuring the substantive relevance of their evaluations and perceptions.
4.1.1. AAI Exposure, Screening, and Sample Validation
Given the conceptual ambiguity between conventional artificial intelligence, generative AI (GenAI), and agentic artificial intelligence (AAI), a structured multi-stage screening procedure was implemented to ensure that respondents possessed a clear understanding of AAI and had meaningful exposure to its distinctive functionalities. This procedure was designed to minimise construct contamination and enhance the internal validity of the study.
For the purposes of this research, minimal exposure to AAI referred to direct experience with at least one of the following functionalities: (1) autonomous workflow execution; (2) AI agents capable of multi-step task completion; (3) delegation of instructional or administrative tasks to AI systems; (4) AI assistants with autonomous decision-support capabilities; or (5) systems capable of independently sequencing and completing interconnected activities with limited real-time human intervention. Examples of such systems included ChatGPT 5 Operator, AutoGPT, Claude Projects and Artifacts, Perplexity Labs, Microsoft Copilot Agents, Gemini Agents, and Manus AI. These examples were illustrative rather than exhaustive and served primarily to help respondents distinguish agentic systems from conventional prompt-response generative AI tools.
The screening protocol consisted of four sequential stages. First, respondents completed an awareness check designed to assess familiarity with AAI concepts. Second, a conceptual differentiation exercise required participants to correctly distinguish AAI systems from conventional generative AI tools based on characteristics such as autonomy, delegation, and multi-step execution. Third, respondents reported their usage status, distinguishing between current users, previous users, and individuals with no direct experience. Finally, behavioural verification items assessed actual engagement with agentic functionalities, including workflow delegation, autonomous execution, and independent task sequencing. Of the 364 initial responses, 26 were excluded because they did not satisfy one or more of these filter requirements, including failures on conceptual differentiation items, inconsistencies between reported awareness and behavioural exposure, or insufficient evidence of meaningful engagement with agentic functionalities. Consequently, the final sample comprised 338 respondents with validated exposure to AAI systems.
Although precise exclusion counts were not retainable for each individual screening stage, all removed cases failed at least one of the predefined criteria relating to awareness, conceptual differentiation, usage status, or behavioural verification. The final analytical sample therefore included only respondents who demonstrated validated exposure to agentic AI through both conceptual understanding and observable engagement with autonomous or semi-autonomous functionalities, thereby ensuring the substantive relevance and internal validity of the study. Respondents were retained only if they correctly identified the defining characteristics of agentic AI, reported either current or prior use, and demonstrated at least one verified behavioural interaction involving autonomous or semi-autonomous task execution. Individuals reporting only intentions to use or no direct experience with agentic AI were excluded from the final analytical sample.
Importantly, respondents who had never used AAI systems or whose experience was limited solely to conventional prompt-based generative AI were excluded from the analytical sample. Similarly, individuals expressing only an intention to use AAI without demonstrable behavioural exposure were not retained for hypothesis testing. This approach ensured that the study captured informed evaluations of pedagogical integration rather than hypothetical attitudes toward unfamiliar technologies. The complete screening protocol, including awareness checks, conceptual differentiation items, and behavioural verification measures, is provided in
Appendix A.
4.1.2. Assessment of Common Method Bias
Given the self-reported nature of the data, potential common method bias (CMB) was addressed through both procedural and statistical remedies (
Podsakoff et al., 2024). Procedurally, respondents were assured of anonymity, and the questionnaire was designed to minimise evaluation apprehension and response tendencies through careful sequencing and separation of measurement items (
Podsakoff et al., 2003). Statistically, a theoretically unrelated marker variable—the primary device used to access AI systems—was incorporated. Although device type may be associated with technology access or usage convenience, it bears no direct theoretical relationship with the value–risk mechanisms, institutional conditions, or pedagogical integration processes examined in the present study. Accordingly, it provides a reasonable diagnostic for detecting potential common method variance without overlapping conceptually with the substantive constructs (
Kock, 2017). Correlations between the marker variable and the focal constructs remained low and non-significant (|r| < 0.15), indicating an absence of systematic shared variance attributable to measurement procedures. In addition, full-collinearity variance inflation factor (VIF) values were below the conservative threshold of 3.3, providing further evidence that common method bias is unlikely to threaten the validity of the findings (
Podsakoff et al., 2003;
Kock, 2017). Collectively, these procedural and statistical safeguards suggest that common method variance does not materially influence the observed relationships.
4.2. Measurement of Constructs
All constructs were operationalised using multi-item reflective scales adapted from established literature and contextualised for the pedagogical use of agentic artificial intelligence (AAI) in higher education. Guided by affordance theory, institutional theory, and automation risk perspectives (
Gibson, 1977;
DiMaggio & Powell, 1983;
Parasuraman & Riley, 1997), the measurement framework captured both the enabling and constraining dimensions associated with delegation-capable AI systems.
Institutional AI Capability (IAC) measured the extent to which institutions provide the infrastructure, governance mechanisms, training opportunities, and implementation support necessary for effective AAI integration (
Erdmann & Toro-Dupouy, 2025;
Denford et al., 2025). Perceived AI Complexity (PAIC) assessed educators’ cognitive and operational challenges in understanding, configuring, and managing autonomous multi-step AI systems (
Cetinkaya & Krämer, 2026;
Xing & Jiang, 2025). Perceived Pedagogical Value (PPV) captured the extent to which educators believed that AAI enhances teaching effectiveness, instructional quality, student engagement, and workload efficiency (
Feng et al., 2025;
Xu et al., 2025).
Perceived Autonomy Risk (PAR) reflected concerns regarding loss of instructional control, execution errors, accountability, and privacy implications arising from the delegation of pedagogical tasks to autonomous systems (
Pikhart & Al-Obaydi, 2025;
Park, 2026). Regulatory Clarity (RC) measured perceptions of the existence of clear governance structures, institutional policies, and legal frameworks guiding the responsible use of AAI (
Colonna, 2026), while AI Legitimacy (AIL) captured the extent to which AAI use was perceived as professionally appropriate and socially accepted within academic environments (
Şen et al., 2026;
Zagami, 2026).
Particular attention was devoted to the conceptualisation of the dependent construct. Rather than measuring simple continuance intention or repeated system use, Pedagogical Integration of AAI (PIA) was conceptualised as the depth, consistency, and instructional embeddedness with which educators incorporated AAI functionalities into teaching, assessment, feedback generation, and broader instructional workflows. Although certain behavioural indicators reflect routinised and sustained engagement, these items capture the institutionalisation and normalisation of agentic AI within pedagogical practices rather than mere intentions to continue using technology. Accordingly, the construct extends beyond traditional continuance perspectives by emphasising integrated, habitual, and pedagogically meaningful utilisation of autonomous AI capabilities (
Silalahi et al., 2026;
Zheng et al., 2025).
Although one indicator is phrased in terms of continued use, it is interpreted within this study as reflecting the institutionalisation of agentic AI within educators’ pedagogical practice rather than a distinct continuance intention. In educational settings, sustained intention to use a teaching technology typically indicates that it has become embedded within routine instructional planning, content delivery, assessment, and learner support. Accordingly, continued use is treated as evidence of routinised pedagogical integration, capturing the habitual incorporation of agentic AI into regular teaching activities rather than merely expressing a future behavioural intention.
All constructs were measured using seven-point Likert scales ranging from 1 (strongly disagree) to 7 (strongly agree) (
Hair & Alamer, 2022;
Becker et al., 2023). Control variables included prior generative AI experience, teaching experience, and academic discipline, while primary device usage was retained as a marker variable for assessing common method bias (
Kock, 2017).
To address concerns regarding coding consistency, all measurement items were aligned such that higher scores consistently represented higher levels of the underlying construct. Reverse-coded items within the Perceived AI Complexity (PAIC) scale are explicitly identified as (R) in
Appendix A.
Pre-Test and Expert Review
A two-stage validation procedure was employed to ensure measurement clarity, contextual relevance, and theoretical consistency (
Hair & Alamer, 2022). First, five experts comprising academics, a practitioner, and specialists in measurement development evaluated the survey instrument with respect to item clarity, construct validity, and alignment with the theoretical foundations of the study (
Becker et al., 2023). Particular emphasis was placed on ensuring that the items appropriately reflected the autonomous, delegatory, and multi-step execution characteristics that distinguish AAI from conventional generative AI systems. Minor revisions relating to wording, sequencing, and contextual framing were subsequently incorporated.
Second, a pilot study involving 30 higher education educators with prior exposure to agentic AI technologies was conducted to assess comprehensibility, reliability, and respondent burden. The pilot participants were selected to ensure familiarity with agentic functionalities such as autonomous workflows, instructional delegation, and AI-assisted task sequencing, thereby enhancing the contextual validity of the instrument.
Reliability diagnostics, including Cronbach’s alpha and composite reliability, exceeded the recommended threshold of 0.70 for all constructs, indicating satisfactory internal consistency (
Hair & Alamer, 2022;
Becker et al., 2023). No substantive issues relating to ambiguity or conceptual overlap emerged during the pilot phase. Minor refinements concerning item phrasing and survey flow were implemented to improve readability and reduce cognitive load.
Importantly, all measurement items were explicitly framed within the context of agentic AI rather than generic artificial intelligence applications. This distinction was maintained throughout the instrument to minimise conceptual contamination between conventional generative AI tools and autonomous, delegation-capable systems. Overall, the pre-test and expert-review procedures strengthened content validity, enhanced measurement precision, and ensured alignment with contemporary best practices in PLS-SEM research (
Becker et al., 2023).
4.3. Data Analysis Technique
The proposed model was evaluated using Partial Least Squares Structural Equation Modelling (PLS-SEM) implemented in SmartPLS 4.1 (
Ringle et al., 2024). This analytical approach is appropriate given the study’s predictive orientation, theory-extension objectives, and the presence of multiple mediation and moderation mechanisms within the proposed framework. PLS-SEM is particularly well suited for complex models that prioritise variance explanation, accommodate reflective measurement specifications, and do not require strict multivariate normality assumptions (
Becker et al., 2023;
Hair & Alamer, 2022).
The analysis followed a two-stage procedure comprising the assessment of the measurement model, followed by the evaluation of the structural model. This approach ensured that the reliability, convergent validity, and discriminant validity of the constructs were established prior to interpreting the hypothesised relationships. Bootstrapping with 5000 resamples was employed to estimate the statistical significance of direct, indirect, and moderation effects. To assess potential endogeneity, the Gaussian copula procedure proposed by
Hult et al. (
2018) was additionally implemented, thereby strengthening the robustness of the findings.
4.3.1. Model Specification
The proposed model examines the influence of Institutional AI Capability and Perceived AI Complexity on the Pedagogical Integration of AAI through two complementary mediating mechanisms: Perceived Pedagogical Value and Perceived Autonomy Risk. These mediators capture the dual evaluative pathways through which educators assess delegation-capable AI systems, reflecting both enabling pedagogical opportunities and concerns associated with autonomy and instructional control.
The model further incorporates Regulatory Clarity and AI Legitimacy as contextual boundary conditions that may strengthen or weaken these relationships. Specifically, regulatory clarity is expected to buffer the positive relationship between perceived complexity and autonomy risk, whereas AI legitimacy is posited to strengthen the translation of pedagogical value into pedagogical integration. Prior generative AI experience, teaching experience, and academic discipline were included as control variables, while primary device usage served as a theoretically unrelated marker variable to address potential common method bias.
4.3.2. Measurement Model Assessment
The measurement model was evaluated using established criteria for reliability and validity. Indicator reliability was assessed through outer loadings, with all items exceeding the recommended threshold of 0.70, indicating adequate representation of their respective constructs. Internal consistency reliability was examined using composite reliability coefficients, all of which surpassed the recommended value of 0.70 (
Hair & Alamer, 2022). Convergent validity was assessed through average variance extracted (AVE), with all constructs exceeding the threshold of 0.50, confirming that the latent variables explained a substantial proportion of the variance in their indicators (
Becker et al., 2023).
Table 2 below shows the measurement model reliability and descriptives.
Discriminant validity was evaluated using both the heterotrait–monotrait (HTMT) ratio and the Fornell–Larcker criterion. All HTMT values remained below the conservative threshold of 0.85 (
Becker et al., 2023), while the square roots of AVE exceeded the corresponding inter-construct correlations, supporting adequate discriminant validity.
To provide additional evidence regarding discriminant validity, bootstrapped HTMT confidence intervals were also examined. None of the 95% confidence intervals included the value of 1.0, providing further support for the empirical distinctiveness of the constructs. Although the relationships among Perceived Pedagogical Value, Perceived Autonomy Risk, and Pedagogical Integration of AAI exhibited comparatively high HTMT values (0.824–0.840), their upper confidence bounds remained below unity (0.872–0.886), suggesting conceptual proximity consistent with the proposed dual-pathway framework rather than problematic construct redundancy. These findings reinforce the adequacy of the measurement model while acknowledging the theoretically interrelated nature of educators’ value, risk, and integration evaluations.
Table 3 below shows the discriminant validity assessment.
Table 4 above shows the bootstrapped HTMT confidence intervals. Although several HTMT values approached the conservative cut-off, particularly among Perceived Pedagogical Value, Perceived Autonomy Risk, and Pedagogical Integration of AAI, these relationships are theoretically expected within the proposed dual-pathway framework. Pedagogical value captures educators’ evaluations of instructional benefits, autonomy risk reflects concerns associated with delegated control, and pedagogical integration represents the behavioural embedding of AAI into instructional practice. While conceptually related, these constructs remain theoretically distinct and correspond to enabling, constraining, and behavioural dimensions, respectively. The observed HTMT values therefore indicate expected conceptual proximity rather than problematic overlap.
Future research may further strengthen discriminant validity assessments by reporting HTMT confidence intervals and comparing constrained alternative measurement specifications. Nevertheless, the combined evidence from HTMT and Fornell–Larcker analyses supports the adequacy of the measurement model.
Multicollinearity was examined using variance inflation factors (VIF), with all values remaining below the conservative threshold of 3.0 (
Kock, 2017), indicating the absence of problematic collinearity. Overall, the measurement model demonstrates satisfactory psychometric properties, supporting the validity and reliability of the constructs employed in the analysis.
4.3.3. Structural Model Assessment
The structural model was evaluated using bootstrapping with 5000 resamples to estimate path coefficients, t-statistics, and bias-corrected confidence intervals (
Becker et al., 2023). This non-parametric procedure provides robust significance testing without imposing distributional assumptions on the data.
The model demonstrated substantial explanatory power across the principal endogenous constructs. The coefficient of determination (R2) indicated that Perceived Pedagogical Value (R2 = 0.621), Perceived Autonomy Risk (R2 = 0.615), and Pedagogical Integration of AAI (R2 = 0.619) all exhibited moderate-to-high levels of explained variance, suggesting that the proposed framework effectively captures the principal determinants of educators’ integration decisions.
Model fit indicators further supported the adequacy of the specification. The SRMR values for both the saturated (0.041) and estimated (0.046) models remained below recommended thresholds, indicating acceptable model fit. The NFI value of 0.887, although slightly below conventional benchmarks, may be more appropriately interpreted as marginal rather than unequivocally acceptable, particularly given the predictive orientation of PLS-SEM. Accordingly, greater emphasis is placed on SRMR and predictive relevance statistics when evaluating model adequacy.
Predictive relevance was assessed using blindfolding procedures. The Q2 values for Perceived Pedagogical Value (Q2 = 0.601), Perceived Autonomy Risk (Q2 = 0.587), and Pedagogical Integration of AAI (Q2 = 0.611) all substantially exceeded zero, indicating strong out-of-sample predictive capability and supporting the model’s predictive validity.
Effect-size analysis (f2) revealed meaningful differences in the substantive importance of the exogenous constructs. Most notably, Institutional AI Capability exhibited an exceptionally large effect on Perceived Pedagogical Value (f2 = 0.964). This finding suggests that institutional infrastructure, governance mechanisms, training opportunities, and organisational support constitute foundational conditions through which educators recognise and realise the pedagogical benefits of AAI. In other words, the educational value of agentic systems appears to depend heavily on the institutional environments within which they are embedded. By contrast, the effects of Perceived Pedagogical Value (f2 = 0.220) and Perceived Autonomy Risk (f2 = 0.213) on Pedagogical Integration of AAI are moderate, indicating that both enabling and constraining evaluations contribute meaningfully to educators’ integration decisions. Regulatory Clarity and AI Legitimacy exhibit comparatively smaller effect sizes, suggesting that they function primarily as contextual boundary conditions rather than principal explanatory drivers.
Mediation analyses were conducted using bootstrapped indirect effects, with significance inferred from confidence intervals that excluded zero. Both specific and serial mediation pathways were examined, with particular attention devoted to the sequential mechanism linking pedagogical value and autonomy risk. The significant indirect effects provide support for the proposed dual-pathway framework, indicating that institutional capability and perceived complexity influence pedagogical integration through interconnected evaluations of value and risk.
Table 5 below shows the main direct and moderation effect results.
Consistent with reporting conventions, all statistically significant paths are reported as p < 0.001 rather than p = 0. Collectively, the structural results support the explanatory and predictive robustness of the proposed framework and demonstrate the importance of value–risk evaluations in understanding the pedagogical integration of delegation-capable AI systems in higher education.
4.3.4. Moderation Effects
Moderation analysis was conducted using the two-stage approach in PLS-SEM to examine whether regulatory clarity and AI legitimacy function as contextual boundary conditions within the proposed value–risk framework. The results indicate that both moderators significantly influence the relationships specified in the conceptual model, although their effects operate by amplifying or buffering existing mechanisms rather than serving as direct determinants of pedagogical integration.
Regulatory clarity significantly moderates the relationship between Perceived AI Complexity and Perceived Autonomy Risk (β = −0.100,
p = 0.008), indicating that the positive association between complexity and autonomy-related concerns becomes weaker when educators perceive clearer institutional and regulatory guidance regarding AAI use. This finding suggests that governance structures, policy frameworks, and explicit operational guidelines may help educators manage the uncertainty associated with complex, delegation-capable AI systems, thereby reducing the extent to which technical complexity translates into perceived risk (
Al-Emran, 2023).
Similarly, AI Legitimacy positively moderates the relationship between Perceived Pedagogical Value and Pedagogical Integration of AAI (β = 0.144,
p < 0.001). This result suggests that when AAI is perceived as professionally appropriate and socially endorsed within educational communities, educators are more likely to translate positive evaluations of pedagogical value into sustained instructional integration. The finding is consistent with legitimacy-based explanations of technology acceptance and institutional conformity, which emphasise the importance of normative support in shaping behavioural commitment (
Suchman, 1995;
Lowry et al., 2025).
Importantly, the direct effects of Regulatory Clarity (β = −0.074, p = 0.062) and AI Legitimacy (β = −0.061, p = 0.054) remain statistically non-significant, indicating that these constructs function primarily as contextual enablers rather than independent drivers of pedagogical integration. Their influence is therefore better understood as shaping the strength of existing relationships within the value–risk framework.
Simple-slope interpretations further reinforce these conclusions. Higher levels of regulatory clarity attenuate the positive association between perceived AI complexity and autonomy-related risk, whereas stronger perceptions of AI legitimacy amplify the positive association between pedagogical value and pedagogical integration. The simple-slope analysis further indicates that the effect of perceived AI complexity on autonomy risk is substantially weaker under conditions of high regulatory clarity than under low regulatory clarity. Likewise, the positive influence of perceived pedagogical value on pedagogical integration becomes considerably stronger when AI legitimacy is perceived to be high, highlighting the role of normative acceptance in translating perceived benefits into instructional practice.
Figure 2 and
Figure 3 present the interaction effects and simple-slope relationships associated with regulatory clarity and AI legitimacy, respectively, thereby providing a more intuitive interpretation of the moderating mechanisms.
4.3.5. Multi-Group Analysis Results
To evaluate the robustness and generalisability of the proposed framework across different respondent groups, measurement invariance and multi-group analyses (MGA) were conducted using the MICOM (Measurement Invariance of Composite Models) procedure and Bootstrap MGA in SmartPLS. Two theoretically relevant subgroup comparisons were examined: (1) Global North versus Global South educators and (2) respondents with low versus high prior generative AI experience.
The MICOM procedure established compositional invariance across virtually all constructs in both comparisons. For the Global North and Global South groups, original correlations ranged from 0.981 to 1.000, with non-significant permutation p-values, thereby satisfying the requirements for compositional invariance. Equality of means and variances was also supported, indicating full measurement invariance across geographic contexts. These findings suggest that respondents from different regional and institutional environments interpreted the focal constructs in comparable ways.
Similarly, for the generative AI experience comparison, compositional invariance was achieved for nearly all constructs, with correlations ranging from 0.923 to 1.000. Although Perceived Autonomy Risk exhibited a significant permutation result (p = 0.013), equality of means and variances remained supported. Accordingly, partial measurement invariance was established, which remains sufficient for conducting meaningful multi-group comparisons within the PLS-SEM framework.
Subsequent Bootstrap MGA revealed no statistically significant differences across any structural paths for either comparison group, with all two-tailed
p-values exceeding the 0.05 threshold. These findings indicate that the proposed dual-pathway mechanism linking institutional capability, perceived complexity, pedagogical value, autonomy risk, and pedagogical integration operates consistently across both geographic contexts and varying levels of prior generative AI experience.
Table 6 below shows the MGA results.
Overall, these results strengthen the external validity and robustness of the proposed framework by demonstrating that the underlying value–risk evaluation process remains structurally stable across diverse educational, cultural, and experiential settings. The consistency of the findings across Global North and Global South contexts further suggests that the pedagogical integration of AAI may be governed by common evaluative mechanisms despite differences in institutional environments and technological maturity.
4.3.6. Endogeneity Assessment
To assess potential endogeneity, the Gaussian copula procedure proposed by
Hult et al. (
2018) was employed. This approach enables the detection of unobserved confounding by examining whether the inclusion of copula terms materially alters the estimated structural relationships.
The results indicate limited evidence of endogeneity within the proposed model. The copula terms associated with Institutional AI Capability → Perceived Autonomy Risk (β = −0.006, p = 0.962), Institutional AI Capability → Perceived Pedagogical Value (β = 0.082, p = 0.370), Perceived AI Complexity → Perceived Autonomy Risk (β = 0.244, p = 0.090), Perceived AI Complexity → Perceived Pedagogical Value (β = −0.201, p = 0.291), Perceived Pedagogical Value → Perceived Autonomy Risk (β = 0.057, p = 0.559), and Perceived Autonomy Risk → Pedagogical Integration of AAI (β = −0.192, p = 0.299) were all statistically non-significant. These findings suggest that unobserved confounding is unlikely to materially affect the majority of the hypothesised relationships.
However, one notable exception concerns the relationship between Perceived Pedagogical Value and Pedagogical Integration of AAI. Although the Gaussian copula term for this relationship was statistically significant (GC β = 0.249, p = 0.001), indicating the possibility of residual endogeneity, the copula-adjusted structural coefficient decreased from the original estimate (β = 0.405) to β = 0.153 and was no longer statistically significant (p = 0.120). This finding suggests that the relationship between perceived pedagogical value and pedagogical integration may be more sensitive to unobserved influences than the remaining structural paths. Accordingly, this association should be interpreted with caution, and future longitudinal or experimental research is needed to examine its causal stability more conclusively.
Overall, the Gaussian copula analysis supports the robustness of the proposed framework, as most structural relationships remained unaffected after accounting for potential endogeneity (
Hult et al., 2018). Nevertheless, the direct relationship between Perceived Pedagogical Value and Pedagogical Integration appears comparatively less robust, indicating that this specific association may be influenced by unobserved factors. While the broader dual-pathway framework remains supported, this relationship should be interpreted with greater caution until validated through longitudinal or experimental research designs.
Table 7 below shows the endogeneity assessment results.
Although most copula-adjusted coefficients remained broadly consistent with the original estimates, the relationship between Perceived Pedagogical Value and Pedagogical Integration showed a noticeable reduction in magnitude and was no longer statistically significant after adjustment. Consequently, this path should be interpreted with greater caution than the remaining structural relationships. Overall, the evidence suggests that endogeneity concerns are limited and do not substantially undermine the explanatory robustness of the proposed model. Nonetheless, future research employing longitudinal designs, experimental approaches, or objective behavioural measures would further strengthen causal inference and reduce the possibility of unobserved confounding in studies of agentic AI integration.
5. Discussion
This study advances the discourse on agentic AAI by shifting the analytical focus from adoption and continuance towards pedagogical integration—the depth, consistency, and instructional embedding of delegation-capable systems in teaching practice—conceptualised as a dynamic value–risk trade-off process (
Feng et al., 2025). In an environment of ambient AI availability, the empirically meaningful variation no longer lies in whether educators continue to use AI but in how they integrate it into instructional decision-making (
Cukurova, 2026;
Silalahi et al., 2026). Unlike ECM-IS and other continuance-centred models that typically treat value and risk as independent determinants, the findings suggest that these dimensions are cognitively interconnected, with perceived pedagogical value being associated with lower autonomy-related risk perceptions (
W. Li, 2025). This perspective positions pedagogical integration as a process-oriented phenomenon rather than a static outcome, offering a more comprehensive understanding of how educators evaluate and adapt to autonomy-enabled systems.
Central to the findings is the identification of a dual-pathway mechanism that captures the simultaneous evaluation of enabling and constraining forces. The results suggest that educators do not assess AAI systems solely on the basis of functional benefits but also consider the implications of delegating instructional responsibilities. This dual evaluation reflects a broader cognitive balancing process in which pedagogical benefits are weighed against concerns related to autonomy, accountability, and uncertainty. By conceptualising AAI use as a delegation-based evaluation problem, the study extends prior research on assistive technologies and highlights the distinctive considerations introduced by systems capable of autonomous, multi-step execution (
Shulner-Tal et al., 2025;
Pikhart & Al-Obaydi, 2025).
A key insight emerging from this study is the role of institutional AI capability as both an enabler and a source of stability within this evaluative process. Institutional support, in the form of infrastructure, training, and governance mechanisms, is positively associated with perceptions of pedagogical value and negatively associated with autonomy-related concerns. This pattern suggests that institutional environments shape not only the availability of technological resources but also how educators interpret and manage issues associated with instructional delegation (
Erdmann & Toro-Dupouy, 2025). In this regard, the findings align with institutional theory by indicating that organisational support may function as a cognitive reference point through which educators interpret both the opportunities and uncertainties associated with agentic technologies.
In contrast, perceived AI complexity emerges as an important constraining factor within the proposed framework. When educators perceive AAI systems as cognitively demanding or difficult to manage, their ability to recognise and enact pedagogical affordances may be reduced, while uncertainty and perceived risk may become more salient. This finding aligns with research on system opacity and unpredictability, which suggests that difficulties in understanding autonomous processes correspond to lower confidence in technology-enabled delegation (
Cetinkaya & Krämer, 2026). Importantly, the influence of complexity appears to operate through evaluative mechanisms rather than through direct behavioural effects, indicating that complexity primarily affects pedagogical integration by shaping how educators perceive value and risk.
One of the notable contributions of this study is the identification of a cross-pathway relationship between perceived pedagogical value and perceived autonomy risk. The findings indicate that higher perceptions of pedagogical value are associated with lower levels of autonomy-related concern, suggesting that educators who recognise stronger instructional benefits may be more inclined to view system autonomy as a manageable pedagogical feature rather than solely as a source of uncertainty (
Singh & Paiva, 2025). Rather than implying a deterministic process, this relationship points to the possibility that positive instructional evaluations correspond to more favourable interpretations of autonomous capabilities. Such a mechanism offers a dynamic perspective on technology evaluation in which perceptions of benefits and concerns are cognitively intertwined.
At the same time, it is important to acknowledge that alternative explanations remain plausible. Although the present framework adopts a theory-driven ordering in which pedagogical value precedes autonomy-risk assessments, the reverse relationship may also exist, whereby higher perceptions of risk influence educators’ evaluations of pedagogical value. The cross-sectional design does not permit empirical verification of temporal ordering, and future longitudinal or experimental studies should compare competing causal structures to examine how value and risk perceptions co-evolve over time in relation to agentic AI use.
Building upon this perspective, the serial mediation results further support the notion of a layered evaluation process. Rather than being associated with isolated factors, pedagogical integration appears to correspond to a sequence of cognitive assessments in which institutional and system-level conditions are linked to perceptions of pedagogical value, which are subsequently associated with autonomy-related concerns and instructional integration. This suggests that engagement with autonomous systems may represent an evolving process shaped by prior evaluations and contextual influences rather than a single adoption decision.
An important clarification concerns the nature of the risk construct examined in this study. The analysis focuses specifically on perceived autonomy risk, reflecting educators’ subjective concerns regarding issues such as instructional control, accountability, privacy, and the implications of delegating tasks to autonomous systems. The study does not measure objective privacy breaches, actual execution errors, or verifiable accountability failures. Consequently, the findings should be interpreted as capturing educators’ perceptions and interpretations of potential risks rather than demonstrating the existence or magnitude of such outcomes. Future research incorporating behavioural, institutional, or system-level data could further examine the relationship between perceived and objectively observable forms of autonomy-related risk.
The role of contextual boundary conditions further refines the proposed mechanism by situating it within broader institutional and normative environments. Regulatory clarity is associated with a weaker relationship between perceived complexity and autonomy-related concerns, suggesting that formal guidance and governance structures may help educators interpret and manage the implications of autonomous instructional systems (
Colonna, 2026). Similarly, AI legitimacy corresponds to a stronger association between perceived pedagogical value and pedagogical integration by reinforcing the social and professional acceptability of agentic technologies (
Şen et al., 2026). These findings indicate that educators’ evaluations of AAI are embedded within broader governance frameworks and normative expectations, highlighting the importance of contextual alignment in supporting pedagogical integration.
Finally, the consistency of the dual-pathway mechanism across different levels of AI experience and geographic contexts provides evidence regarding the robustness of the proposed framework. The absence of significant structural differences between Global North and Global South samples suggests that the underlying value–risk evaluation process may operate similarly across diverse institutional environments (
Silalahi et al., 2026). Nevertheless, this observation should be interpreted cautiously given the purposive sampling strategy and uneven subgroup sizes. While contextual differences may influence the strength or manifestation of specific relationships, the overall pattern of findings aligns with the proposition that educators across regions engage in comparable cognitive evaluations when considering the pedagogical integration of AAI.
Overall, the findings position the value–risk trade-off as a mechanism-oriented framework for understanding pedagogical integration of AAI, offering an alternative perspective to traditional adoption models that pay limited attention to autonomy and delegation dynamics. By integrating affordance theory, institutional theory, and automation risk perspectives, the study provides a process-oriented explanation of how educators navigate the opportunities and concerns associated with autonomy-enabled systems. Rather than depicting pedagogical integration as a simple consequence of perceived usefulness or satisfaction, the framework emphasises the interplay among value perceptions, autonomy-related concerns, and contextual influences, thereby offering a theoretically grounded lens for examining the evolving role of agentic AI in higher education.
7. Limitations and Future Research
Despite its contributions, this study is subject to several limitations that provide avenues for future research. First, the cross-sectional design constrains causal inference and does not permit empirical verification of the temporal ordering proposed in the dual-pathway framework. Although the hypothesised relationships are grounded in affordance theory, institutional theory, and automation risk perspectives, alternative explanations remain plausible. In particular, perceptions of autonomy-related risk may also influence educators’ evaluations of pedagogical value rather than following the sequencing proposed in the present model. Future longitudinal, panel, and experimental studies are needed to examine how value and risk perceptions co-evolve over time and to compare competing causal structures underlying the pedagogical integration of agentic AI (
Hair & Alamer, 2022). Certain disciplinary and regional subgroups, including Business educators (n = 14) and respondents from the MENA region (n = 6), were relatively small and should therefore be interpreted with caution. Future research should seek broader representation across underrepresented educational and geographic contexts.
Second, the purposive online sampling strategy may have introduced self-selection bias by attracting educators who are more digitally confident, interested in artificial intelligence, or predisposed towards experimenting with emerging educational technologies. Although this approach ensured meaningful exposure to delegation-capable AI systems, the findings should be interpreted as reflecting informed evaluations among educators with validated AAI experience rather than the broader higher education population. Future studies could employ probability-based sampling approaches and investigate institutions characterised by different levels of digital maturity and AI readiness to strengthen the external validity and contextual applicability of the findings (
Podsakoff et al., 2024).
Third, while the sample encompassed respondents from diverse Global North and Global South contexts, certain regional and disciplinary subgroups remained relatively small, particularly those from the MENA region and business education. Consequently, subgroup-specific interpretations should be treated with caution, and broader contextual generalisations remain tentative. Future research should pursue more balanced cross-cultural and disciplinary samples to examine how institutional environments, governance arrangements, and professional norms influence the pedagogical integration of agentic AI across heterogeneous educational settings (
DiMaggio & Powell, 1983).
Fourth, the study relied exclusively on self-reported perceptions of pedagogical value, autonomy-related risk, and pedagogical integration. Accordingly, the findings capture educators’ subjective evaluations rather than directly observed instructional practices or objectively measured educational outcomes. Moreover, the study examines perceived autonomy risk, including concerns regarding accountability, privacy, and instructional control, rather than actual occurrences of execution failures, data breaches, or governance violations. Future research should incorporate classroom observations, learning analytics, digital trace data, institutional usage records, and student performance indicators to investigate the relationship between perceived and objective manifestations of pedagogical integration within agentic AI environments (
Parasuraman & Riley, 1997).
Finally, the present study adopts an educator-centred perspective and therefore does not incorporate the views of students, institutional leaders, instructional designers, or educational technology developers. The pedagogical integration of agentic AI is inherently a multi-actor and institutionally embedded process involving interactions among diverse stakeholders operating within broader systems of governance, legitimacy, and professional norms. Future studies should therefore employ multi-level and multi-stakeholder designs to investigate how different actors collectively shape perceptions of pedagogical value, autonomy-related risk, legitimacy, and instructional delegation. Such efforts would contribute to a more holistic understanding of how agentic AI becomes embedded within contemporary higher education ecosystems (
Suchman, 1995).