1. Introduction
We report a 12-week quasi-experimental field study with a total analysed sample of N = 187 employees across three Yemeni organisations (AI-Adaptive: n = 94; Control: n = 93), comparing an AI-adaptive cybersecurity awareness training platform against traditional instructor-led training (ILT). The remainder of this Introduction sets out the knowledge–behaviour gap that motivates this work.
Organisational cybersecurity hinges not only on the sophistication of technical controls but also on the daily decisions of individual employees. A single mis-click on a phishing email, a reused password, or an unlocked workstation can lead to substantial security breaches despite existing technical controls. Yet decades of security awareness training have exposed a stubborn paradox: employees who correctly identify phishing on a knowledge test routinely fall victim to simulated attacks minutes later [
1,
2]. This knowledge–behaviour gap—the systematic failure of intention to predict action in security contexts—remains among the most consequential unsolved problems in applied information security. Traditional instructor-led training (ILT) addresses knowledge but rarely the psychological mechanisms that convert knowledge into protective behaviour: self-efficacy, perceived threat, social norms, and perceived ease of compliance [
3,
4,
5]. The unresolved question, then, is how to bridge this persistent gap by directly engaging the psychological levers that ILT typically leaves untouched.
The key innovation of our deployment is the use of interpretable P(G) and P(S) Bayesian Knowledge Tracing signals as an auditable pedagogical decision layer that triggers PMT/TPB-aligned interventions, rather than as mere mastery estimates.
AI-driven adaptive learning systems offer a theoretically promising alternative. By tracking individual mastery in real time via algorithms such as Bayesian Knowledge Tracing (BKT) [
6], they can deliver content calibrated to each learner’s Zone of Proximal Development while simultaneously building competency and self-efficacy. Within the wider knowledge-tracing literature, classic BKT remains attractive when interpretability, robustness with modest sample sizes, and deployment transparency matter, even as more complex alternatives such as Deep Knowledge Tracing (DKT) have broadened the modelling landscape [
7,
8]. For the specific goals of this study, BKT’s interpretability is key to our theoretical contribution because it allows the deployment of an auditable, theory-grounded pedagogical decision layer that operationalises and tests PMT/TPB-aligned responses, rather than treating personalisation as a black box. The key innovation of our BKT deployment is therefore the explicit use of interpretable P(G) and P(S) signals as an auditable pedagogical decision layer to guide PMT/TPB-aligned responses, beyond typical mastery estimation. This transparency also fosters organisational trust by justifying the trigger of each intervention: high P(S) flags an attention-related slip and routes the learner to attention-to-detail and execution-control content, whereas high P(G) flags a likely informed guess or misconception and routes the learner to explicit rule-recall and self-efficacy scaffolds. Adaptive decisions, therefore, remain transparent rather than black-boxed—an attribute increasingly demanded in security-relevant deployments.
Recent cybersecurity-training publications also point toward growing interest in generative-AI-enabled and adaptive training architectures, but much of that literature is still conceptual, review-based, or prototype-oriented rather than field-validated [
9,
10,
11]. Whereas those works propose frameworks or simulated benefits, the present study provides the first formal empirical test, in a live organisational setting, of objective IT-verified incident outcomes and rigorous mediation analysis of PMT/TPB constructs. The novelty of the present design, therefore, lies less in proposing an entirely new knowledge-tracing algorithm than in the empirical validation of an integrated system in which our custom-configured BKT model—designed so that interpretable P(G) and P(S) signals function as an auditable pedagogical decision layer—serves as the theory-grounded decision engine for a PMT/TPB-aligned instructional layer. Importantly, unlike prior studies that may have measured PMT/TPB constructs only as outcome variables, this work integrates them directly into the adaptive decision-making logic itself, transforming them into actionable levers for behavioural change rather than passive endpoints. Standard BKT libraries such as pyBKT [
12] provide robust parameter estimation but do not natively offer such a theory-linked pedagogical decision layer, which underscores the architectural novelty of the integrated system evaluated here.
Methodologically, field evaluations face a further challenge: when participants cannot be randomly assigned—because departmental scheduling or organisational constraints preclude it—baseline groups often differ dramatically. Standard ANCOVA produces biased estimates when baseline differences exceed approximately 1 SD [
13]; yet many published evaluations report Cohen’s d > 1.0 in baseline characteristics without applying advanced causal-inference methods. A validated analytical playbook for this common situation is conspicuously absent from the information-systems security literature.
5. Discussion
5.1. Theoretical Implications: Mediation as the Bridge
The central theoretical claim of this study is that AI-adaptive training reduces the knowledge–behaviour gap not merely by increasing information transfer, but by actively modulating the psychological mediators posited by PMT and TPB. Our mediation analysis provides the first empirical test of this claim in a real-world organisational setting. The indirect effects via coping self-efficacy (0.843, 95% CI [0.52, 1.19]) and PBC (0.524, 95% CI [0.28, 0.81]) were both significant and together accounted for 66.4% of the total training effect on policy compliance behaviour.
As anticipated in the integrated PMT–TPB framework, the results point to a specific design principle: adaptive systems should prioritise mastery experiences and controllability cues first, then use threat-salience material as a secondary reinforcement rather than as the primary motivational lever.
Critically, the indirect path through threat appraisal alone was non-significant—replicating the pattern identified by Herath and Rao [
2] and confirmed in Bada et al.’s meta-analysis [
1], where coping appraisal consistently outperforms threat appraisal as a proximal predictor of compliance. This finding carries a direct practical implication: fear-based messaging—still dominant in many corporate security awareness campaigns—is insufficient to drive sustained compliance when coping resources remain underdeveloped. The competitive advantage of AI-adaptive training lies in providing individualised mastery experiences that build coping capacity at scale, not in amplifying threat salience.
Three complementary explanations help interpret the non-significant indirect path via threat appraisal. First, by design, the adaptive logic used threat-salience vignettes only as secondary reinforcement once P(L) > 0.70, while the dominant pedagogical action under low-to-moderate mastery was coping-oriented scaffolding; the intervention, therefore, delivered a deliberately small “dose” of threat content. Second, baseline threat-severity scores were already high across both groups (AI M = 4.12, SD = 0.58; Control M = 4.08, SD = 0.61, on a 5-point scale), implying a ceiling effect that mechanically constrains observable change in threat appraisal. Third, the observed pattern replicates the consistent finding in the PMT literature [
1,
2] that coping appraisal outperforms threat appraisal as a proximal predictor of compliance—fear-based messaging, in particular, is insufficient where coping resources remain underdeveloped. Taken together, the null indirect threat-appraisal path is not a design anomaly but a theoretically expected pattern: AI-adaptive training appears to bridge the knowledge–behaviour gap principally by building coping capacity rather than by amplifying perceived threat.
Limits of mediation interpretation (added in revision). We interpret the mediation results as evidence of plausible explanatory pathways rather than definitive causal proof. While indirect effects via coping self-efficacy and PBC were statistically significant and survived FDR correction, the quasi-experimental design does not permit the inference that the BKT-driven adaptive mechanism itself caused the observed psychological changes. Alternative explanations remain possible—for example, increased time-on-task, differential engagement with adaptive content, novelty effects, or unmeasured differences in instructor enthusiasm in the ILT arm could each contribute. The convergence of the four-method triangulation, Rosenbaum Γ = 2.1, and E-values ≥ 3.4 collectively constrain the magnitude of any such alternative explanation, but they do not eliminate it. Future randomised, dismantling-style experiments will be required to isolate the causal contribution of the BKT-driven adaptation layer itself.
Rosenbaum Γ = 2.1 means that any unmeasured confounder would need to more than double a participant’s odds of being assigned to the AI-adaptive condition to overturn the observed effect. E-values ≥ 3.4 mean that such a confounder would, in addition, need to be associated with the outcome by a risk ratio of at least 3.4. These bounds do not eliminate hidden bias, but they make it implausible that small, ordinary, or trivial confounders alone would suffice to explain the observed association.
These findings extend PMT beyond its traditional application in static “fear appeal” communications [
3,
16] to a dynamic, adaptive context where the training system functions as a continuous coping-appraisal enhancement mechanism. The BKT-driven adaptive algorithm operationalises Bandura’s [
21] concept of mastery experiences as the most potent source of self-efficacy—now applied automatically and at scale to enterprise cybersecurity.
5.3. Contextual Implications: Developing-Country Specificity and Boundary Conditions
Three contextual factors specific to the Yemeni setting likely shape the findings. First, the exceptionally low baseline compliance rates in the control group (12.9% prior training) mean that effect sizes may be larger than would be observed in populations with higher baseline competency. Second, the offline-first architecture validated here is directly transferable to analogous resource-constrained settings—public utilities, civil-society entities, and NGOs in Sub-Saharan Africa, South Asia, and conflict-affected regions—where CSAT adoption is lowest and cyber-incident risk is highest. Third, high power-distance culture [
28]—in which deference to authority figures is strong—may have inflated short-term compliance in authority-delivered ILT workshops.
Generalisability caveats: findings should be extrapolated cautiously to high-income populations with mature security cultures and large IT departments. The same adaptive logic remains directly extensible to higher-maturity use cases: the BKT engine can adaptively coach administrators on least-privilege principles, track adherence to ISO 27001, GDPR, HIPAA, or PCI DSS controls, reinforce secure-coding mastery, and shorten incident-reporting latency.
Boundary conditions for transfer (added in revision). Three boundary conditions are worth drawing out explicitly. First, the magnitude of “extreme baseline imbalance” observed here (Cohen’s d > 2.0 on prior training and baseline knowledge) is partly a product of the cross-organisational mix in Yemen: a large public utility, a private bank with more mature security policies, and a humanitarian NGO. In high-income enterprises with uniform onboarding training, baseline imbalances on these covariates may be substantially smaller; conversely, in cross-sector partnerships or supply-chain training programmes that bridge SMEs and large enterprises, imbalances may be comparable or greater. Second, the coping-self-efficacy mediation pathway is expected to remain dominant across cultures (it is well replicated in PMT meta-analyses [
1]), but the magnitude of the PBC pathway is plausibly moderated by national power-distance and by the formality of organisational compliance regimes [
28]. Third, in highly regulated industries (GDPR-bound EU finance, HIPAA-bound US healthcare), the avoided-incident cost component will dominate the cost–benefit equation, raising the NPV figures reported here by an order of magnitude. We, therefore, recommend that the present results be read as a lower-bound effect-size envelope and as a transferable methodological template, rather than as a point-estimate forecast for any specific high-income setting.
A fourth driver of the larger-than-benchmark effect sizes (d = 0.66–0.89 vs. the meta-analytic benchmark d ≈ 0.33) is plausibly contextual rather than methodological. Baseline cybersecurity awareness in developing-country and conflict-affected workforces is systematically lower than in the predominantly OECD samples that populate the Bada et al. [
1] meta-analysis: in our control group, only 12.9% of employees had received any prior structured cybersecurity training, compared with 60–80% in typical published OECD samples. The mechanical consequence is a larger ceiling-of-improvement: the same absolute knowledge gain translates into a larger standardised effect size because the baseline is further from the ceiling. We, therefore, expect that effective AI-adaptive interventions will continue to outperform the published benchmark in developing-country and capacity-building contexts, but that the standardised effect envelope will compress toward the d ≈ 0.40–0.55 range as baseline awareness rises. The absolute risk-reduction value (IT-verified incidents avoided, reportable breaches prevented) may nonetheless remain operationally large because the per-incident financial cost is much higher in highly regulated settings.
Cross-cultural information-security behaviour research suggests that high power-distance environments [
28] and highly formalised compliance regimes can amplify the short-term behavioural impact of authority-delivered ILT while simultaneously dampening the marginal contribution of perceived behavioural control (PBC) in the AI-adaptive condition; conversely, in low-power-distance settings with weaker formal compliance norms, PBC tends to carry a larger share of the mediation pathway because voluntary, self-directed protective behaviour is the dominant route to compliance. This moderation pattern is consistent with Hofstede’s [
28] cultural-dimensions framework and should be formally tested in cross-national replication.
In organisations with mature security cultures and uniformly high baseline knowledge, the AI-adaptive system’s standardised effect on knowledge and self-reported behaviour may be smaller than the d = 0.66–0.89 envelope reported here, because of a ceiling-of-improvement effect. However, the absolute reduction in high-cost incidents (e.g., reportable breaches under GDPR or HIPAA) may remain substantial in monetary terms even when standardised effect sizes compress, because the per-incident financial impact is much larger. The portable contribution of this study is, therefore, the four-method analytical framework and the BKT→PMT/TPB design logic, rather than the specific effect-size magnitudes, both of which can be replicated directly in higher-maturity contexts.
6. Conclusions
This study provides robust, theoretically grounded, and practically significant evidence that AI-adaptive cybersecurity training outperforms traditional ILT on every primary outcome evaluated. The consensus effect sizes (d = 0.66–0.89), verified across four complementary causal-inference methods, exceed the meta-analytic benchmark for security awareness training by a factor of two to three, and the objective institutional outcomes—48.9% reduction in IT-verified security incidents (IRR = 0.51 [0.38, 0.68]) and 75% reduction in phishing click-rates (χ2(1) = 8.74, p = 0.003)—demonstrate that theoretical gains translate into measurable operational risk reduction.
Three principal contributions emerge. First, this study provides, to our knowledge, the first theory-driven empirical test of PMT/TPB constructs as mediators of AI-adaptive training effects, demonstrating that coping self-efficacy and PBC—not threat appraisal—are the active psychological mechanisms (combined mediation: 66.4% of total effect). This insight redirects training design away from fear-based messaging toward efficacy-building, personalised mastery experiences. Second, the four-method triangulation protocol (ANCOVA, PSM, DiD, mixed-effects), with full statistical diagnostics (c-stat = 0.89; HL p = 0.51; parallel trends β = 0.42, p = 0.64; Rosenbaum Γ = 2.1; E-values ≥ 3.4), provides a validated analytical playbook for the highly common situation of field evaluation with extreme baseline imbalances. Third, the validated offline-first platform, tested in a conflict-affected, resource-constrained context, establishes a replicable technological model for global cybersecurity capacity building in underserved environments.
Methodological diagnostics—what each one confirms (added in revision). The four-method triangulation is accompanied by full statistical diagnostics that, taken together, justify confidence in the result: c-statistic = 0.89 indicates good PSM discrimination; Hosmer–Lemeshow p = 0.51 indicates no evidence of propensity-score model mis-fit; the parallel-trends β = 0.42, p = 0.64 confirms that the central DiD assumption is satisfied; Rosenbaum Γ = 2.1 implies that any unmeasured confounder would need to more than double the odds of treatment assignment to overturn the result and E-values ≥ 3.4 quantify the minimum strength such a confounder would need to explain the effect away. Sustained behaviour change in cybersecurity is itself a long-horizon problem; we, therefore, stress that the priority next step is a 12–24-month follow-up evaluation with decay-aware knowledge tracing and continuous adaptive refreshers.
Taken together, these findings reposition AI-adaptive CSAT from a marginal enhancement of awareness toward a measurable lever for organisational risk reduction, and the PMT/TPB-mediated pathways provide a theoretical explanation of why—moving beyond “black-box” observations of adaptive system performance to actionable insights for future training design.
Author Contributions
Conceptualisation, M.M.A.-G., M.A. and I.A.-B.; methodology, M.M.A.-G. and M.A.; software, M.M.A.-G.; validation, M.M.A.-G., M.A. and I.A.-B.; formal analysis, M.M.A.-G.; investigation, M.M.A.-G.; resources, M.M.A.-G. and I.A.-B.; data curation, M.M.A.-G.; writing—original draft preparation, M.M.A.-G.; writing—review and editing, M.A. and I.A.-B.; visualisation, M.M.A.-G.; supervision, M.A. and I.A.-B.; project administration, M.M.A.-G.; funding acquisition, M.M.A.-G. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Board of AL-Hikma University (protocol code IRB-2024-CS-041; date of approval: 7 February 2026).
Informed Consent Statement
Written informed consent was obtained from all participants involved in the study.
Data Availability Statement
The de-identified, aggregated datasets and the R(4.3.1)/Stata(18) analysis scripts that support the findings of this study (ANCOVA, PSM, DiD, mixed-effects, mediation) are available from the corresponding author upon reasonable request, subject to the data-sharing agreement with the participating Yemeni organisations and the institutional ethical approval (IRB-2024-CS-041), which restricts the release of any participant-level identifiable data. A subset of fully anonymised analysis-ready files and the BKT parameter-estimation code will be archived in a public repository (GitHub) upon acceptance, and the corresponding DOI will be provided in the final published version.
Acknowledgments
The authors thank the IT, HR, and security teams at the three participating Yemeni organisations for their cooperation during data collection, the two postgraduate-qualified instructors who delivered the standardised control-condition workshops, and the independent IT auditors who blindly classified Tier 2–3 incidents. The authors also thank the third-party vendor that administered the blinded phishing simulations.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Bada, M.; Sasse, A.M.; Nurse, J.R.C. Cyber security awareness campaigns: Why do they fail to change behaviour? arXiv 2019, arXiv:1901.02672. [Google Scholar] [CrossRef] [Scilit]
- Herath, T.; Rao, H.R. Encouraging information security behaviors in organizations: Role of penalties, pressures and perceived effectiveness. Decis. Support Syst. 2009, 47, 154–165. [Google Scholar] [CrossRef] [Scilit]
- Boss, S.R.; Galletta, D.F.; Lowry, P.B.; Moody, G.D.; Polak, P. What do systems users have to fear? Using fear appeals to engender threats and coping appraisals for system security. MIS Q. 2015, 39, 837–864. [Google Scholar] [CrossRef] [Scilit]
- Ifinedo, P. Understanding information systems security policy compliance: An integration of the theory of planned behavior and the protection motivation theory. Comput. Secur. 2012, 31, 83–95. [Google Scholar] [CrossRef] [Scilit]
- Bulgurcu, B.; Cavusoglu, H.; Benbasat, I. Information security policy compliance: An empirical study of rationality-based beliefs and information security awareness. MIS Q. 2010, 34, 523–548. [Google Scholar] [CrossRef] [Scilit]
- Corbett, A.T.; Anderson, J.R. Knowledge tracing: Modeling the acquisition of procedural knowledge. User Model. User-Adapt. Interact. 1995, 4, 253–278. [Google Scholar] [CrossRef] [Scilit]
- Khajah, M.; Lindsey, R.V.; Mozer, M.C. How deep is knowledge tracing? In Proceedings of the 9th International Conference on Educational Data Mining, Raleigh, NC, USA, 29 June–2 July 2016; pp. 94–101. [Google Scholar]
- Abdelrahman, G.; Wang, Q.; Nunes, B. Knowledge tracing: A survey. ACM Comput. Surv. 2023, 55, 1–37. [Google Scholar] [CrossRef] [Scilit]
- Zhdanov, D.; Caldera, T.M.; Califf, M.E. Using generative AI for cybersecurity awareness training in healthcare. In Proceedings of the 19th Pre-ICIS Workshop on Information Security and Privacy, Bangkok, Thailand, 15 December 2024. [Google Scholar]
- Ahmed, A.M.; Mejri, M. Generative AI for cybersecurity awareness training. In Proceedings of the IECON 2025—51st Annual Conference of the IEEE Industrial Electronics Society, Madrid, Spain, 14–17 October 2025. [Google Scholar]
- Sengupta, S.; Varma, U.; Islam, T. Empowering cybersecurity education: A review of adaptive learning paradigms and practical implications. In Proceedings of the IS-EUD 2025: 10th International Symposium on End-User Development, CEUR Workshop Proceedings; CEUR-WS: Bonn, Germany, 2025; Volume 3978. [Google Scholar]
- Badrinath, A.; Wang, F.; Pardos, Z.A. pyBKT: An accessible Python library of Bayesian knowledge tracing models. In Proceedings of the 14th International Conference on Educational Data Mining, Online, 29 June–2 July 2021. [Google Scholar]
- Shadish, W.R.; Cook, T.D.; Campbell, D.T. Experimental and Quasi-Experimental Designs for Generalized Causal Inference; Houghton Mifflin: Boston, MA, USA, 2002. [Google Scholar]
- Lebek, B.; Uffen, J.; Neumann, M.; Hohler, B.; Breitner, M.H. Information security awareness and behavior: A theory-based literature review. Manag. Res. Rev. 2014, 37, 1049–1092. [Google Scholar] [CrossRef] [Scilit]
- McCormac, A.; Zwaans, T.; Parsons, K.; Calic, D.; Butavicius, M.; Pattinson, M. Individual differences and information security awareness. Comput. Hum. Behav. 2017, 69, 151–156. [Google Scholar] [CrossRef] [Scilit]
- Rogers, R.W. A protection motivation theory of fear appeals and attitude change. J. Psychol. 1975, 91, 93–114. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ajzen, I. The theory of planned behavior. Organ. Behav. Hum. Decis. Process. 1991, 50, 179–211. [Google Scholar] [CrossRef] [Scilit]
- Pardos, Z.A.; Heffernan, N.T. Modeling individualization in a Bayesian networks implementation of knowledge tracing. In Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2010; Volume 6075, pp. 255–266. [Google Scholar] [CrossRef] [Scilit]
- Angrist, J.D.; Pischke, J.-S. Mostly Harmless Econometrics: An Empiricist’s Companion; Princeton University Press: Princeton, NJ, USA, 2009. [Google Scholar]
- van de Sande, B. Properties of the Bayesian knowledge tracing model. J. Educ. Data Min. 2013, 5, 1–10. [Google Scholar] [CrossRef] [Scilit]
- Bandura, A. Self-Efficacy: The Exercise of Control; Freeman: New York, NY, USA, 1997. [Google Scholar]
- Egelman, S.; Peer, E. Scaling the security wall: Developing a security behavior intentions scale (SeBIS). In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems (CHI ‘15), Seoul, Republic of Korea, 18–23 April 2015; pp. 2873–2882. [Google Scholar] [CrossRef] [Scilit]
- Vandenberg, R.J.; Lance, C.E. A review and synthesis of the measurement invariance literature: Suggestions, practices, and recommendations for organizational research. Organ. Res. Methods 2000, 3, 4–70. [Google Scholar] [CrossRef] [Scilit]
- National Institute of Standards and Technology. Computer Security Incident Handling Guide; NIST Special Publication 800-61 Rev. 2; NIST: Gaithersburg, MD, USA, 2012. Available online: https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-61r2.pdf (accessed on 24 June 2026).
- Austin, P.C. Optimal caliper widths for propensity-score matching when estimating differences in means and differences in proportions in observational studies. Pharm. Stat. 2011, 10, 150–161. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hayes, A.F. Introduction to Mediation, Moderation, and Conditional Process Analysis, 3rd ed.; Guilford Press: New York, NY, USA, 2022. [Google Scholar]
- Rosenbaum, P.R.; Rubin, D.B. The central role of the propensity score in observational studies for causal effects. Biometrika 1983, 70, 41–55. [Google Scholar] [CrossRef]
- Hofstede, G. Culture’s Consequences: Comparing Values, Behaviors, Institutions, and Organizations Across Nations, 2nd ed.; Sage: Thousand Oaks, CA, USA, 2001. [Google Scholar]
Figure 1.
Integrated PMT–TPB Framework with proposed AI-adaptive training mediation pathways. Solid arrows represent theorised mediation paths, later empirically validated as significant (coping self-efficacy and PBC) or non-significant (threat appraisal) in
Section 4.6; the dashed arrow represents the direct BKT reinforcement loop. ***
p < 0.001; ns: not significant. Blue-toned boxes denote intervention and mediator constructs; red boxes denote outcome constructs; solid blue arrows indicate mediation paths; dashed orange arrow represents the direct BKT reinforcement loop.
Figure 2.
CONSORT-style participant flow diagram. N = 200 enrolled; n = 187 analysed (93.5% retention); differential attrition: χ
2(1) = 0.14,
p = 0.71 (non-significant). PSM-matched subsample: n = 81 pairs per group. Note. Boxes with bold borders denote key sample-size milestones. Diagram designed for greyscale-robust printing (distinct line styles and light/dark contrasts rather than hue alone). Source: this study; n’s reported in
Section 3.1.
Figure 3.
Security-incident trends by group. (A) Pre-intervention parallel-trends verification (β = 0.42, SE = 0.89, p = 0.64—non-significant interaction confirms parallel trends). (B) Full 12-week study period showing a 48.9% post-intervention incident reduction in the AI group. Solid line = AI-Adaptive; dashed line = Control—IRR (incidence-rate ratio) is the ratio of post-intervention incident rates between the AI-adaptive and control groups; an IRR < 1 indicates a relative reduction.
Figure 4.
Simulated phishing click-rates at three time points (Weeks 0, 6, and 12). AI-Adaptive: 8.8% → 5.1% → 2.1% (75% relative reduction). Control: 9.0% → 8.6% → 8.4% (no meaningful change). Identical stimuli were administered by a blinded third-party vendor. Any highlight arrow in the figure is explanatory only and marks the post-intervention decline in the AI-Adaptive group; no result depends on colour alone. χ2(1) = 8.74, p = 0.003, φ = 0.22.
Figure 5.
Kaplan–Meier time-to-competency survival curves (80% mastery threshold). AI-Adaptive median = 4.2 weeks (95% CI: [3.6, 4.8]); Control did not reach threshold within 12 weeks. Log-rank χ2(1) = 47.3, p < 0.001.
Figure 6.
Mediation path diagram (PROCESS Model 4; N = 187; 5000 bootstrap draws). a1 = 1.24 *** (SE = 0.18); b1 = 0.68 *** (SE = 0.12); Indirect1 = 0.843 [0.52, 1.19] ***. a2 = 0.97 *** (SE = 0.21); b2 = 0.54 *** (SE = 0.14); Indirect2 = 0.524 [0.28, 0.81] ***. a3 = 0.31 ns (SE = 0.24); b3 = 0.19 ns (SE = 0.16); Indirect3 = 0.059 [−0.07, 0.21] ns. Direct effect c′ = 0.63 * (SE = 0.28). Total c = 2.06 ***; proportion mediated = 66.4%. *** p < 0.001; * p < 0.05; ns = non-significant. Significant versus non-significant paths are distinguished by line style and contrast, not by colour alone.
Figure 7.
Effect-size summary across primary outcomes (four-method consensus). All outcomes significantly exceed the meta-analytic benchmark (d = 0.33) [
1]. Rosenbaum bounds Γ = 2.1; E-values ≥ 3.4; FDR-corrected q < 0.005 for all five outcomes.
Table 1.
Summary comparison: this study versus representative recent AI-adaptive cybersecurity training studies.
| Study | System Focus | Empirical Status | Mediator Testing | Objective Incident Data | Context |
|---|
| Zhdanov et al. [9] | Generative-AI healthcare awareness training | Design/pilot-stage evaluation | No | No | US healthcare |
| Ahmed and Mejri [10] | GAI-driven adaptive CSAT framework | Framework/architecture paper | No | No | Conference prototype |
| Sengupta et al. [11] | Review of adaptive cybersecurity education | Systematic/structured review | No | No | Cross-study |
| Present study | Custom-configured 4-parameter BKT model with interpretable P(G) and P(S) signals serving as an auditable pedagogical decision layer that directly triggers PMT/TPB-aligned interventions (rather than only estimating mastery) | 12-week field evaluation | Yes | Yes | Yemen (developing country, offline-first) |
Table 2.
Baseline characteristics by group (key variables).
| Characteristic | AI-Adaptive (n = 94) | Control (n = 93) | Std. Diff. Pre-PSM | Std. Diff. Post-PSM |
|---|
| Age (Mean, SD) | 31.2 (6.4) | 38.7 (8.1) | 1.02 (large) | 0.06 |
| Female (%) | 38.3% | 41.9% | 0.07 | 0.03 |
| University degree (%) | 71.3% | 44.1% | 0.58 (moderate) | 0.07 |
| Prior cyber training (%) | 82.4% | 12.9% | 2.14 (extreme) | 0.08 |
| Technical literacy (0–10) | 6.8 (1.4) | 4.2 (1.9) | 1.55 (large) | 0.09 |
| Baseline knowledge (0–100) | 64.3 (11.2) | 41.8 (13.4) | 1.82 (extreme) | 0.10 |
| Baseline compliance (0–10) | 5.9 (1.7) | 3.8 (2.1) | 1.10 (large) | 0.09 |
Table 3.
Conceptual mapping between BKT diagnostic signals and PMT/TPB constructs.
| BKT Signal | Inferred Cognitive State | PMT Construct | TPB Construct | Triggered InterVention | Refs. |
|---|
| High P(S)—slip | Execution failure | Response cost/response effort | Perceived Behavioural Control (PBC) | Checklist and slower review items | [2,5,16,17] |
| High P(G)—informed-guess | Misconception or low confidence | Self-efficacy and coping | Self-efficacy and attitude | Misconception-correction + rule recall | [2,5,21] |
| P(L) < 0.50 | Unfamiliarity | Response efficacy | PBC | Worked example + scaffold + recovery drill | [6,18,21] |
| 0.40 ≤ P(L) ≤ 0.80 | Consolidation (ZPD band) | Coping appraisal | Attitude | Progressively harder items | [6,12,20] |
| P(L) > 0.70 | Mastery | Threat appraisal (booster) | Subjective norms | Brief threat-salience vignette | [1,3,16] |
Table 4.
Training content, module coverage, and time-on-task by group.
| Module/Security Concept | AI | Control | AI Avg. Time (Min) | Control (Min/Session) |
|---|
| 1. Phishing Recognition and Avoidance | ✓ | ✓ | 18.4 | 30 |
| 2. Password Security and MFA | ✓ | ✓ | 14.2 | 20 |
| 3. Social Engineering Awareness | ✓ | ✓ | 16.1 | 20 |
| 4. Malware and Ransomware Prevention | ✓ | ✓ | 13.8 | 15 |
| 5. Secure Email Practices | ✓ | ✓ | 11.9 | 15 |
| 6. Data Handling and Classification | ✓ | ✓ | 15.3 | 20 |
| 7. Physical Security and Clean Desk | ✓ | ✓ | 8.7 | 10 |
| 8. Incident Reporting Procedures | ✓ | ✓ | 12.4 | 15 |
| 9. Remote Work Security | ✓ | ✓ | 9.6 | 10 |
| 10. Security Policy Compliance | ✓ | ✓ | 14.2 | 15 |
| TOTAL (Mean across 12 weeks) | — | — | 134.6 min (~12.8 h ± 2.4) | 120 min (~12.0 h) |
Table 5.
Descriptive statistics and four-method effect-size estimates by primary outcome.
| Outcome | AI M(SD) Wk12 | Ctrl M(SD) Wk12 | ANCOVA d [95% CI] | PSM d [95% CI] | DiD β (SE) | ME [95% CI] | FDR-adj q |
|---|
| Knowledge Score (0–100) | 82.4 (8.6) | 67.1 (11.3) | 0.89 [0.71, 1.07] | 0.84 [0.63, 1.05] | 16.8 (2.3) *** | 17.2 [14.1, 20.3] | <0.001 (original p < 0.001) |
| Behavioural Intentions (SeBIS 1–7) | 5.82 (0.71) | 5.16 (0.87) | 0.76 [0.57, 0.95] | 0.72 [0.52, 0.92] | 0.71 (0.12) *** | 0.68 [0.47, 0.89] | <0.001 (original p < 0.001) |
| Policy Compliance (0–10) | 7.84 (1.18) | 6.12 (1.54) | 0.78 [0.59, 0.97] | 0.74 [0.53, 0.95] | 1.72 (0.28) *** | 1.68 [1.20, 2.16] | <0.001 (original p < 0.001) |
| Security Incidents (Tier 2–3) | 11 | 22 | IRR = 0.51 [0.38, 0.68] | IRR = 0.49 [0.36, 0.66] | −0.49 (0.09) *** | IRR = 0.51 [0.38, 0.68] | <0.001 (original p < 0.001) |
| Phishing Click-Rate (%) | 2.1% | 8.4% | φ = 0.22 [0.09, 0.35] | φ = 0.21 [0.08, 0.34] | −0.063 (0.019) *** | OR = 0.22 [0.07, 0.67] | 0.004 (original p = 0.003) |
Table 6.
Mediation analysis: indirect effects of training on policy compliance (5000 bootstrap draws).
| Pathway | a-Path b [SE] | b-Path b [SE] | Indirect Effect b | 95% Boot CI [Lo, Hi] | Significance |
|---|
| Training → Coping Self-Efficacy → Compliance | 1.24 [0.18] | 0.68 [0.12] | 0.843 | [0.52, 1.19] | Significant *** |
| Training → PBC → Compliance | 0.97 [0.21] | 0.54 [0.14] | 0.524 | [0.28, 0.81] | Significant *** |
| Training → Threat Appraisal → Compliance | 0.31 [0.24] ns | 0.19 [0.16] ns | 0.059 | [−0.07, 0.21] | Non-significant |
| Direct Effect (c′): Training → Compliance | — | — | 0.630 | [0.08, 1.18] | Significant * |
| Total Effect (c): Training → Compliance | — | — | 2.056 | [1.58, 2.54] | Significant *** |
| Proportion Mediated (Self-Efficacy + PBC) | — | — | 66.4% | — | — |
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |