Highlights
What are the main findings?
- AAIGF-E provides 111 controls across seven lifecycle phases with a dual-narrative architecture serving both AI governance stakeholders and OT engineering teams, mapped to NERC CIP, NIST AI RMF, ISA/IEC 62443, and MITRE ATLAS.
- Scenario-driven validation against three operationally relevant failure modes (adversarial DER dispatch manipulation, unauthorized agentic AI action in OT, and silent model drift) confirms structural coverage and tests prompt-injection pathways through operational data feeds, showing that DI-006 provides technical runtime ingestion validation while a residual semantic-manipulation gap remains before downstream actuation controls, alongside other implementation and assurance limitations.
What are the implications of the main findings?
- Compliance with existing instruments does not equal operational assurance: pre-operational assurance gates, rather than post-deployment detection, must carry the governance load for AI coupled to physical grid infrastructure.
- Utilities, municipal operators, and regulators gain a phased implementation pathway connecting AI lifecycle risk management to existing compliance cycles, extensible to other smart-city infrastructure domains.
Abstract
AI is being deployed in smart-grid and smart-city energy infrastructure faster than governance frameworks can validate it. Existing instruments—NERC CIP, NIST AI RMF 1.0, and ISA/IEC 62443—address cybersecurity compliance and AI trustworthiness principles but do not provide operational control architecture for AI systems influencing grid stability, distributed energy resource (DER) dispatch, or real-time reliability decisions. Compliance with these instruments is not equivalent to operational assurance: their requirements do not by themselves establish validated input integrity controls, adversarial robustness testing, or bounded autonomy architecture for every AI-enabled operational use case. This paper presents the AI and Agentic Intelligence Governance Framework, Electric Sector (AAIGF-E): 111 controls across seven lifecycle phases (Govern, Design, Implement, Assure, Monitor, Respond, Recover), mapped to existing electric-sector compliance obligations. A defining structural feature of AAIGF-E is its dual-narrative control architecture, in which each control carries both an AI governance rationale and a parallel electric-sector operational translation, enabling the framework to serve AI governance stakeholders and operational technology (OT) engineering teams within the same control set without requiring cross-discipline translation. Scenario-driven validation tests the framework against three operationally relevant failure modes: adversarial manipulation of DER dispatch AI, unauthorized autonomous action execution by an agentic AI system in an OT environment, and silent model drift in grid stability forecasting. Validation confirms structural coverage across prevention, detection, and response phases and identifies residual implementation and assurance gaps while identifying a residual prompt-injection pathway in which technically valid operational content may manipulate agent task scope or apparent authority before downstream actuation controls activate. The 111-control register is further characterized through quantitative analysis of standards coverage, lifecycle and governance-domain distribution, cross-framework convergence, and sector-specific translation. A separate independent coding reproducibility check samples 35 controls across five reference frameworks, producing 175 binary mapping decisions; 155 decisions agree with the internal coding key (88.6% raw agreement; Cohen’s ). To reduce reliance on author-led validation, a representative operational scenario is also evaluated by independent AI governance, cybersecurity, and ICS/OT experts who derive their judgments before reviewing the AAIGF-E reference assessment. Their domain selections and governance recommendations show strong overall convergence with the framework assessment. AAIGF-E is designed for electric utilities and smart-city operators within North American regulatory frameworks who need governance architecture that connects AI lifecycle risk management to existing compliance cycles rather than replacing them.
1. Introduction
AI is already operating in smart grid and smart-city energy infrastructure. AI-assisted optimization, real-time load balancing, and increasingly autonomous operational functions are already being deployed across the electric sector [1,2]. NERC has identified AI-driven control systems as an emerging reliability consideration [1]. The U.S. Department of Energy has identified opportunities for AI across grid planning, permitting, operations and reliability, and resilience [2]. Such deployments can proceed without structured governance, lifecycle assurance, or bounded autonomy controls appropriate for systems with direct physical consequences.
This distinction matters architecturally, not just operationally. A failure in an IT system typically degrades information: a wrong recommendation, a stale dashboard, a delayed report. A failure in an OT-connected AI system can degrade a physical process: a voltage excursion, an unplanned islanding event, an equipment trip. AI governance frameworks built for IT risk, model bias in a recommendation engine, and explainability for a lending decision do not transfer cleanly to a domain where a wrong output has a physical stability consequence measured in operating cycles, not information-quality terms. This is the structural reason a sector-specific control architecture is necessary rather than a direct application of general-purpose AI governance principles.
The gap is not missing principles. NIST AI RMF 1.0 defines trustworthiness but provides no control architecture to enforce it: application of its subcategories does not itself establish a sector-specific pre-operational assurance gate. ISO/IEC 42001 provides a management system structure but does not specify what assurance activities must occur before an AI system goes live. NERC CIP governs bulk-electric-system cybersecurity. ISA/IEC 62443 governs deterministic industrial control systems. These frameworks provide critical governance foundations. NERC CIP established enforceable reliability standards where none existed. NIST AI RMF provides a systematic approach to AI risk across the full system lifecycle. ISA/IEC 62443 defines security architecture for industrial control environments that remains the reference standard for OT practitioners. These instruments do not collectively provide a sector-specific control architecture for adaptive AI operating within physical OT stability constraints; that boundary reflects their respective scopes, which AAIGF-E is designed to complement rather than replace. That architecture requires a structural shift toward pre-operational gating, emphasizing assurance before deployment rather than relying primarily on post-deployment monitoring. This philosophy underpins the framework’s Assure phase and reflects a significant departure from traditional electric-sector compliance practice. The missing piece is operational: a sector-specific control architecture that connects AI lifecycle risk management to the compliance cycles and physical stability constraints that electric sector operators already work within.
This paper presents the AI and Agentic Intelligence Governance Framework, Electric Sector (AAIGF-E). The framework spans seven lifecycle phases and eleven domain clusters, mapped to existing compliance obligations that electric sector operators already work within. It is designed not to replace existing instruments but to operationalize them for AI and, critically, for the emerging class of agentic AI systems that extend autonomous action beyond decision support into action execution. While the compliance alignment logic is explicitly scoped to North American regulatory frameworks (NERC CIP, NIST AI RMF), the underlying control architecture—including lifecycle phase structure, domain clustering, dual-narrative design, and assurance-first philosophy—is designed for reuse under alternative regulatory mappings with targeted gap analysis.
This paper argues that the remaining gap does not arise from an absence of governance principles but from the limited degree to which existing instruments specify sector-specific, pre-operational AI assurance controls for cyber-physical electric-sector deployment. AAIGF-E addresses this gap by making assurance gates a load-bearing part of the control architecture, with governance documentation supporting rather than substituting for operational assurance. In this respect, AAIGF-E occupies a different design position from the frameworks and standards against which it is mapped. NIST AI RMF provides cross-sector AI risk-management outcomes and trustworthiness characteristics; ISO/IEC 42001 establishes organizational requirements for an AI management system; NERC CIP defines enforceable cybersecurity obligations for bulk-electric-system assets; ISA/IEC 62443 provides cybersecurity requirements and architectural principles for industrial automation and control systems; and MITRE ATLAS provides a threat-informed knowledge base for adversarial techniques against AI systems. AAIGF-E does not duplicate these functions. Its contribution is to integrate their relevant governance, cybersecurity, and adversarial-assurance objectives into an electric-sector AI control architecture that specifies lifecycle activation, pre-operational assurance gates, runtime data-trust controls, bounded AI-to-OT actuation, monitoring and response obligations, and operational translation for grid environments. The novelty therefore lies in the control-level orchestration and sector-specific operationalization of these complementary instruments rather than in claiming replacement or functional equivalence with any of them. This paper makes three contributions. First, it defines and quantitatively characterizes a 111-control architecture across seven lifecycle phases, with 78 controls activating at the Assure phase, and evaluates control-level mapping convergence and reproducibility across the reference frameworks and standards. Second, it introduces a dual-narrative control architecture that expresses each control in both AI-governance and OT-operational terms, enabling governance requirements to be interpreted within electric-sector operational contexts. Third, it provides scenario-driven validation against three operationally relevant smart-grid and smart-city failure modes, confirming structural coverage while identifying residual implementation and assurance limitations. For prompt-injection and data-layer trust risks, the validation distinguishes DI-006 technical runtime ingestion validation from OP-009 downstream actuation mediation and identifies a residual semantic-manipulation gap where otherwise technically valid operational content can influence task scope or apparent authority before the actuation gate. This preserves the v1.0 control boundaries while identifying a concrete requirement for future control formalization and empirical adversarial testing.
This study contributes to smart-city research by shifting AI governance from principle-based trustworthiness toward operational assurance for cyber-physical urban infrastructure. While the validation focuses on electric-sector smart-city infrastructure, the framework’s lifecycle structure may also be applicable to other AI-mediated urban systems [3,4]. The paper therefore advances smart-city governance scholarship by providing a control-level architecture for AI systems whose decisions affect physical urban operations.
The paper is structured as follows. Section 2 reviews the literature across smart-city cyber-physical systems resilience, smart-grid cybersecurity, AI governance, human oversight, agentic AI, threat-informed assurance, and digital twins, and identifies the structural governance gap. Section 3 describes the Design Science Research methodology, framework architecture, control translation logic, scenario validation design, and limitations. Section 4 presents scenario-driven validation results across three failure modes. Section 5 discusses implications, unresolved gaps, and practitioner constraints. Section 6 concludes with future work directions.
2. Literature Review
2.1. Smart Cities as Cyber-Physical Critical Infrastructure
Smart cities increasingly operate as tightly coupled cyber-physical systems in which urban services depend on continuous coordination across energy, mobility, buildings, communications, and public safety platforms. Within this ecosystem, the smart grid functions as a resilience backbone due to the electrification of transport, distributed energy resource (DER) proliferation, and integration of intelligent control layers. Lee identified that CPS integrated computation and physical processes through continuous feedback, arguing that conventional object-oriented software abstractions were inadequate for the safety and reliability requirements this coupling introduced [5]. Amin characterized the growing complexity and data intensity of interconnected power systems, showing how tight coupling across sensing, communication, and control layers amplified systemic risk at grid scale [6]. Cardenas et al. framed CPS security as requiring three things: understanding the consequences of attacks, identifying the properties that distinguish CPS from conventional IT security, and building prevention, detection, and resilience mechanisms for failures that propagate from digital control into physical processes [7].
Gungor et al. surveyed the communication technologies and standards enabling the smart grid, mapping the state of the art and open research issues across the sector [8]. Fang et al. structured the smart-grid landscape into three constituent systems, smart infrastructure, smart management, and smart protection, and surveyed enabling technologies within each [9]. Mohsenian-Rad and Leon-Garcia proposed an energy consumption scheduling algorithm combined with a real-time price predictor, showing reductions in both user electricity costs and peak-to-average load demand [10]. While these applications improve operational efficiency and sustainability, they introduce new governance challenges, as AI systems increasingly influence load balancing, DER coordination, and storage dispatch decisions that operate within constrained physical tolerances of voltage and frequency stability. The literature thus establishes that smart-city energy infrastructures are not conventional IT systems; they are mission-critical CPS environments where algorithmic decisions may have systemic physical consequences. The governance challenge extends beyond the grid itself. Smart-city deployments now include AI systems managing traffic signal coordination, building energy management across district heating and cooling networks, public safety platforms with autonomous anomaly detection, and smart lighting infrastructure responding to real-time occupancy and environmental data [3,4]. Each of these systems operates within physical tolerances, influences real-time resource allocation, and in many deployments shares data pipelines with the grid-connected energy layer. An AI system managing district building energy loads issues dispatch signals that directly influence DER aggregation and grid demand response programs. A failure in that system is not isolated to the building management domain. It propagates into grid operations. The governance boundary between smart-city OT and smart-grid OT is not a clean line. It is a coupling point. AAIGF-E addresses this coupling directly: the framework governs AI systems influencing grid-connected energy infrastructure, which in smart-city environments includes building management systems, EV charging networks, and demand response aggregators operating alongside traditional grid assets [11].
2.2. Smart-Grid Cybersecurity Foundations and the AI Governance Gap
Smart-grid cybersecurity governance is relatively mature compared to many other domains. NISTIR 7628 provides analytical guidance for smart grid cybersecurity architecture, risk assessment, and system-specific mitigation strategies [12]. The electric sector is further governed by enforceable reliability standards under the North American Electric Reliability Corporation (NERC) Critical Infrastructure Protection (CIP) framework [13,14]. Industrial control system security is structured by the ISA/IEC 62443 series, which defines security levels, zones and conduits, and lifecycle requirements for industrial automation and control systems [15]. Information security management for these environments is further supported by ISO/IEC 27001, which provides the baseline security controls and audit framework against which OT-adjacent IT systems are assessed [16]. NIST SP 800-53 provides a comparable, US-federal-oriented catalog of security and privacy controls that informs several of AAIGF-E’s operational security domain controls as a supplementary reference alongside the five principal mapped frameworks [17].
However, these instruments were not designed to govern adaptive AI systems, model drift, data poisoning risks, or semi-autonomous decision-making architectures. They assume deterministic or rule-based control environments. As AI systems are introduced into smart-grid and smart-city contexts, a governance gap emerges: cybersecurity compliance exists, but structured AI lifecycle assurance and bounded autonomy controls are underdeveloped. This gap aligns with broader observations in AI governance research that domain-neutral risk management frameworks require sector-specific operationalization to become actionable [18]. Consequently, AAIGF-E treats this sector-specific operationalization as its central design problem: translating cybersecurity compliance obligations already in force into the AI-specific lifecycle assurance and bounded autonomy controls these instruments do not provide.
2.3. AI Governance and Lifecycle Risk Management
The NIST AI Risk Management Framework (AI RMF 1.0) formalizes governance, mapping, measurement, and management functions across the AI lifecycle, emphasizing trustworthiness characteristics including safety, security, accountability, and explainability. However, it intentionally remains sector-agnostic [18]. Bovens defined accountability narrowly as a relationship between an actor and a forum, in which the actor must explain and justify conduct, the forum can question and pass judgment, and the actor may face consequences, and offered evaluative perspectives, democratic, constitutional, and learning, for assessing accountability arrangements [19]. Floridi et al. synthesized five ethical principles for AI and translated them into twenty concrete recommendations for assessing, developing, incentivizing, and supporting AI, directed at policymakers and other stakeholders [20]. Building on the governance and cyber-physical systems literature reviewed above, this paper treats traceability, responsibility allocation, and escalation logic as architectural properties that lifecycle governance controls should satisfy rather than administrative afterthoughts. In the context of smart-grid and smart-city infrastructures, lifecycle governance must extend beyond documentation and policy alignment to define operational triggers, assurance checkpoints, and bounded authority mechanisms consistent with infrastructure stability constraints.
Recent scholarship has sharpened the governance challenge for urban AI deployments specifically. Cugurullo et al. [3] document the emergence of AI urbanism as a distinct trajectory from smart-city urbanism: where smart-city systems optimize existing urban infrastructure, AI urbanism introduces autonomous decision-making agents that actively reconfigure urban services, governance structures, and physical space. City traffic management platforms, autonomous building energy controllers, and predictive public safety systems are operating across multiple urban domains simultaneously, often without governance architectures capable of bounding their autonomy or auditing their decisions. These forms of urban AI are already being deployed across contemporary smart-city environments [3,4]. The governance gap AAIGF-E addresses in electric-sector contexts mirrors this broader urban AI governance deficit: operational AI systems are ahead of the frameworks designed to govern them, and the frameworks that exist were not built for adaptive, multi-domain AI operating within physical infrastructure constraints [4].
2.4. Human Oversight and Automation in Critical Systems
Human-automation interaction research provides foundational grounding for meaningful oversight in autonomous systems. Parasuraman and Riley [21] identified automation bias as a systemic risk in high-reliability environments. Sheridan [22] and Endsley [23] further articulated supervisory control and situational awareness principles necessary for safe human–machine interaction. In safety-critical domains, oversight must be operationally meaningful: humans must retain evaluative authority, defined intervention points, and adequate system transparency to maintain control [24]. This aligns with contemporary regulatory developments, including high-risk AI oversight requirements under the EU AI Act [25], which mandate human-in-the-loop safeguards for systems impacting safety and critical infrastructure. Recent work on embodied AI in critical infrastructure similarly argues that resilience depends on bounded autonomy within a hybrid governance architecture, proposing oversight modes mapped to sectors by task complexity, risk level, and consequence severity [26]. While AAIGF-E is scoped to North American regulatory contexts, the oversight architecture principles informing its human authority controls—confidence thresholds, intervention triggers, and escalation logic—are consistent with the EU Act’s human-in-the-loop requirements and draw on the same foundational human-automation literature.
For smart-grid and smart-city systems, oversight must accommodate reduced intervention latency and distributed decision-making. As AI autonomy deepens, oversight mechanisms must be architecturally embedded through confidence thresholds, drift detection, rollback logic, and structured escalation. AAIGF-E operationalizes this requirement directly: PG-006 and OP-009 embed confidence thresholds and actuation gatekeeping as architectural properties of the control set itself, not supplementary guidance layered on top of it.
2.5. Agentic AI and Distributed Decision Authority
Multi-agent systems and distributed control research demonstrates increasing use of agent-based architectures for energy coordination. Rahimi and Ipakchi positioned demand response, distributed generation, and distributed storage collectively as distributed energy resources (DERs) and argued for treating demand response as a market resource in transmission and wholesale operations [27]. Pipattanasomporn et al. designed and implemented a FIPA-compliant multi-agent system for a microgrid, demonstrating autonomous transition to island mode when upstream outages are detected [28]. Multi-agent systems enable decentralized coordination across DER clusters, microgrids, and demand response programs. Glavic et al. reviewed applications of reinforcement learning to electric power system decision and control problems across multiple time scales and assessed RL’s prospects given emerging communication and instrumentation technologies [29]. Zhang et al. reviewed deep reinforcement learning fundamentals and surveyed its applications across power-system energy management, demand response, electricity markets, and operational control [30]. While these architectures enhance adaptability and scalability, they distribute decision authority and reduce centralized human visibility.
Agentic AI systems, characterized by iterative planning, tool usage, and autonomous action loops, expand this autonomy depth. Yao et al. proposed ReAct, interleaving reasoning traces with task-specific actions so that language models can induce, track, and update action plans while interfacing with external tools and environments [31]. Xi et al. surveyed the use of large language models as a general foundation for building AI agents, synthesizing a previously fragmented body of agent research into a unified framework [32]. Their integration into infrastructure control environments may amplify governance risk by increasing autonomy and distributing decision authority beyond what earlier, narrower agent architectures required. Distributed autonomy introduces challenges in traceability, responsibility attribution, and cascading error propagation. A comprehensive recent survey formalizes agentic AI for smart grids specifically, proposing autonomy levels from advisory analytics through fully coordinated autonomous operation, and identifying governance, safety, and human-oversight requirements as central open challenges alongside the underlying reinforcement learning and multi-agent architectures [33]. However, this and related surveys catalog technical capability and autonomy taxonomies without specifying the lifecycle control architecture that would bind that autonomy to sector-specific safety thresholds. Consequently, the literature lacks sector-specific governance models that bind agentic autonomy to infrastructure stability thresholds and lifecycle safety controls.
2.6. Threat-Informed Assurance and Adversarial Risk
AI-enabled smart-grid systems are exposed to adversarial manipulation, data poisoning, model tampering, and adversarial input attacks. Huang et al. established adversarial machine learning as a field of study, providing an early taxonomy of attacks against online learning algorithms [34]. Biggio and Roli provided a decade-spanning retrospective of the field, reviewing threat models and evasion and poisoning attacks and discussing the limitations of current defenses [35]. MITRE ATLAS formalizes adversarial tactics and techniques for AI systems and is increasingly used to inform threat modeling [36]. Recent prompt-injection research shows that indirect instructions embedded in untrusted external content can influence LLM-based agents and become especially consequential when those agents are connected to privileged tools or operational systems [37]. In OT-connected settings, this creates a credible pathway by which compromised context could influence consequential actions, reinforcing the need to separate untrusted content from trusted instructions and authorization state. In CPS environments, adversarial ML risks intersect with operational reliability constraints. Threat-informed assurance thus becomes a necessary complement to governance lifecycle models. Without adversarial validation gates embedded into operational workflows, AI systems in smart-city energy contexts may fail under stress or manipulation.
2.7. Digital Twins and Scenario-Based Assurance
Digital-twin architectures are increasingly adopted in smart-city and smart-grid environments to simulate infrastructure behavior, perform predictive maintenance, and stress-test operational configurations. Tao et al. reviewed the state of the art in digital-twin research, covering key components and major industrial applications [38]. Djebali et al. surveyed digital-twin design and modeling processes specifically for smart-grid applications, including a proposed digital-twin operations management design [39]. From a governance perspective, digital twins provide a structured environment for scenario-driven validation of AI behavior under stress conditions. They enable simulation of cascading failure scenarios, drift conditions, adversarial manipulation, and rollback procedures without exposing live infrastructure. However, the literature on digital twins primarily emphasizes modeling and optimization rather than governance gating and autonomy boundaries. AAIGF-E addresses this by treating governance gating as a lifecycle property enforced independently of the simulation environment, so digital-twin validation, when adopted, tests the framework’s existing control activation logic rather than substituting for it. Future validation iterations of AAIGF-E are designed to be compatible with digital-twin simulation environments, enabling live scenario testing of control activation sequences and escalation logic without exposing production grid infrastructure to governance validation workloads.
2.8. Identified Gap
Across CPS resilience research, smart-grid AI applications, governance frameworks, human oversight scholarship, and adversarial AI literature, a structural gap emerges. No currently published governance instrument provides a sector-specific operational architecture that:
- Binds AI autonomy to grid stability constraints;
- Embeds lifecycle gating aligned with OT realities;
- Integrates adversarial assurance into operational workflows;
- Provides scenario-driven escalation logic for agentic AI;
- Aligns with existing energy-sector compliance ecosystems.
This gap motivates the development of AAIGF-E: a sector-specific governance architecture for smart-grid and smart-city systems that operationalizes lifecycle risk management into bounded autonomy and scenario-driven control mechanisms. Table 1 presents the authors’ structured assessment of existing framework limitations relative to AAIGF-E design objectives, based on primary review of published framework documentation rather than a formal crosswalk methodology.
Table 1.
Comparative analysis of existing governance frameworks relative to AAIGF-E design objectives.
3. Materialsand Methods
Microsoft Word and Microsoft Excel were used for manuscript preparation, tabulation, and basic quantitative analysis. No specialized experimental equipment or analytical software was used.
3.1. Research Design
This paper applies Design Science Research (DSR) methodology to develop and validate a sector-specific AI governance artifact for smart-grid and smart-city environments. Hevner et al. established design science as a research paradigm distinct from behavioral science, providing a conceptual framework and evaluation guidelines for information systems artifacts [41]. Peffers et al. built on this foundation with a concrete nominal process model for conducting and presenting Design Science Research [42]. DSR is appropriate when the research objective is to construct a novel artifact, such as a framework, model, or architecture, and evaluate its utility against a defined problem class. The problem class here is the governance gap identified in Section 2: no existing instrument binds AI autonomy to grid stability constraints while integrating adversarial assurance and lifecycle gating aligned with OT realities.
The artifact produced is AAIGF-E: a structured control architecture of 111 controls organized across seven lifecycle phases and eleven domain clusters. DSR evaluation is conducted through scenario-driven control mapping, in which realistic smart-city and smart-grid failure scenarios are used to trace control activation, identify coverage gaps, and validate architectural coherence. This approach is consistent with DSR practice in security and governance research where proprietary operational data are unavailable, but real-world failure modes are well documented in the incident literature and regulatory guidance [43]. This study performs analytical validation through scenario-grounded control activation tracing rather than empirical deployment testing. Structural coverage, sequencing coherence, and gap surfacing are the evaluation criteria. Empirical validation against live operational deployments is identified as the primary future work direction in Section 6.
Rationalefor Analytical Scenario Validation
Analytical scenario validation is appropriate at this stage for four reasons. First, live utility data are rarely available for public research because operational technology environments are safety-critical and security-sensitive. Second, critical infrastructure testing cannot be safely conducted in production systems when scenarios involve adversarial data manipulation, unsafe agentic action, or stability forecasting drift. Third, DSR permits demonstration-stage artifact evaluation when the research goal is to test structural coherence and practical utility rather than measure operational performance. Fourth, the objective of this validation is coverage analysis: whether the right controls exist at the right lifecycle phases and whether control activation sequences follow a prevention–detection–response logic.
Future validation should extend the completed preliminary expert assessment through a larger panel and additional scenarios, followed by digital-twin simulation and pilot implementation with a utility or municipal operator. Expanded expert review would test whether the observed convergence generalizes across different domain combinations, consistent with established design science evaluation approaches [41,42,43]. Digital-twin simulation would test control activation without exposing production infrastructure [38,39]. Pilot implementation would evaluate implementation cost, procurement constraints, audit artifacts, and organizational learning in a live institutional setting.
3.2. Framework Architecture
AAIGF-E is intended for AI-enabled or agentic systems whose outputs, recommendations, automated decisions, or actions can materially influence electric-sector cyber-physical operations, operational decision-making, safety, reliability, security, or recovery. The framework is not intended to imply that all 111 controls apply uniformly to every AI use case. Control applicability depends on factors including the system’s operational role, degree of autonomy, potential physical consequence, data and system dependencies, lifecycle stage, deployment architecture, and applicable regulatory context. Control applicability should therefore be interpreted in relation to the characteristics and risk context of the specific implementation rather than as uniform activation of all 111 controls.
AAIGF-E is structured around a seven-phase AI lifecycle aligned to the operational realities of electric-sector environments. The Govern (G) phase addresses policy, accountability structures, regulatory alignment, and decision authority. The Design (D) phase covers AI system architecture, risk classification, and security-by-design requirements. The Implement (I) phase encompasses deployment controls, access enforcement, and integration security. The Assure (A) phase governs validation, testing, adversarial assurance, and pre-operational gating. The Monitor (M) phase covers operational telemetry, drift detection, behavioral monitoring, and anomaly detection. The Respond (Rs) phase addresses incident classification, containment, and escalation. The Recover (Rc) phase governs controlled re-enablement of trusted AI capability, re-validation, and lessons learned documentation.
Controls span eleven domain clusters: Policy and Governance (PG, 10 controls), Risk Assessment (RA, 3), System Integrity (SI, 18), Model Integrity (MI, 16), Testing and Evaluation (TE, 10), Data Integrity (DI, 7), Operational Control (OP, 10), Output and Decision Integrity (OM, 8), Supply Chain (SC, 9), Privacy (PR, 12), and Incident Response (IR, 8). Most controls are intentionally multi-phase: a model integrity control activates during Design when the architecture is specified, during Implement when the model is deployed, during Assure when integrity verification is conducted, and during Monitor when drift or tampering is detected. Governance obligations do not reset at deployment; they persist and evolve across the system lifecycle.
The complete AAIGF-E control register, including all 111 controls with Control ID, Name, Lifecycle Phase, Domain, Control Objective, AI Narrative, Electric-Sector Narrative, and Reference Standards mappings, is documented in [11]. Table 2 and Table 3 present representative controls across seven domains.
Table 2.
Representative AAIGF-E controls across seven domains–AI Governance Narratives (part A).
Table 3.
Representative AAIGF-E controls across seven domains–Electric-Sector Narratives (part B).
3.3. Control Translation Logic
Each AAIGF-E control was constructed through a four-step translation process. First, a governance outcome was defined: what the control must achieve in terms of trustworthiness, safety, or operational integrity, derived from the gap analysis in Section 2 to ensure every control traces to at least one identified structural deficiency. Second, the outcome was translated into an operational requirement specific to smart-grid and smart-city deployment contexts, including voltage and frequency stability constraints, DER coordination boundaries, OT network segmentation requirements, and human override authority thresholds. Third, the control was mapped to applicable reference standards. NIST AI RMF 1.0 is referenced in 102 of 111 controls. ISO/IEC 42001 is referenced in 101 controls. MITRE ATLAS is actively referenced in 59 controls; the remainder are marked not applicable given their operational focus. NERC CIP standards are referenced in 70 controls. ISA/IEC 62443 is referenced in 25 controls. This layered mapping ensures controls are not invented from scratch but are grounded in existing compliance expectations that electric-sector operators already recognize.
Fourth, sector-specific narratives were authored for each control: an AI narrative explaining the governance rationale and an electric-sector narrative translating that rationale into grid operational language. This dual-narrative structure allows the framework to serve both AI governance stakeholders and OT engineering teams without requiring translation across technical vocabularies.
The resulting control distribution reflects deliberate design choices. Across the seven lifecycle phases, 78 controls activate at Assure, 64 at Implement, 52 at Design, 45 at Monitor, 36 at Respond, 35 at Govern, and 8 at Recover. These figures sum to 318 activation points—not 111—because most controls are intentionally multi-phase. The concentration at Assure and Implement reflects OT’s zero-tolerance for post-deployment failure: in grid-coupled systems, prevention and pre-operational gating carry the governance load because the cost of runtime failure is disproportionate to the cost of assurance. Domain concentration follows the same principle: SI (18 controls), MI (16), and OP (10) dominate because system integrity, model integrity, and actuation authority and operational control are the proximate failure surfaces for AI in electric-sector environments.
As a worked example, DI-006 (Runtime Data Validation at AI Inference Ingestion Points) was derived as follows. The governance outcome (Step 1) was defined as closing the gap identified in Section 2: no existing framework specifies pre-inference data validation as an architectural requirement rather than a general data-quality recommendation. The operational requirement (Step 2) translated that outcome into a concrete constraint: SCADA values, PMU measurements, DER set points, and market signals must fall within physically plausible ranges before reaching the inference engine. The standards mapping (Step 3) tied this requirement to NIST AI RMF’s Map function, ISO/IEC 42001’s operational planning provisions, and NERC CIP’s data integrity requirements. The dual narrative (Step 4) then produced the AI governance and electric-sector operational language shown in Table 3. This same four-step process was applied uniformly across all 111 controls.
3.4. Independent Coding Reproducibility Check
Mapping decision criteria. Each sampled AAIGF-E control–framework pair was coded as a binary YES/NO decision. A mapping was coded YES when the AAIGF-E control’s stated objective and required governance or technical action were substantively addressed by an identifiable requirement, control objective, safeguard, or technique in the reference framework. Terminological similarity alone was insufficient; the referenced provision had to address the same underlying governance or control function. A mapping was coded NO when no such substantive correspondence could be identified or when similarity was limited to broad thematic relevance without an equivalent control function. Because the five reference sources differ in structure and normative status, correspondence was assessed on substantive control intent rather than identical terminology, numbering, or one-to-one structural equivalence. A positive mapping therefore indicates substantive alignment and does not imply equivalence, certification, or compliance with the referenced framework.
To assess the reproducibility of the standards-mapping judgments independently of the framework authors, a bounded coding exercise was conducted on a 35-control sample from the 111-control register. For each sampled control, an independent coder assigned a binary applicability judgment (YES/NO) against each of the five principal reference frameworks used in the quantitative analysis: NIST AI RMF, ISO/IEC 42001, NERC CIP, ISA/IEC 62443, and MITRE ATLAS. This produced 175 independent binary coding decisions (35 controls × 5 frameworks). The independent coder was provided with the coding file containing the sampled controls and framework-assessment fields and completed the YES/NO judgments independently. The completed file was returned to the authors, after which the coder’s judgments were compared with the internal mapping key for agreement analysis. The coder was not asked to reproduce the authors’ evidence annotations, confidence ratings, or qualitative rationale; the purpose of the exercise was specifically to test reproducibility of the binary mapping decision under the time-bounded review design.
The completed coding sheet was compared decision-by-decision with the internal coding key. Raw percentage agreement was calculated as the number of identical YES/NO decisions divided by the 175 total decisions. Cohen’s was additionally calculated across the pooled binary decisions to account for agreement expected by chance. Framework-specific raw agreement was also calculated to identify where mapping judgments were more or less reproducible. Disagreements were retained as observed rather than recoded to the internal key. Because only one independent coder completed the bounded exercise and no qualitative rationale was collected, the result is interpreted as a preliminary reproducibility check of the mapping process, not as independent validation of the AAIGF-E framework or proof of correctness of individual mappings.
3.5. Operational Parameterization of Selected Controls
AAIGF-E specifies governance objectives and required control capabilities rather than universal engineering set points. Operational parameterization is therefore deployment-specific: implementing organizations translate each control objective into measurable thresholds, decision rules, evidence requirements, and escalation actions appropriate to the AI use case, grid function, consequence severity, and validated operating envelope. This distinction prevents the framework from prescribing brittle values that would not transfer across utilities or operational contexts, while still requiring that the parameters be defined, approved, testable, and auditable before operational reliance.
For DI-006 (Runtime Data Validation at AI Inference Ingestion Points), parameterization occurs at the inference boundary. Implementers define acceptable schema and type constraints, value-range limits, timestamp validity and synchronization tolerances, completeness requirements, consistency checks, and rules for missing, stale, or contradictory telemetry. The applicable values depend on the source and operational function: PMU, SCADA, DER, and market inputs operate at different update rates and physical tolerances. The governance requirement is therefore not a universal freshness interval or numeric range; it is that these limits are established from the validated operating context, enforced before inference, tested during assurance, and linked to a defined reject, quarantine, fallback, or human-verification action when violated.
For OP-009 (AI-to-OT Command Mediation and Actuation Gatekeeping), parameterization defines the boundary between AI-generated recommendations and executable OT actions. Relevant parameters include the classes of commands the AI may request, assets and zones within scope, magnitude or consequence thresholds requiring human authorization, permitted operating states, interlock and engineering-limit checks, rate or sequence constraints, and conditions that force blocking or fail-safe execution. High-consequence actions such as protective-setting changes, switching, or DER curtailment above organization-defined thresholds can therefore require explicit approval even when lower-consequence actions are permitted under bounded authority. The control objective is the presence of a technically enforced mediation layer whose authorization rules are established during design, verified before deployment, and logged so that an AI-generated command cannot bypass the approved actuation boundary.
For OM-003 (Model Behavior Monitoring and Drift Detection), parameterization translates continuous monitoring into use-case-specific behavioral limits. Implementers define the monitored indicators, reference baselines, observation windows, drift or calibration thresholds, persistence criteria, alert severity, and the actions triggered when the validated envelope is exceeded. Thresholds should reflect physical consequence as well as statistical deviation: a stability-assessment model may require tighter limits and faster escalation than a lower-consequence forecasting application. The framework therefore requires explicit and testable monitoring thresholds without prescribing a single drift metric or universal value. Exceedance must be tied to predetermined actions such as enhanced human review, revalidation, restricted reliance, rollback, or temporary suspension of AI influence.
Together, these examples illustrate how AAIGF-E moves from control intent to operational enforcement without conflating governance requirements with equipment- or utility-specific engineering settings. Parameter values remain organization-defined, but their derivation, approval, testing, monitoring, evidence retention, and associated response actions are governed by the control architecture. This creates an auditable bridge between framework-level requirements and implementation while preserving portability across different electric-sector environments.
3.6. Scenario Validation Design
To evaluate the AAIGF-E against realistic failure conditions, three scenario types were selected based on operationally relevant smart-grid and smart-city failure modes and threat intelligence. The three scenarios were deliberately scoped to adversarial manipulation, unauthorized autonomy, and silent drift, the failure classes with the most immediate consequences for physical safety and grid operations. Supply chain, privacy, and foundational system-integrity domains were not included as core validation scenarios; the rationale for this scope boundary, and these domains’ expected non-activation, is addressed explicitly in Section 4.5. Scenario A reflects recognized adversarial risks involving manipulation or corruption of data used by AI/ML systems in operational environments, including poisoning and evasion techniques identified in the AI threat-modeling literature [1]. Scenario B examines the risk of consequential or unauthorized actions when an agentic AI system is permitted to influence operational processes without sufficiently enforced autonomy, authorization, and actuation boundaries. Scenario C examines the risk of silent model degradation when operational conditions shift beyond the distribution represented during model development and validation.
Three evaluation criteria were established prior to scenario execution: (1) structural coverage, assessing whether controls addressing the failure mode exist at the correct lifecycle phases; (2) sequencing coherence, assessing whether control activation follows a logical prevention–detection–response order and whether trigger conditions, control dependencies, decision gates, and escalation or fail-safe paths can be traced without critical dependencies being bypassed; and (3) gap surfacing, assessing whether the analysis reveals unresolved gaps with explicit rationale for why they cannot be addressed within current governance infrastructure. A framework would fail evaluation if scenarios revealed failure modes for which no controls existed, if critical trigger or decision points lacked an enforceable control response, if dependencies permitted consequential actions to bypass required assurance or authorization gates, or if identified gaps were dismissed without explicit deferral rationale.
Scenario A: Compromised DER Dispatch AI. An AI system managing distributed energy resource dispatch receives tampered input data, producing dispatch commands that destabilize local voltage profiles. This scenario stress-tests controls in the MI (Model Integrity), DI (Data Integrity), OM (Output and Decision Integrity), and IR (Incident Response) domains. The scenario assumes a tampered telemetry feed producing dispatch commands that push local voltage profiles outside nominal operating tolerance across several consecutive dispatch cycles before anomaly detection triggers.
Scenario B: Agentic AI with Unauthorized OT Action Execution. An agentic AI system authorized for anomaly investigation autonomously executes remediation actions across OT network segments without human approval, propagating configuration changes that trigger protective relay operations. This scenario stress-tests PG (governance authority boundaries), OP (actuation authority and operational control), and IR (containment and escalation) controls and directly evaluates whether the framework’s human oversight architecture is operationally meaningful or symbolic. The scenario assumes the agent holds read/write access to two adjacent OT network segments and executes an unapproved configuration change within a single operational cycle, without an intervening human approval step.
Scenario C: AI Model Drift in Grid Stability Forecasting. A machine learning model used for real-time stability assessment drifts silently due to distribution shift in load patterns following a large-scale DER integration event. Operators continue trusting outputs until a near-miss event triggers manual investigation. This scenario stress-tests TE (Testing and Evaluation), OM (drift detection), and RA (continuous risk monitoring) controls. The scenario assumes drift accumulates over several weeks following a large-scale DER integration event that shifts the model’s input distribution beyond its validated training envelope, with no automated drift-detection trigger firing during that interval.
3.7. Limitations
Five limitations bound the scope and generalizability of this work. First, the framework has not yet been evaluated through live utility deployment or controlled cyber-physical experimentation; accordingly, the quantitative mapping analysis, reproducibility assessment, scenario-based structural evaluation, and preliminary expert elicitation reported in this study should not be interpreted as empirical evidence of implementation-specific control effectiveness or operational risk reduction. Second, the control mapping prioritizes the North American regulatory context. Utilities operating under ENTSO-E, IEC network codes, or national grid operator frameworks will need mapping adjustments before direct application. Three elements are jurisdiction-specific: the NERC CIP reference mappings, the FERC-jurisdiction scoping assumptions, and the evidence artifacts tied to CIP audit cycles. Four elements are designed to be transferable subject to jurisdiction-specific validation and adaptation: the seven-phase lifecycle architecture, the eleven-domain-cluster structure, the dual-narrative control design, and the assurance-first philosophy. Third, agentic AI governance controls address autonomy boundaries and escalation logic as currently understood. Agentic AI architectures are not yet stable enough for these controls to be treated as comprehensive; the framework should therefore be treated as a versioned artifact requiring scheduled review cycles rather than a fixed compliance baseline. Fourth, the completed independent expert assessment provides preliminary content-validation evidence but is bounded by a small purposive panel (N = 5) and a single operational scenario; it does not support statistical generalization. Fifth, the independent coding exercise assessed mapping reproducibility for a sampled subset rather than independent reproduction of the full control-derivation process. Accordingly, the reported agreement and Cohen’s should be interpreted as evidence of mapping-process reproducibility, not as proof that an independent research team would derive the identical 111-control architecture by applying the same control-development methodology. Broader multi-coder replication and independent derivation studies would provide stronger evidence of methodological reproducibility. Independent testing in live utility environments and expanded expert assessment across additional scenarios remain future validation steps.
Figure 1 presents the framework’s structural architecture. Figure 2 complements this with a detailed Scenario B decision flow showing trigger conditions, control dependencies, decision gates, escalation conditions, and fail-safe paths for selected controls. The figure illustrates one operational application of the control architecture rather than a replacement for the seven AAIGF-E lifecycle phases.
Figure 1.
AAIGF-E architecture.
Figure 2.
Triggered control-assurance matrix for Scenario B: Agentic AI with Unauthorized OT Action Execution. The figure traces the v1.0 governance, privilege, boundary, actuation, override, and response controls used in the Scenario B analysis and explicitly identifies the residual semantic-manipulation pathway that is not closed by an existing v1.0 control. The decision sequence is a scenario-evaluation flow and is distinct from the seven AAIGF-E lifecycle phases (Govern, Design, Implement, Assure, Monitor, Respond, Recover). Bold: control IDs and decision outcomes; italic: control names.
4. Results: Scenario-Driven Control Validation
4.1. Validation Approach
Control validation was conducted by tracing each scenario against the AAIGF-E control set, identifying which controls activate at each lifecycle phase, assessing whether activated controls are sufficient to prevent or contain the failure, and surfacing sequencing gaps where control activation arrives too late or depends on a prerequisite control that was bypassed.
The scenario analysis is interpreted as a structural evaluation of control coverage, sequencing coherence, and gap surfacing under defined failure conditions rather than as evidence of implementation-specific operational effectiveness.
4.2. Scenario A: Adversarial Manipulation of DER Dispatch AI
Failure mode: A DER dispatch AI system receives manipulated sensor feeds at its inference ingestion point. Input validation controls are absent or misconfigured. The model produces dispatch commands that push voltage profiles outside stability tolerances before operators identify anomalous behavior.
Control activation sequence: Three Design-phase controls gate this failure before it reaches the model. DI-004 requires documented approval for every data source supplying inference data; no ungoverned sensor feed reaches the model in a compliant deployment. MI-003 requires the model to detect, reject, or safely degrade under tampered inputs; it spans Design through Monitor, making input validation an architectural property, not a post-deployment configuration. MI-006 bounds what the model can output under any condition, embedding voltage and frequency limits in the operational specification before go-live.
At the Assure phase, AAIGF-E-DI-006 requires schema conformance, value-range integrity, timestamp validity, completeness, and consistency checking at ingestion, serving as the last preventive gate before technically invalid or inconsistent data reach the model at runtime. AAIGF-E-MI-009 requires validation against scenarios outside typical training data, which in this context include sensor spoofing and feed injection cases. AAIGF-E-TE-007 requires adversarial input testing before production deployment.
At the Monitor phase, AAIGF-E-OM-002 monitors live ingestion for distribution shifts, timing gaps, and replay patterns. AAIGF-E-OM-007 detects when dispatch command outputs deviate from established operational baselines. At the Respond phase, AAIGF-E-IR-003 mandates rapid isolation of AI influence, and AAIGF-E-MI-010 enables restoration to a validated baseline dispatch state.
Coverage finding: Prevention coverage is strong if Design and Assure controls are fully implemented before operational deployment. The critical dependency is sequencing: MI-003 and DI-006 must be active before the system goes live. The identified gap is detection latency. OM-002 and OM-007 are reactive controls; if ingestion anomaly detection thresholds are not tuned to DER dispatch sensitivity, manipulated feeds may generate dispatch commands for multiple control cycles before deviation from output baselines becomes statistically detectable. This gap motivates a tighter coupling between DI-006 (ingestion validation) and OM-007 (output anomaly detection) as a compensating control pair.
The counterfactual is direct. Without DI-006 enforcing runtime ingestion validation and MI-003 requiring manipulation resilience as architectural properties, the compromised feed reaches the model unchallenged. OM-007 may eventually surface anomalous dispatch output, but only after multiple control cycles have executed against corrupted inputs. With AAIGF-E controls in place, DI-006 rejects or flags the anomalous telemetry at ingestion. The compromised data never reach the inference engine. The dispatch command is never generated. The failure is preventable at the data layer, not the response layer.
4.3. Scenario B: Agentic AI with Unauthorized OT Action Execution
Failure mode: An agentic AI system deployed for grid anomaly investigation autonomously executes remediation actions—including configuration changes across OT network segments—without human approval. The agent interprets its investigation mandate as authorization to act. Propagated configuration changes trigger protective relay operations, causing an unplanned islanding event. This scenario is the highest-consequence test case in this validation set, directly probing whether the AAIGF-E’s human oversight architecture is operationally enforceable or merely procedural.
For Scenario B, the structural evaluation additionally assesses whether authority, privilege, zone-boundary, actuation-mediation, operator-override, and coordinated-response controls prevent or contain unauthorized OT action. It also tests whether technically valid but semantically manipulative operational content can alter an agent’s interpretation of task scope or apparent authority before those downstream controls activate; where no v1.0 control closes that pathway, the condition is recorded explicitly as a residual gap rather than treated as covered.
Control activation sequence: At the Govern phase, AAIGF-E-PG-002 requires agentic systems with OT actuation capability to be approved under elevated risk tiers. AAIGF-E-PG-006 defines authority thresholds for high-consequence AI decisions, requiring human approval before execution when defined consequence thresholds are met. AAIGF-E-PG-001 ensures accountability for the agentic system is formally assigned before deployment or material changes to its remit.
At the Design and Implement phases, AAIGF-E-OP-001 maps AI-specific privileged actions to approved roles with enforced separation; agentic systems must have explicit privilege assignments rather than inherited or assumed authority. AAIGF-E-OP-007 constrains where AI services may operate across OT, DMZ, and IT zones. AAIGF-E-OP-009 then enforces mediation and authorization before any AI-influenced command reaches OT control functions, providing the architectural enforcement point for consequential actuation. AAIGF-E-OP-006 provides pre-engineered, auditable mechanisms for operators to pause or override AI influence.
AAIGF-E-DI-006 provides upstream runtime validation of ingested operational data for schema conformance, value-range integrity, timestamp validity, completeness, and consistency. These checks can reject or flag technically invalid inputs before inference, but they do not determine whether otherwise valid content is semantically manipulating the agent’s interpretation of task scope or apparent authority.
At the Respond phase, AAIGF-E-IR-006 establishes joint playbooks between cybersecurity, operations engineering, and control-room leadership when an agentic failure crosses OT network boundaries or affects physical grid configuration.
Coverage finding: AAIGF-E v1.0 provides layered governance and operational constraints for unauthorized OT action: PG-002, PG-006, and PG-001 establish risk tiering, authority, and accountability; OP-001 and OP-007 constrain privilege and zone boundaries; OP-009 provides the mandatory actuation-mediation boundary; OP-006 enables operator override; and IR-006 coordinates response. DI-006 contributes technical runtime ingestion validation upstream of this chain. However, Scenario B exposes a residual data-layer prompt-injection pathway that these controls do not close: operational content can satisfy DI-006’s schema, range, timestamp, completeness, and consistency checks while remaining semantically manipulative in a way that alters task scope, apparent authority, or a human approver’s judgment before OP-009 is reached. The scenario therefore identifies a genuine v1.0 control gap rather than treating DI-006 as a semantic content-trust control. Section 5.2 records the corresponding control direction as deferred future work.
4.4. Scenario C: Silent Model Drift in Grid Stability Forecasting
Failure mode: A machine learning model used for real-time stability assessment drifts silently following a large-scale DER integration event that shifts load patterns outside the model’s training distribution. Operators continue trusting model outputs. No drift detection alert is generated. A near-miss event triggers manual investigation, which reveals the model has been operating outside its validated performance envelope for weeks.
Control activation sequence: At the Design and Assure phases, AAIGF-E-TE-002 requires training data to adequately represent the full range of expected operational conditions, including rare events; a large-scale DER integration event should be explicitly represented or flagged as a known distribution boundary in the model’s evaluation evidence. AAIGF-E-TE-005 requires predefined performance thresholds and decision gates, including sensitivity to distribution shift conditions, not only point-in-time accuracy metrics. AAIGF-E-TE-008 requires testing under realistic edge cases with defined safety constraints. AAIGF-E-RA-001 requires documented risk assessment before deployment; the known distributional sensitivity of the stability model to DER penetration levels is a risk that must be documented and monitored.
At the Monitor phase, AAIGF-E-OM-003 is the primary detection control, requiring continuous monitoring for drift, degradation, and calibration changes with automated triggers. AAIGF-E-OM-007 establishes statistical baselines against which output deviations are measured. AAIGF-E-RA-003 requires ongoing monitoring of risk signals with defined response actions when thresholds are exceeded. AAIGF-E-TE-010 requires documentation of model limitations and distributional assumptions, enabling operators to understand when operational conditions have moved outside the model’s validated envelope.
This scenario illustrates an assurance gap that can arise even where an organization satisfies applicable cybersecurity and governance obligations. Compliance with such requirements does not by itself establish that an AI model remains within its validated performance envelope as operational conditions change. Without continuous drift monitoring, defined behavioral baselines, and revalidation triggers, degraded outputs may remain operationally plausible and therefore continue to influence human or automated decision-making without timely detection. Without OM-003 requiring continuous drift monitoring with defined triggers and TE-005 establishing predefined acceptance criteria, a degraded model may continue producing outputs that appear plausible, giving operators no mechanistic reason to question continued reliance. With these controls implemented, threshold exceedance can trigger review and revalidation before continued operational reliance. The scenario therefore demonstrates the governance function of connecting model-behavior monitoring to explicit assurance and revalidation decisions rather than treating initial compliance or validation as sufficient for continued use.
Coverage finding: The framework’s coverage is structurally sound, but this scenario exposes a threshold calibration problem. OM-003 and OM-007 are only effective if drift detection thresholds are set relative to grid stability sensitivity, not generic statistical drift metrics. A model serving stability forecasting in a high-DER environment requires tighter drift thresholds than a model serving demand forecasting because the physical consequences of degraded output are asymmetric. The AAIGF-E correctly requires drift detection but does not prescribe threshold methodology; this is an appropriate design choice for a governance framework but means implementing organizations must develop domain-specific threshold guidance as an operational complement to the control set.
4.5. Cross-Scenario Coverage Summary
Across all three scenarios, 47 of 111 AAIGF-E controls activated at one or more lifecycle phases. Table 4 summarizes domain coverage by scenario.
Table 4.
AAIGF-E domain coverage by validation scenario.
Three structural findings emerge from the cross-scenario analysis. First, the IR domain activates in every scenario at the Respond phase, confirming that incident response architecture is a universal requirement regardless of failure type. IR-003 (Containment and Safe Isolation) and IR-006 (Coordinated Cyber and Operations Response) are the highest-frequency controls across all three scenarios. Second, no single scenario activates the full control set; the framework is a lifecycle model, not a scenario checklist, and non-activated domains address failure modes outside the three scenario classes tested. Third, Scenario B shows that DI-006 technical runtime ingestion validation and OP-009 actuation gatekeeping provide important upstream and downstream safeguards but do not close the residual semantic-manipulation pathway identified in the scenario; Scenario C identifies threshold calibration as an implementation-support gap. The Scenario B finding requires future control formalization and empirical adversarial testing, while the Scenario C finding requires sector-specific implementation guidance.
4.6. Quantitative Control-Mapping Analysis
To examine the structural composition and standards integration of AAIGF-E, a quantitative analysis was conducted across the framework’s 111 controls. The analysis examined alignment with the five principal reference frameworks used in the control mapping, distribution across AAIGF-E lifecycle anchors and governance domains, cross-framework convergence, and electric-sector translation. The analysis characterizes the architecture and standards alignment of the framework; it does not establish control equivalence, certification, compliance with the referenced standards, or empirical effectiveness (Table 5).
Table 5.
Quantitative characteristics of the AAIGF-E control architecture.
4.6.1. Standards Coverage
Of the 111 AAIGF-E controls, 102 (91.9%) map to the NIST AI Risk Management Framework and 101 (91.0%) to ISO/IEC 42001. Seventy controls (63.1%) map to NERC Critical Infrastructure Protection requirements, 25 (22.5%) to ISA/IEC 62443, and 59 (53.2%) have active mappings to MITRE ATLAS. An additional 52 controls were assessed as not applicable to MITRE ATLAS and were therefore not counted as active ATLAS mappings.
The high coverage of NIST AI RMF and ISO/IEC 42001 reflects the framework’s grounding in general AI governance and management-system requirements, while NERC CIP and ISA/IEC 62443 introduce electric-sector and operational-technology considerations. The NIST AI RMF and ISO/IEC 42001 mappings intersect on 92 controls, while their union covers all 111 controls.
4.6.2. Lifecycle and Governance Distribution
The 111 controls contain 318 lifecycle activation points because individual controls may apply at more than one lifecycle stage. Assure has the largest number of control activations (78), followed by Implement (64), Design (52), Monitor (45), Respond (36), Govern (35), and Recover (eight). This distribution indicates that AAIGF-E is not concentrated solely on pre-deployment governance; controls extend across implementation, assurance, monitoring, response, and recovery activities.
Control distribution across the 11 active governance domains is similarly heterogeneous. System Integrity (SI) contains 18 controls, followed by Model Integrity (MI) with 16, Privacy (PR) with 12, Policy and Governance (PG) with 10, Testing and Evaluation (TE) with 10, Operational Control (OP) with 10, Supply Chain (SC) with nine, Output and Decision Integrity (OM) with eight, Incident Response (IR) with eight, Data Integrity (DI) with seven, and Risk Assessment (RA) with three. These differences reflect variation in the governance functions represented by each domain rather than an assumption that domains should contain equal numbers of controls.
4.6.3. Cross-Framework Convergence
Cross-framework convergence was assessed by counting how many of the five principal reference frameworks—NIST AI RMF, ISO/IEC 42001, NERC CIP, ISA/IEC 62443, and active MITRE ATLAS mappings—were associated with each AAIGF-E control. One control mapped to a single framework, 25 controls mapped to two, 44 to three, 31 to four, and 10 controls mapped across all five. The mean was 3.22 principal frameworks per control (Figure 3).
Figure 3.
Cross-framework convergence across AAIGF-E controls. Eighty-five of 111 controls (76.6%) map to at least three of the five principal reference frameworks.
Accordingly, 85 of the 111 controls (76.6%) mapped to at least three of the five principal reference frameworks. This concentration indicates substantial convergence among AI-governance, cybersecurity, electric-sector, operational-technology, and adversarial-AI concerns within the AAIGF-E control architecture. Convergence does not imply equivalence among source requirements. Rather, it indicates that individual AAIGF-E controls frequently bring together governance objectives addressed from different perspectives across multiple standards and frameworks.
4.6.4. Electric-Sector Translation
Seventy-three controls (65.8%) contain explicit mapping to NERC CIP and/or ISA/IEC 62443 and are classified in this analysis as sector-specific or operationally translated controls. The remaining 38 controls (34.2%) retain primarily generic AI-governance applicability.
This distribution provides quantitative evidence of AAIGF-E’s intended translation function: general AI-governance requirements are combined with electric-sector cybersecurity and operational-technology considerations within a common control architecture. The result is not a claim that AAIGF-E replaces or establishes compliance with NERC CIP or ISA/IEC 62443; rather, the mapping demonstrates where general AI-governance requirements intersect with sector-specific operational and cybersecurity concerns.
Taken together, the quantitative results show broad alignment with established AI-governance frameworks, lifecycle-spanning control coverage, substantial cross-framework convergence, and explicit electric-sector translation. These characteristics provide structural evidence that AAIGF-E operates as an integrated governance architecture rather than merely reproducing requirements from individual source frameworks. The analysis remains a structural and standards-based assessment; claims regarding practical effectiveness remain bounded by the scenario-based evaluation and preliminary independent expert content validation. Empirical evaluation of implementation-specific outcomes therefore remains a distinct validation stage requiring deployment-based or controlled experimental evidence.
4.7. Independent Coding Reproducibility Results
Across the 175 binary mapping decisions, the independent coder agreed with the internal coding key on 155 decisions and disagreed on 20, yielding 88.6% raw agreement. Cohen’s was 0.742, indicating that the observed agreement remained substantial after accounting for chance agreement. Agreement was not uniform across the five reference frameworks: NIST AI RMF achieved 35/35 agreement (100.0%), NERC CIP 32/35 (91.4%), MITRE ATLAS 31/35 (88.6%), ISO/IEC 42001 30/35 (85.7%), and ISA/IEC 62443 27/35 (77.1%). Table 6 summarizes the results.
Table 6.
Independent coder agreement with the internal standards-mapping key.
The result provides preliminary evidence that the binary standards-mapping procedure can be reproduced independently, while also showing that reproducibility varies by source framework. The comparatively lower agreement for ISA/IEC 62443 identifies the strongest area of mapping ambiguity in this sample and should therefore be treated conservatively in interpretation. Because the exercise collected only binary judgments, it does not establish why individual disagreements occurred or whether the internal or independent judgment was substantively preferable. The coding result is consequently reported as process-reproducibility evidence rather than as confirmation that the underlying mappings are objectively correct. This exercise evaluated the reproducibility of the framework-mapping procedure rather than independent reproduction of the original 111-control derivation process.
4.8. Preliminary Independent Expert Elicitation
To address the self-validation concern raised in peer review, AAIGF-E’s assessment of Validation Case Study 02 (AI-Assisted Equipment Shutdown Recommendation) was examined through a preliminary independent expert elicitation. A staged, independent-first evaluation instrument was distributed to practitioners with relevant domain expertise, in which each respondent first derived their own governance domain selection and recommendation from the operational scenario and evidence alone, before being shown AAIGF-E’s own assessment for comparison. The exercise is reported as preliminary expert elicitation and content-validation evidence rather than conclusive or statistically generalizable validation, consistent with the demonstration-phase scope of the broader Design Science Research evaluation described in Section 3.
4.8.1. Panel Composition
Six responses were received. One respondent selected no domain-relevant area of expertise (checking only “Other”) and was excluded from analysis under a pre-specified eligibility criterion requiring at least one of: AI Governance, Cybersecurity, ICS/OT, Electric Utilities, Risk Management, AI Assurance, or Standards Development. This yielded five eligible responses. Table 7 summarizes the panel composition.
Table 7.
Preliminary expert-elicitation panel composition (N = 5 eligible responses).
Respondents were permitted to select multiple areas of expertise; every eligible respondent selected at least two.
4.8.2. Domain Selection Convergence
AAIGF-E’s own assessment of Case Study 02 activated seven domains: Policy and Governance (PG), Risk Assessment (RA), Data Integrity (DI), Model Integrity (MI), Operational Control (OP), Incident Response (IR), and Output and Decision Integrity (OM). Table 8 reports how frequently each domain was independently selected by the eligible panel.
Table 8.
Independent domain selection frequency against the AAIGF-E reference assessment (N = 5).
The six domains common to both the reference assessment and majority independent selection (DI, MI, PG, RA, OP, IR) showed strong convergence, each selected by at least 80% of the panel. Two divergences are notable. First, OM (Output and Decision Integrity) was independently selected by only two of five respondents despite being part of the reference assessment; the evidence table’s “model drift status: Unknown” and “recent automated controller activity: Unknown” rows are the OM-relevant items, and their under-selection suggests some experts weighted these as Model Integrity considerations rather than Output and Decision Integrity considerations, a reasonable alternative categorization given the conceptual overlap between the two domains. Second, SC (Supply Chain) was independently selected by three of five respondents despite not appearing in the reference assessment, plausibly triggered by the evidence table’s “vendor assurance status: Current” row; while vendor assurance is evidence relevant to the scenario, the reference assessment treats it as a resolved, non-triggering condition rather than an active governance concern. Neither divergence indicates disagreement with the underlying evidence; both reflect reasonable differences in domain categorization at the margins of AAIGF-E’s domain boundaries.
4.8.3. Recommendation Convergence
Of the five eligible expert responses, four selected VERIFY FIRST, consistent with the AAIGF-E reference assessment. One response selected RELY/PROCEED; however, its accompanying written rationale explicitly recommended VERIFY FIRST and independently identified the same evidence gaps as the reference assessment, including unverified sensor timestamp synchronization, unknown model drift status, and an unknown status for recent automated controller activity. This discrepancy most plausibly reflects a response-selection error rather than a substantive disagreement, as the written rationale directly contradicts the selected option. Conservatively, recommendation convergence based on the selected option was 4/5 (80%). Interpreting the response according to its written rationale would yield 100% recommendation convergence (5/5). Accordingly, both the conservative (selected response) and rationale-based interpretations are reported to maintain transparency and preserve the integrity of the original dataset.
The four unambiguous VERIFY FIRST rationales independently converged on the same underlying reasoning present in the reference assessment. Unresolved conditions, including model drift status, sensor timestamp synchronization, and recent controller activity, are individually insufficient to justify withholding action but collectively warrant verification before approval. This recommendation is further supported by the fact that independent protection systems remained active and the mandatory protection-trip threshold was not reached.
4.8.4. Framework Quality Assessment
Table 9 summarizes the panel’s evaluation of AAIGF-E on seven framework attributes, rated on a five-point scale (1 = very poor to 5 = excellent).
Table 9.
Framework quality ratings from the preliminary expert-elicitation panel (N = 5, mean of 1–5 scale).
All seven attributes received mean ratings at or above 4.0 (Good), with no attribute falling below the midpoint of the scale. Overall completeness and value to critical infrastructure received the lowest mean ratings (4.0), reflecting the same domain-boundary observations discussed above rather than a distinct quality concern.
4.8.5. Limitations of This Validation
This validation is preliminary rather than conclusive, and three limitations should inform its interpretation. First, the eligible panel size (N = 5) is small; it does not support statistical generalization, and the findings should be read as convergence evidence from a purposively selected panel rather than a representative sample. Second, the panel was purposively selected based on direct professional relationships and domain expertise rather than randomly sampled, which is standard practice for expert elicitation studies of this kind but bounds the claims that can be made about the broader population of AI governance and OT security practitioners. Third, validation was conducted against a single scenario (Case Study 02); the domain-selection convergence and divergence patterns reported here are specific to this scenario’s evidence structure and may not generalize to scenarios activating different domain combinations. Expansion to a larger panel and additional validated scenarios is identified as future work in Section 6.
5. Discussion
5.1. What the Scenarios Reveal About the Framework
The three scenarios confirm a tiered architecture: prevention is front-loaded into Design and Assure phases, detection depends on threshold calibration decisions the framework correctly delegates rather than prescribes, and recovery is minimal by design. Organizations that compress or skip the early phases do not inherit a weaker version of the framework. They inherit a detection-only posture, which is structurally inadequate when the AI system is coupled to physical grid infrastructure.
Scenario A makes this explicit. MI-003 and DI-006 are architectural prevention controls. If they are not implemented before go-live, OM-002 and OM-007 become the first line of defense: reactive controls carrying a load they were not designed to bear alone. The phase distribution confirmed by scenario analysis is not incidental. AAIGF-E is intentionally assurance-heavy: 78 of 111 controls activate at the Assure phase because pre-operational gating is the primary risk control in environments where post-deployment correction is operationally costly and physically consequential. Based on the comparative mapping presented in this study, NIST AI RMF 1.0, ISO/IEC 42001, and ISA/IEC 62443 provide strong governance or security foundations but contain fewer explicit pre-operational AI assurance requirements. NIST AI RMF’s closest analog is its Manage function, which operates at the organizational process level rather than specifying technical pre-operational gates, a limitation also evident in the function’s companion implementation guidance [18,44]. ISO/IEC 42001 similarly defines AI management-system obligations without mandating the specific assurance activities that must occur before deployment [40]. ISA/IEC 62443 provides lifecycle cybersecurity requirements for industrial automation and control systems but does not directly address model assurance or AI behavioral drift [15]. AAIGF-E inverts that emphasis deliberately: assurance gates are the load-bearing structure, and governance documentation supports them rather than substituting for them.
This distribution also differs structurally from existing governance frameworks. Unlike AI governance frameworks that primarily define organizational principles or management-system requirements, AAIGF-E demonstrates control-level convergence across multiple governance and sector-specific reference frameworks while maintaining a sector-oriented operational architecture. These findings indicate that the framework translates governance objectives into operational controls intended for direct implementation rather than organizational guidance alone.
The distinction is one of design purpose rather than superiority: the mapped instruments retain their respective roles in organizational AI governance, management systems, electric-sector cybersecurity, industrial-control security, and adversarial threat characterization, while AAIGF-E provides the lifecycle and operational control layer needed to coordinate those objectives for AI systems that can influence physical grid behavior.
5.2. The Agentic AI Finding Is the Most Consequential
Scenario B surfaces a consequential data-layer threat: prompt injection through operational feeds in OT-connected agentic systems. An agentic AI consuming live SCADA data, historian feeds, or event logs can be manipulated through those feeds before an action request reaches OP-009’s actuation gate. Existing v1.0 controls provide important but incomplete protection. PG-002, PG-006, and PG-001 establish authority and accountability; OP-001 and OP-007 constrain privilege and zone boundaries; OP-009 mediates consequential OT action; OP-006 provides operator override; IR-006 coordinates response; and DI-006 provides technical runtime validation of ingested operational data.
The residual failure mode is semantic rather than syntactic or numerical. Content can pass DI-006’s schema-conformance, value-range, timestamp, completeness, and consistency checks while still manipulating an agent’s interpretation of task scope, apparent authority, or a human approver’s judgment. OP-009 can block an unauthorized consequential command at the actuation boundary, but it does not itself validate the semantic trustworthiness of the operational content that shaped the preceding reasoning or approval context. The human in the loop can therefore also become part of the attack surface if compromised context influences the approval decision.
This finding is consistent with, but more operationally specific than, recent work on governing autonomous AI in critical infrastructure. Sharma and Pursiainen argue that resilience in critical-infrastructure AI depends on bounded autonomy within a hybrid governance architecture, with oversight modes calibrated to task complexity, risk, and consequence severity [26]. AAIGF-E’s existing v1.0 controls operationalize bounded authority and actuation constraints, but Scenario B demonstrates that bounded actuation alone is insufficient when semantically manipulative content can influence the reasoning context upstream of the command gate.
The corresponding control direction is therefore treated as deferred future work rather than retrofitted into DI-006. A future extension of OP-009 should require agentic systems consuming live OT data feeds to apply source attestation and content-plausibility validation before feed content is permitted to update task scope or authorization context. Formalization is deferred pending further empirical adversarial testing and development of OT-grade content-trust methods for agentic systems, including alignment with relevant MITRE ATLAS and ISA/IEC 62443 work. Until such a requirement is formalized, organizations may use compensating measures such as trusted-source restrictions, network-path validation, certificate checks, correlation across independent sensors, and mandatory human verification, but these measures should not be represented as an existing AAIGF-E v1.0 control.
5.3. Drift Threshold Calibration Is an Implementation Gap, Not a Framework Gap
Scenario C’s finding—that OM-003 and OM-007 are only effective if thresholds are calibrated to grid stability sensitivity rather than generic statistical drift—should not be read as a framework deficiency. Governance frameworks that prescribe threshold values for specific operational contexts become brittle and jurisdiction-specific. The AAIGF-E correctly requires drift detection and continuous risk monitoring without mandating any methodology. What the scenario reveals is an implementation support gap: electric utilities implementing OM-003 need domain-specific guidance on drift threshold methodology for stability-critical AI, connecting statistical drift metrics to physical consequence severity. NERC, NIST, or an industry working group is better positioned to produce that guidance than a governance framework. The framework’s role is to require the capability and create the accountability structure around it.
5.4. Compliance Coverage Does Not Equal Operational Assurance
Seventy of 111 AAIGF-E controls reference NERC CIP standards; this coverage is intentional because electric-sector operators already operate within NERC CIP compliance cycles. However, satisfying NERC CIP compliance requirements does not by itself close the operational assurance gap identified in this paper. NERC CIP was developed for cybersecurity governance of Bulk-Electric-System assets and does not explicitly address model drift, adaptive AI behavior, or agentic action telemetry [13,45]. Its audit artifacts map reasonably well to AAIGF-E’s Govern and Implement phase controls. They map poorly to the Monitor and Respond phase controls that govern adaptive AI behavior. Drift detection thresholds, model behavioral baselines, and agentic action logs do not have natural homes in current CIP audit frameworks. Implementing organizations will need to extend their evidence management practices to cover AI-specific telemetry, and regulators will eventually need to update audit expectations accordingly.
5.5. The Framework’s Practical Constraint
The AAIGF-E was constructed without validation against live operational deployments. Scenario-driven validation tests structural coherence—whether the right controls exist in the right phases for operationally relevant failure modes—but does not test implementation friction, control cost, or whether utilities with legacy OT environments can realistically implement Design-phase controls before deploying AI systems that were procured as vendor-integrated packages.
A significant portion of AI deployed in smart-city energy environments arrives embedded in vendor systems, not built in-house, bypassing the Design phase almost entirely. SC domain controls (SC-001 through SC-009) address supply chain governance, but their effectiveness depends on procurement authority and contractual leverage that many municipal and utility procurement environments do not currently exercise. This is where the framework’s next validation phase must go: testing control applicability against vendor-integrated AI deployment pathways, where the implementing organization is a consumer rather than a builder.
In practice, adoption of AAIGF-E does not require a utility to replace its existing governance or cybersecurity programs. An implementing organization would first determine the applicable AI use case, operational assets, autonomy level, and regulatory context; activate the relevant AAIGF-E controls across the applicable lifecycle phases; translate organization-defined control requirements into measurable technical or procedural parameters; assign control ownership and required evidence; and integrate resulting assurance, monitoring, escalation, and incident-response requirements into existing operational and compliance processes. Where the AI capability is vendor-integrated or technically opaque, the same control objectives are applied through procurement requirements, contractual evidence obligations, interface-level validation, runtime monitoring, and bounded actuation authority rather than relying on access to model internals.
The framework’s smart-city applicability extends beyond the grid but requires explicit mapping. AAIGF-E was validated against grid-specific failure modes: DER dispatch manipulation, agentic OT action execution, and stability forecasting drift. Smart-city AI systems operating outside direct grid coupling, including traffic management platforms, building automation controllers, and public safety anomaly detection systems, share the same structural governance problem [3,4]. They operate within physical tolerances. They influence real-time resource allocation. They may be deployed without pre-operational assurance gates. The seven-phase lifecycle and eleven domain cluster architecture of AAIGF-E may provide a transferable governance structure for these systems, subject to domain-specific validation and mapping: an AI system managing district traffic signal coordination requires the same Design-phase input integrity controls (MI-003, DI-006), the same Assure-phase robustness testing gates (TE-007, MI-009), and the same Monitor-phase drift detection obligations (OM-003) as a DER dispatch AI. What changes is the OT narrative layer, not the control architecture. Smart-city operators outside the electric utility sector should treat the grid-specific OT narratives in AAIGF-E as implementation templates, not constraints [4]. The compliance alignment columns require jurisdictional remapping for non-utility contexts, but the assurance-first philosophy and lifecycle gating structure may be transferable, subject to domain-specific validation and adaptation.
5.6. Governance and Policy Implications in the AI Era: Institutional Settings and Action Plans
AI governance in smart-city infrastructure is not a technical problem with a policy wrapper. It is an institutional problem with a technical dimension [46,47,48]. The question is not whether AI systems are accurate or efficient. It is whether the institutions responsible for urban infrastructure, municipal governments, utilities, regulators, vendors, and control room operators have the capacity to authorize, supervise, audit, and stop AI systems that affect physical urban operations [49,50]. Existing smart-city governance arrangements may also inadequately address citizen accountability, equity, and legitimacy [51]. AAIGF-E is intended to translate these broader governance concerns into operational control requirements for grid-connected infrastructure.
For implementation purposes, AAIGF-E proposes five complementary institutional roles rather than claiming a universal allocation of legal responsibility. Municipal governments can provide public-accountability oversight for public value, citizen safety, equity, privacy, and legitimacy; utilities and infrastructure operators can translate governance objectives into engineering constraints, safety thresholds, fallback mechanisms, and incident-response procedures; regulators and standards bodies can define or recognize evidence expectations for AI-enabled critical infrastructure; vendors can be required through procurement and contractual mechanisms to disclose relevant AI functions and provide assurance evidence where the operating organization lacks direct access to model internals; and citizens remain legitimacy stakeholders where smart-city AI affects mobility, energy access, public safety, service allocation, or privacy [47,51]. These role allocations are design recommendations of the framework and must be adapted to the applicable legal, regulatory, contractual, and organizational context.
The governance architecture described above can be translated into a sequenced implementation pathway. Table 10, Table 11 and Table 12 are illustrative, non-prescriptive implementation proposals derived from AAIGF-E; their timeframes, institutional assignments, and adoption levels have not been empirically validated and should be adapted to organizational and jurisdictional conditions.
Table 10.
Illustrative phased policy action plan for AI-era smart governance.
Table 11.
Illustrative minimum action plan by institutional actor.
Table 12.
Illustrative implementation pathway for AAIGF-E adoption.
The shift from principle-based trustworthiness to institutionalized operational assurance is the work that remains. Principles such as transparency, accountability, fairness, safety, and explainability are necessary. They are not sufficient. In cyber-physical urban infrastructure, a governance failure can propagate beyond the governance layer into service, safety, or legitimacy consequences. AAIGF-E provides the control architecture for that shift. AI systems influencing physical urban infrastructure must be governed through lifecycle accountability, pre-operational assurance, bounded autonomy, continuous monitoring, incident readiness, and recovery-based learning. Within AAIGF-E, these capabilities constitute the proposed minimum institutional capacity for governing AI deployed in critical operations.
6. Conclusions
The governance problem is not unsolved: it is under-operationalized. Existing instruments define what AI trustworthiness should look like but stop short of specifying the control architecture that enforces it before a grid-influencing AI system goes live. AAIGF-E fills that gap with 111 controls across seven lifecycle phases, structured through a dual-narrative architecture that makes the same control set legible to AI governance stakeholders and OT engineering teams without requiring translation between them. The design principle is single and non-negotiable: prevention and assurance carry the load, because recovery from AI-induced grid failure is not a routine operational procedure. AI systems influencing grid stability, DER dispatch, or real-time reliability decisions should not be permitted to operate under governance models that lack pre-operational enforcement. Compliance with current instruments does not satisfy that requirement [45].
Three additional evaluation components strengthen the framework assessment. Quantitative characterization of the frozen 111-control register demonstrates standards coverage, lifecycle and governance-domain distribution, cross-framework convergence, and sector-specific translation. The broader mapping analysis shows substantial convergence across the five principal reference frameworks and standards, with 85 of 111 controls mapping to at least three. An independent reproducibility check of 175 binary mapping decisions produced 88.6% raw agreement with the internal key and Cohen’s , providing preliminary evidence that the sampled mapping procedure can be reproduced while identifying ISA/IEC 62443 as the least consistent mapping boundary in the sample. Independent experts in AI governance, cybersecurity, and ICS/OT also evaluated a representative operational scenario after first deriving their governance judgments without viewing the AAIGF-E reference assessment. Their domain selections and recommendation rationales showed strong overall convergence with the framework assessment. Given the small purposive panel and single-scenario design, these findings constitute preliminary independent content validation rather than conclusive empirical validation.
Scenario-driven validation confirms the framework’s structural coherence. Prevention controls in the Design and Assure phases address adversarial input manipulation, unauthorized autonomy, and silent model drift before operational deployment. Detection and response controls in the Monitor and Respond phases provide containment and recovery architecture when prevention fails. The control sequencing is intentional: organizations that compress or skip early lifecycle phases inherit a detection-only posture that is insufficient for grid-coupled AI systems operating within physical stability tolerances.
Three findings from validation deserve direct attention from practitioners and standards bodies. First, agentic AI in OT environments requires actuation gatekeeping at the architectural level, not the policy level. OP-009 must be a hard enforcement boundary, not a documented requirement. Scenario B demonstrates that prompt-injection risk requires layered data and actuation safeguards but also exposes a residual semantic-manipulation gap. DI-006 provides technical runtime ingestion validation, and OP-009 provides the downstream actuation gate; neither v1.0 control should be interpreted as a semantic content-trust control for task-scope or authority manipulation. Second, drift detection controls are only as effective as their threshold calibration. OM-003 and OM-007 create the accountability structure; implementing organizations and sector regulators must develop the domain-specific threshold methodology that makes those controls operational in high-DER grid environments. Third, NERC CIP compliance cycles provide a deployment pathway for Govern and Implement phase controls, but they do not yet accommodate the AI-specific telemetry artifacts that Monitor and Respond phase controls generate. Closing that audit gap requires regulatory engagement, not just framework adoption.
The vendor integration pathway remains the framework’s most significant unvalidated constraint. Vendor-integrated AI represents an important deployment pathway in smart-city energy environments and may arrive embedded in procured systems, limiting the operating organization’s direct influence over Design-phase activities. Supply chain controls SC-001 through SC-009 address this structurally, but their effectiveness depends on procurement authority and contractual leverage that many utilities and municipal operators do not currently exercise.
This research is scoped to the North American electric sector. The compliance alignment logic, evidence architecture, and regulatory references reflect NERC CIP, NIST AI RMF 1.0, and FERC-jurisdiction assumptions that do not transfer directly to other regions without gap analysis. In future work, we will extend AAIGF-E validation to other regulatory contexts, including ENTSO-E network codes, IEC-governed grid environments, and national AI governance frameworks in regions with high DER penetration and active smart-city programs. The core questions are applicability and suitability: which control domains are transferable without substantive change, which require jurisdictional remapping or adaptation, and whether the assurance-first design philosophy holds under different regulatory incentive structures. That work requires live operational partners in those regions, not just framework comparison.
Future work should pursue four directions. First, live operational validation against deployed AI systems in utility environments would test implementation friction and control cost in ways scenario analysis cannot. Second, empirical adversarial testing and sector-specific implementation guidance for agentic AI controls—particularly the residual prompt-injection pathway identified in Scenario B, multi-agent trust hierarchies, and tool-use authorization boundaries in OT contexts—are warranted given the pace of agentic deployment in infrastructure-adjacent systems. Third, a compliance bridge document connecting AAIGF-E Monitor and Respond phase controls to NERC CIP audit evidence frameworks would accelerate practitioner adoption without requiring regulatory change as a precondition. Fourth, future work will examine the economic and operational implications of implementing AAIGF-E controls in smart-grid deployments; this may include applying option-valuation and stochastic-planning approaches to assess how governance requirements affect investment flexibility under endogenous and exogenous uncertainty [52]. Further research may also extend the framework to reinforcement-learning-based controllers in isolated renewable microgrids and hydrogen-storage systems, where autonomous control must remain bounded by physical and operational constraints [53]. Related planning models for storage technologies under decision-dependent innovation uncertainty may provide a basis for evaluating how governed AI and storage investments alter long-term system risk and flexibility [54]. The control architecture exists. The implementation experience does not. Until AAIGF-E is tested against live utility deployments and regulatory bodies formalize AI-specific audit expectations, the problem this paper opened with remains: satisfying today’s obligations does not guarantee tomorrow’s safe operation. That is the next problem.
Author Contributions
Conceptualization, S.A.R.; methodology, S.A.R.; investigation, S.A.R. and D.K.; validation, S.A.R., D.K. and S.M.; writing—original draft preparation, S.A.R. and D.K.; writing—review and editing, S.A.R., D.K. and S.M.; visualization, S.A.R. and D.K.; supervision, S.M. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Participants were informed of the purpose of the expert elicitation, the intended academic use of their responses, and the confidentiality of individual responses. Participation was voluntary, and completion of the evaluation instrument constituted consent to participate.
Data Availability Statement
The complete AAIGF-E control register, including all 111 controls with Control ID, Name, Lifecycle Phase, Domain, Control Objective, AI Narrative, Electric-Sector Narrative, and Reference Standards mappings, is documented in Rana and Bodungen [11].
Acknowledgments
The authors used ChatGPT (GPT-5.6 Sol, OpenAI, San Francisco, CA, USA) primarily for language, grammar, and editorial review, and Claude (Opus 5.5, Anthropic, San Francisco, CA, USA) solely as an additional review tool. The authors reviewed and approved all resulting changes and remain fully responsible for the content of the manuscript.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| AAIGF-E | AI and Agentic Intelligence Governance Framework, Electric Sector |
| AI | Artificial Intelligence |
| CIP | Critical Infrastructure Protection |
| CPS | Cyber-Physical System |
| DER | Distributed Energy Resource |
| DSR | Design Science Research |
| EU | European Union |
| ICS | Industrial Control System |
| ML | Machine Learning |
| NERC | North American Electric Reliability Corporation |
| NIST | National Institute of Standards and Technology |
| OT | Operational Technology |
| PMU | Phasor Measurement Unit |
| SCADA | Supervisory Control and Data Acquisition |
References
- North American Electric Reliability Corporation. Artificial Intelligence and Machine Learning in Real-Time System Operations: White Paper Revision 1; NERC: Washington, DC, USA, 2024; Available online: https://www.nerc.com/globalassets/our-work/reports/white-papers/whitepaper-ai-and-ml-in-real-time-system-operations.pdf (accessed on 1 March 2025).
- U.S. Department of Energy, Office of Critical and Emerging Technologies. AI for Energy: Opportunities for a Modern Grid and Clean Energy Economy; DOE: Washington, DC, USA, 2024. Available online: https://www.energy.gov/sites/default/files/2024-04/AI%20EO%20Report%20Section%205.2g(i)_043024.pdf (accessed on 27 July 2026).
- Cugurullo, F.; Caprotti, F.; Cook, M.; Karvonen, A.; McGuirk, P.; Marvin, S. The rise of AI urbanism in post-smart cities: A critical commentary on urban artificial intelligence. Urban Stud. 2024, 61, 1168–1182. [Google Scholar] [CrossRef] [Scilit]
- Wolniak, R.; Stecuła, K. Artificial intelligence in smart cities: Applications, barriers, and future directions: A Review. Smart Cities 2024, 7, 1346–1389. [Google Scholar] [CrossRef] [Scilit]
- Lee, E.A. Cyber physical systems: Design challenges. In Proceedings of the 11th IEEE International Symposium on Object Oriented Real-Time Distributed Computing (ISORC), Orlando, FL, USA, 5–7 May 2008; pp. 363–369. [Google Scholar]
- Amin, S.M. Smart grid: Overview, issues and opportunities. Advances and Challenges in Sensing, Modeling, Simulation, Optimization and Control. Eur. J. Control 2011, 17, 547–567. [Google Scholar] [CrossRef] [Scilit]
- Cardenas, A.A.; Amin, S.; Sinopoli, B.; Giani, A.; Perrig, A.; Sastry, S. Challenges for securing cyber physical systems. In Proceedings of the Workshop on Future Directions in Cyber-Physical Systems Security, Newark, NJ, USA, 22–24 July 2009. [Google Scholar]
- Gungor, V.C.; Sahin, D.; Kocak, T.; Ergut, S.; Buccella, C.; Cecati, C.; Hancke, G.P. Smart grid technologies: Communication technologies and standards. IEEE Trans. Ind. Inform. 2011, 7, 529–539. [Google Scholar] [CrossRef] [Scilit]
- Fang, X.; Misra, S.; Xue, G.; Yang, D. Smart grid: The new and improved power grid: A survey. IEEE Commun. Surv. Tutor. 2012, 14, 944–980. [Google Scholar] [CrossRef] [Scilit]
- Mohsenian-Rad, A.H.; Leon-Garcia, A. Optimal residential load control with price prediction in real-time electricity pricing environments. IEEE Trans. Smart Grid 2010, 1, 120–133. [Google Scholar] [CrossRef] [Scilit]
- Rana, S.A.; Bodungen, C. Adaptive AI Governance Framework for the Electric Sector. SSRN 2026. Available online: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6584459 (accessed on 18 May 2026).
- NISTIR 7628 Rev. 1; Guidelines for Smart Grid Cybersecurity. National Institute of Standards and Technology (NIST): Gaithersburg, MD, USA, 2014.
- North American Electric Reliability Corporation. CIP Standards. Available online: https://www.nerc.com/pa/Stand/Pages/CIPStandards.aspx (accessed on 1 January 2025).
- North American Electric Reliability Corporation. Reliability Standards. Available online: https://www.nerc.com/pa/Stand/Pages/default.aspx (accessed on 1 January 2025).
- ISA/IEC 62443; Series of Standards. International Society of Automation (ISA): Research Triangle Park, NC, USA, 2018.
- ISO/IEC 27001:2022; Information Security, Cybersecurity and Privacy Protection: Information Security Management Systems—Requirements. International Organization for Standardization (ISO): Geneva, Switzerland, 2022.
- NIST SP 800-53 Rev. 5; Security and Privacy Controls for Information Systems and Organizations. National Institute of Standards and Technology (NIST): Gaithersburg, MD, USA, 2020.
- NIST AI 100-1; AI Risk Management Framework: AI RMF 1.0. National Institute of Standards and Technology (NIST): Gaithersburg, MD, USA, 2023.
- Bovens, M. Analysing and assessing accountability: A conceptual framework. Eur. Law J. 2007, 13, 447–468. [Google Scholar] [CrossRef] [Scilit]
- Floridi, L.; Cowls, J.; Beltrametti, M.; Chatila, R.; Chazerand, P.; Dignum, V.; Luetge, C.; Madelin, R.; Pagallo, U.; Rossi, F.; et al. AI4People—An ethical framework for a good AI society: Opportunities, Risks, Principles, and Recommendations. Minds Mach. 2018, 28, 689–707. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Parasuraman, R.; Riley, V. Humans and automation: Use, misuse, disuse, abuse. Hum. Factors 1997, 39, 230–253. [Google Scholar] [CrossRef] [Scilit]
- Sheridan, T.B. Telerobotics, Automation, and Human Supervisory Control; MIT Press: Cambridge, MA, USA, 1992. [Google Scholar]
- Endsley, M.R. Toward a theory of situation awareness in dynamic systems. Hum. Factors 1995, 37, 32–64. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.D.; See, K.A. Trust in automation: Designing for appropriate reliance. Hum. Factors 2004, 46, 50–80. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- European Parliament. Regulation (EU) 2024/1689 of the European Parliament and of the Council on Artificial Intelligence (AI Act); Official Journal of the European Union: Brussels, Belgium, 2024.
- Sharma, P.; Pursiainen, C.H. Resilience meets autonomy: Governing embodied AI in critical infrastructure. arXiv 2026, arXiv:2603.15885. [Google Scholar] [CrossRef] [Scilit]
- Rahimi, F.; Ipakchi, A. Demand response as a market resource under the smart grid paradigm. IEEE Trans. Smart Grid 2010, 1, 82–88. [Google Scholar] [CrossRef] [Scilit]
- Pipattanasomporn, M.; Feroze, H.; Rahman, S. Multi-agent systems in a distributed smart grid: Design and implementation. In Proceedings of the IEEE PES Power Systems Conference and Exposition, Seattle, WA, USA, 15–18 March 2009. [Google Scholar]
- Glavic, M.; Fonteneau, R.; Ernst, D. Reinforcement learning for electric power system decision and control: Past Considerations and Perspectives. IFAC-PapersOnLine 2017, 50, 6918–6927. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Zhang, D.; Qiu, R.C. Deep reinforcement learning for power system applications: An overview. CSEE J. Power Energy Syst. 2020, 6, 213–225. [Google Scholar] [CrossRef] [Scilit]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. ReAct: Synergizing reasoning and acting in language models. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar]
- Xi, Z.; Chen, W.; Guo, X.; He, W.; Ding, Y.; Hong, B.; Zhang, M.; Wang, J.; Jin, S.; Zhou, E.; et al. The rise and potential of large language model based agents: A survey. arXiv 2023, arXiv:2309.07864. [Google Scholar]
- Kiasari, M.; Aly, H. Agentic Artificial Intelligence for Smart Grids: A Comprehensive Review of Autonomous, Safe, and Explainable Control Frameworks. Energies 2026, 19, 617. [Google Scholar] [CrossRef] [Scilit]
- Huang, L.; Joseph, A.D.; Nelson, B.; Rubinstein, B.I.P.; Tygar, J.D. Adversarial machine learning. In Proceedings of the 4th ACM Workshop on Security and Artificial Intelligence, Chicago, IL, USA, 21 October 2011; pp. 43–58. [Google Scholar]
- Biggio, B.; Roli, F. Wild patterns: Ten years after the rise of adversarial machine learning. Pattern Recognit. 2018, 84, 317–331. [Google Scholar] [CrossRef] [Scilit]
- MITRE. ATLAS: Adversarial Threat Landscape for Artificial-Intelligence Systems. Available online: https://atlas.mitre.org (accessed on 1 January 2025).
- Gulyamov, S.; Gulyamov, S.; Rodionov, A.; Khursanov, R.; Mekhmonov, K.; Babaev, D.; Rakhimjonov, A. Prompt Injection Attacks in Large Language Models and AI Agent Systems: A Comprehensive Review of Vulnerabilities, Attack Vectors, and Defense Mechanisms. Information 2026, 17, 54. [Google Scholar] [CrossRef] [Scilit]
- Tao, F.; Zhang, H.; Liu, A.; Nee, A.Y.C. Digital twin in industry: State-of-the-art. IEEE Trans. Ind. Inform. 2019, 15, 2405–2415. [Google Scholar] [CrossRef] [Scilit]
- Djebali, S.; Guerard, G.; Taleb, I. Survey and insights on digital twins design and smart grid’s applications. Future Gener. Comput. Syst. 2024, 153, 234–248. [Google Scholar] [CrossRef] [Scilit]
- ISO/IEC 42001:2023; Information Technology: Artificial Intelligence: Management System. International Organization for Standardization (ISO): Geneva, Switzerland, 2023.
- Hevner, A.R.; March, S.T.; Park, J.; Ram, S. Design science in information systems research. MIS Q. 2004, 28, 75–105. [Google Scholar] [CrossRef] [Scilit]
- Peffers, K.; Tuunanen, T.; Rothenberger, M.A.; Chatterjee, S. A design science research methodology for information systems research. J. Manag. Inf. Syst. 2007, 24, 45–77. [Google Scholar] [CrossRef] [Scilit]
- vom Brocke, J.; Hevner, A.; Maedche, A. Introduction to design science research. In Design Science Research. Cases; Springer: Cham, Switzerland, 2020; pp. 1–13. [Google Scholar]
- National Institute of Standards and Technology. NIST AI RMF Playbook; NIST: Gaithersburg, MD, USA, 2023. Available online: https://airc.nist.gov/airmf-resources/playbook/ (accessed on 21 September 2026).
- Rana, S.A. Operationalizing AI Governance in Bulk Electric Systems: A Control-Level Gap Analysis of NERC CIP Using AAIGF-E. SSRN 2026. Available online: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6711358 (accessed on 19 May 2026).
- Kitchin, R. The real-time city? Big data and smart urbanism. GeoJournal 2014, 79, 1–14. [Google Scholar] [CrossRef] [Scilit]
- Jasanoff, S. (Ed.) States of Knowledge: The Co-Production of Science and the Social Order; Routledge: London, UK, 2004. [Google Scholar]
- Kitchin, R.; Dodge, M. Code/Space: Software and Everyday Life; MIT Press: Cambridge, MA, USA, 2011. [Google Scholar]
- OECD. Recommendation of the Council on Artificial Intelligence; OECD/LEGAL/0449; OECD: Paris, France, 2019. [Google Scholar]
- ISO/IEC 23894:2023; Information Technology, Artificial Intelligence, Guidance on Risk Management. International Organization for Standardization (ISO): Geneva, Switzerland, 2023.
- Cardullo, P.; Kitchin, R. Smart urbanism and smart citizenship: The neoliberal logic of citizen-focused smart cities in Europe. Environ. Plan. C Polit. Space 2019, 37, 813–830. [Google Scholar] [CrossRef] [Scilit]
- Giannelos, S. Option Valuation of Smart Grid Technology Projects Under Endogenous and Exogenous Uncertainty. Doctoral Dissertation, Imperial College London, London, UK, 2016. [Google Scholar]
- Zhang, T.; Giannelos, S.; Pudjianto, D.; Strbac, G. Performance evaluation of reinforcement learning for hydrogen integration in renewable microgrids. IET Conf. Proc. 2026, 2025, 877–883. [Google Scholar] [CrossRef] [Scilit]
- Giannelos, S.; Konstantelos, I.; Strbac, G. A new class of planning models for option valuation of storage technologies under decision-dependent innovation uncertainty. In Proceedings of the 2017 IEEE Manchester PowerTech, Manchester, UK, 18–22 June 2017; pp. 1–6. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.


