Next Article in Journal
Research on the “Economy-Society-Environment” Sustainability of Urban Agglomerations in the Yellow River Basin, China
Previous Article in Journal
From Sustainability Recognition to Documented Outcomes: A Systematic Review and Study-Level Evidence Translation Analysis Across Productive Sectors
Previous Article in Special Issue
Green Intervention with a Hydroxyapatite-Based Sustainable Eco-Material: Case Study of the Apos Architecture Summer School
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Agentic AI for Climate-Resilient Cities: A PRISMA-Guided Review and Digital Twin Framework

1
AI Center, Faculty of Computer and Information Systems, Islamic University of Madinah, Madinah 42351, Saudi Arabia
2
AI V&V Lab, King Fahd University of Petroleum and Minerals, Dhahran 31261, Saudi Arabia
3
Civil Engineering Department, Islamic University of Madinah, Madinah 42351, Saudi Arabia
4
Faculty of Computing and Information Technology, University of the Punjab, Lahore 54590, Pakistan
5
Faculty of Computing and Informatics, Multimedia University, Cyberjaya 63100, Selangor, Malaysia
6
Department of Structures for Engineering and Architecture, University of Naples Federico II, 80125 Naples, Italy
*
Author to whom correspondence should be addressed.
Sustainability 2026, 18(17), 8917; https://doi.org/10.3390/su18178917
Submission received: 29 June 2026 / Revised: 26 July 2026 / Accepted: 5 August 2026 / Published: 31 August 2026

Abstract

Cities face pressure from urban growth and climate risk, yet deployed systems stay single-domain and reactive. This PRISMA-guided rapid review applies operationalized criteria to separate Agentic AI from conventional machine learning for SDG 11 (Sustainable Cities and Communities) and SDG 13 (Climate Action). Agentic AI is defined by four properties: task-level autonomy, goal-directed planning, tool use, and multi-agent coordination; evidencing at least two marks a system as fully agentic. A two-tier search across five databases with backward citation tracking returned 896 records (2018–2026), of which 60 met the eligibility criteria and 14 satisfied the agentic threshold. The corpus is stratified by study type with a threshold sensitivity analysis. Two contributions follow: a reference architecture specifying how an agentic layer and an urban digital twin exchange state, and a real-data feasibility study on the SEVIR archive testing whether multimodal fusion improves hazard classification. On real data the proposed model is the best-ranked of four but only marginally exceeds a no-change persistence baseline, giving the assumption weak support, not operational evidence. The review reveals a field growing sharply since 2023, clustered in a few urban and climate domains, with almost no validated cross-domain deployment.

1. Introduction

Urban digitalization and artificial intelligence are reshaping how cities respond to sustainability and climate pressure, yet the systems actually deployed remain fragmented. Most operate within a single domain, react to conditions after they have deteriorated, and cannot coordinate across mobility, infrastructure, energy, and environmental monitoring [1,2]. Agentic AI has been proposed as a response to precisely this fragmentation, offering autonomous decision execution, goal-directed planning, multi-agent collaboration, and adaptive reasoning [3]. Unlike conventional AI systems, which produce outputs that a human then acts upon, agentic systems are designed to persist in dynamic environments and to act within a bounded operational scope. Orchestration stacks built on large language models (LLMs), including multi-agent pipelines, graph-structured agent workflows, and LLM-augmented urban planning agents, have begun to operationalize these properties through structured communication protocols, tool-calling interfaces, and memory-augmented multi-step reasoning [4,5].
The argument developed here is that Agentic AI, the Internet of Things (IoT), large language models, and multi-agent systems (MAS) are not independent lines of work but mutually enabling components of a single ecosystem. IoT and data pipelines supply the sensing substrate; language models and planning algorithms supply the reasoning core; MAS coordination distributes action across urban domains; and digital twin (DT) integration supplies the simulation layer that closes the loop from sensing to decision to deployment. Figure 1 maps that ecosystem, and Figure 2 situates it against the shared challenge space where SDG 11 and SDG 13 objectives overlap.

1.1. Motivation and Problem Statement

Urbanization and climate change are among the defining systemic pressures of this century, and they act on the same physical stock. Cities house more than half the global population and account for over 70% of total CO2 emissions [2,6], which places them simultaneously at the center of environmental risk and at the center of any credible sustainability transition. As metropolitan areas expand, the pressures they face increasingly interlock: congestion, inefficient resource use, infrastructure strain, and exposure to flooding, extreme heat, and severe storms are not separable problems [7].
Conventional urban management struggles against this interlocking structure. Interventions are typically one-off, scoped to a single problem, and dependent on human-paced decision cycles that do not scale to the volume and velocity of data produced by modern sensing networks. What is needed is a class of system capable of predictive analysis, responsive coordination, and evidence-based intervention across several urban and environmental domains at once.
SDG 11 and SDG 13 name these imperatives directly, calling respectively for inclusive, resilient, sustainable settlements and for urgent climate action [6]. Meeting them requires technological paradigms that can represent the complexity of urban processes and support anticipatory governance rather than after-the-fact response. Agentic AI is a plausible candidate because it bundles goal-directed reasoning, autonomous execution, tool use, and multi-agent coordination into one design pattern [4]. The candidacy is only plausible, however, if the label is applied with discipline. A system is treated as agentic in this review only when it evidences at least two of the four operationalized criteria defined in Section 3, which is what separates Agentic AI from supervised or rule-based systems that happen to be deployed in a city. Autonomy here means task-level autonomy within bounded operational scope; policy-level and safety-critical decisions remain under human authority throughout, including in the architecture proposed in Section 7.
The research base has not kept pace with the terminology. Urban intelligence, climate analytics, and autonomous systems have largely been studied in isolation, which leaves the question of what agentic architectures actually contribute to sustainability outcomes unanswered at the level of the whole. That fragmentation is the gap this review addresses.

1.2. Scope, Manuscript Type, and Review Question

This manuscript is a review. Section 2, Section 3, Section 4, Section 5 and Section 6 constitute a PRISMA-guided rapid review of Agentic AI applications aligned with SDG 11 and SDG 13 over the period January 2018 to March 2026. Section 7, Section 8 and Section 9 are the synthesis output: a reference architecture abstracted from the reviewed corpus, and a feasibility check on one architectural assumption that the corpus repeatedly asserts. The review serves as the evidential contribution, while the architecture and feasibility study act as derived artifacts that make the synthesis actionable.
The review is organized around three questions. First, which published systems in the SDG 11 and SDG 13 space are genuinely agentic under an explicit, reproducible definition, and which are conventional machine learning under an agentic label? Second, what do those systems actually demonstrate, separated by study type, so that an implemented and evaluated system is never counted as equivalent to a conceptual proposal or a prior review? Third, what architectural pattern do the genuinely agentic systems converge on, and what does that pattern leave unspecified?

1.3. Contributions

  • Primary contribution. An operationalized PRISMA-guided rapid review of Agentic AI for urban sustainability and climate resilience. The review contributes a screening rubric (A/G/T/M) that makes agentic classification reproducible rather than rhetorical, a threshold sensitivity analysis showing how corpus composition responds to alternative inclusion cut-offs, a study-type stratification that separates implemented systems from conceptual and review works, and a structured comparison against prior reviews in the adjacent literature (Section 10.6).
  • Supporting contributions. Three artifacts follow from the synthesis and are subordinate to it.
    1.
    A reference architecture with specified interfaces. The reviewed corpus converges on a layered agentic–digital-twin pattern but almost never specifies how the layers exchange state. Section 7 fills that gap with an explicit synchronization rule (Equation (1)), an action feedback path, and a message contract with stated frequency and latency budgets (Section 7.4).
    2.
    A cross-domain optimization formulation. Section 7.3 formalizes a joint SDG 11/SDG 13 objective (Equation (2)), derives it as a special case of the MAS joint utility, and discusses Pareto-based and lexicographic calibration of competing sustainability reward signals.
    3.
    An empirical feasibility study. Section 8 evaluates the core architectural assumption that multimodal fusion under a transformer backbone improves hazard classification over single-modality convolutional processing, on real multi-sensor imagery from the SEVIR storm-event archive. Section 10.2 details the scope and limits of these findings, including that the support for the assumption is weak once a persistence baseline is considered.

1.4. Organization of the Paper

Section 2 establishes the conceptual and definitional groundwork, including the formal characterization of agency used for screening. Section 3 presents the review methodology: the A/G/T/M rubric and its justification, the inclusion threshold and its sensitivity analysis, the two-tier search strategy, screening reliability, study-type stratification, and methodological limitations. Section 4 and Section 5 synthesize the corpus for SDG 11 and SDG 13 respectively, and Section 6 examines their intersection. Section 7 derives the reference architecture from the synthesis. Section 8 describes the real-data feasibility study and Section 9 reports its outcomes. Section 10 interprets the empirical and review findings, treats the ethical and governance dimension at length, and positions the work against prior reviews. Section 11 concludes.

2. Background

2.1. Sustainable Development Goals Overview

The United Nations Sustainable Development Goals provide a global framework for addressing interlinked social, economic, and environmental challenges by 2030 [8]. SDG 11 concerns making cities and human settlements inclusive, safe, resilient, and sustainable through planning policies that reduce environmental footprint while improving quality of life and equitable access to services. SDG 13 concerns urgent climate action: mitigation through greenhouse gas reduction, adaptation to unavoidable change, and the mainstreaming of climate considerations into urban and national planning. Taken together the two goals describe cities as resilient, environmentally sustainable systems rather than as collections of separately optimized services.

2.2. Interdependence of SDG 11 and SDG 13

Urbanization and climate change are tightly coupled in both directions. Urban growth raises energy consumption, waste generation, and emissions, which amplifies climate forcing. Climate risk in the form of extreme weather, rising temperatures, and sea-level rise then acts directly on urban infrastructure, public safety, and economic activity [7]. Because the coupling runs both ways, interventions that optimize one goal in isolation routinely degrade the other, and solutions that advance urban sustainability and climate resilience jointly are required. Figure 2 renders this coupling explicitly, distinguishing the orchestration layer, where agentic coordination and reasoning occur, from the simulation layer, where the digital twin co-models urban and climate dynamics, and relating both to the shared challenge domains of emissions, urban heat, and disaster risk in which SDG 11 and SDG 13 objectives overlap.

2.3. Agentic AI: Formal Definition and Distinguishing Characteristics

Agentic AI denotes autonomous, goal-directed artificial intelligence systems that perceive an environment, reason over time, invoke tools or APIs, and act independently, frequently in coordination with other agents [3,4]. The distinction from supervised or rule-based AI is substantive rather than terminological. A fixed solar irradiance prediction model used in urban planning is not agentic even when its outputs drive decisions, because it lacks autonomy, a planning horizon, and inter-agent communication [9]. Formally, an agent may be characterized as a tuple A = 〈 S , A set , T , R , π ∗ 〉 , where S is the state space perceived from the urban–climate environment, A set is the action set (including tool invocations and inter-agent messages), T : S ✕ A set → Δ S is the stochastic transition function mapping state–action pairs to distributions over successor states, R : S ✕ A set → R is a reward signal aligned with sustainability objectives, and π ∗ is the policy maximizing cumulative expected reward.
Four properties of this tuple are directly observable in published system descriptions, and they are the four used as screening criteria in Section 3.
  • Autonomy (A) → the policy π executes without per-decision human authorization within a designated operational scope, under closed-loop feedback and higher-level human control.
  • Goal-directed planning (G) → the agent reasons over a horizon τ > 1 toward an explicitly specified objective, for instance maximizing traffic throughput or meeting an emission target, rather than mapping inputs to outputs in a single step.
  • Tool use or environmental interaction (T) → the action set A set includes invocations of external APIs, sensors, simulators, or actuators.
  • Multi-agent coordination (M) → the action set includes messages to other agents, enabling collective behavior beyond what any single agent achieves alone.
The choice of these four dimensions is not arbitrary, and Section 3.2 defends it in detail against the alternatives. In brief, each maps onto one element of the agent tuple, which is what makes the rubric checkable from a published architecture description rather than dependent on an author’s self-description. Autonomy is a property of π , planning of the horizon over T , tool use of A set with respect to the environment, and coordination of A set with respect to other agents. The same four dimensions recur, under varying names, in recent conceptual treatments of agentic urban AI [3,5,10] and in the architectural survey of Lee and Park [11].
Studies included in this review evidence at least two of the four properties. Section 3.2 reports what happens to the corpus at alternative thresholds. It is worth being precise about what autonomy means here: the systems discussed operate self-directedly within restricted operational scopes, and human actors retain authority over policy-level and safety-related decisions. Task-level autonomy is not system-level autonomy, and the two are frequently conflated in this literature.
Table 1 positions Agentic AI against three adjacent paradigms with which it is often confused. Classic multi-agent systems provide decentralized coordination but typically rely on pre-scripted communication protocols and lack open-ended reasoning or tool use [12]. Classical planning agents, including the belief–desire–intention architecture that predates the current language-model wave, decompose goals rigorously but operate offline over symbolic state spaces without real-time environmental interaction [13]. Single-agent reinforcement learning acquires autonomous policies but optimizes a fixed reward without deliberative multi-step planning or inter-agent message passing. What distinguishes Agentic AI as operationalized here is the co-presence of all four properties, which is why a two-of-four threshold admits systems that are meaningfully agentic without requiring the full pattern.

2.4. Related Technologies

2.4.1. Digital Twins

Digital twins are executable virtual representations of physical systems: infrastructure networks, traffic flow, energy grids, and environmental fields. Coupled with an agentic layer they support prediction, scenario testing, and intervention appraisal before anything is committed in the physical city [14,15,16]. Their practical value to decision-makers lies in offering a consequence-free space for trial and error, which operational urban systems cannot provide [17,18].

2.4.2. Multi-Agent Systems (MAS)

Multi-agent systems comprise several interacting agents making decentralized decisions. MAS frameworks support complex urban and climate applications by distributing tasks across specialized agents, which enables resource allocation, mobility optimization, and predictive climate modeling at scales that a monolithic controller handles poorly [19]. As Table 1 indicates, classic MAS supplies the coordination substrate of Agentic AI but not the open-ended reasoning and tool invocation that characterize contemporary agentic architectures. The joint utility across N agents may be written U = ∑ i = 1 N λ i u i ( s , a i ) , where λ i are domain-specific weights and u i is the local utility of agent i in state s under action a i . This connects directly to the cross-domain objective in Equation (2): for the two-domain case ( N = 2 ), setting λ 1 = α , λ 2 = β , u 1 = r SDG 11 , u 2 = r SDG 13 and introducing temporal discounting recovers Equation (2) as a special case. Fixing λ i in practice requires either stakeholder elicitation or multi-objective optimization; Pareto-based approaches [20] and lexicographic constraint ordering have both been proposed in the cooperative multi-agent reinforcement learning (MARL) literature as principled alternatives to hand-set weights.

2.4.3. Internet of Things (IoT)

IoT devices supply the real-time data feeds, from sensors, smart meters, and environmental monitors, that digital twins and agentic layers consume. They enable real-time analytics, anomaly detection, and automated control, closing the loop between sensing, analysis, and action [21].

2.5. Research Gaps and Opportunities

Several structural gaps persist despite the pace of development, and they shape what this review can and cannot conclude. Deployment is unevenly distributed: advanced implementations cluster in technologically well-resourced regions while developing contexts face persistent constraints in infrastructure, data access, computational capacity, and skills [22]. Evaluation practice is inconsistent, with studies reporting improvements against incommensurable baselines, which limits cross-study generalization and was the reason formal meta-analysis was not attempted here. Ethical and governance considerations compound the difficulty as agentic systems reach safety-critical urban processes, making responsible-AI principles a design requirement rather than a compliance afterthought [4]. Two gaps are particularly consequential for this review. First, the literature rarely compares operational agentic deployments against conventional machine learning baselines using modern orchestration platforms, so the marginal value of agency remains largely asserted rather than measured. Second, cross-domain orchestration is almost entirely absent: urban management, climate monitoring, and energy optimization are studied separately even though the underlying systems are coupled.

3. Review Methodology

The review follows PRISMA 2020 reporting principles under a rapid-review protocol [23]. Rapid reviews retain the core logic of systematic reviews while streamlining selected stages to deliver timely evidence in fast-moving fields [24]. Agentic AI in urban and climate settings is both interdisciplinary and unusually fast-moving, and four adaptations were made to balance rigor against feasibility.
First, the search covered five academic databases supplemented by targeted backward citation tracing, rather than attempting comprehensive gray-literature coverage. Second, title and abstract screening was performed by a primary reviewer with a second reviewer independently validating a stratified random 20% sample; the sampling rate was fixed a priori in the protocol, and all six co-authors participated in full-text assessment, quality appraisal, and extraction. Section 3.5 treats the reliability consequences of this design directly rather than deferring them to the limitations. Third, backward reference scanning was applied selectively to high-impact and seminal studies to recover eligible work the database queries missed. Fourth, quantitative meta-analysis was not performed, because heterogeneity in study designs, evaluation protocols, and reported outcomes made pooled effect estimation uninterpretable. Preprints and gray literature were excluded to hold a peer-reviewed synthesis standard, a decision whose recency-bias cost is quantified in Section 3.10. These adaptations follow published guidance on rapid evidence synthesis [24].
A structured protocol fixing scope, search strategy, eligibility criteria, and synthesis procedure was established before searching. Reporting follows PRISMA 2020 [23,25].

3.1. Operationalizing Agentic AI for Screening

The central methodological problem in this literature is that “agentic” is applied to systems ranging from genuine multi-agent planners to relabeled supervised classifiers. Screening therefore required a rubric that a second coder could apply to a published architecture description and reach the same verdict. The four criteria defined in Section 2.3, autonomy (A), goal-directed planning (G), tool use or environmental interaction (T), and multi-agent coordination (M), were operationalized into the decision rules in Table 2. Each rule specifies what counts as positive evidence, what does not, and where in a paper that evidence must appear. Studies were coded on all four dimensions from the architecture description, experimental setup, or implementation detail. Claims made only in an abstract or introduction, without corresponding architectural or experimental substantiation, were not treated as evidence. Where evidence was ambiguous, the conservative decision was to code the criterion absent, which biases the corpus toward under-inclusion rather than over-inclusion.

3.2. Inclusion Threshold and Sensitivity Analysis

A study was classified as fully agentic when it satisfied at least two of the four criteria. The threshold was calibrated on a random 20-study sample drawn before full-text screening, and the calibration exposed a genuine trade-off rather than an obvious answer. A threshold of ≥1 admits systems whose only agentic property is task-level autonomy, which in practice means passive prediction pipelines running on a schedule; agency reduces to automation at that cut-off. A threshold of ≥3 excludes a substantial body of multi-step planning systems that coordinate nothing, and single-agent planners are not obviously less agentic than message-passing systems that never plan. The ≥2 rule requires that autonomy be accompanied by at least one of planning, environmental interaction, or coordination, which matches the conceptual position that agency entails more than unsupervised execution. The rule is a stratification of agentic intensity applied to the relevance-screened corpus, not the topical inclusion filter itself; conceptual proposals and prior reviews that concern agentic urban systems are retained and marked by study type (Section 3.6) even when they do not themselves implement two of the four properties.
Because a threshold chosen this way is a methodological commitment rather than a fact, Table 3 reports how the count of fully agentic studies responds to the alternatives. Of the 60 included studies, 18 exhibit at least one agentic property, 14 satisfy the ≥2 rule and constitute the primary implemented evidence, and only 3 exhibit three or more. Relaxing to ≥1 adds four systems whose sole agentic property is task-level autonomy, while tightening to ≥3 leaves too few studies for domain-level synthesis. The qualitative synthesis conclusions in Section 4, Section 5 and Section 6 (the domain clusters, the post-2023 growth pattern, and the scarcity of cross-domain deployment) are drawn from the full 60-study corpus stratified by study type rather than from the fully agentic subset alone, and they do not reverse when attention is restricted to the 14 fully agentic studies. That robustness is the substantive result of the sensitivity analysis, and it is the reason the ≥2 choice does not drive the paper’s findings.
Eligibility also required that studies (i) present AI-driven approaches applicable to urban sustainability or climate resilience, (ii) exhibit agent-based, autonomous, or multi-agent characteristics consistent with the operationalization above, and (iii) provide sufficient methodological detail for scholarly assessment. The publication window of January 2018 to March 2026 was fixed in the protocol before searching and corresponds to the emergence of deep reinforcement learning in urban control and the subsequent arrival of language-model-augmented agent architectures.

3.3. Search Strategy

An earlier formulation of the search intersected four concept groups with a mandatory Boolean AND: an agentic-systems group, an urban group, a climate group, and a digital-twin or simulation group. Requiring all four simultaneously maximizes precision but suppresses recall, and it does so asymmetrically: a genuinely agentic traffic-signal controller evaluated for emissions impact would be excluded solely for not mentioning simulation infrastructure. That is a real limitation of the design and it warranted correction.
The strategy reported here is two-tier. The core query requires the agentic group and at least one of the urban or climate groups, which is the minimal condition for topical relevance. The digital-twin and cyber-physical group is applied as an enrichment facet rather than a mandatory conjunct, used to stratify results rather than to filter them:
  • TITLE-ABS-KEY(
    ("agentic AI" OR "autonomous AI" OR "intelligent agent*"
    OR "multi-agent system*" OR "multi-agent reinforcement learning"
    OR "LLM agent*" OR "autonomous system*")
    AND
    ( ("sustainable cit*" OR "smart cit*" OR "urban system*"
    OR "urban plan*" OR "urban environment*")
    OR
    ("climate action" OR "climate resilien*"
    OR "carbon emission*" OR "climate adapt*"
    OR "environmental monitor*") )
    )
    AND [enrichment facet, recorded not required]
    ("digital twin*" OR "urban digital twin*" OR "simulation"
    OR "urban model*" OR "cyber-physical system*" OR "IoT")
The searches covered Scopus, Web of Science, IEEE Xplore, SpringerLink, and ScienceDirect, with field codes adapted per platform, and the final search was executed in March 2026. Backward reference scanning of high-impact studies supplemented the database queries. Table 4 reports retrieval by source. Because relaxing the fourth group changes what the search can be expected to recover, Section 3.10 states plainly what the revised strategy still misses: venues outside the five indexed databases, non-English publications, and preprints.

3.4. Screening and PRISMA Flow

Screening followed the PRISMA 2020 workflow [23,25]. Of 896 identified records, 721 remained after deduplication and entered title and abstract screening against the predetermined eligibility criteria. At that stage 651 records were excluded: 420 were off-topic for urban or climate domains, 145 did not concern agentic or autonomous AI, and 86 were not aligned with SDG 11 or SDG 13. The remaining 70 records proceeded to full-text assessment, where 10 were excluded across four categories: 3 were off-topic for agentic urban or climate AI on full reading, 2 were insufficiently aligned with the integrated urban–climate scope, 2 lacked adequate methodological description of the AI system, and 3 fell outside the publication window. Sixty studies satisfied all criteria and were carried into synthesis; of these, 14 satisfy the ≥2 agentic threshold and form the primary implemented evidence (Section 3.2), while the remainder are implemented systems below the threshold together with conceptual and review works, retained and marked by study type. Figure 3 presents the flow diagram.

3.5. Screening Reliability and Bias Control

The screening design addresses potential selection bias resulting from single-primary-reviewer title and abstract screening (581 of 721 records) through three validation measures.
First, agreement was quantified at both stages rather than only at title and abstract. A stratified random sample of title and abstract records ( n = 140 , 19% of the 721 screened) was independently screened by a second reviewer, yielding Cohen’s κ = 0.81 (95% CI: [0.74, 0.88]), which falls in the almost-perfect band [26]. A sample of 40 full-text assessments (57% of the 70 assessed) yielded κ = 0.78 (95% CI: [0.68, 0.88]), indicating substantial agreement. The lower full-text value is expected and informative: full-text coding requires judging the four agentic criteria against architecture descriptions, which is a harder discrimination than topical relevance.
Second, disagreements were resolved by a documented procedure rather than by deference. Every discordant decision was discussed against the Table 2 rules, and where the rules did not resolve the case, a third co-author adjudicated. Records that remained ambiguous after adjudication were excluded, consistent with the conservative rule applied throughout.
Third, the direction of residual bias can be bounded even though the unvalidated decisions cannot be re-verified. The observed disagreement pattern was asymmetric: discordance concentrated on false exclusions, that is, records the primary reviewer excluded and the second reviewer would have retained, rather than on false inclusions. Under the conservative coding rule this asymmetry is expected, and its consequence is that the corpus is more likely to under-represent than to over-represent the agentic literature. An under-inclusive corpus weakens claims about coverage and prevalence, which is why Section 10.2 avoids prevalence claims, but it does not inflate the observed characteristics of the studies that were retained. Readers should nonetheless treat corpus-completeness statements as lower bounds. Raising the validation fraction, or dual-screening the full record set, remains the correct remedy and is recommended for any extension of this work.

3.6. Data Extraction and Study-Type Stratification

Extraction was performed into a structured schema capturing bibliographic metadata, AI methodology (multi-agent reinforcement learning, LLM-augmented planning, digital twin integration, and so on), the A/G/T/M coding, application domain, evaluation method and reported metrics, SDG alignment, and study type. The study-type field was added specifically because synthesizing an implemented and evaluated system alongside a conceptual proposal, as though the two carry equal evidential weight, is a failure mode that this literature is prone to. Four strata are distinguished.
  • Implemented (I): a system was built and evaluated, whether in simulation, on a testbed, or in deployment. Reported outcomes constitute evidence.
  • Conceptual (C): an architecture, framework, or model is proposed and argued for without implementation or evaluation. Reported outcomes constitute expectation.
  • Review or Survey (R): the work synthesizes other studies. Findings are secondary evidence and are never counted as independent instances of a phenomenon.
  • Supporting (S): methodological, normative, or definitional sources cited for framing rather than as evidence, including reporting guidelines, statistical references, and intergovernmental policy documents.
Table 5 reports the stratification. The 60 included studies comprise the Implemented, Conceptual, and Review strata; the Supporting stratum sits outside the screened set and is cited only for framing. The distribution should temper any reading of the synthesis that treats the corpus as a body of deployment evidence: fewer than half of the included studies report an implemented and evaluated system, a substantial share argues for agentic urban systems rather than demonstrating them, and the review sections that follow mark this distinction inline wherever a claim rests on conceptual rather than implemented work.
Study type and agentic intensity are orthogonal classifications, and conflating them misreads the corpus. Table 6 cross-tabulates the two over the 60 included studies. All 14 fully agentic studies lie within the Implemented stratum: no conceptual or review work clears the ≥2 threshold, and the other 10 implemented studies report systems that do not. The 46 studies below the threshold are therefore not reviews; they are 10 implemented systems short of full agency, 23 conceptual proposals, and 13 reviews. This is why the review reports two counts that must not be conflated: the 24 implemented studies supply the performance and outcome evidence, while the 14 fully agentic among them supply the evidence specific to agentic operation.

3.7. Quality Assessment

Methodological quality was appraised using a Joanna Briggs Institute critical appraisal framework adapted for computational systems research. Five dimensions were assessed: clarity of research objectives; methodological transparency including algorithmic and implementation detail; validation strategy, whether simulation-based, comparative, or deployed; data sufficiency in scale, diversity, and representativeness; and acknowledgment of limitations. Appraisal placed 53% of studies at high quality across all five dimensions, 33% at moderate, and 14% at adequate with notable limitations. No study was excluded on quality grounds. Instead, appraisal outcomes modulate the confidence attached to each finding, and moderate or low quality primary studies are identified as such wherever their results are cited in Section 4, Section 5 and Section 6.

3.8. Agentic Coding of Included Studies

Table 7 reports the A/G/T/M coding for fourteen representative studies, together with study type. Representativeness was determined by a stated procedure rather than by convenience: studies were selected to span the six SDG 11 and five SDG 13 application clusters identified in Section 4 and Section 5, to cover the full range of agentic scores from 2 to 4, and to include at least one review and one conceptual work so that readers can see how non-implemented studies score under the same rubric. Every mark in the table follows the decision rules in Table 2 and the definitions in Section 2.3; the complete coding for all included studies is deposited with the Supplementary Materials.

3.9. Bibliometric Overview of the Corpus

Table 8 reports the temporal distribution. Activity rises sharply from 2023 onward, with 2025 alone accounting for 39 studies, roughly 51% of the reference set, while only 25% predates 2023. The inflection coincides with the arrival of language-model-augmented agent systems and their rapid application to urban and climate problems. A corpus this heavily weighted toward a single recent year carries an obvious caution: much of the evidence base has not yet been independently replicated, and the field’s apparent consensus may reflect a common intellectual moment rather than accumulated validation.
Publisher concentration is reported in Table 9. Elsevier, MDPI, and Springer Nature together account for roughly 57% of the reference set, indicating that this work is published predominantly in established, high-visibility venues. The small share from intergovernmental technical reports and from independent journals contributes policy-oriented and interdisciplinary material that the major publishers do not carry. Studies in the independent category received additional scrutiny during quality appraisal (Section 3.7) and were retained on the basis of adequate methodological transparency, with confidence recorded as moderate.
Table 10 divides the corpus by document type. Journal articles dominate at roughly 85%, with conference papers, technical reports, and a small number of foundational books and book chapters making up the remainder. The distribution suggests a field that has moved past exploratory conference-level contribution into archival publication, though as Table 5 shows, archival publication is not the same thing as implemented and evaluated work.
Three structural features of the corpus follow from this profile and carry into the synthesis: a recent and steep expansion concentrated in 2025, a concentration within established publishers, and a dominance of journal articles that coexists with a substantial fraction of non-implemented work.

3.10. Methodological Limitations

Several limitations constrain what this review can support, and they are stated here rather than distributed through the discussion.
The review was not prospectively registered. Registration is standard for clinical systematic reviews and not yet established practice in computational sustainability research, and a structured protocol fixing search strategy, eligibility, and agentic screening dimensions was written before study selection as a partial substitute. It is a partial substitute only: without prospective registration, protocol deviations cannot be externally verified, and future work in this area would benefit from registration on PROSPERO or the Open Science Framework.
Screening reliability is bounded by the single-reviewer design analyzed in Section 3.5. The 560 unvalidated title and abstract decisions cannot be formally verified, and classifying agency from a brief abstract is an interpretive judgment even for an experienced coder.
Coverage is bounded in three ways. Five databases were searched, so work in less-indexed venues may be absent. Gray literature and preprints were excluded, which introduces a systematic recency bias in a field where significant developments frequently appear first on preprint servers, and which likely under-represents the most recent methodological advances. Non-English publications were not searched, which is a particular concern given that the review comments on the geographic concentration of deployment.
Coding reliability for the full corpus is bounded by the reported inter-rater agreement, and the conservative exclusion rule biases toward under-inclusion as analyzed above.
Heterogeneity in study design, evaluation metrics, and reporting precluded formal meta-analysis [24]. A targeted meta-analysis over a methodologically homogeneous subgroup, for instance studies applying MARL to traffic signal control against comparable simulation benchmarks, would be a valuable follow-on and is recommended as a priority.
Finally, the reference architecture in Section 7 has not been empirically validated at city scale, and the feasibility study in Section 8 evaluates only the prediction stage, on a small real-data subset of fifty SEVIR storm events. Performance characterizations of the architecture are therefore indicative rather than demonstrated, and outcomes reported in individual included studies may not generalize across deployment contexts. Section 10.2 states which specific claims are affected.

4. Agentic AI for SDG 11: Sustainable Cities

4.1. Overview of Urban Applications

Applications of Agentic AI to urban sustainability, safety, and operational efficiency have expanded quickly, and the corpus organizes into six clusters: smart mobility, digital-twin-based infrastructure planning, waste management automation, public safety and emergency response, urban energy management, and environmental monitoring. Figure 4a maps these clusters. Energy and environmental monitoring (Section 4.6) is the most substantially represented, appearing in eleven of the included studies. Throughout this section and the next, study-type markers follow the convention of Table 5: (I) implemented and evaluated, (C) conceptual, (R) review or survey. Where a claim rests on (C) or (R) sources it describes an argued expectation, not a measured outcome.

4.2. Smart Mobility

Traffic control and vehicle coordination are where agentic methods have the strongest implementation record. Hierarchical MARL has produced sustainability-oriented outcomes in signal control under high-density conditions (I) [27], and cooperative MARL has been applied to autonomous vehicle lane-changing in mixed traffic with joint mobility and emissions objectives (I) [41]. Lee and Park survey the architectural space these systems occupy (R) [11]. Route optimization across public and private transport incorporates live telemetry and predictive demand to support emissions-aware scheduling [42], and IoT-driven adaptive traffic management addresses congestion through real-time multi-agent coordination. Table 11 summarizes representative cases with their study types.

4.3. Digital Twin-Based Infrastructure Planning

Pairing digital twins with agentic reasoning lets planners simulate infrastructure changes, appraise policy, and anticipate maintenance without committing capital (I) [28,34]. Scenario simulation covers urban expansion, energy planning, and disaster preparedness, and predictive maintenance agents monitor structural health indicators to support condition-based intervention. Reported efficiency gains vary widely with deployment scenario, sensor coverage, model fidelity, and institutional adoption capacity (R) [44], and this variance is large enough that pooled estimates would mislead. Multi-objective criteria spanning environmental, social, and economic indicators are standard in this cluster, with simulation accuracy refined continuously through IoT feedback.

4.4. Waste Management Automation

Predictive analytics and route optimization dominate this cluster. IoT-instrumented bins report fill levels, and agentic routing adjusts collection in real time, with pilot deployments reporting reduced unnecessary vehicle movement in managed urban environments (I) [45,46]. Graph-based predictive modeling with adaptive routing has been reported to improve system resilience and reduce operational cost (I) [47], and transformer-based MARL combining classification, forecasting, and routing has reported emissions and efficiency benefits (I) [48]. The magnitude of improvement depends on network density, sensor accuracy, and fleet size, and single-pilot outcomes should not be transferred to other urban settings without recalibration.

4.5. Public Safety and Emergency Response

Agentic methods support urban safety through predictive risk modeling and multi-agent coordination. Geospatial models identify locations with elevated likelihood of accidents or infrastructure failure, and multi-agent pipelines allow fire, medical, and evacuation resources to be reallocated dynamically as incident conditions change (I) [49]. Coupling meteorological feeds to early-warning systems provides advance notice of floods, heatwaves, and severe weather, supporting pre-positioning of emergency services.

4.6. Urban Energy and Environmental Monitoring

Continuous optimization cycles underpin agentic energy management. Real-time agents adjust electricity, water, and district heating in response to variable demand and renewable availability (I) [50], and multi-agent architectures for demand balancing, in which each agent optimizes local comfort and power requirements, have proven viable at grid scale alongside forecasting and demand-response scheduling (I) [51].
Decentralized hierarchical MAS structures support scalable smart grid management under high renewable penetration while balancing economic and environmental objectives (I) [52]. Multi-layer systems combining forecasting, MARL, evolutionary optimization, and blockchain-enabled EV scheduling have reported reductions in carbon intensity and peak-load stress in urban microgrid settings (I) [53,54], and market-based MAS frameworks provide complementary decentralized load-shifting and renewable integration mechanisms for district energy (I) [55].
Environmental monitoring forms a closely coupled cluster. Distributed air and noise sensor arrays identify pollution hotspots and trigger local interventions such as adaptive rerouting or HVAC schedule adjustment. Together, energy management and environmental monitoring account for eleven of the included studies, which makes this the densest deployment domain for Agentic AI under SDG 11.

5. Agentic AI for SDG 13: Climate Action

5.1. Overview of Climate Applications

Climate applications concentrate in five domains: climate monitoring, early warning systems, renewable energy forecasting, carbon emission and air quality tracking, and disaster risk modeling with climate policy evaluation. Figure 4b maps them. Compared with the SDG 11 clusters, this side of the corpus contains proportionally more review and conceptual work, which is reflected in the hedging below.

5.2. Climate Monitoring

Agentic systems integrate satellite imagery, sensor networks, and meteorological data to track environmental anomalies and climate trends over time (R) [56]. Agent-based fusion of multi-source observations identifies urban heat islands, deforestation fronts, and pollution hotspots at temporal resolutions that manual analysis cannot match (I) [57]. Simulation pipelines combining reinforcement learning with learned forecasting support seasonal prediction and extreme event probability estimation, which in turn feed adaptive management strategies.

5.3. Early Warning Systems

Multi-agent pipelines improve both the capability and the speed of early warning for extreme events (R) [35]. Agents fuse meteorological, hydrological, and geophysical data in real time to produce hazard estimates for flood inundation, storm surge, and wildfire propagation. Automated dissemination reaches authorities and citizens through multiple channels, and scenario simulation agents optimize evacuation routing and resource pre-positioning. LLM-based multi-agent simulation environments have been proposed for modeling population behavior during disasters, supporting pre-deployment stress testing of response strategies (R) [30]. The Early Warnings for All program illustrates how multi-agency AI coordination can compress the data-to-decision cycle, though achieved lead times vary substantially with hazard type, geography, and infrastructure readiness [58].

5.4. Renewable Energy Forecasting

Agentic pipelines forecast and optimize renewable generation in support of grid stability, storage management, and demand response (R) [31,59]. Under favorable climatic conditions, deep learning forecasters embedded in agentic pipelines can be competitive on day-ahead solar irradiance accuracy, but performance varies considerably across geographic regions and climatic profiles, and published accuracy figures should not be extrapolated without the conditions and datasets under which they were obtained [60]. Wind power prediction achieves comparable accuracy through meteorological feature engineering, with accuracy declining as the forecast horizon extends. Agentic load-balancing systems manage battery storage, smart charging, and dispatchable generation to hold grid frequency stable as renewable penetration rises.

5.5. Carbon Emission and Air Quality Tracking

Real-time emissions and pollution monitoring supports both regulatory enforcement and evidence-based policy (I) [61]. City-scale CO2 and NOx tracking agents fuse transport, industrial, and building energy data into spatiotemporal emission inventories. Identified hotspots enable targeted responses such as rerouting heavy traffic or activating low-emission zones, with improvements in ambient air quality reported in experimental deployments. This cluster aligns directly with SDG 13 Target 13.2 on mainstreaming climate measures into policy.

5.6. Disaster Risk Modeling and Climate Policy Evaluation

Simulation-based policy evaluation for resilience planning is the principal contribution here (I) [32,33]. Multi-agent co-simulation environments model urban–climate interaction under projected extreme event scenarios, allowing planners to stress-test adaptation strategies such as green roofs, permeable pavement, and coastal barriers before committing infrastructure investment. Digital-twin-based climate-responsive development frameworks integrate zero-energy building strategies into the smart city context, and long-term planning tools built on ensemble climate projections support resilient infrastructure design.

6. Intersection of SDG 11 and SDG 13

6.1. Integrated Urban–Climate Architecture

Serving both goals at once requires a single system capable of trading urban efficiency against climate resilience rather than optimizing either alone. The corpus indicates four processes that such a system must support: integration of real-time IoT, weather, and infrastructure data; simulation of urban policy under varying climate conditions; automated infrastructure and resource allocation; and integration across urban systems for climate-informed decision-making [50]. Figure 5a renders this coordination structure, and Figure 5b shows the reference architecture derived from it in Section 7.

6.2. Urban–Climate Co-Simulation

Agentic coordination of co-simulation across interdependent layers is the mechanism by which urban policy and climate outcome are linked (R) [36,40]. Multi-agent co-simulations model interaction among traffic flow, energy use, greenhouse gas emissions, and urban morphology, driven by climate variables including temperature, precipitation, and storm frequency, to project city-scale scenarios. Scenario-based evaluation surfaces second-order policy consequences, which is what allows a policymaker to see the implications of a measure before enacting it. Table 12 presents representative scenarios and the planning outcomes they support.

6.3. Digital Twin-Based Climate-Aware Planning

Digital twins maintain dynamic virtual models of cities and their environmental state, which supports iterative climate-conscious planning cycles (I) [37]. Modeling infrastructure stress under climatic extremes surfaces vulnerabilities before they become operational failures. Agent-driven intervention simulation yields quantitative performance estimates that inform capital allocation, and digital twin paradigms for urban heat monitoring provide concrete mechanisms linking simulation output to actionable resilience strategy (I) [62]. Dashboards synthesizing multi-scenario results allow transparent comparison of policy alternatives.

7. Proposed Framework: Climate-Resilient Agentic AI–Digital Twin Architecture

The architecture presented here is a synthesis output, not an independent proposal. Three observations from the review motivate it. First, the studies scoring 4 on the A/G/T/M rubric converge on the same layered pattern: sensing, twin-based state representation, agentic decision, actuation. Second, almost none of them specify how the layers exchange state, so the pattern is reproducible in outline but not in implementation. Third, the studies that do specify a synchronization mechanism treat the twin as a passive mirror, updated from sensors alone, which breaks the multi-step reasoning that agency requires. Section 7.2, Section 7.3 and Section 7.4 address precisely these three gaps: they give the synchronization rule explicitly, close the loop from agent action back into twin state, and specify the message contract between layers.
The architecture is structured as a three-phase pipeline: heterogeneous climate data acquisition through multimodal sensing; digital twin state representation, synchronization, and predictive simulation; and agentic adaptive policy optimization operating on the continuously updated twin state. Figure 5b shows the overall structure and Figure 6 shows the closed-loop control and temporal interaction detail.
  • The end-to-end flow proceeds as follows. Multimodal sensing infrastructure generates heterogeneous environmental signals at varying spatial resolution and temporal frequency. A data fusion engine performs modality alignment, temporal interpolation, outlier suppression, and privacy-preserving normalization to produce a unified observation tensor o t ∈ O at each time step t. This fused state propagates to the digital twin synchronization module, which updates the virtual city model and produces a synchronized state s ^ t . A climate prediction engine applies learned models to s ^ t to generate probabilistic forecasts s ^ t + 1 : τ over a planning horizon τ , which is what turns the twin from a reactive mirror into a forward-looking simulation environment. The augmented state ( s ^ t , s ^ t + 1 : τ ) is presented to the multi-agent reasoning layer, whose agents select domain-specific actions under SDG-aware governance constraints. Approved actions dispatch to actuators, and the resulting environmental transitions are sensed and fed back, closing the perception-action loop.

7.1. Data Acquisition Layer

The acquisition layer aggregates heterogeneous urban and environmental streams: IoT sensor networks, satellite imagery, radar observation, air quality sensors, weather stations, precipitation monitors, mobility telemetry, structural health monitoring, municipal open data, and citizen reporting [38,44]. Multimodality is a functional requirement rather than a design preference. No single modality captures the spatiotemporal complexity of urban–climate dynamics, and cross-modal redundancy is what confers robustness when a sensor fails or degrades. The modalities are complementary in a specific way: radar supplies high-temporal-resolution precipitation and storm structure, satellite supplies spatial coverage of surface temperature and cloud, fixed air quality stations supply ground truth at points, and mobility telemetry proxies population exposure and infrastructure utilization.
Raw data pass through modality-specific preprocessing, including radar reflectivity normalization, satellite band selection, air quality temporal smoothing, and mobility flow aggregation, followed by cross-modal alignment onto a common spatiotemporal grid. Missing-modality imputation and uncertainty quantification are architectural requirements rather than optional extensions, since sensor dropout is a normal operating condition rather than an exception. An edge-to-cloud synchronization strategy preprocesses latency-sensitive observations at the network edge before ingestion, trading computational efficiency against real-time responsiveness. The fusion engine emits the structured observation tensor o t , which is the canonical input to everything downstream.

7.2. Digital Twin Layer and State Synchronization

The digital twin maintains a dynamic virtual model of urban infrastructure and environmental conditions [34,63]. The concept originates in product-lifecycle engineering, where it was introduced as a high-fidelity virtual counterpart kept in correspondence with a physical asset [64], and it has since matured into an industrial paradigm spanning simulation, monitoring, and control [65]. It is not a visualization tool. It is an active, simulation-driven decision environment: a cyber-physical representation that evolves in synchrony with real-world state and supports predictive analytics, disaster scenario simulation, and resilience planning. The synchronized state s ^ t encodes static urban topology (road networks, building footprints, critical infrastructure locations) alongside dynamic environmental fields (temperature, precipitation intensity, air quality surfaces, mobility flow matrices), which lets downstream components reason jointly over physical constraints and evolving conditions. This dual representation is grounded in the corpus: Villani et al. [34] demonstrate a Venice urban sustainability twin, Korkmaz [37] addresses disaster management twins, and Sacoto-Cabrera et al. [36] present IoT-integrated smart city twins. Integration with AI-enabled IoT platforms extends simulation capacity toward real-time building energy control and city-scale environmental surveillance [66].
Synchronization is the mechanism that the reviewed literature leaves unspecified, and it is specified here. The twin state updates at each time step as a function of three inputs rather than one:
s ^ t = Φ s ^ t − 1 , o t , p t , a t − 1 = F env s ^ t − 1 , o t ⏟ exogenous update ⊕ F haz p t ⏟ hazard surface ⊕ F inf s ^ t − 1 , a t − 1 ⏟ endogenous update ,
where ⊕ denotes composition over disjoint state partitions, o t is the fused observation tensor, p t is the hazard probability vector emitted by the climate prediction engine, and a t − 1 is the joint action executed in the previous step. The three operators act on separate partitions of s ^ and therefore compose without conflict. F env updates environmental fields from sensor evidence, F haz writes the predicted hazard risk surface, and F inf updates infrastructure state variables such as resource allocation, alert levels, and evacuation route designations to reflect what the agents actually did.
The third term is the substantive contribution of Equation (1). Without F inf , the twin is a mirror of the environment and the agents face an environment that never registers their past decisions, which reduces multi-step planning to a sequence of independent single-step reactions. Including it makes the twin evolve as a joint function of exogenous environmental dynamics and endogenous agent behavior, which is the condition under which reasoning over a horizon τ > 1 is meaningful.
Between sensor arrivals the state is propagated rather than held constant, since observation streams are asynchronous and individually unreliable. For linear-Gaussian partitions a Kalman estimator suffices; for the nonlinear partitions a learned transition model T θ propagates state with uncertainty. Denoting by Δ m the nominal inter-arrival interval of modality m and by δ t the elapsed time since its last update, the confidence weight applied to that modality’s contribution decays as ω m ( t ) = exp ( − δ t / Δ m ) , so that stale observations lose influence smoothly instead of being trusted until they are discarded. The twin simulates storm development and track, flood inundation, temperature anomaly and heat island effects, pollutant propagation, mobility disruption under extreme weather, and emergency response parameters including resource availability and population exposure.
The climate prediction engine applies learned models to s ^ t to produce probabilistic forecasts over the horizon τ . Storm classification models assign severity class distributions from observed radar patterns, flood models propagate precipitation forecasts through digital terrain models to estimate inundation probability, and air quality models project pollutant dispersal under forecast wind fields. The choice of a transformer backbone with multimodal fusion at this stage is an architectural hypothesis, and Section 8 tests it on real SEVIR data. The resulting predictive outputs augment s ^ t to produce the forward-looking state ( s ^ t , s ^ t + 1 : τ ) on which all agentic reasoning is conditioned.

7.3. Agentic AI Layer and Optimization Objective

The agentic layer supplies autonomous goal-directed reasoning over the twin state [5]. Climate response agents receive projections of the augmented state ( s ^ t , s ^ t + 1 : τ ) as local observations, select actions within their domains (issuing a storm alert, pre-positioning emergency resources, adjusting signal timing, rebalancing grid load), and announce intended actions to peers through the coordination protocol. The twin’s role is what makes the layer work: it supplies a simulation-rich synchronized environment, and the agentic layer returns adaptive decisions that the twin then absorbs through F inf in Equation (1). That return path is the agent-to-twin feedback mechanism, and it is what distinguishes this architecture from the open-loop pattern common in the corpus. At the governance level every agent action is subject to human oversight; the system is not fully autonomous, consistent with the task-level autonomy definition of Section 2.3. Grounding studies include Cao et al. [27] on hierarchical MARL for signal optimization, Tiggeloven et al. [30] on agentic pipelines for early warning, and Cho et al. [32] on AI-driven climate policy modeling.
Agents operate under centralized training with decentralized execution (CTDE). Constraint satisfaction, covering safety thresholds on emergency response violation rates and minimum service-level requirements, is enforced through Lagrangian relaxation with domain-specific multipliers ( λ Storm , λ Flood , λ Evacuation ) adapted during training to balance reward maximization against safety compliance. Emergency coordination protocols impose priority ordering over action spaces under time-critical conditions, so that life-safety objectives such as evacuation and alert issuance take precedence over efficiency objectives such as demand compliance when the environmental state is severe. An SDG-aware optimization module aggregates agent actions, evaluates joint reward, applies the weighting structure ( w 1 , w 2 , w 3 , w 4 ) reflecting institutional priorities, tracks the Pareto front in real time, and generates decision-support artifacts (predicted risk levels, recommended actions, estimated SDG contributions) for human decision-makers at the governance layer.
The cross-domain optimization objective is:
π ∗ = arg max π E ∑ t = 0 T γ t α · r SDG 13 ( s t , a t ) + β · r SDG 11 ( s t , a t ) ,
where π ∗ is the optimal cross-domain policy, γ ∈ ( 0 , 1 ] is the discount factor, r SDG 13 and r SDG 11 are reward signals for climate action and urban sustainability, and α , β ≥ 0 are weighting coefficients. As a special case of the MAS joint utility of Section 2, Equation (2) corresponds to N = 2 with λ 1 = α , λ 2 = β , u 1 = r SDG 13 , u 2 = r SDG 11 and explicit temporal discounting. Prototype operationalizations of r SDG 13 include CO2 reduction (tonnes per planning period), early warning lead time (hours ahead of threshold), renewable penetration, and adaptation plan coverage; r SDG 11 may combine traffic delay reduction, energy demand compliance, waste collection trip reduction, and air quality improvement.
Calibrating α and β is an open problem rather than an implementation detail. Fixed scalar weights are vulnerable to reward magnitude disparity and rarely reflect stakeholder trade-offs. Pareto-optimal weighting [20] identifies a set of non-dominated policies across the ( α , β ) space and leaves stakeholders to select an operating point without committing to a scalarization in advance. Lexicographic goal ordering [67] imposes a priority structure instead, so that SDG 13 resilience requirements must be satisfied before SDG 11 objectives are optimized, which prevents critical goals from being traded away for marginal gains elsewhere.

7.4. Inter-Layer Interfaces and Communication Protocol

Specifying what each layer computes is insufficient for reproducibility if the contract between layers is left implicit, which is the condition of most architectures in the corpus. Table 13 states that contract: the payload crossing each boundary, its nominal frequency, the latency budget within which it must arrive to remain useful, and the failure behavior when it does not.
Transport is publish-subscribe rather than request-response, because the layers operate at different and independently varying rates and because a blocking call from the agentic layer to the twin would couple decision latency to sensor latency. Each interface in Table 13 corresponds to a topic; producers publish on their own schedule and consumers subscribe with the stated latency budget. A broker supporting quality-of-service levels and last-will notification, of the kind standard in industrial IoT deployment, satisfies the requirement; the architecture does not depend on a specific broker implementation, but it does depend on three properties: per-message timestamping, at-least-once delivery for the governance and actuation topics, and explicit staleness signaling rather than silent reuse of old values.
Two failure modes deserve attention because they are specific to the closed loop rather than to distributed systems generally. The first is divergence between twin state and physical state when the a t feedback path degrades: the twin continues to believe resources are allocated as the agents intended while the physical actuation failed. The divergence counter in Table 13 exists for this reason, and exceeding a threshold should trigger twin resynchronization from sensor evidence alone with the infrastructure partition reset. The second is oscillation, where agents respond to a forecast that their own actions invalidate. Rate-limiting on the infrastructure partition and hysteresis on alert-level transitions bound this behavior, at the cost of some responsiveness.

7.5. Multi-Agent Coordination

The multi-objective character of climate-resilient urban management, with competing goals across transportation, energy, environmental protection, and population safety, requires domain-specific agents coordinating through structured protocols [68,69]. The layer is grounded in the corpus: Cao et al. [27] demonstrate hierarchical MARL coordination for traffic optimization, Dragomir and Dragomir [52] show decentralized hierarchical MAS for grid management, and Burger [39] models agent-based governance for equitable mobility. Coordination employs shared knowledge representations, negotiation protocols, and consensus algorithms including distributed constraint optimization (DCOP) and auction-based allocation. Concretely, a coordination round proceeds as follows: agents publish intent declarations on the peer topic; conflicts over shared resources (emergency vehicles, road capacity, grid headroom) are resolved by auction where resources are divisible and by DCOP where they are coupled through constraints; the resolved joint action is published to governance for approval and to the twin for state update. The 2 s peer latency budget in Table 13 bounds the round, and agents that fail to respond within it are excluded from that round and fall back to local policy, which trades coordination quality for bounded decision time. Figure 6b shows the full temporal sequence.

7.6. Extended SDG 13 Climate Action Integration

Three SDG 13 sub-targets map onto distinct architectural components. The twin’s disaster scenario simulation and the agentic layer’s autonomous response coordination together address SDG 13.1 on strengthening resilience and adaptive capacity to climate hazards. The governance architecture addresses SDG 13.2 on integrating climate measures into policy by keeping decision-makers in the loop, generating decision-support artifacts legible to planners, and maintaining an audit trail of agent decisions. Scenario visualization from the twin supports SDG 13.3 on education, awareness, and institutional capacity by letting policymakers explore modeled futures without specialist computational skills.
Resilience, mitigation, and adaptation are operationalized across layers. Resilience covers disaster preparedness, flood and storm robustness, and emergency adaptation, supported by the twin’s real-time probabilistic risk surfaces. Mitigation covers emission-aware optimization, renewable coordination, and resource management, supported by conditioning agent actions on carbon intensity signals in the emission reward. Adaptation covers evacuation route planning, infrastructure response sequencing, and resilience planning under uncertain futures, supported by the twin’s scenario-conditioned counterfactual simulation.
The SDG 13 reward in Equation (2) is therefore a multi-dimensional aggregation:
r SDG 13 ( s t , a t ) = w 1 r emission ( s t , a t ) + w 2 r resilience ( s t , a t ) + w 3 r warning ( s t , a t ) + w 4 r adaptation ( s t , a t ) ,
where w 1 , w 2 , w 3 , w 4 ≥ 0 satisfy ∑ k = 1 4 w k = 1 under normalized formulations. r emission captures carbon mitigation as deviation from a reference emission trajectory, operationalizing SDG 13.2. r resilience measures contribution to infrastructure robustness and population safety under hazard exposure, formulated as a monotonically decreasing function of a composite vulnerability index over flood inundation extent, structural exposure, and population density, supporting SDG 13.1. r warning rewards timely accurate alerts, increasing with the lead time at which a storm or flood threshold is predicted and actioned, which incentivizes the agentic layer to exploit the twin’s predictive capacity. r adaptation evaluates adaptive response effectiveness across evacuation routing, resource pre-positioning, and infrastructure protection, encoding SDG 13.3 planning objectives by rewarding decisions that measurably reduce expected harm under simulated disaster scenarios. Equation (3) substitutes directly into Equation (2), with the w k treated as inner-level hyperparameters subject to the same Pareto or lexicographic calibration.

7.7. Framework Alignment with SDG Targets

Table 14 maps each architectural layer to specific SDG 13 and SDG 11 targets and associated performance indicators.

7.8. Architectural Limitations and Failure Modes

The architecture is a reference specification, and four limitations bound what it can be claimed to offer. Computational scalability is the first: high-fidelity twin co-simulation of a large metropolitan area carries a heavy computational burden, and while edge-cloud orchestration mitigates it, the mitigation introduces latency trade-offs in exactly the time-critical emergency scenarios the architecture is meant to serve. Data heterogeneity and interoperability follow: legacy urban infrastructure frequently lacks standardized APIs and schemas, so a federated data layer with semantic interoperability is a precondition for instantiation rather than an enhancement [70]. Security and adversarial vulnerability are third: policies and policy inputs are susceptible to adversarial perturbation, and multi-agent systems operating over networked sensor infrastructure are exposed to spoofing and denial-of-service attacks that could corrupt actuation decisions. The message contract in Table 13 bounds the blast radius of such failures but does not prevent them. Fourth, while full city-scale deployment remains an ongoing objective for the field, Section 8 provides only a bounded real-data feasibility evaluation of the prediction component, and the multi-agent response layer remains specified but unvalidated. Longitudinal field testing against operational smart city deployments represents a natural progression for future production deployment.

8. Feasibility Study: Real-Data Demonstration on the SEVIR Archive

This section empirically probes the key architectural assumption drawn from the review: that multimodal fusion under attention-based backbones improves hazard classification over single-modality convolutional processing. The reference architecture of Section 7 relies on this mechanism at the climate prediction stage. The study below tests this assumption on real multi-sensor meteorological imagery from the Storm EVent ImageRy (SEVIR) archive [71], under a leakage-controlled, event-level nowcasting protocol with an explicit persistence reference. It is a bounded feasibility check on one architectural assumption, not a claim of operational performance, and the outcome is reported together with its statistical limits, including the finding that on real data the proposed variant only marginally exceeds a trivial persistence baseline.

8.1. Data: Real Multi-Sensor SEVIR Imagery

The data are drawn from the Storm EVent ImageRy (SEVIR) archive [71], a curated collection of spatiotemporally aligned National Weather Service storm events over the continental United States. Two complementary modalities instantiate the multimodal sensing of Section 7: NEXRAD vertically integrated liquid (VIL) as the radar channel, and two GOES-16 channels (the 6.9 μ m water-vapor band, IR069, and the 10.7 μ m clean infrared window, IR107) as the satellite channels. A bounded subset of fifty fully radar–satellite-aligned storm events (Hail and Thunderstorm Wind categories) was retrieved directly from the public archive, each event a sequence of frames at five-minute spacing. Frames were resampled to a common grid, and the data were partitioned at the level of whole events, so that no storm contributes frames to more than one of the training, validation, and test splits; this event-level partitioning is what prevents the train–test leakage that overlapping frames from a single storm would otherwise introduce.
The labeling deserves emphasis, because it is what makes the multimodal comparison fair. A severity label read from the radar field at the same instant is trivially recoverable from that field and therefore cannot test whether the satellite channels add information. The task is instead posed as nowcasting: the input is the multimodal imagery at time t, and the label is a three-class storm-intensity category (nominal, moderate, severe) derived from the moderate-VIL areal coverage approximately one hundred minutes later. Because storms grow and decay over that horizon, a single frame does not determine its own future class, so the satellite channels, whose water-vapor and cloud-top signals lead radar in developing convection, have a genuine opportunity to contribute. A no-change persistence forecast, which predicts the current intensity class as the future one, is reported throughout as the reference floor that any learned model must beat.
Four variants isolate the two factors of interest, modality and architecture, in a 2 × 2 design: Radar CNN (single modality, convolutional), Multimodal CNN (multimodal, convolutional), ViT Single (single modality, transformer), and ViT Multimodal (multimodal, transformer). All four are trained and evaluated under the identical event-level protocol of Section 8.5, with the persistence forecast as the external reference.

8.2. Processing Pipeline

The pipeline implements the sensing-to-prediction portion of the Section 7 data flow. Real SEVIR observations are read per event and time step, normalized channel-wise using training-split statistics, and fused by concatenating the VIL radar tensor with the two GOES-16 satellite channels along the channel dimension. The fused representation passes to the climate prediction engine, which emits storm severity class probabilities p t ; in the reference architecture these propagate to the twin synchronization module, which updates the hazard risk surface via F haz in Equation (1). The feasibility study evaluates this prediction stage in isolation; the downstream agentic layer that would consume p t and close the loop through F inf is specified in Section 7.3 but is not evaluated here, and its empirical validation is deferred to future work (Section 10.8).

8.3. Model Descriptions

8.3.1. Radar CNN (Baseline)

Successive convolution and pooling operations over radar reflectivity learn local spatial pattern representations. This is the single-modality convolutional reference. It lacks the global context modeling needed to reason across the full spatial extent of large storm systems.

8.3.2. Multimodal CNN (Fusion)

The baseline extended with the two GOES-16 satellite channels (water vapor, IR069; clean infrared window, IR107) concatenated to the radar channel at input, allowing convolutional layers to learn cross-modal as well as spatial relationships. The design targets cases where radar alone is insufficient, such as developing convection whose cloud-top signature precedes its VIL response, and measures the marginal value of satellite sensing within a computationally tractable architecture.

8.3.3. ViT Single (Radar Only)

A Vision Transformer over radar observations alone, partitioning the radar grid into non-overlapping patches processed by multi-head self-attention. The global receptive field captures long-range spatial dependencies such as spiral precipitation bands, stratiform-convective transitions, and mesoscale convective system extent, which locally constrained convolutional filters represent poorly. This variant isolates the contribution of attention independent of multimodal fusion.

8.3.4. ViT Multimodal (Proposed)

Global attention combined with multimodal fusion, incorporating the two GOES-16 satellite channels as additional input channels alongside the radar patches. Self-attention jointly attends to radar spatial structure and satellite context at each patch location, yielding representations that reflect both large-scale storm organization and the water-vapor and cloud-top conditions that modulate near-term evolution.

8.4. Reference Baseline

A four-way internal ablation establishes the relative contribution of modality and architecture, but it is reasonable to ask what the learned models should be measured against. For a nowcasting task the primary answer is not another supervised architecture but persistence: the no-change forecast that predicts the storm-intensity class one hundred minutes ahead to equal the current class. Persistence is the standard and most demanding reference in short-range convective nowcasting precisely because storms are strongly temporally autocorrelated; the operationally meaningful question is not whether one network outranks another by a few points but whether any learned model anticipates change that a no-change forecast misses. Reporting persistence is therefore what prevents a model that merely reproduces the current state from being mistaken for one that forecasts, and every learned variant is compared against it on the identical held-out test events. Against a static classification benchmark this reference would be trivial; against a nowcasting benchmark it is the baseline most capable of exposing an overclaim, which is why it is the appropriate one here.
Comparison against larger trained architectures (residual networks such as ResNet-18 and hierarchical transformers such as Swin-T or ConvNeXt-T) is deliberately deferred rather than reported on this subset, for a statistical reason rather than an evasive one. Those models carry roughly eleven to twenty-eight million parameters, whereas the real subset supplies on the order of a few hundred training samples drawn from fifty storm events; a from-scratch fit is thus under-constrained by several orders of magnitude and would memorize the training events rather than generalize across them. A leaderboard assembled from such fits would rank overfitting propensity, not architectural merit, and would mislead precisely the reader the comparison is meant to inform. The compact variants evaluated here are small enough to train meaningfully at this scale, and the external-baseline comparison is specified as the first step of the scaled protocol in Section 10.8, where an enlarged corpus makes it statistically defensible (Section 10.2).

8.5. Experimental Protocol and Statistical Treatment

The protocol uses ten independent seeds per configuration, each varying network initialization and augmentation while holding the event-level data partition fixed. Results report mean ± standard deviation of macro-averaged F1 across seeds on the held-out test events, with accuracy, severe-class recall, and expected calibration error (ECE) as secondary metrics. Macro-F1 is the primary metric because the future-class distribution on the small test set is imbalanced, and accuracy alone would reward a majority-class predictor; severe-class recall and ECE are reported because Section 10.8 identifies both as operationally necessary: the first because a missed severe storm is the costly error, the second because the reward term r warning consumes calibrated probabilities. Differences between the proposed variant and each internal variant are assessed by Wilcoxon signed-rank test over paired per-seed scores, with Holm correction for multiple comparisons; given the modest seed count and the small number of test events the test has limited power, so p-values are interpreted conservatively and differences that do not reach significance are described as unresolved rather than as wins. The persistence floor is a deterministic reference and is reported without a seed distribution.

8.6. Hyperparameter Specifications

Table 15 reports key hyperparameters for the four internal classification variants. Complete configuration files and implementation code are available in the repository cited in the data availability statement.

9. Feasibility Study Results

Results are reported for the storm-classification prediction stage, which is the component this feasibility study evaluates on real data. Metrics are computed on held-out test events under the event-level nowcasting protocol of Section 8.5, and every learned variant is read against the persistence floor. The multi-agent response layer of the architecture is not evaluated here; its empirical validation is deferred to future work (Section 10.8).

9.1. Storm Classification Performance

Figure 7 shows ten-seed macro-F1 across the four internal variants on the held-out SEVIR test events, and Table 16 reports the corresponding numbers against the persistence floor.
The absolute performance is far below what the same architectures reach on easy synthetic data, which is the expected consequence of moving to a real, leakage-controlled nowcasting task at a one-hundred-minute lead. The single-modality Radar CNN reaches macro-F1 0.578 ± 0.044 , the lowest of the four. Adding the satellite channels within the convolutional family lifts the Multimodal CNN only marginally, to 0.580 ± 0.097 , and with a much wider spread, indicating that convolutional fusion does not reliably exploit the satellite signal on a dataset this small. Attention helps more than convolutional fusion: ViT Single reaches 0.612 ± 0.057 on radar alone. The proposed ViT Multimodal, combining attention with satellite fusion, reaches the highest mean at 0.692 ± 0.077 (accuracy 0.762 ). The ordering Radar CNN ≲ Multimodal CNN < ViT Single < ViT Multimodal is directionally consistent with the architectural hypothesis of Section 7.2, but the margins are small relative to the seed variance.
Three points follow, and they are deliberately conservative. First, the proposed variant is the best-ranked of the four and is the only one whose advantage over the weakest variant reaches significance ( p = 0.029 versus Radar CNN); its advantage over ViT Single is not resolved at this sample size ( p = 0.160 ), so the specific claim that multimodal fusion beats attention alone is unsupported on this subset. Second, and most important, the proposed variant only marginally exceeds the persistence floor (macro-F1 0.692 versus 0.676 ; accuracy 0.762 versus 0.750 ). On real data the net value of the learned model over a no-change forecast is small, which the synthetic testbed entirely concealed. Third, the wide seed-to-seed spread, particularly for the multimodal CNN, reflects the difficulty of training from scratch on fifty events, and is itself a finding relevant to anyone instantiating the architecture on limited data. Taken together, these results support the architectural hypothesis only weakly and do not license a deployment claim.
Two secondary metrics in Table 16 qualify this picture in the proposed variant’s favor. On severe-class recall, the operationally critical quantity since a missed severe storm is the costly error, every learned model exceeds the persistence floor (0.88 to 0.96 versus 0.81), so the models add value on the severe class even where their macro-F1 does not cleanly separate from persistence. And on calibration, the proposed ViT Multimodal attains the lowest expected calibration error (0.105) of the four; this matters because the hazard classifier drives the reward term r warning in Equation (3) through its class probabilities rather than its top-1 label, so a well-ranked but poorly calibrated model (such as ViT Single here, at ECE 0.180) would misinform that reward.

9.2. Robustness Checks

Two checks probe whether the ordering in Table 16 is an artifact of the single fixed split or the single lead time.
  • Cross-validation. The ten-seed protocol varies only initialization on one event-level split, whose test set is a single partition of seven events. Rotating the split over five event-level folds (three seeds per fold) yields mean macro-F1 of 0.592 ± 0.109 (Radar CNN), 0.620 ± 0.120 (Multimodal CNN), 0.643 ± 0.165 (ViT Single), and 0.664 ± 0.124 (ViT Multimodal), against a persistence mean of 0.588 ± 0.181 . The ordering is preserved and the proposed variant’s margin over persistence widens ( + 0.076 against + 0.016 on the fixed split), but the fold-to-fold spread is large, with one fold collapsing to near 0.42 for every method, which reinforces that fifty events is too small for a tight estimate.
  • Lead-time sensitivity.Figure 8 sweeps the nowcasting lead, and the behavior is non-monotonic and physically interpretable. At a forty-minute lead the task is nearly trivial: persistence ( 0.847 ) matches the best learned model, because storms change little over that horizon. At one hundred minutes the picture inverts in the way the multimodal hypothesis predicts: both multimodal models beat persistence (Multimodal CNN 0.688 , ViT Multimodal 0.720 against 0.676 ) while both radar-only models fall below it (Radar CNN 0.530 , ViT Single 0.615 ). At the horizon where nowcasting is neither trivial nor hopeless, the satellite channels are precisely what lets the model improve on a no-change forecast. At one hundred and forty minutes every model degrades to or below persistence, the horizon having outrun the predictive content of a single frame. This sweet-spot structure is a stronger argument for multimodal fusion than any single operating point, and it is consistent with water-vapor and cloud-top signals leading radar in convective evolution.
Scaling the event count and adding trained external baselines remain the immediate steps needed to tighten the comparisons that are currently underpowered; Section 10.8 specifies that protocol.

9.3. Multi-Agent Response Layer: Status and Future Validation

The constrained multi-agent response layer specified in Section 7.3, namely the Lagrangian CTDE formulation with per-domain multipliers ( λ Storm , λ Flood , λ Evacuation ) and the SDG-decomposed reward of Equation (3), is presented in this paper as a specified design rather than as an empirically validated system. A credible evaluation of it requires two components that are outside the scope of this feasibility study: a response environment driven by real hazard dynamics rather than a hand-built simulator, and a genuine multi-agent learning implementation benchmarked against established constrained and unconstrained baselines under a multi-seed protocol. Assembling and validating that stack is a substantial undertaking in its own right, and we deliberately decline to report multi-agent performance numbers here rather than report figures from a simulator that does not yet meet that bar. The empirical contribution of this feasibility study is therefore confined to the prediction stage evaluated above. Section 10.8 specifies the protocol under which the response layer should be validated, and Section 10.2 keeps the paper’s claims within the boundary of what has actually been demonstrated.

10. Key Findings and Discussion

Across the corpus, the direction of travel is consistent: urban intelligence is moving toward integrated, autonomous, simulation-driven ecosystems. Table 17 compares representative studies, with study type reported so that the evidential weight of each entry is visible.

10.1. Emerging Trends

The literature converges on distributed urban intelligence, pairing multi-agent systems with digital twin platforms for scenario exploration, policy stress-testing, and anticipatory planning. The bibliometric profile (Table 8) shows the inflection at 2023, coinciding with the arrival of language-model-based agent frameworks. Reinforcement learning and adaptive optimization dominate in traffic control, energy load balancing, and emergency logistics. Explainability and ethical governance appear with increasing frequency, and task-level autonomy under human supervision has emerged as the prevailing design posture rather than full autonomy.

10.2. What the Evidence Demonstrates, and What It Does Not

This subsection explicitly demarcates the scope of empirical and review findings to ensure rigorous interpretation of all results.
  • Demonstrated by the review. That a distinguishable class of agentic urban and climate systems exists and can be identified reproducibly through the A/G/T/M rubric. That the literature clusters into six SDG 11 and five SDG 13 application domains, with energy and environmental monitoring densest. That publication activity rose sharply from 2023. That cross-domain orchestration spanning both goals is almost absent from the corpus. That a substantial fraction of the literature argues for agentic urban systems rather than demonstrating them (Table 5). That these findings are robust to tightening the inclusion threshold (Table 3).
  • Demonstrated by the feasibility study, within a bounded real-data scope. That the prediction stage of the architecture is implementable on real multi-sensor SEVIR imagery under a leakage-controlled, event-level nowcasting protocol. That on this fifty-event subset the proposed multimodal transformer is the best-ranked of four variants, while its advantage over an attention-only radar model is not statistically resolved and its margin over a no-change persistence baseline is small. No claim is made about the multi-agent response layer, which is specified but not empirically evaluated (Section 9.3).
  • Scope and Operational Boundaries. The findings presented reflect the systematic synthesis of published literature and a bounded real-data feasibility evaluation of the prediction stage. Operational deployment across full-scale municipal infrastructure represents a broader field-wide objective, and the benchmark results in Table 16 establish real-data prediction-stage performance within controlled experimental parameters, bounded by the persistence reference reported alongside them.
Prevalence and magnitude claims are absent from this paper by design. An under-inclusive corpus assembled under a conservative coding rule can support statements about what exists and what patterns recur; it cannot support statements about how common something is in the field as a whole.

10.3. Observed Impacts

Reported impacts cluster in operational efficiency, infrastructure responsiveness, and climate preparedness [72]. Simulation-based decision support lets planners test interventions before commitment, which reduces uncertainty in high-stakes decisions. Agent-based coordination in mobility is reported to improve traffic flow and reduce emissions [27,41], and predictive analytics in climate risk management to identify hazards earlier and support faster response [30,58]. Renewable forecasting is reported to improve grid stability and demand-aligned scheduling [59], and LLM-based multi-agent architectures are beginning to support evidence-based planning [18]. Every impact in this paragraph is a reported capability rather than an independently verified outcome, and most derive from pilot-scale or simulated deployment. Section 10.2 governs how far these should be carried.

10.4. Ethical, Social, and Governance Dimensions

Agentic systems in urban governance differ from analytical AI in a way that changes the ethical calculus rather than merely intensifying it: they act. A biased forecast misleads a planner who may catch the error; a biased agent with actuation authority reallocates resources before anyone reviews the reasoning. The dimensions below are therefore treated as design constraints on the architecture in Section 7, not as an appendix to it. Each carries genuine benefit alongside genuine risk, and treating either side alone produces bad policy.

10.4.1. Privacy and Surveillance

The sensing substrate that makes agentic urban management possible is the same infrastructure that makes pervasive surveillance possible, and the distinction is a matter of governance rather than technology. Mobility telemetry sufficient to optimize evacuation routing is sufficient to reconstruct individual movement histories; air quality sensing at building resolution reveals occupancy patterns. The benefit is real: fine-grained exposure data are what allow heat and flood interventions to reach the households most at risk rather than the households most legible to a planning department. The risk is equally real and falls unevenly, since populations with insecure housing or immigration status bear disproportionate cost from data that others treat as innocuous. Architectural mitigations exist and belong in the design rather than in the deployment agreement: spatial and temporal aggregation before the twin ingests mobility data, differential privacy on citizen-reported streams, on-edge processing so that raw records never enter the central twin, and retention limits enforced in the fusion layer of Section 7. None of these is free, and each degrades the resolution available to the agents, which is the trade-off that must be made explicitly rather than by default.

10.4.2. Algorithmic Bias and Spatial Equity

Bias in urban agentic systems is primarily spatial, which makes it harder to detect than the demographic bias that dominates fairness literature. Sensor networks are denser in wealthier districts, historical incident data reflects historical response patterns rather than historical need, and a reward function that maximizes aggregate traffic throughput will systematically favor arterial corridors over neighborhood streets. An agent trained on such data does not merely reproduce inequity; it operationalizes it at machine speed and then generates the data that confirms it. The benefit case is that explicit reward specification makes distributional choices auditable in a way that discretionary municipal decision-making never has been: the weights w k in Equation (3) state a priority ordering that can be contested. Realizing that benefit requires disaggregated evaluation as standard practice, reporting outcomes by district and by exposed population rather than in aggregate, and treating sensor coverage disparity as a measured quantity rather than an assumption. The corpus does not currently do this, and its absence is among the more consequential gaps this review identifies.

10.4.3. Accountability and Liability

When an autonomous system delays a flood warning, responsibility is genuinely unclear, and the ambiguity is structural rather than a gap in current law. Candidates include the municipality that deployed the system, the vendor that built the policy, the operator who approved the action, and the agent designer who specified the constraint threshold. Distributed multi-agent architectures worsen this, since a harmful joint action may emerge from a negotiation in which no individual agent’s behavior was faulty. The architecture’s response is the audit trail described in Section 7.6 together with the governance interface in Table 13: every joint action is logged with the state that produced it, the constraint multipliers active at the time, and the human approval record. This makes post-hoc reconstruction possible, which is a precondition for accountability but not the same thing as it. The genuine benefit is that agentic systems can be made far more auditable than the human decision processes they augment, since the reasoning is recorded by construction. The genuine risk is that auditability without an assigned locus of responsibility produces sophisticated post-hoc explanation of harms that nobody is answerable for.

10.4.4. Transparency, Explainability, and Human Oversight

Task-level autonomy under human oversight is the design posture adopted throughout this paper, and it is coherent only if the oversight is substantive. Oversight degrades in two directions. Automation bias leads operators to approve recommendations they have not evaluated, particularly under time pressure, which is precisely the condition in which emergency response operates. Alert fatigue produces the opposite failure, with operators dismissing warnings after repeated false positives. Both convert human-in-the-loop from a safeguard into a formality that transfers liability without transferring control. Mitigations are partly architectural, including the decision-support artifacts of Section 7.3 that present predicted risk and estimated SDG contribution alongside a recommendation, and partly institutional, including staffing levels adequate for genuine review and periodic audit of approval rates as an oversight-quality indicator. The EU trustworthy AI guidelines [73] treat human agency and oversight as a requirement rather than an aspiration, and an approval rate near unity should be read as evidence that oversight has failed.

10.4.5. Environmental Cost and Distributional Justice

A system justified by emissions reduction should account for its own emissions. Continuous twin co-simulation at city scale, transformer inference over multimodal streams at minute resolution, and multi-agent training are computationally substantial, and the corpus almost never reports the energy cost of the systems it proposes. Net benefit is plausible but is currently assumed rather than demonstrated, and lifecycle accounting should be a reporting requirement in this literature. The distributional dimension compounds it: computational infrastructure and the expertise to operate it concentrate in high-income cities, so the cities facing the most severe climate exposure are frequently least able to deploy these systems [19,22]. Lightweight architectures, open data standards, and federated deployment models are not secondary equity concerns but conditions for the paradigm to serve the goals it claims. The compact scale of the models evaluated in Section 9 is pertinent here, though a full efficiency case would require the parameter and energy accounting this feasibility study does not yet report.

10.5. Challenges and Structural Gaps

Several structural barriers stand between the reviewed work and large-scale implementation. Integrating heterogeneous data from legacy urban infrastructure remains difficult in the absence of widely adopted ontologies [22,70]. Scalability and energy efficiency are unresolved in high-frequency data environments. Geographic disparity complicates adoption, particularly where infrastructure and financing capacity are constrained. Table 18 summarizes the principal gaps and the opportunities they define.

10.6. Comparison with Prior Reviews

Novelty in a review is a claim about coverage and method, and it should be demonstrated rather than asserted. Table 19 positions this work against seven prior reviews spanning the four adjacent areas: SDG-oriented AI, smart cities, digital twins, and agentic or multi-agent systems.
Three differences are substantive rather than incremental. First, no prior review in this space applies an operationalized definition of agency as a screening criterion. Sharifi et al. [40] map smart city and SDG co-benefits systematically but do not distinguish agentic systems from conventional AI; Sacoto-Cabrera et al. [36] cover IoT, AI, and digital twin integration without SDG framing or agency criteria; Lee and Park [11] survey agentic architectures but without a systematic protocol. The A/G/T/M rubric with its threshold sensitivity analysis is what allows this review to state which systems are agentic and to show that the answer does not depend on where the threshold is set. Second, the four adjacent literatures have been reviewed separately but not jointly, and their intersection is where cross-domain urban–climate orchestration would live; the finding that this intersection is nearly empty is only visible from a review scoped across all four. Third, the reviews above identify architectural patterns without specifying interfaces. Section 7.4 converts the recurring pattern into a specification with a synchronization rule, a feedback path, and a message contract, which is what makes it reproducible.
The comparison also clarifies what this review does not improve upon. Its corpus is smaller than the broad AI-for-SDG surveys, its search is narrower than reviews scoped to a single technology, and its empirical component is a small real-data feasibility study where none of the comparators attempt an empirical component at all. Depth on the agentic question was purchased with breadth elsewhere.

10.7. Policy and Governance Implications

Realizing the potential identified here requires institutional readiness, regulatory clarity, and cross-sector cooperation alongside technical advance [22]. The EU High-Level Expert Group guidelines on trustworthy AI [73] supply a normative framework covering human agency, robustness, transparency, fairness, and accountability that maps directly onto the design constraints in Section 10.4, and agentic deployment in safety-critical urban services should be conditioned on demonstrated compliance rather than declared alignment. Hybrid governance models combining adaptive AI with participatory energy management and community-informed data are institutionally viable paths [29]. Regulatory frameworks should require transparency and accountability in AI-supported planning, and capacity building deserves priority, since municipal authorities cannot exercise meaningful oversight over systems they cannot interpret. Citizen sensing, community co-design, and deliberative participatory governance align with SDG 11 Target 11.3 and offer a practical route toward legitimate agentic governance [39].

10.8. Real-Data Validation: Completed Steps and Remaining Protocol

The feasibility study in Section 9 already executes the first step of a real-data validation for the prediction stage. This section states what that step establishes and specifies the audited protocol that remains.
The classification component has been partially validated. The study above uses real NEXRAD vertically integrated liquid and GOES-16 satellite imagery from the SEVIR archive [71] under an event-level split and a nowcasting label, and it confirms two things the synthetic testbed could not: that absolute performance falls sharply on real data, and that the proposed variant’s advantage is marginal once a persistence floor is reported. Three extensions remain before the variant ordering can be treated as established. First, scaling from the fifty-event subset to a corpus large enough to power the comparisons that are currently unresolved, in particular the proposed-versus-attention-only contrast. Second, adding trained external baselines (residual and modern convolutional networks and hierarchical transformers), which a fifty-event subset cannot support without overfitting. Third, hardening the partition into a temporal protocol, training on earlier seasons and testing on later ones with a second split by geographic region, and extending the metrics to per-class recall on the severe class and calibration error, since a hazard classifier that is well ranked but poorly calibrated cannot drive the reward term r warning in Equation (3).
The multi-agent component transfers less readily and should not be forced. A full real-data instantiation requires an operational city partner with actuation authority, which is a multi-year undertaking. Two prerequisites precede even a simulated evaluation: a genuine multi-agent learning implementation of the Lagrangian CTDE formulation, benchmarked against established constrained and unconstrained algorithms under a multi-seed protocol; and a response environment whose hazard stream is driven by real radar-derived severity sequences rather than a hand-built generator. Constructing that environment from the SEVIR severity sequences already assembled for the classification study, while keeping the response side simulated, is the tractable step recommended for the immediate next stage of this work.

10.9. Future Research Directions

Integrating Agentic AI with digital twin ecosystems remains the most promising direction for intelligent urban governance [66]. Emerging LLM-based orchestration platforms should be evaluated explicitly against the A/G/T/M criteria and compared with MARL architectures for urban deployment, since the two families are currently discussed in separate literatures. Six priorities follow from this review, ordered by the gap they close: scaling the real-data validation begun here, from the fifty-event meteorological feasibility subset of Section 9 to larger multi-region corpora, to real urban datasets, and to the multi-agent response layer, following the protocol in Section 10.8; disaggregated equity evaluation reporting outcomes by district and exposed population rather than in aggregate; cross-domain orchestration spanning both goals, which the corpus shows to be nearly absent; shared benchmarking protocols including standardized MARL baselines for constrained urban optimization and meteorological benchmarks for hazard classification; lifecycle energy accounting for proposed systems, so that net climate benefit is demonstrated rather than assumed; and longitudinal impact assessment across environmental, resilience, and equity dimensions. Advances in decentralized coordination and explainable agentic reasoning cut across all six.

11. Conclusions

This review examined 60 studies on Agentic AI for SDG 11 and SDG 13, screened between January 2018 and March 2026 and characterized against four operationalized criteria: autonomy, goal-directed planning, tool use, and multi-agent coordination, of which 14 satisfy the ≥2 agentic threshold and form the primary implemented evidence [23]. The methodological contribution is the rubric itself, which makes the agentic label checkable from published architecture descriptions rather than accepted on an author’s word, together with the sensitivity analysis showing that the synthesis conclusions hold under stricter thresholds.
What the corpus shows is a field expanding rapidly since 2023, concentrated in a small number of urban and climate domains, and organized around a recurring layered pattern that couples agentic reasoning to twin-based simulation. What it does not show is validated performance at city scale. A substantial share of the literature argues for agentic urban systems rather than demonstrating them, cross-domain orchestration spanning both goals is nearly absent, and evaluation practice is heterogeneous enough to preclude pooled estimation. These are the findings, and they are deliberately modest.
Two artifacts follow from the synthesis. The reference architecture of Section 7 specifies what the reviewed pattern leaves implicit: a synchronization rule in which the twin updates from agent actions as well as sensor evidence (Equation (1)), and an inter-layer message contract with stated frequencies, latency budgets, and degradation behavior (Table 13). The cross-domain objective (Equation (2)) connects the joint SDG formulation to the MAS utility and offers Pareto and lexicographic calibration in place of fixed scalar weighting. The feasibility study establishes that the prediction stage is implementable on real multi-sensor SEVIR imagery, and finds that the proposed multimodal transformer is the best-ranked of four variants but only marginally exceeds a no-change persistence baseline on a fifty-event subset; the support for the architectural assumption is real but weak, and explicitly not a demonstration of operational performance.
For practitioners, the implications differ by role. Urban planners gain a co-simulation posture for stress-testing infrastructure across mobility, energy, and flood risk, though the evidence supporting it remains largely pilot-scale. System designers gain a specification standard: the ≥2-of-4 checklist in Table 1 prevents a dashboard or rule-based scheduler from being classified as agentic, and Table 13 states what an implementation must actually provide. Policymakers should treat alignment with trustworthy AI principles [73] as a precondition for deployment in safety-critical services and should read a near-unity human approval rate as evidence that oversight has lapsed rather than that the system is performing well. For developing regions, lightweight architectures and open data standards are the condition under which this paradigm serves the 2030 Agenda [8] rather than widening the gap it is meant to close.
Three key priorities define the frontier for future research. Extending real-data validation from the fifty-event feasibility subset to multi-region operational datasets, and to the multi-agent response layer, remains the central empirical gap specified in Section 10.8. Systematic screening protocols continue to evolve as new literature emerges. And incorporating spatial equity metrics into agentic reward specifications represents a vital design imperative. Agentic AI offers a promising architectural direction for resilient, low-carbon cities; the evidence assembled here, a review synthesis together with a bounded real-data feasibility study, supports that direction more than it demonstrates operational performance, which remains the work still to be done.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/su18178917/s1, Figure S1: PRISMA 2020 flow diagram of study selection; Table S1: Two-tier database search strings and full query syntax for the five databases searched; Table S2: PRISMA screening summary with record counts at each stage; Table S3: Curated database of all 76 references and the 60 included studies, with the complete A/G/T/M agentic coding and study-type classification; Table S4: Inter-rater reliability calculations (Cohen’s κ ) for title/abstract and full-text screening; Table S5: Document type distribution statistics for the reference corpus.

Author Contributions

Conceptualization, T.A.S. and A.A.; software, A.A.; formal analysis, A.A.; validation, M.T.N. and D.H.; investigation, M.T.N. and D.H.; resources, S.K. and A.F.; data curation, S.K. and A.F.; writing—original draft preparation, A.A.; writing—review and editing, T.A.S., A.A., M.T.N., D.H., S.K. and A.F.; visualization, A.A.; supervision, S.K. and A.F.; project administration, S.K. and A.F. All authors have read and agreed to the published version of the manuscript.

Funding

This research is supported by Multimedia University (MMU) through its Article Page Charge (APC) Sponsorship Scheme.

Data Availability Statement

All data generated or analyzed during this study are publicly available in a Zenodo repository at: https://doi.org/10.5281/zenodo.21563406. The repository includes the curated study dataset, detailed search strategies, PRISMA screening records, the flow diagram, document classification statistics, inter-rater reliability calculations, the complete A/G/T/M coding and study-type classification for all included studies, model hyperparameter configuration files, and implementation code. The reproducibility repository containing the experimental implementation, reinforcement learning environment, training configuration, evaluation pipeline, and supporting source code is available at: https://github.com/aliakarma/agentic-weather-rl (accessed on 30 July 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Batty, M.; Axhausen, K.W.; Giannotti, F.; Pozdnoukhov, A.; Bazzani, A.; Wachowicz, M.; Ouzounis, G.; Portugali, Y. Smart cities of the future. Eur. Phys. J. Spec. Top. 2012, 214, 481–518. [Google Scholar] [CrossRef] [Scilit]
  2. Rolnick, D.; Donti, P.L.; Kaack, L.H.; Kochanski, K.; Lacoste, A.; Sankaran, K.; Ross, A.S.; Milojevic-Dupont, N.; Jaques, N.; Waldman-Brown, A.; et al. Tackling Climate Change with Machine Learning. ACM Comput. Surv. 2022, 55, 42. [Google Scholar] [CrossRef] [Scilit]
  3. Tiwari, A. Conceptualising the emergence of agentic urban AI: From automation to agency. Urban Inform. 2025, 4, 13. [Google Scholar] [CrossRef] [Scilit]
  4. Liu, J.; Chu, C.; Zhao, Y.; Aoki, G.; Xiao, Z. Agentic AI for sustainable development: Leveraging large language model-enhanced agent-based modeling for complex policy strategies. Emerg. Media 2025, 3, 401–413. [Google Scholar] [CrossRef] [Scilit]
  5. Liu, R.; Zhe, T.; Peng, Z.R.; Catbas, N.; Ye, X.; Wang, D.; Fu, Y. Urban planning in the age of agentic AI: Emerging paradigms and prospects. ACM SIGKDD Explor. Newsl. 2025, 27, 35–42. [Google Scholar] [CrossRef] [Scilit]
  6. United Nations. Sustainable Development Goals: Goal 11 and Goal 13; Technical Report; United Nations Department of Economic and Social Affairs: New York, NY, USA, 2024. [Google Scholar]
  7. Hossain, I.; Haque, A.K.M.M.; Rana, M.S.; Al Masud, A. Assessing urban environmental sustainability using SDG aligned indices in major city corporations of Bangladesh. Discov. Cities 2026, 3, 9. [Google Scholar] [CrossRef] [Scilit]
  8. United Nations. Transforming Our World: The 2030 Agenda for Sustainable Development; Technical Report; United Nations: New York, NY, USA, 2015. [Google Scholar]
  9. Vinuesa, R.; Azizpour, H.; Leite, I.; Balaam, M.; Dignum, V.; Domisch, S.; Felländer, A.; Langhans, S.D.; Tegmark, M.; Fuso Nerini, F. The role of artificial intelligence in achieving the Sustainable Development Goals. Nat. Commun. 2020, 11, 233. [Google Scholar] [CrossRef] [Scilit]
  10. Tiwari, A. Beyond automation: The emergence of agentic urban AI. Automation 2025, 6, 29. [Google Scholar] [CrossRef] [Scilit]
  11. Lee, Y.; Park, E. Toward sustainable agentic AI systems: A survey of architectures and methodologies. Sustain. Dev. 2026. [Google Scholar] [CrossRef] [Scilit]
  12. Wooldridge, M. An Introduction to MultiAgent Systems, 2nd ed.; John Wiley & Sons: Chichester, UK, 2009. [Google Scholar]
  13. Rao, A.S.; Georgeff, M.P. BDI Agents: From Theory to Practice. In Proceedings of the First International Conference on Multi-Agent Systems (ICMAS-95); AAAI: Washington, DC, USA, 1995; pp. 312–319. [Google Scholar]
  14. Ali, Z.A.; Zain, M.; Hasan, R.; Pathan, M.S.; Al Salman, H.; Almisned, F.A. Digital twins: Cornerstone to circular economy and sustainability goals. Environ. Dev. Sustain. 2025. [Google Scholar] [CrossRef] [Scilit]
  15. Choi, S.; Yoon, S. AI agent-based intelligent urban digital twin (I-UDT): Concept, methodology, and case studies. Smart Cities 2025, 8, 28. [Google Scholar] [CrossRef] [Scilit]
  16. Ye, X.; Du, J.; Han, Y.; Newman, G.; Retchless, D.; Zou, L.; Ham, Y.; Cai, Z. Developing human-centered urban digital twins for community infrastructure resilience: A research agenda. J. Plan. Lit. 2023, 38, 187–199. [Google Scholar] [CrossRef] [Scilit]
  17. Riaz, K.; McAfee, M.; Gharbia, S.S. Management of climate resilience: Exploring the potential of digital twin technology, 3D city modelling, and early warning systems. Sensors 2023, 23, 2659. [Google Scholar] [CrossRef] [Scilit]
  18. Syed, T.A.; Akarma, A.; Alatify, A.; Naqash, M.T.; Alqurashi, A. Agentic AI-enhanced digital twins for Smart City civil infrastructure: A secure, autonomous and auditable management framework. PLoS ONE 2026, 21, e0353610. [Google Scholar] [CrossRef] [Scilit]
  19. Syed, T.A.; Khan, S.; Jan, S.; Ali, G.; Nauman, M.; Akarma, A.; Ali, A. Agentic AI Framework for Cloudburst Prediction and Coordinated Response. arXiv 2025, arXiv:2511.22767. [Google Scholar] [CrossRef] [Scilit]
  20. Kim, I.Y.; de Weck, O.L. Adaptive weighted-sum method for bi-objective optimization: Pareto front generation. Struct. Multidiscip. Optim. 2005, 29, 149–158. [Google Scholar] [CrossRef] [Scilit]
  21. Waqar, A.; Barakat, T.A.H.; Almujibah, H.R.; Alshehri, A.M.; Alyami, H.; Alajmi, M. Analytical approach to smart and sustainable city development with IoT. Sci. Rep. 2025, 15, 23617. [Google Scholar] [CrossRef] [Scilit]
  22. Alahi, M.E.E.; Sukkuea, A.; Tina, F.W.; Nag, A.; Kurdthongmee, W.; Suwannarat, K.; Mukhopadhyay, S.C. Integration of IoT-enabled technologies and artificial intelligence (AI) for smart city scenario: Recent advancements and future trends. Sensors 2023, 23, 5206. [Google Scholar] [CrossRef] [Scilit]
  23. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit]
  24. Devane, D.; Hamel, C.; Gartlehner, G.; Nussbaumer-Streit, B.; Griebler, U.; Affengruber, L.; Saif-Ur-Rahman, K.M.; Garritty, C. Key concepts in rapid reviews: An overview. J. Clin. Epidemiol. 2024, 175, 111518. [Google Scholar] [CrossRef] [Scilit]
  25. Helms Andersen, T.; Marcussen, T.M.; Termannsen, A.D.; Lawaetz, T.W.H.; Nørgaard, O. Using artificial intelligence tools as second reviewers for data extraction in systematic reviews: A performance comparison of two AI tools against human reviewers. Cochrane Evid. Synth. Methods 2025, 3, e70036. [Google Scholar] [CrossRef] [Scilit]
  26. Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
  27. Cao, Q.; Li, J.; Trucco, P. Sustainability-oriented urban traffic system optimization through a hierarchical multi-agent deep reinforcement learning framework. Sustainability 2026, 18, 1606. [Google Scholar] [CrossRef] [Scilit]
  28. Yang, S.; Macatulad, E.G.; Biljecki, F.; Dane, G. Can urban digital twins support the realization of Sustainable Development Goal 11? Identifying key social and technical challenges. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2025, X-4/W7-2025, 129–136. [Google Scholar] [CrossRef] [Scilit]
  29. White, G.; Zink, A.; Codecá, L.; Clarke, S. A digital twin smart city for citizen feedback. Cities 2021, 110, 103064. [Google Scholar] [CrossRef] [Scilit]
  30. Tiggeloven, T.; Pfeiffer, S.; Matanó, A.; van den Homberg, M.; Thalheimer, L.; Reichstein, M.; Torresan, S. The role of artificial intelligence for early warning systems: Status, applicability, guardrails, and ways forward. iScience 2025, 28, 113689. [Google Scholar] [CrossRef] [Scilit]
  31. Algburi, S.; Al Kareem, S.S.A.; Sapaev, I.B.; Mukhitdinov, O.; Hassan, Q.; Khalaf, D.H.; Jabbar, F.I. The role of artificial intelligence in accelerating renewable energy adoption for global energy transformation. Unconv. Resour. 2025, 8, 100229. [Google Scholar] [CrossRef] [Scilit]
  32. Cho, H.; Ackom, E. Artificial intelligence (AI)-driven approach to climate action and sustainable development. Nat. Commun. 2025, 16, 1228. [Google Scholar] [CrossRef] [Scilit]
  33. Magazzino, C.; Zoundi, Z. Enhancing climate action evaluation using artificial neural networks: An analysis of SDG 13. Sustain. Futur. 2025, 9, 100439. [Google Scholar] [CrossRef] [Scilit]
  34. Villani, L.; Gugliermetti, L.; Barucco, M.A.; Cinquepalmi, F. A digital twin framework to improve urban sustainability and resiliency: The case study of Venice. Land 2025, 14, 83. [Google Scholar] [CrossRef] [Scilit]
  35. Ghaffarian, S.; Taghikhah, F.R.; Maier, H.R. Explainable artificial intelligence in disaster risk management: Achievements and prospective futures. Int. J. Disaster Risk Reduct. 2023, 98, 104123. [Google Scholar] [CrossRef] [Scilit]
  36. Sacoto-Cabrera, E.J.; Pérez-Torres, A.; Tello-Oquendo, L.; Cerrada, M. IoT, AI, and digital twins in smart cities: A systematic review for a thematic mapping and research agenda. Smart Cities 2025, 8, 175. [Google Scholar] [CrossRef] [Scilit]
  37. Korkmaz, M.; Akyildiz, Y.E.; Demirkesen, S.; Toprak, S.; Nowak, P.; Ciftci, B. A digital twin approach to sustainable disaster management: Case of Cayirova. Sustainability 2025, 17, 9626. [Google Scholar] [CrossRef] [Scilit]
  38. Vitanova, L.; Petrova-Antonova, D.; Shirinyan, E. Urban digital twin for assessing and understanding urban heat island impacts. Urban Clim. 2025, 62, 102530. [Google Scholar] [CrossRef] [Scilit]
  39. Burger, K. Towards equitable, smart, and sustainable urban mobility: Governance archetypes and their relations. Transp. Res. Part D Transp. Environ. 2025, 145, 104797. [Google Scholar] [CrossRef] [Scilit]
  40. Sharifi, A.; Allam, Z.; Bibri, S.E.; Khavarian-Garmsir, A.R. Smart cities and sustainable development goals (SDGs): A systematic literature review of co-benefits and trade-offs. Cities 2024, 146, 104659. [Google Scholar] [CrossRef] [Scilit]
  41. Louati, A.; Louati, H.; Kariri, E.; Neifar, W.; Hassan, M.K.; Khairi, M.H.H.; Farahat, M.A.; El-Hoseny, H.M. Sustainable smart cities through multi-agent reinforcement learning-based cooperative autonomous vehicles. Sustainability 2024, 16, 1779. [Google Scholar] [CrossRef] [Scilit]
  42. Khamis, A. Smart mobility education and capacity building for sustainable development: A review and case study. Sustainability 2025, 17, 7999. [Google Scholar] [CrossRef] [Scilit]
  43. Chong, Y.W.; Villanueva-Libunao, K.; Chee, S.Y.; Alvarez, M.J.; Yau, K.L.A.; Keoh, S.L. Artificial intelligence policies to enhance urban mobility in Southeast Asia. Front. Sustain. Cities 2022, 4, 824391. [Google Scholar] [CrossRef] [Scilit]
  44. Zhu, M.; Jin, J. Data-driven urban digital twins and critical infrastructure under climate change: A review of frameworks and applications. Urban Plan. 2025, 10, 10109. [Google Scholar] [CrossRef] [Scilit]
  45. Fang, B.; Yu, J.; Chen, Z.; Osman, A.I.; Farghali, M.; Ihara, I.; Hamza, E.H.; Rooney, D.W.; Yap, P.S. Artificial intelligence for waste management in smart cities: A review. Environ. Chem. Lett. 2023, 21, 1959–1989. [Google Scholar] [CrossRef] [Scilit]
  46. Belyamani, I. Artificial intelligence in waste management systems: Applications, challenges, and prospects. Waste Manag. Bull. 2025, 3, 100269. [Google Scholar] [CrossRef] [Scilit]
  47. Anitha, R.; Parthiban, A. AI-IoT-graph synergy for smart waste management: A scalable framework for predictive, resilient, and sustainable urban systems. Front. Sustain. 2025, 6, 1675021. [Google Scholar] [CrossRef] [Scilit]
  48. Patel, R.; Kimsanovich, I.A.; Waiker, V.; Muniyandy, E.; Naidu, S.M.; Mahamatov, N.; Shahin, O.R. Transformer driven multi-agent reinforcement learning framework for integrated waste classification forecasting and adaptive routing. Int. J. Adv. Comput. Sci. Appl. 2025, 16, 743–755. [Google Scholar] [CrossRef] [Scilit]
  49. Syed, T.A.; Muhammad, M.A.; AlShahrani, A.A.; Hammad, M.; Naqash, M.T. Smart water management with digital twins and multimodal transformers: A predictive approach to usage and leakage detection. Water 2024, 16, 3410. [Google Scholar] [CrossRef] [Scilit]
  50. Bibri, S.E.; Huang, J. AI and AI-powered digital twins for smart, green, and zero-energy buildings: A systematic review of leading-edge solutions for advancing environmental sustainability goals. Environ. Sci. Ecotechnol. 2025, 28, 100628. [Google Scholar] [CrossRef] [Scilit]
  51. Malik, M.M.; Altamimi, A.; Kazmi, S.A.A.; Khan, Z.A.; Ansari, M.W.; Mujahid, K.; Gao, J. A full-fledged, multi-agent system representing the architecture of smart cities by balancing energy with optimal electricity forecasting, integrating individual comfort, and extracting financial gains. IEEE Access 2024, 12, 172280–172296. [Google Scholar] [CrossRef] [Scilit]
  52. Dragomir, O.E.; Dragomir, F. A decentralized hierarchical multi-agent framework for smart grid sustainable energy management. Sustainability 2025, 17, 5423. [Google Scholar] [CrossRef] [Scilit]
  53. Khanna, A.; Srivastava, D.; Sah, A.; Dangi, S.; Sharma, A.; Tiang, S.S.; Tiang, J.J.; Lim, W.H. AI-driven multi-agent energy management for sustainable microgrids: Hybrid evolutionary optimization and blockchain-based EV scheduling. Computation 2025, 13, 256. [Google Scholar] [CrossRef] [Scilit]
  54. Jan, S.; Razzaqi, H.A.; Akarma, A.; Belgaum, M.R. A blockchain-monitored agentic AI architecture for trusted perception–reasoning–action pipelines. In Proceedings of the 2025 IEEE International Conference on Computing and Applications (ICCA); IEEE: Piscataway, NJ, USA, 2025. [Google Scholar] [CrossRef] [Scilit]
  55. Fritz, J.M.; Riebesel, L.; Xhonneux, A.; Müller, D. MASSIVE: A scalable framework for agent-based scheduling of micro-grids using market mechanisms. Energy Inform. 2025, 8, 101. [Google Scholar] [CrossRef] [Scilit]
  56. Causevic, A.; Causevic, S.; Fielding, M.; Barrott, J. Artificial intelligence for sustainability: Opportunities and risks of utilizing Earth observation technologies to protect forests. Discov. Conserv. 2024, 1, 2. [Google Scholar] [CrossRef] [Scilit]
  57. Shaamala, A.; Tilly, N.; Yigitcanlar, T. Leveraging urban AI for high-resolution urban heat mapping: Towards climate resilient cities. Environ. Plan. B Urban Anal. City Sci. 2025, 52, 2251–2266. [Google Scholar] [CrossRef] [Scilit]
  58. International Telecommunication Union (ITU); World Meteorological Organization (WMO); United Nations Office for Disaster Risk Reduction (UNDRR); International Federation of Red Cross and Red Crescent Societies (IFRC). Global: AI-Powered Early-Warning Systems Under the Early Warnings for All (EW4All) Initiative; Technical Report; United Nations Office for Disaster Risk Reduction (UNDRR): Geneva, Switzerland, 2025. [Google Scholar]
  59. El-Shabrawy, M.; Gaber, K.S.; Eid, M.M.; Alhussan, A.A.; Khafaga, D.S.; El-Kenawy, S. Explainable AI for intelligent green energy forecasting: Deep learning with iHow optimization algorithm (iHOW). Sci. Rep. 2025, 15, 41158. [Google Scholar] [CrossRef] [Scilit]
  60. Smart, E.E.; Olanrewaju, L.O.; Usman, J.; Otaru, K.; Muhammad, D.U.; Amalu, P.N.; Popoola, E.T. Artificial intelligence (AI) in renewable energy forecasting and optimization. World J. Adv. Eng. Technol. Sci. 2025, 15, 1100–1112. [Google Scholar] [CrossRef] [Scilit]
  61. Wang, Q.; Li, Y.; Pata, U.K.; Li, R. Artificial intelligence and global carbon inequality: Addressing the challenges and opportunities for SDG 10, SDG 12, and SDG 13. Geosci. Front. 2025, 16, 102072. [Google Scholar] [CrossRef] [Scilit]
  62. Hossain, M.I.; Hossan, M.R.; Shaon, Z.H.; Ferdous, M.N. Linking digital twin paradigm for urban heat monitoring and policy integration to building smart city climate resilience. Discov. Cities 2026, 3, 1. [Google Scholar] [CrossRef] [Scilit]
  63. Huzzat, A.; Anpalagan, A.; Khwaja, A.S.; Woungang, I.; Alnoman, A.A.; Pillai, A.S. A comprehensive review of digital twin technologies in smart cities. Digit. Eng. 2025, 4, 100040. [Google Scholar] [CrossRef] [Scilit]
  64. Grieves, M.; Vickers, J. Digital Twin: Mitigating Unpredictable, Undesirable Emergent Behavior in Complex Systems. In Transdisciplinary Perspectives on Complex Systems; Springer: Cham, Switzerland, 2017; pp. 85–113. [Google Scholar] [CrossRef] [Scilit]
  65. Tao, F.; Zhang, H.; Liu, A.; Nee, A.Y.C. Digital Twin in Industry: State-of-the-Art. IEEE Trans. Ind. Inform. 2019, 15, 2405–2415. [Google Scholar] [CrossRef] [Scilit]
  66. Raihan, A. Synergistic Integration of Digital Twins and Artificial Intelligence for Sustainable Energy and Environmental Systems: A Comprehensive Review. Sustain. Cities Soc. Adv. 2026, 2, 100024. [Google Scholar] [CrossRef] [Scilit]
  67. Xue, B.; Lin, X.; Zhang, X.; Zhang, Q. Multiple trade-offs: An improved approach for lexicographic linear bandits. Proc. AAAI Conf. Artif. Intell. 2025, 39, 21850–21858. [Google Scholar] [CrossRef] [Scilit]
  68. Clemen, T.; Ahmady-Moghaddam, N.; Lenfers, U.A.; Ocker, F.; Osterholz, D.; Ströbele, J.; Glake, D. Multi-agent systems and digital twins for smarter cities. In Proceedings of the 2021 ACM SIGSIM Conference on Principles of Advanced Discrete Simulation; ACM: New York, NY, USA, 2021; pp. 45–55. [Google Scholar] [CrossRef] [Scilit]
  69. Lu, Q.; Parlikad, A.K.; Woodall, P.; Ranasinghe, G.D.; Xie, X.; Liang, Z.; Konstantinou, E.; Heaton, J.; Schooling, J. Developing a digital twin at building and city levels: Case study of West Cambridge campus. J. Manag. Eng. 2020, 36, 05020004. [Google Scholar] [CrossRef] [Scilit]
  70. Tan, Y.R.; Hofmeister, M.; Phua, S.Z.; Brownbridge, G.; Rustagi, K.; Akroyd, J.; Mosbach, S.; Bhave, A.; Kraft, M. Beyond connected digital twins—Can digital twins really deliver sustainable cities? Sustain. Cities Soc. 2025, 131, 106596. [Google Scholar] [CrossRef] [Scilit]
  71. Veillette, M.; Samsi, S.; Mattioli, C. SEVIR: A Storm Event Imagery Dataset for Deep Learning Applications in Radar and Satellite Meteorology. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2020; Volume 33, pp. 22009–22019. [Google Scholar]
  72. Aleke, C.U.; Ovie, E.O.; Asere, J.B.; Oriakhi, E.A.; Igah, G.C.; Ajibade, O.E.; Egbuna, I.K. Digital twins for climate-resilient infrastructure: Simulating environmental impact on buildings. Asian J. Geogr. Res. 2025, 8, 132–141. [Google Scholar] [CrossRef] [Scilit]
  73. High-Level Expert Group on Artificial Intelligence. Ethics Guidelines for Trustworthy AI; Technical Report; European Commission: Luxembourg, 2019. [Google Scholar]
  74. Yigitcanlar, T.; Kamruzzaman, M.; Buys, L.; Ioppolo, G.; Sabatini-Marques, J.; da Costa, E.M.; Johnson, J.J. Can cities become smart without being sustainable? A systematic review of the literature. Sustain. Cities Soc. 2019, 45, 348–365. [Google Scholar] [CrossRef] [Scilit]
  75. Shahat, E.; Hyun, C.T.; Yeom, C. City digital twin potentials: A review and research agenda. Sustainability 2021, 13, 3386. [Google Scholar] [CrossRef] [Scilit]
  76. Zhou, Y.; Han, Q.; Sarabi, S.; de Vries, B. Multi-agent systems in climate-resilient land-use planning: A review. Int. J. Digit. Earth 2025, 18, 2487051. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Agentic AI ecosystem for climate-resilient cities. Urban and climate challenges (congestion, energy demand, flood risk, heat) are linked through enabling technologies (IoT, LLMs, knowledge graphs) to an agentic core, an integrated digital twin, and SDG 11/SDG 13 application domains.
Figure 1. Agentic AI ecosystem for climate-resilient cities. Urban and climate challenges (congestion, energy demand, flood risk, heat) are linked through enabling technologies (IoT, LLMs, knowledge graphs) to an agentic core, an integrated digital twin, and SDG 11/SDG 13 application domains.
Sustainability 18 08917 g001
Figure 2. Coupling of SDG 11 and SDG 13 through shared state. An orchestration layer (agentic coordination and reasoning) and a simulation layer (digital twin co-modeling) act on a shared state space of emissions, urban heat, and disaster risk, with bidirectional feedback between policy learning and state estimation.
Figure 2. Coupling of SDG 11 and SDG 13 through shared state. An orchestration layer (agentic coordination and reasoning) and a simulation layer (digital twin co-modeling) act on a shared state space of emissions, urban heat, and disaster risk, with bidirectional feedback between policy learning and state estimation.
Sustainability 18 08917 g002
Figure 3. PRISMA 2020 flow diagram of study selection [23]. Of 896 identified records, 721 were screened after deduplication, 70 underwent full-text assessment, and 60 met all criteria for the final synthesis.
Figure 3. PRISMA 2020 flow diagram of study selection [23]. Of 896 identified records, 721 were screened after deduplication, 70 underwent full-text assessment, and 60 met all criteria for the final synthesis.
Sustainability 18 08917 g003
Figure 4. Application maps for the two goals. (a) Domain-specific agents for mobility, energy, waste, infrastructure, and safety interacting through a shared urban state under centralized coordination and reward-driven learning. (b) Observation, prediction, decision, and execution layers with feedback, spanning renewable forecasting, hazard prediction, and adaptive policy optimization.
Figure 4. Application maps for the two goals. (a) Domain-specific agents for mobility, energy, waste, infrastructure, and safety interacting through a shared urban state under centralized coordination and reward-driven learning. (b) Observation, prediction, decision, and execution layers with feedback, spanning renewable forecasting, hazard prediction, and adaptive policy optimization.
Sustainability 18 08917 g004
Figure 5. From observed pattern to reference architecture. (a) Co-optimization structure abstracted from the corpus: data acquisition, state estimation, predictive modeling, policy learning, and constraint-based coordination with feedback. (b) The proposed architecture specifying the interfaces (a) leaves implicit: multimodal sensing, data fusion, climate prediction, agentic optimization, and SDG-aware orchestration under human governance.
Figure 5. From observed pattern to reference architecture. (a) Co-optimization structure abstracted from the corpus: data acquisition, state estimation, predictive modeling, policy learning, and constraint-based coordination with feedback. (b) The proposed architecture specifying the interfaces (a) leaves implicit: multimodal sensing, data fusion, climate prediction, agentic optimization, and SDG-aware orchestration under human governance.
Sustainability 18 08917 g005aSustainability 18 08917 g005b
Figure 6. Layer interaction mechanism. (a) Closed-loop control: environmental sensing updates the synchronized twin state, which feeds transformer-based forecasting and constraint-aware multi-agent orchestration under human-governed SDG objectives. (b) Temporal sequence implementing Equation (1) and the message contract specified in Section 7.4: fusion, synchronization, prediction, optimization, policy evaluation, actuation, and feedback.
Figure 6. Layer interaction mechanism. (a) Closed-loop control: environmental sensing updates the synchronized twin state, which feeds transformer-based forecasting and constraint-aware multi-agent orchestration under human-governed SDG objectives. (b) Temporal sequence implementing Equation (1) and the message contract specified in Section 7.4: fusion, synchronization, prediction, optimization, policy evaluation, actuation, and feedback.
Sustainability 18 08917 g006
Figure 7. Storm-severity nowcasting on the real SEVIR test events: macro-F1 (mean ± SD over ten seeds). The proposed ViT Multimodal variant attains the highest mean but, with a wide seed-to-seed spread, only marginally exceeds the no-change persistence forecast (Table 16).
Figure 7. Storm-severity nowcasting on the real SEVIR test events: macro-F1 (mean ± SD over ten seeds). The proposed ViT Multimodal variant attains the highest mean but, with a wide seed-to-seed spread, only marginally exceeds the no-change persistence forecast (Table 16).
Sustainability 18 08917 g007
Figure 8. Lead-time sensitivity of storm-severity nowcasting on the real SEVIR test events (macro-F1, mean ± SD over five seeds). At short lead the persistence floor (dashed) matches the best model; at the one-hundred-minute mid-horizon the multimodal models exceed persistence while radar-only models do not; at long lead all models fall to the floor.
Figure 8. Lead-time sensitivity of storm-severity nowcasting on the real SEVIR test events (macro-F1, mean ± SD over five seeds). At short lead the persistence floor (dashed) matches the best model; at the one-hundred-minute mid-horizon the multimodal models exceed persistence while radar-only models do not; at long lead all models fall to the floor.
Sustainability 18 08917 g008
Table 1. Agentic AI versus related paradigms across the four operationalized criteria. ✓ = present; ∼ = partial or situational; ✕ = not evidenced. The rubric in Table 2 specifies the evidence required for each mark.
Table 1. Agentic AI versus related paradigms across the four operationalized criteria. ✓ = present; ∼ = partial or situational; ✕ = not evidenced. The rubric in Table 2 specifies the evidence required for each mark.
ParadigmAutonomy
(A)
Goal-Directed
Planning (G)
Tool Use/
Env. Interaction (T)
Multi-Agent
Coordination (M)
Supervised/Rule-based AI✕✕✕✕
RL Control (single agent)✓∼∼✕
Classic Multi-Agent Systems✓∼✕✓
Planning Agents (STRIPS/HTN)∼✓✕✕
Agentic AI (this review)✓✓✓✓
Table 2. Coding rubric applied to every screened study. Each criterion is scored 1 or 0; evidence must appear in the architecture, experimental setup, or implementation, not in abstract claims alone, and ambiguous cases are coded 0.
Table 2. Coding rubric applied to every screened study. Each criterion is scored 1 or 0; evidence must appear in the architecture, experimental setup, or implementation, not in abstract claims alone, and ambiguous cases are coded 0.
CriterionCoded 1 When the Paper EvidencesCoded 0 Despite Superficial Resemblance
A (Autonomy)A policy or controller that executes actions without per-decision human authorization inside a stated operational scope, with a closed feedback loop from outcome to next decision.Decision-support dashboards; systems whose outputs a human must approve before every action; open-loop forecasters.
G (Goal-directed planning)Explicit optimization or search over a horizon greater than one step toward a stated objective: an RL return, a planning objective, a declared multi-step goal decomposition.Single-step input–output prediction; static optimization solved once offline; objective functions that are only loss functions for supervised fitting.
T (Tool use/env. interaction)Invocation of external APIs, simulators, sensors, retrieval systems, or actuators as part of the action set, with the returned result affecting subsequent behavior.Reading a static dataset; preprocessing pipelines; visualization of results in an external tool.
M (Multi-agent coordination)Two or more decision-making entities exchanging messages, bids, constraints, or shared state, where the joint outcome depends on the exchange.Ensembles of models; parallel independent predictors; modular pipelines whose stages do not negotiate.
Table 3. Sensitivity of the fully agentic count to the classification threshold, over the 60 included studies. The ≥2 row is the operative threshold for primary implemented evidence; synthesis conclusions drawn from the full corpus are unchanged under stricter thresholds.
Table 3. Sensitivity of the fully agentic count to the classification threshold, over the 60 included studies. The ≥2 row is the operative threshold for primary implemented evidence; synthesis conclusions drawn from the full corpus are unchanged under stricter thresholds.
ThresholdStudies Meeting ItChange vs. ≥2Effect on the Agentic Subset
≥118 + 4 Adds systems whose only agentic property is task-level autonomy; the distinction from conventional ML collapses.
≥214—Operative threshold. Requires autonomy plus at least one of planning, interaction, or coordination; these studies form the primary implemented evidence.
≥33 − 11 Retains only multi-property architectures; too few for domain-level synthesis, though the qualitative conclusions persist.
=42 − 12 Full-pattern architectures only; used in Section 7 to identify the recurring architectural pattern.
Table 4. Database search strategy and record retrieval by source. Counts reflect pre-deduplication retrieval under the two-tier query of Section 3.3.
Table 4. Database search strategy and record retrieval by source. Counts reflect pre-deduplication retrieval under the two-tier query of Section 3.3.
SourceField Code AppliedRecords Retrieved
ScopusTITLE-ABS-KEY315
Web of ScienceTS (Topic Search)228
IEEE XploreFull Text & Metadata146
SpringerLinkAll Content102
ScienceDirectAll Fields87
Backward citation trackingReference scanning of seminal studies18
Total—896
Table 5. Study-type stratification of the reference corpus, with evidential weight decreasing down the table. The Implemented, Conceptual, and Review strata form the 60 included studies; Supporting sources are cited only for framing. Percentages are of the full reference set.
Table 5. Study-type stratification of the reference corpus, with evidential weight decreasing down the table. The Implemented, Conceptual, and Review strata form the 60 included studies; Supporting sources are cited only for framing. Percentages are of the full reference set.
StratumCount%Role in Synthesis
Implemented (I)2431.6Built and evaluated systems; all performance and outcome claims trace here. The 14 that also clear the ≥2 agentic threshold (Table 6) are the primary evidence for agentic operation.
Conceptual (C)2330.3Architectural patterns and design rationale only; no outcome claims.
Review/Survey (R)1317.1Contextual positioning and gap identification; never counted as independent instances.
Supporting (S)1621.1Reporting guidelines, statistical references, normative and intergovernmental documents, foundational agent and digital-twin sources, comparative prior-review citations, and the SEVIR dataset citation used in the feasibility study.
Included (I + C + R)
Total references
60
76
78.9
100
Screened synthesis set.
Table 6. Cross-tabulation of study type against agentic intensity (the A/G/T/M count) for the 60 included studies. The fully agentic set (≥2, the ≥3 and =2 columns) totals 14, all within the Implemented stratum; no conceptual or review work clears the threshold.
Table 6. Cross-tabulation of study type against agentic intensity (the A/G/T/M count) for the 60 included studies. The fully agentic set (≥2, the ≥3 and =2 columns) totals 14, all within the Implemented stratum; no conceptual or review work clears the threshold.
Study Type≥3=2=1=0Row Total
Implemented (I)3112824
Conceptual (C)0022123
Review/Survey (R)0001313
Total31144260
Table 7. Agentic coding for 14 representative studies, per the rubric in Table 2. A = Autonomy; G = Goal-directed planning; T = Tool use or environmental interaction; M = Multi-agent coordination. ✓ = evidenced; — = not evidenced. Type: I = Implemented, C = Conceptual, R = Review.
Table 7. Agentic coding for 14 representative studies, per the rubric in Table 2. A = Autonomy; G = Goal-directed planning; T = Tool use or environmental interaction; M = Multi-agent coordination. ✓ = evidenced; — = not evidenced. Type: I = Implemented, C = Conceptual, R = Review.
StudyDomainTypeAGTMCount
Cao et al. [27]Traffic signal optimizationI✓✓✓✓4
Yang et al. [28]Infrastructure planningI✓✓✓—3
White et al. [29]Smart city citizen engagementI—✓✓—2
Tiggeloven et al. [30]Climate early warningR✓✓✓—3
Algburi et al. [31]Renewable energy AIR—✓✓—2
Cho et al. [32]Climate policy evaluationI—✓✓—2
Magazzino et al. [33]Climate action evaluationI—✓✓—2
Villani et al. [34]Urban digital twin sustainabilityI✓✓✓—3
Ghaffarian [35]Disaster risk managementR✓✓——2
Sacoto-Cabrera et al. [36]IoT–digital twin integrationR✓—✓✓3
Korkmaz [37]Resilience digital twinI✓✓✓—3
Vitanova et al. [38]Urban climate modelingI✓✓✓—3
Burger [39]Mobility governanceC—✓✓✓3
Sharifi et al. [40]Smart city–SDG synthesisR—✓—✓2
Table 8. Temporal distribution of the reference corpus by publication year. Pre-2020 sources are aggregated and are chiefly Supporting-stratum references. Cumulative percentages are rounded down.
Table 8. Temporal distribution of the reference corpus by publication year. Pre-2020 sources are aggregated and are chiefly Supporting-stratum references. Cumulative percentages are rounded down.
YearBefore 20202020202120222023202420252026Total
Studies103425739676
Cumul. %13172225314092100—
Table 9. Distribution of publications by major publisher category; minor publishers are aggregated under “Others”.
Table 9. Distribution of publications by major publisher category; minor publishers are aggregated under “Others”.
Publisher CategoryCountPercentageExample Venues
Elsevier1621.1%Cities, iScience, Sustainable Cities and Society
MDPI1418.4%Sustainability, Sensors, Smart Cities
Springer/Springer Nature1317.1%Nature Communications, npj Urban Sustainability, Transdisciplinary Perspectives
Other Academic Publishers2735.5%ACM, SAGE, IEEE, NeurIPS, Taylor & Francis, Frontiers, Wiley, etc.
Technical Reports (UN/Intl.)33.9%UN SDGs, UNDRR EW4All
Independent/Misc. Journals33.9%WJAETS, EJSMT, AJGR
Total76100%
Table 10. Distribution of publications by document type, from the submitted reference list.
Table 10. Distribution of publications by document type, from the submitted reference list.
Document TypeBibTeX TypeCountPercentage
Journal Articles@article6585.5%
Conference Papers@inproceedings56.6%
Technical Reports@techreport45.3%
Books/Book Chapters@book/@incollection22.6%
Total 76100%
Table 11. Representative Agentic AI applications in smart mobility (SDG 11). Criteria (A/G/T/M) follow Table 7; Type: I = Implemented, C = Conceptual, R = Review. For C and R entries, “Key Outcome” is an argued position rather than a measured result.
Table 11. Representative Agentic AI applications in smart mobility (SDG 11). Criteria (A/G/T/M) follow Table 7; Type: I = Implemented, C = Conceptual, R = Review. For C and R entries, “Key Outcome” is an argued position rather than a measured result.
StudyTypeAI ParadigmUrban ContextKey OutcomeLimitation
Cao et al. [27]IHierarchical MARL (A,G,T,M)Urban traffic signal controlSustainability-oriented traffic optimizationSimulation-based; real-world validation needed
Khamis [42]RMaaS integration AI (G,T,M)Smart transit planningImproved modal shift equityLimited rural applicability
Burger [39]CAgent-based governance (G,T,M)Policy simulationEquitable mobility archetypesNormative framing required
Chong et al. [43]IAI policy analysis (G,T)Southeast Asian citiesEnhanced policy alignmentCross-context generalizability
Table 12. Representative urban–climate co-simulation scenarios and the planning outcomes they support.
Table 12. Representative urban–climate co-simulation scenarios and the planning outcomes they support.
ScenarioSimulation FocusPlanning Outcome
Urban SprawlTraffic + Emissions + HeatwavesIdentification of high-risk urban heat zones; targeted cooling intervention strategies
Renewable IntegrationEnergy Demand + Climate VariabilityOptimal spatial allocation of storage assets and smart grid scheduling
Disaster PreparednessFlood + Storm + Population DensityEmergency response prioritization and pre-positioned resource allocation
Green InfrastructureLand Cover + Urban Temperature + RunoffCost-benefit ranking of nature-based adaptation interventions
Table 13. Inter-layer message contract for the reference architecture. Frequencies are nominal city-scale design targets; latency budgets give the interval beyond which a message is treated as stale; degradation behavior specifies the required response to a missed or late message.
Table 13. Inter-layer message contract for the reference architecture. Frequencies are nominal city-scale design targets; latency budgets give the interval beyond which a message is treated as stale; degradation behavior specifies the required response to a missed or late message.
InterfacePayloadFrequencyLatency BudgetDegradation Behavior
Sensing → FusionPer-modality raw records with timestamp and sensor IDModality-specific
(1 Hz to 15 min)
—Modality flagged missing; imputation invoked
Fusion → TwinObservation tensor o t with per-modality confidence ω m ( t ) 1/min10 sTwin propagates via T θ ; confidence decays per ω m
Twin → PredictionSynchronized state s ^ t 1/min5 sLast valid s ^ reused; staleness flag set
Prediction → TwinHazard probability vector p t over τ 1/min20 sHazard surface held; forecast horizon truncated
Twin → AgentsAugmented state ( s ^ t , s ^ t + 1 : τ ) , per-agent local projection1/min5 sAgent falls back to reactive single-step policy
Agent ↔ AgentIntent declarations, bids, constraint messages (DCOP/auction)Event-driven2 sNon-responding agent excluded from round; local policy applied
Agents → GovernanceJoint action a t , predicted risk, SDG contribution estimatesPer decision1 sAction withheld pending human review
Agents → TwinExecuted joint action a t (feeds F inf , Equation (1))Per decision5 sInfrastructure partition not advanced; divergence counter incremented
Table 14. Mapping of the proposed framework layers to SDG 13 and SDG 11 targets and associated performance indicators.
Table 14. Mapping of the proposed framework layers to SDG 13 and SDG 11 targets and associated performance indicators.
Framework LayerPrimary FunctionSDG 13 ContributionSDG 11 Contribution
Data AcquisitionReal-time sensing and multi-source integrationCity-scale climate monitoringInfrastructure efficiency monitoring
Digital TwinScenario simulation and stress testingRisk prediction and adaptationUrban planning and resilience
Agentic AIAutonomous goal-directed coordinationDisaster early warningSmart mobility optimization
Multi-Agent LayerDistributed resource allocation and negotiationEmergency managementPublic safety and equity
Table 15. Hyperparameters for the four storm-classification variants on the real SEVIR subset. ViT variants use a compact hybrid convolutional–transformer configuration matched to the 32 × 32 grid. Seeds follow the fixed list of Section 8.5.
Table 15. Hyperparameters for the four storm-classification variants on the real SEVIR subset. ViT variants use a compact hybrid convolutional–transformer configuration matched to the 32 × 32 grid. Seeds follow the fixed list of Section 8.5.
ParameterRadar/Multimodal CNNViT Single/ViT MultimodalNote
Input channels1 / 31 / 3radar (VIL); + IR069, IR107 (GOES-16)
Input grid32 × 3232 × 32resampled SEVIR frames
Architecture3 conv blocks + BNconv stem + 2–3 transformer layerspatch tokens 8 × 8
Attention headsN/A4embedding dim 128
OptimizerAdamWAdamWlr = 1 × 10−3, weight decay 10−3
LR schedulecosine annealcosine anneal—
Batch size3232—
Training epochs1515best-validation checkpoint
Loss functionCross-entropyCross-entropy—
Random seeds { 13 , 42 , 77 , 101 , 256 , 512 , 777 , 1024 , 1337 , 2024 } 10 independent runs
Table 16. Storm-severity nowcasting across ten seeds on the real SEVIR test events (fifty aligned events, event-level split, one-hundred-minute lead). Macro-F1 and accuracy are mean ± SD; severe recall and ECE (expected calibration error, lower is better) are means over seeds. p-values are Holm-corrected Wilcoxon signed-rank against the proposed variant. Persistence is a deterministic no-change reference floor (ECE undefined). Bold marks the best trained variant per column.
Table 16. Storm-severity nowcasting across ten seeds on the real SEVIR test events (fifty aligned events, event-level split, one-hundred-minute lead). Macro-F1 and accuracy are mean ± SD; severe recall and ECE (expected calibration error, lower is better) are means over seeds. p-values are Holm-corrected Wilcoxon signed-rank against the proposed variant. Persistence is a deterministic no-change reference floor (ECE undefined). Bold marks the best trained variant per column.
ModelModalityMacro-F1AccuracySevere RecallECEp
Radar CNN (baseline)Radar 0.578 ± 0.044 0.625 ± 0.056 0.881 0.134 0.029
Multimodal CNN (fusion)Multimodal 0.580 ± 0.097 0.682 ± 0.049 0.962 0.136 0.074
ViT Single (radar only)Radar 0.612 ± 0.057 0.657 ± 0.072 0.931 0.180 0.160
ViT Multimodal (proposed)Multimodal0.692 ± 0.0770.762 ± 0.045 0.919 0.105 —
Persistence (no-change)reference 0.676 0.750 0.812 ——
Table 17. Comparison of representative studies on Agentic AI for sustainable, climate-resilient cities. Criteria follow Table 7; Type: I = Implemented, C = Conceptual, R = Review. For R and C entries, reported outcomes are argued positions rather than measured results.
Table 17. Comparison of representative studies on Agentic AI for sustainable, climate-resilient cities. Criteria follow Table 7; Type: I = Implemented, C = Conceptual, R = Review. For R and C entries, reported outcomes are argued positions rather than measured results.
StudyTypeAI ParadigmDigital TwinDomainStrengthLimitation
Lee et al. [11]RAgentic AI surveyPartialSustainability architecturesComprehensive architecture taxonomySurvey; no experimental validation
Yang et al. [28]IAgent-based DT (A,G,T)YesInfrastructure planningSDG 11 target alignmentHigh infrastructure cost
Tiggeloven et al. [30]RDeep learning EWS (G,T)NoClimate early warningBroad status assessmentLimited interpretability; not an implemented system
White et al. [29]IDT citizen platform (G,T)YesCitizen engagementParticipatory DT governanceLimited autonomous decision-making
Algburi et al. [31]REnergy-AI review (G,T)NoRenewable energy adoptionBroad policy and technology coverageReview; no empirical system
Cho et al. [32]IAI policy model (G,T)PartialClimate policy evaluationSDG interlinkage mappingCausal inference limitations
Table 18. Research gaps and future opportunities in Agentic AI for urban sustainability and climate resilience.
Table 18. Research gaps and future opportunities in Agentic AI for urban sustainability and climate resilience.
Research AreaIdentified GapFuture Opportunity
Smart MobilityIsolated domain optimization modelsIntegrated multi-agent orchestration across transport modes
Climate ForecastingLimited policy feedback linkageAI-driven policy simulation with causal inference
Digital TwinsHigh infrastructure and data costScalable federated cloud-based twin architectures
Urban GovernanceAbsence of operational ethical AI frameworksResponsible AI governance with participatory design
Cross-domain AIFragmented single-domain deploymentsUnified urban intelligence platforms spanning multiple SDGs
Equity and AccessAI concentrated in high-income citiesLightweight architectures for resource-constrained regions
Evaluation PracticeIncommensurable baselines and metricsShared benchmarks for constrained urban optimization
Table 19. Structured comparison with prior reviews across the four adjacent literatures, along six dimensions: agentic criteria, digital twins (DT), multi-agent coordination (MAS), SDG framing, documented protocol with inter-rater agreement ( κ ), and an empirical component. ✓ = present; ∼ = partial; ✕ = absent.
Table 19. Structured comparison with prior reviews across the four adjacent literatures, along six dimensions: agentic criteria, digital twins (DT), multi-agent coordination (MAS), SDG framing, documented protocol with inter-rater agreement ( κ ), and an empirical component. ✓ = present; ∼ = partial; ✕ = absent.
ReviewYearPrimary ScopeAgentic CriteriaDTMASSDG FramingProtocol + κ Empirical
Vinuesa et al. [9]2020AI across all 17 SDGs✕✕✕✓∼✕
Rolnick et al. [2]2022ML for climate change✕∼∼∼✕✕
Yigitcanlar et al. [74]2019Smart city sustainability✕✕✕∼✓✕
Shahat et al. [75]2021City digital twins✕✓✕✕∼✕
Sharifi et al. [40]2024Smart cities and SDGs✕✕∼✓✓✕
Sacoto-Cabrera et al. [36]2025IoT, AI, DT in smart cities✕✓∼✕✓✕
Huzzat et al. [63]2025DT technologies in cities✕✓✕✕∼✕
Zhou et al. [76]2025MAS in land-use planning∼∼✓∼∼✕
Lee and Park [11]2026Agentic AI architectures∼∼✓∼✕✕
This review2026Agentic AI for SDG 11 + 13✓✓✓✓✓∼
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Syed, T.A.; Akarma, A.; Naqash, M.T.; Hameed, D.; Kamal, S.; Formisano, A. Agentic AI for Climate-Resilient Cities: A PRISMA-Guided Review and Digital Twin Framework. Sustainability 2026, 18, 8917. https://doi.org/10.3390/su18178917

AMA Style

Syed TA, Akarma A, Naqash MT, Hameed D, Kamal S, Formisano A. Agentic AI for Climate-Resilient Cities: A PRISMA-Guided Review and Digital Twin Framework. Sustainability. 2026; 18(17):8917. https://doi.org/10.3390/su18178917

Chicago/Turabian Style

Syed, Toqeer Ali, Ali Akarma, Muhammad Tayyab Naqash, Danial Hameed, Shahid Kamal, and Antonio Formisano. 2026. "Agentic AI for Climate-Resilient Cities: A PRISMA-Guided Review and Digital Twin Framework" Sustainability 18, no. 17: 8917. https://doi.org/10.3390/su18178917

APA Style

Syed, T. A., Akarma, A., Naqash, M. T., Hameed, D., Kamal, S., & Formisano, A. (2026). Agentic AI for Climate-Resilient Cities: A PRISMA-Guided Review and Digital Twin Framework. Sustainability, 18(17), 8917. https://doi.org/10.3390/su18178917

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop