1. Introduction
Cities are increasingly on the front line of compounding shocks and stresses, including heatwaves, flooding, bushfires, infrastructure failures, and cascading disruptions across interconnected systems. Synthesised evidence shows that climate risks are intensifying and that adaptation and resilience-building must move rapidly from aspiration to implementation, particularly at the urban and local government levels [
1]. At the same time, global resilience agendas, such as disaster risk reduction and “build back better” commitments, increasingly emphasise governance, investment choices, and institutional capacity rather than focusing solely on individual projects [
2]. Yet “urban resilience” remains contested and inconsistently operationalised, making it difficult for Australian local governments to translate resilience intent into clear, repeatable, and accountable delivery routines [
3].
Over the past decade, smart city technologies, including sensors, platforms, analytics, dashboards, and digital twins, have been promoted as a route to more efficient, sustainable, and “data-driven” urban management. However, the smart city agenda has attracted sustained criticism: it can become overly technology-led, fragmented across departments and vendors, and weakly connected to public value and long-term outcomes [
4]. In practice, many initiatives are implemented as discrete technology deployments rather than as part of a coherent governance pathway that links investment choices, service delivery, and resilience outcomes over time.
A related implementation failure is the “pilot-forever” pattern: cities initiate demonstrations, living labs, and experimental projects but struggle to scale successful initiatives, discontinue unsuccessful ones, and embed learning into routine decision-making. Urban experimentation is widely documented as a dominant mode of innovation in cities, yet the governance arrangements required to convert experiments into durable systems, standards, funding mechanisms, accountability, and capability are often under-specified [
5]. Where initiatives scale, they may do so unevenly, reinforcing spatial and social inequities or creating new risks related to surveillance, data ownership, and platform dependence. This is not unique to Australian local government: internationally, cities have often used living labs and pilot projects to trial digital innovation, yet have struggled to secure the governance alignment, interoperability, procurement pathways, and long-term operational funding needed for wider institutional uptake [
6,
7].
These concerns sharpen when equity and legitimacy are treated as add-ons rather than core design requirements. Critical smart city scholarship shows that participation and “citizen-focused” narratives can mask unequal power, uneven benefits, and exclusion, particularly when platformisation and data-driven governance are not matched by robust public safeguards [
8]. Similarly, policy work highlights recurring gaps in data governance, resources, and institutional capacity that can hinder safe scaling and restrict smart initiatives to small, non-transferable pilots [
9].
Together, these dynamics reveal a clear gap in both research and practice: a shortage of integrated, governance-ready frameworks that (i) sequence smart city efforts into a realistic delivery pathway, (ii) align technology choices with budgeting, asset management, procurement, and service delivery routines, (iii) treat equity, ethics, and data governance as core, non-negotiable design and governance requirements, and (iv) make evaluation and learning continuous and decision-relevant rather than post hoc. While collaborative and multi-actor governance models provide useful foundations for understanding how public agencies coordinate across boundaries, they rarely translate into step-by-step implementation architectures that specify decision points, evidence requirements, and the conditions for scaling [
10].
In this paper, urban resilience is operationalised through a service continuity lens: the capacity of local governments and their partners to maintain or rapidly restore essential services, such as local infrastructure, drainage, waste, and community-facing operations, under shock, stress, and disruption. This lens is important because it shifts the focus from technology uptake alone to whether digital interventions strengthen governable, equitable, and accountable continuity in public service delivery.
This paper addresses that gap by introducing the Smart Urban Resilience Framework (SURF), a phase-gated, tier-aware governance architecture for institutionalising smart urban resilience in local government. The SURF is not proposed as another resilience indicator set, smart city maturity model, or pilot-learning framework. Rather, it is presented as an institutionalisation model for how smart resilience initiatives can move from experimentation to business-as-usual under public sector conditions. In this sense, the SURF bridges broad normative claims about ‘smart’ or ‘resilient’ cities and the practical realities of project delivery in local government. It does so by linking public sector decision cycles, minimum evidence gates, concurrent action tracks, and tier-aware implementation pathways into a practical governance architecture for moving from fragmented pilots to embedded and measurable resilience practice.
Against this background, this paper makes three linked contributions. First, it offers a theoretical contribution by conceptualising smart urban resilience as a problem of institutionalisation, showing how digital initiatives become resilience-relevant only when they are embedded in governance, budgeting, operational integration, and learning routines rather than treated as stand-alone pilots. Second, it makes a methodological contribution by providing an auditable evidence-to-design logic that translates findings from the systematic review, tiered comparative analysis, Sydney case study, and stakeholder input into specific framework components. Third, it makes a practical contribution by specifying phase logic, gate criteria, action tracks, and evidence pack tools that can support more transparent decisions to fund, scale, refine, or retire initiatives.
Positioning SURF: Theoretical Contribution and Distinction from Existing Approaches
“Smart resilience” practice is very often assembled from separate policy areas and toolkits, reflecting its cross-sectoral bases. Accordingly, the SURF is positioned here against four common areas of work that shape how councils design and govern initiatives for “smart resilience” (
Table 1). These are (i) resilience frameworks, (ii) smart city maturity models, (iii) data governance, and (iv) living labs for experimentation and innovation. Resilience frameworks and indicator sets help define what resilience means and what should be measured, but they often under-specify how local governments move from planning to implementation, particularly how evidence informs decisions to fund, scale, refine, or discontinue initiatives. Smart city maturity models and capability roadmaps assist cities in diagnosing digital preparedness, although these can drift towards “tech maturity” and remain weakly connected to resilience outcomes, equity safeguards, legitimacy, and long-term operational sustainability. Data governance and ethics guidance provide essential principles (privacy, security, stewardship, transparency) and role clarity, but they often lack decision processes, evidence standards, and checkpoints that prevent premature lock-in. Finally, the living labs and experimentation literature shows how pilots can foster learning under uncertainty, yet it often leaves unresolved the transition from pilots to business-as-usual operations, procurement, sustainable funding, and demonstrable results.
The SURF builds on these diverse fields by integrating them within a single operational delivery architecture. It provides a phase-gated process aligned with public sector decision schedules; specifies two evidence gates that discipline investment decisions (Validate & Fund; Scale/Refine/Retire); organises delivery through four concurrent action tracks so that governance, piloting, integration, and learning progress together; and incorporates tier-aware guidance to support both high- and lower-capacity councils through minimum viable foundations, shared services, and staged scaling options.
Taken together, these comparisons position the SURF as an institutionalisation framework rather than an assessment framework, capability maturity model, or experimentation toolkit. Its central concern is not simply what to measure, how digitally mature a council is, or how to run pilots, but how public organisations convert promising digital initiatives into durable, governable, and accountable resilience practice over time.
Section 2 outlines the conceptual foundations that underpin these design choices and explains why this governance architecture is needed for smart urban resilience.
3. Materials and Methods
This section describes the methodological approach used to develop the Smart Urban Resilience Framework (SURF). It details the overall study design, the evidence base and data sources, the analytical steps used to convert evidence into framework components, and the procedures adopted to enhance trustworthiness, expert review and iterative refinement, and ethical rigour.
3.1. Study Design and Overall Approach
The SURF was developed through an evidence-informed qualitative design that treats smart urban resilience as a socio-technical and governance challenge, rather than a purely technical optimisation problem. The study used a multi-source case strategy that combined: (i) a PRISMA-guided systematic literature review (SLR) supplemented with a targeted synthesis of practice-oriented frameworks and guidance; (ii) a comparative Tier 1–Tier 2 document analysis; (iii) an in-depth Sydney case study; and (iv) semi-structured interviews with stakeholders involved in planning, resilience, and technology delivery. This design is well suited to examining complex, context-dependent “how” and “why” questions and tracing how governance arrangements, institutional capacity, and implementation routines shape outcomes over time [
21]. Rather than treating each evidence stream as a standalone study, the methods were structured as an integrated evidence-to-design pipeline: evidence was synthesised into themes and design requirements, translated into SURF components, and then refined through expert feedback to strengthen usability, clarity, and decision relevance.
Throughout the evidence-to-design process, service continuity was treated as the organising resilience outcome, allowing the analysis to focus on how digital initiatives could be embedded within governance, planning, and delivery systems to maintain or restore essential local government functions under stress.
The empirical and conceptual inputs to the SURF were drawn from four linked evidence streams: the systematic review, the Tier 1–Tier 2 comparative analysis, the Sydney case study, and stakeholder interviews. The article’s original contribution lies not in re-reporting these streams as standalone studies, but in showing how they were synthesised, translated, and operationalised into an auditable framework construction process and its resulting governance architecture. In this sense, the article’s contribution is the derivation of SURF itself: a phase-gated, tier-aware model for institutionalising smart urban resilience through explicit decision points, concurrent action tracks, enabling layers, and evidence-based scaling logic.
Relationship to Prior Publications
The four evidence streams informing the SURF are documented in companion outputs with distinct purposes. The systematic review established the conceptual and empirical links between digital technologies, governance, and resilience capacities [
22]. The Tier 1–Tier 2 comparison examined how governance conditions, resourcing, and institutional variation shape implementation possibilities across Australian local government contexts [
23]. The Sydney case study examined how these dynamics play out in a detailed metropolitan setting [
24]. Building on these inputs, the present article contributes the cross-stream synthesis and framework-building process through which this evidence base was translated into SURF’s phases, gates, action tracks, enabling layers, and implementation tools.
3.2. Evidence Base and Data Sources
- (1)
Systematic literature review and targeted synthesis
The academic evidence base was established through a systematic literature review reported in accordance with PRISMA 2020 [
25]. Searches were conducted in Web of Science and Scopus, supplemented by manual searches (e.g., reference list checks) to reduce database- and indexing-related bias. Screening progressed from title/abstract review to full-text assessment, resulting in a final corpus of 115 articles.
To ensure the synthesis was implementation-relevant, the SLR used a hybrid analytical approach: (i) qualitative thematic synthesis in NVivo 14 and (ii) bibliometric keyword co-occurrence mapping in VOSviewer 1.6.20 to identify dominant clusters and research fronts [
26,
27].
In addition, a targeted synthesis incorporated practice-oriented frameworks, standards, and guidance that are often underrepresented in peer-reviewed collections but strongly shape how councils and governments operationalise “smart” and “resilient” agendas (e.g., data governance guidance, digital ethics and accountability, resilience strategy tools, and indicator frameworks). This targeted synthesis was used specifically to translate academic themes into governance and delivery design choices aligned with Australian institutional contexts.
The full SLR methods and results are reported in the companion review [
22]; here, the review’s synthesised themes are re-expressed as implementation design requirements for SURF.
- (2)
Comparative Tier 1–Tier 2 document analysis
A structured document analysis was conducted to compare Tier 1 and Tier 2 Australian cities to assess how institutional capacity and governance influence possible implementation pathways. This review examined 22 official strategy documents from Tier 1 cities (Sydney, Melbourne, Brisbane, Adelaide) and Tier 2 cities (Geelong, Newcastle, Hobart, Sunshine Coast). The analysis employed a two-phase qualitative method: (i) profiling each city and (ii) identifying common patterns across cities, such as technological integration, community engagement, resilience goals, and funding approaches.
Documents were drawn from official strategies, council reports, policy documents, and publicly accessible plans related to digital innovation and resilience; where available, supporting materials (e.g., meeting minutes, media releases, publicly available indicators) were used to contextualise and cross-check claims [
28].
The Tier 1–Tier 2 comparison (22 official strategy documents; two-phase qualitative analysis) is reported in [
23]. In this paper, the tiered findings inform the SURF’s tier-aware guidance, feasibility assumptions, and scaling expectations.
- (3)
Sydney case study
Sydney served as a high-capacity (Tier 1) case to explore how smart-enabled resilience is put into practice amid complex multi-level governance, enhanced institutional capacity, and extensive digital programmes. The study employed a phased approach involving document review and project analysis, which included (i) evaluating two key strategies, the Smart City Strategic Framework and the Resilient Sydney Strategy; (ii) providing a typology of 36 smart city initiatives carried out from 2015 to 2023; and (iii) conducting detailed examinations of three major initiatives, NSW Spatial Digital Twin, Land iQ, and SIMPaCT, to understand their delivery methods, partnerships, governance conditions, and contributions to resilience.
The Sydney case analysis (including an assessment of key strategies and major initiatives such as the NSW Spatial Digital Twin, Land iQ, and SIMPaCT) is reported in [
24]. Here, it is used to stress-test governance and scaling conditions and to calibrate the SURF’s gate-evidence expectations.
- (4)
Semi-structured interviews
Semi-structured interviews elicited practitioner accounts of decision-making, barriers to scaling beyond pilots, institutional incentives, and the feasibility of alternative governance arrangements. Interviews were conducted between July and November 2023 with stakeholders across Tier 1 and Tier 2 settings. The interview sample comprised 10 participants in decision-making or advisory roles related to local governance, technology implementation, and urban resilience.
Interviews followed an open-ended guide and typically lasted ~45 min. They were conducted online, audio-recorded with consent, transcribed, de-identified, and securely stored. Participants were also provided with a short summary of key points to confirm accuracy prior to final coding (member checking) [
29]. Recruitment used a purposeful sampling strategy focused on participants with direct experience in smart city and resilience implementation and evaluation [
30].
The interview sample was intentionally oriented towards strategic and advisory perspectives rather than frontline operational roles. It included participants across Tier 1 and Tier 2 contexts, spanning cities such as Sydney, Melbourne, Geelong, and Hobart, and covered a mix of policy, planning, industry, and community-facing expertise. In the broader study design, this role mix included policymakers, urban planners/resilience specialists, private-sector technology providers, and a community-focused NGO representative. This composition strengthened the study’s insight into governance, implementation, and scaling conditions, but it also limited the representativeness of the interview evidence for frontline operational practice and resident experience.
To protect participant confidentiality in a small, role-specific sample spanning identifiable city contexts, the manuscript does not provide a more detailed tier-by-tier disaggregation of interviewees. The sample nevertheless covered both Tier 1 and Tier 2 settings and was used primarily to inform a governance-oriented synthesis rather than for statistical comparison.
Table 2 summarises the distinct analytical contribution of each evidence stream to SURF development. Presenting these source-specific findings separately clarifies how the framework was grounded in multiple forms of evidence before those findings were consolidated through cross-stream synthesis.
3.3. Data Handling, Cross-Stream Synthesis, and Framework Derivation
All textual materials, including documents and interview transcripts or notes, were coded through thematic analysis to identify recurring constraints, enabling conditions, and decision requirements [
26]. Coding was supported by qualitative data software where appropriate, to maintain a transparent audit trail, manage cross-case retrieval, and run structured comparison queries [
31]. As summarised in
Table 2, each evidence stream contributed distinct findings relevant to SURF development.
Coding was led by the first author in NVivo, using an evolving codebook that combined deductive domains derived from the study objectives with inductive refinement as new sub-themes emerged. Coding proceeded at the level of meaningful segments (sentences and paragraphs) across documents and interview materials. Analytic memos were maintained throughout to document coding decisions, emerging interpretations, and the rationale for merging, splitting, or refining codes. Because the study employed a reflexive thematic and cross-stream synthesis approach, formal intercoder reliability statistics were not used; instead, consistency was supported through iterative codebook refinement, memo-based audit trails, and repeated comparison across evidence streams.
Appendix A Table A1 summarises the step-by-step transformation of coded evidence into synthesis themes, requirement statements, and SURF constructs.
Beyond identifying themes within individual evidence streams, the analysis was designed to make the progression from evidence to framework design explicit and traceable. Cross-stream synthesis therefore proceeded by consolidating overlapping and complementary findings across the systematic review, Tier 1–Tier 2 comparison, Sydney case evidence, and interviews. Recurring issues were then translated into requirement statements specifying what a council-ready smart urban resilience framework would need to enable in practice. These requirements were written to be implementation-relevant rather than purely conceptual. Where the evidence indicated context-specific differences in capability, governance complexity, or feasible delivery pathways, the requirements were retained as conditional or tier-aware. The resulting requirement statements were then organised into SURF’s major structural elements: phases and evidence gates, concurrent action tracks, enabling layers, and tier-aware implementation pathways. This process provided the main audit trail linking evidence source, synthesis theme, design requirement, and framework feature.
Analysis proceeded through four linked steps:
Within-stream coding: Each evidence stream (SLR, tier documents, Sydney case evidence, interviews) was coded to identify themes relevant to implementation and governance, including data governance, capability, funding, equity, and evaluation.
Cross-stream synthesis: Themes were consolidated into a cross-source evidence matrix linking (a) the recurrent issue, (b) supporting evidence by source, (c) how the issue manifested across tiers and in Sydney, and (d) its implications for the SURF.
Translation into design requirements: Consolidated issues and themes were translated into explicit requirement statements specifying what a council-ready smart urban resilience framework would need to enable in practice.
Framework construction: Requirement statements were operationalised into the SURF’s phases and gates, concurrent action tracks, enabling layers, tier-aware pathways, and associated evidence expectations, using a case-study pattern to locate decision points and learning loops [
21].
In practical terms, the translation from evidence to framework design followed the logic of recurrent implementation problems. For example, evidence of fragmented pilots, weak monitoring, and limited post-pilot learning informed the inclusion of Gate 2 and the Learning, Evaluation and Impact track. Recurring concerns about unclear ownership, sponsorship, and cross-agency coordination informed the Governance and Strategy track and Gate 1 criteria relating to decision rights and governability. Persistent barriers around data stewardship, privacy, cybersecurity, and interoperability informed the Digital Foundations and Enabling Systems layer. Differences between Tier 1 and Tier 2 settings, particularly in workforce capacity, funding stability, and platform readiness, informed SURF’s tier-aware pathways rather than separate frameworks for different city types.
Two brief examples illustrate this translation process. First, across the systematic review, comparative policy analysis, and interviews, a recurring finding was that pilots often lacked robust monitoring and evaluation, making it difficult to decide whether they should be scaled, refined, or retired. This indicated that the evaluation needed to function as an ongoing decision-support mechanism rather than as end-of-project reporting. It was therefore translated into a design requirement for explicit monitoring, learning, and evidence thresholds, which was operationalised in SURF through the Learning, Evaluation and Impact action track, the IOOI measurement logic, and the Gate 2 evidence pack. Second, across the Tier 1–Tier 2 comparison, Sydney case evidence, and interviews, recurring barriers included unclear ownership, fragmented cross-agency coordination, and short-term or uncertain funding pathways. These findings indicated that smart resilience initiatives struggle to institutionalise when governance arrangements and sponsorship remain under-specified. This was translated into a design requirement for early clarification of decision rights, accountability, and delivery governance, which was operationalised in the SURF through Phase 0 (Readiness and Business Case), the Governance and Strategy action track, and Gate 1 criteria relating to ownership, sponsorship, and governability.This evidence-to-framework workflow is summarised in
Figure 1.
3.4. Trustworthiness, Expert Review and Iterative Refinement, and Ethics
Trustworthiness was strengthened through (i) triangulation across data types and perspectives; (ii) systematic matrix-based comparison across tiers to identify both common mechanisms and context-sensitive differences; and (iii) expert review of the draft framework [
32,
33].
For expert review and iterative refinement, an outreach process was conducted in early April 2025, inviting 20 individuals, including former interviewees and additional experts, to review a concise 2–3-page SURF summary and provide written feedback or meet briefly online. Nine participants ultimately contributed feedback: five provided detailed written comments and four participated in online meetings lasting approximately 30–45 min. These contributors represented a cross-section of roles relevant to smart urban resilience, including local government practitioners, state-linked stakeholders, private-sector delivery partners, and academic experts. Feedback was used not simply to confirm the framework, but to refine its practical logic, feasibility, and clarity. In particular, the review process led to several targeted revisions, including reframing SURF from a more linear roadmap into a phase-gated governance pathway; making Phase 0 (Readiness and Business Case) explicit; strengthening the treatment of leadership continuity, internal capability, and organisational ownership within the enabling conditions; and broadening the evaluation logic from narrow KPIs towards impact measurement, evidence curation, and organisational learning. Changes were documented to preserve traceability between reviewer feedback and subsequent refinement of the framework.
Table 3 summarises examples of expert feedback themes and the resulting refinements made to SURF.
Ethical protocols were adhered to throughout, including obtaining informed consent, maintaining confidentiality, handling data securely, and de-identifying participants.
4. Results
4.1. What SURF Delivers and How to Read It
This section presents the Smart Urban Resilience Framework (SURF) as a practical, governance-embedded architecture for moving from isolated smart city initiatives towards more institutionalised and service-continuity-oriented resilience practice. SURF is designed for application across Tier 1 and Tier 2 Australian cities, while recognising differences in starting capacity, delivery models, and infrastructure maturity.
To make the Results Section easy to follow, the findings are presented in four steps:
SURF architecture: what the SURF is and how its components fit together (
Figure 2).
Phases and decision gates: how the SURF sequences investment, learning, and escalation decisions over time.
Action tracks and enabling layers: how workstreams progress in parallel to avoid pilot drift and under-governed scaling.
Evidence-to-design traceability: how recurrent implementation issues from the evidence base were translated into specific SURF features (
Table 4), followed by the operationalisation of gate criteria and evidence packs (
Table 5), and a worked example.
Across these steps, the results present the SURF as a decision spine (phases and gates) supported by concurrent workstreams (action tracks) and foundational enablers that are intended to reduce common implementation failures, including pilot proliferation, unclear ownership, weak evaluation, and a loss of momentum after early trials.
4.2. SURF Architecture: A Decision Spine Plus Concurrent Delivery Tracks and Enabling Layers
Figure 2 presents the SURF as an integrated architecture with three linked elements. First, the decision spine provides a temporal pathway through clearly defined phases, punctuated by formal evidence gates that determine whether initiatives should proceed, be refined, or be scaled. Second, four concurrent action tracks organise delivery work in parallel: Governance and Strategy, Pilots and Innovation, Infrastructure and Integration, and Learning, Evaluation and Impact, which must be progressed together, so that pilots do not outpace governability, operational integration, or decision-relevant evaluation. Third, three enabling layers span all phases and tracks: Guiding Principles, Governance Enablers, and Digital Foundations and Enabling Systems.
The SURF is intentionally designed so that local governments are not required to complete one action track before beginning another. Instead, the framework treats resilience outcomes as the result of co-evolution across governance, technology, community engagement, and measurement, guided by staged decisions that allocate resources in proportion to evidence and readiness.
A key structural feature of the SURF is that the enabling layers are treated as prerequisites for durable implementation rather than optional add-ons. Guiding Principles establish the normative basis for action, including equity and inclusion, public legitimacy, trust, ethics/privacy, and participation. Governance Enablers provide the organisational conditions for implementation, including procurement, partnerships, capacity, and funding rhythms. Digital Foundations and Enabling Systems provide the technical conditions required for safe and scalable delivery, including interoperability, cybersecurity, stewardship, analytics, and platform readiness. Measurement and adaptation are operationalised through the Learning, Evaluation and Impact action track and reinforced through the learning-and-adaptation loop shown in
Figure 2.
4.3. Phases: From Business Case Readiness to Institutionalisation
The SURF is implemented as a staged pathway that aligns with budgeting and policy cycles, recognising that electoral turnover and changing political priorities may affect continuity and the pace of progression between phases. In the SURF diagram, the Year 0 readiness and validation period comprises two linked phases: Phase 0 (Readiness and Business Case, pre-Year 0) and Phase 1 (Discover, Co-design and Plan, Year 0), culminating in an evidence-based Gate 1 decision. Throughout these early phases, assessment and progression are led by a designated local government sponsoring unit and overseen by a cross-functional steering group, with input from relevant internal stakeholders and external delivery partners.
Phase 0: Readiness and Business Case (pre-Year 0)
This phase assesses whether there is a credible case to proceed and if the foundational conditions are in place to responsibly design a pilot. Key outputs include a cross-departmentally agreed baseline and an initial outcomes logic (including equity considerations); an evaluation of data readiness, governance constraints, and implementation risks; and a feasible governance and funding pathway. These outputs are typically produced by the sponsoring unit with input from strategy, finance, procurement, digital/data, infrastructure, and community-facing functions. The aim is to build shared understanding and a well-founded case before committing to a funded pilot.
Phase 1: Discover, Co-design and Plan (Year 0)
Phase 1 translates the readiness case into an implementable pilot design. This phase involves co-designing the pilot with internal stakeholders and external delivery partners whose roles are material to implementation, governance, service delivery, data stewardship, or community impact; clarifying roles, responsibilities, and decision rights; defining KPIs (or, where appropriate, an IOOI [input–output–outcome–impact] logic); and compiling the Gate 1 evidence pack. In practice, participating actors may include strategy and resilience staff, finance and procurement teams, digital/data units, operational service areas, community-engagement staff, relevant state agencies, utilities, vendors, research partners, and affected community representatives, depending on the initiative. The purpose is to de-risk the initiative, demonstrate fit with governance and ethics requirements, and ensure the pilot is fundable and measurable.
Phase 2: Pilot, Test and Learn (Years 1–2)
In this phase, pilots are conducted in a structured manner and reviewed by the sponsoring unit and steering group against agreed criteria covering performance, cost, operational feasibility, and equity impacts. The emphasis is on generating credible evidence of learning, rather than merely delivering a demonstration, so that Gate 2 decisions are based on results, transferability, and institutional fit.
Phase 3: Scale and Institutionalise (Years 3–5)
Validated solutions are embedded in policy, procurement, asset, and operational plans, as well as in ongoing monitoring systems. The purpose is institutionalisation, making outcomes durable through routine governance, resourcing, and continuous improvement rather than one-off projects.
4.4. Decision Gates: Evidence-Based Progression and Prevention of Pilot Drift
The SURF proposes two decision gates to ensure that movement through these three phases is evidence-based:
Gate 1: Validate & Fund (end of Phase 1).
Gate 1 determines whether the city should invest in a pilot. Passing Gate 1 requires evidence that the initiative is sufficiently defined and governable to be tested responsibly. This includes (at a minimum) clear ownership and decision rights, a resourcing and funding pathway, data and ethical safeguards, stakeholder and cross-agency alignment, and a feasible evaluation design.
Gate 2: Scale/Refine/Retire (end of Phase 2).
Gate 2 determines whether the pilot should be scaled, refined and re-tested, or retired. Passing Gate 2 requires evidence of demonstrated benefits relative to costs and risks, transferability beyond the pilot setting, operational feasibility, and a pathway to embed the initiative within ongoing institutional arrangements (budgeting, procurement, governance, capability, and monitoring).
The gates are designed to prevent two common implementation failures: (1) pilots that proceed without adequate readiness and become stalled, contested, or under-governed; and (2) pilots that generate activity but insufficient credible evidence of learning, public value, and operational feasibility, such that decisions about scaling are driven more by political sponsorship, short-term visibility, or institutional momentum than by systematic evaluation.
4.5. Action Tracks and Enabling Layers: What Must Progress in Parallel
As shown in
Figure 2, SURF is organised around a decision spine, four concurrent action tracks, and three enabling layers that provide cross-cutting foundations across all phases. The decision spine provides the staged pathway (readiness → planning → pilot/test → scale/institutionalise) and the two evidence gates (Validate & Fund; Scale/Refine/Retire) that structure escalation decisions. Running alongside this spine are four concurrent action tracks that organise delivery work, while the enabling layers provide the normative, organisational, and technical conditions needed to support progress across phases and tracks. This distinction separates decision timing (spine), delivery work (tracks), and cross-cutting prerequisites (enabling layers).
In operational terms, the four action tracks generate evidence for both gates, with different emphasis at each stage. At Gate 1, the key question is whether an initiative is sufficiently defined, governable, and measurable to proceed responsibly. Evidence is drawn primarily from governance readiness, ownership and sponsorship, pilot scope and co-design, data and interoperability readiness, and baseline and evaluation planning. At Gate 2, the key question is whether the pilot has produced evidence strong enough to justify scaling, refinement, or retirement. Evidence is drawn primarily from pilot results, implementation lessons, institutional fit, integration feasibility, lifecycle and operating implications, and monitoring evidence on performance, equity, and transferability. Gate outcomes then feed back into the action tracks by promoting refinement of design, governance arrangements, integration pathways, and learning priorities, while the broader learning loop carries these lessons into the next cycle of decision-making.
SURF specifies four concurrent action tracks that run across all phases. While each track’s content is tailored to local context, the Results Section consolidates delivery into four workstreams:
Governance and Strategy: decision rights, cross-department coordination, partnership/accountability arrangements, procurement alignment, and policy integration.
Pilots and Innovation: living labs, stakeholder co-design, structured experimentation, and iterative refinement based on user needs and operational realities.
Infrastructure and Integration: integration of validated digital interventions into physical systems, service delivery, asset management, and place-based solutions (including nature-based options where relevant).
Learning, Evaluation and Impact: indicators, baselines, evaluation design, ongoing monitoring, and feedback routines that keep learning decision-relevant beyond the pilot period.
Three enabling layers support all four tracks across all phases:
Guiding Principles: equity and inclusion, public legitimacy, trust, ethics/privacy, and participation.
Governance Enablers: procurement, partnerships, capacity, and funding rhythms.
Digital Foundations and Enabling Systems: data availability, interoperability, platform strategy (including shared services), cybersecurity and privacy safeguards, fit-for-purpose analytics, and supporting digital infrastructure.
The decision spine and gates integrate these action tracks and enabling layers into auditable decisions. At Gate 1 (Validate & Fund), evidence is drawn across the four action tracks, with particular emphasis on Governance and Strategy and Pilots and Innovation, to assess whether the enabling conditions are sufficient to proceed responsibly. At Gate 2 (Scale/Refine/Retire), evidence again draws on the action tracks, with particular emphasis on Infrastructure and Integration and Learning, Evaluation and Impact, to determine whether scaling is feasible and justified, or whether refinement or retirement is warranted. In this way, the action tracks describe the work, the enabling layers describe the cross-cutting conditions for implementation, and the decision spine governs when evidence is assessed and decisions are taken.
4.6. Tier-Aware Application: Common Structure and Different Default Delivery Modes
The SURF is designed to be structurally consistent across city tiers while recognising distinct starting conditions.
For Tier 1 cities, SURF supports larger portfolios and earlier integration of advanced capabilities because institutional capacity, funding scale, and cross-agency infrastructure often enable multi-track progress at pace. In this context, the gates serve as a portfolio discipline, prioritising initiatives that are evidence-ready and aligned with resilience outcomes.
For Tier 2 cities, the SURF emphasises minimum viable foundations first (particularly in data/platform strategy, governance ownership, and measurement design). Where bespoke systems are unrealistic, the SURF explicitly accommodates delivery through shared services, regional collaborations, vendor-managed platforms, and open standards, provided gate evidence demonstrates governance control, data safeguards, and credible outcome measurement. The gate thresholds remain evidence-based; what changes is the feasible pathway to meeting them.
Across Tier 1 and Tier 2 settings, several requirements remain non-negotiable because they underpin governability and public legitimacy. These include clear ownership and decision rights, data stewardship arrangements, privacy and ethical safeguards, a baseline logic for intended outcomes, and a credible evaluation pathway. What varies by tier is not whether these conditions must be met, but how they are met in practice. Tier 1 councils may meet them through in-house capability, enterprise platforms, and larger delivery portfolios, whereas Tier 2 councils may rely more on shared services, regional collaborations, vendor-managed platforms, staged evidence packs, and more incremental staffing or delivery models. In this sense, the SURF maintains common gate disciplines across tiers while allowing for flexibility in the delivery pathways through which those disciplines are satisfied.
4.7. Evidence-to-Design Traceability: How the Evidence Base Shaped the SURF
Across the evidence base, a recurring implementation gap is evident: initiatives often begin with technology procurement or short-term pilots, whereas governance, data foundations, capabilities, and evaluation readiness lag. The SURF responds to this gap by embedding (i) a staged decision spine, (ii) explicit gate criteria, and (iii) concurrent workstreams that support alignment between delivery activity and learning evidence.
To avoid overstating quantification,
Table 4 is presented as a recurrent thematic alignment rather than a frequency ranking. Where relevant, issues are expressed as cross-cutting conditions (e.g., “ownership and accountability clarity”, “interoperability and shared platforms”, “evaluation and learning infrastructure”, and “cross-agency coordination”) and translated into concrete SURF features (e.g., Gate 1 evidence pack requirements, Gate 2 scale/refine/retire decisions, measurement track requirements, and the digital foundations layer).
Building on the source-specific contributions summarised in
Table 2,
Table 4 provides a condensed evidence-to-design audit trail showing how recurrent implementation issues identified across the evidence base were translated into design requirements and then operationalised as specific SURF features. Rather than presenting these issues as descriptive findings alone, the table makes explicit how particular governance and implementation problems informed the SURF’s phases, gates, action tracks, enabling layers, and tier-aware pathways.
4.8. Operationalising the Gates: Minimum Criteria and Evidence Packs
Table 5 sets out the minimum gate criteria and the associated evidence packs required to inform decision-making. These minimum criteria apply across both Tier 1 and Tier 2 settings; what varies by context is the form, scale, and delivery pathway through which councils assemble the required evidence.
At Gate 1, the evidence pack serves as a readiness dossier. It typically includes: problem definition and intended resilience outcomes; governance, ownership, and decision rights; resourcing and funding pathway; data sources and stewardship arrangements; risk, privacy, and ethical safeguards; stakeholder engagement plan; and an evaluation plan with indicators and a baseline logic.
At Gate 2, the evidence pack serves as a scaling dossier. It typically includes pilot results (outcomes, costs, risks, and implementation lessons); evidence of transferability beyond the pilot context; operational feasibility and lifecycle costs; delivery and procurement implications; platform and data governance arrangements for scaling; capability requirements; and a monitoring plan for sustained outcome tracking.
These evidence packs are structured to support three legitimate Gate 2 outcomes: scale, refine, or retire, so that “retire” is treated as responsible governance (when benefits are weak or risks are excessive), not as failure.
4.9. Worked Pathway Examples
The worked examples below illustrate how the SURF can be used both as an operational pathway for prospective planning and as a retrospective interpretive lens for assessing real initiatives. The first example shows how Gate 1 limits pilot selection to initiatives with minimal readiness, how Phase 2 is explicitly framed as learning and evaluation, and how Gate 2 links scaling to evidence and institutional embedding rather than enthusiasm or short-term funding availability.
Worked example of progression through SURF (illustrative)
Context (Tier 2 council): A mid-sized council is experiencing recurring flash flooding and heat stress in several suburbs. Internal analytics capacity is limited, asset data are fragmented, and delivery relies heavily on vendors.
Phase 0, Readiness and Business Case (pre-Year 0):
The council establishes a baseline (incident logs, response times, hotspot locations, and vulnerability indicators) and confirms data readiness (quality of the stormwater asset register, IoT feasibility, data stewardship responsibilities, and cybersecurity requirements). A governance and funding pathway is defined, including executive sponsorship, decision rights across infrastructure, environment, and community services, and a feasible budget pathway (capital works allocation plus external grant).
Phase 1, Discover, Co-design and Plan (Year 0):
A pilot is co-designed with internal teams and external partners (e.g., a water authority, local community groups, and a technology provider). The pilot scope and intended resilience outcomes are defined (e.g., faster hazard detection, reduced disruption, improved service response equity). Evaluation is specified using IOOI (input–output–outcome–impact) logic, with baseline measures and KPIs. The Gate 1 evidence pack is compiled, including: problem definition and outcomes logic; delivery and procurement approach; data inventory and stewardship plan; privacy/ethics and cybersecurity assessment; roles and decision rights; stakeholder engagement plan; and an evaluation plan.
Gate 1, Validate & fund:
A go/no-go decision is made based on problem–solution fit, governability (ownership and decision rights), data/ethics readiness, equity screening, and the feasibility of delivery and evaluation.
Phase 2, Pilot, Test and Learn (Years 1–2):
The pilot is implemented in priority hotspots (e.g., sensors, alerts, and a dashboard for operations teams). Performance, cost, and equity impacts are monitored against the baseline. Learning evidence is documented: what worked, what failed, implementation conditions, data quality issues, and whether outcomes are transferable beyond the pilot area.
Gate 2, Scale/Refine/Retire:
The council decides to scale, refine, or retire based on evidence strength (benefits vs. costs/risks), lifecycle cost and operational feasibility, equity impacts, scalability/replicability, and an integration pathway into business-as-usual systems (asset plans, procurement, governance, monitoring).
Phase 3, Scale and Institutionalise (Years 3–5):
Validated solutions are embedded in routine operations: procurement and vendor management arrangements, asset and maintenance plans, data governance and cybersecurity controls, staff capability development, and ongoing monitoring dashboards. The learning loop is used to update priorities and improve subsequent initiatives.
The example is illustrative; sequencing and duration vary across initiatives and funding cycles.
Retrospective illustration: applying SURF to SIMPaCT
As a brief retrospective illustration, the SURF can also be read against SIMPaCT (Smart Irrigation Management for Parks and Cool Towns), one of the Sydney initiatives examined in the case study. SIMPaCT uses environmental monitoring, IoT, and AI-enabled optimisation to improve water efficiency and reduce heat impacts in public parks, linking digital innovation to climate adaptation and place-based resilience outcomes.
Viewed through SURF, Phase 0 (Readiness and Business Case) would require the initiative to define the service and risk problem clearly, establish baseline conditions, confirm data and sensor readiness, and clarify operational ownership, funding responsibilities, and governance oversight. Gate 1 would then depend on whether the initiative was sufficiently governable and measurable to proceed responsibly, including whether data quality, accountability arrangements, and a feasible evaluation pathway were in place.
In Phase 2 (Pilot and Learn), the SURF would treat SIMPaCT not simply as a technology demonstration, but as a structured pilot requiring evidence on technical performance, operational readiness, maintenance feasibility, and equity of benefit distribution. This is especially important because the Sydney case suggests that the long-term resilience value of SIMPaCT depends not only on site-level outcomes but also on sustained funding, ownership arrangements, and the capacity to replicate the model across multiple locations and councils. Gate 2 would therefore ask whether the pilot had generated sufficient evidence to justify scaling, refinement, or retirement.
Under the SURF, progression to Phase 3 (Scale and Institutionalise) would require embedding the initiative beyond the pilot itself through routine governance oversight, monitoring, maintenance responsibilities, and a credible pathway for replication and long-term operation. This retrospective illustration is intended to show how the SURF can be applied analytically to a real smart resilience initiative; it does not constitute prospective validation of the framework.
5. Discussion
The SURF is designed to address a practical implementation problem: Australian councils are investing in smart technologies, yet benefits often remain fragmented, difficult to scale beyond pilots, and loosely linked to measurable resilience outcomes [
7,
9,
34]. In this paper, those outcomes are understood through a service continuity lens, focusing on whether digital initiatives can be embedded in governance and delivery systems in ways that support the continuity of essential local government services under stress. SURF responds by integrating a phase-gated pathway, two explicit decision gates, concurrent action tracks, and an IOOI (input–output–outcome–impact) measurement chain, positioning monitoring as decision support rather than end-of-project reporting.
A central argument emerging from the results is that “smart urban resilience” is not achieved through technology adoption alone. It depends on whether technology is coupled with governance arrangements, operational routines, skills, data stewardship, evaluation capability, and budget pathways that enable initiatives to move credibly from exploration to institutionalisation. The SURF makes that coupling explicit and testable through staged work, evidence thresholds, and learning loops.
5.1. Theoretical Contribution: Bridging Socio-Technical Transitions and Governance
The first theoretical contribution of the SURF is to reframe smart urban resilience as a socio-technical transition. In transition theory, change occurs not because technology becomes available, but because technologies become embedded within institutions, standards, capabilities, funding pathways, and shared expectations, often across multiple levels of government and over long time horizons [
35]. The SURF translates this transition logic into a governance-ready pathway for local government, specifying how councils can move from readiness and planning to piloting, scaling, and institutionalisation through structured stages and decision checkpoints.
The SURF is best understood primarily as a governance and implementation architecture for socio-technical transition in local government, rather than as a maturity model, indicator set, or evaluation framework alone. Its distinct contribution lies in specifying how digital initiatives move from experimentation to institutionalisation under public sector conditions. This matters because digital systems introduce governance problems that are not fully captured by general resilience or collaboration frameworks alone, including interoperability burdens across legacy and shared systems, contested data rights and stewardship responsibilities, heightened cybersecurity and privacy exposure, dependence on vendors and external platforms, risks of platform lock-in, and uneven organisational data capability. The SURF addresses these socio-technical challenges by embedding them within the same phase-gated decision architecture that governs ownership, evidence requirements, scaling decisions, and long-term operational integration.
Read in relation to adjacent debates, the SURF sits at the intersection of smart urbanism, urban experimentation, and multi-scalar governance. Smart city research from an urban context perspective has shown that the field often remains bounded by system implementation and insufficiently engaged with broader urban planning concerns [
36]. Research on urban living laboratories further shows that experimentation can generate learning, but that its wider value depends on how it is governed, translated, and sustained beyond demonstration settings, including how questions of power, participation, and inclusion are handled in collaborative arrangements [
37,
38]. Cross-scale urban systems research likewise shows that governance challenges are shaped by interactions across spatial, temporal, and functional dimensions rather than by single projects in isolation [
39]. The SURF builds on these insights, but differs from them by translating them into a council-ready governance architecture that links experimentation, evidence gates, operational integration, and institutionalisation around service continuity.
Second, the SURF contributes to multi-level and collaborative governance scholarship by operationalising ideas that are often presented descriptively. Multi-level governance highlights that authority and capacity are dispersed across levels and actors, creating coordination challenges that no single organisation can resolve alone [
15]. Collaborative governance further suggests that cross-sector collaboration requires institutional design, leadership, incentives, and capability, not just goodwill [
10]. The SURF makes these governance ideas actionable through three design moves:
A phase cadence that aligns with public sector decision and budget cycles;
Concurrent tracks that require governance, delivery, data foundations, capability, and evaluation to progress together; and
Gate decisions that clarify who authorises progress and what evidence justifies escalation.
Third, the SURF strengthens the link between the people-centred smart city agenda and implementation practice. While guidance increasingly emphasises legitimacy, inclusion, transparency, and accountability as foundational, these principles often sit alongside rather than within the mechanisms that determine funding and scaling decisions [
40]. SURF operationalises people-centred commitments by embedding equity checks and explicit evidence requirements into the same decision architecture that governs whether a pilot proceeds (Gate 1) and whether scaling is justified (Gate 2).
5.2. Practical Contribution: Minimising Fragmentation and Facilitating Auditable Scaling Choices
On the practice side, the SURF targets two persistent implementation failures: fragmentation and “pilot-forever” dynamics. Fragmentation occurs when initiatives are deployed as isolated departmental projects, procured as one-off solutions, and evaluated inconsistently [
34]. “Pilot forever” occurs when pilots show promise but fail to transition into operational ownership and funded, business-as-usual delivery, particularly in AI/IoT initiatives where governance, data stewardship, and evaluation requirements are substantial [
7]. In this sense, the SURF addresses not only fragmentation in general, but also the specific governance burdens created by digital systems, including data stewardship, interoperability, cybersecurity assurance, vendor dependence, and the organisational work required to embed digital capability in routine service delivery.
The SURF responds by reshaping delivery governance through four concurrent action tracks that reduce the risk of technology outpacing institutional capacity: Governance and Strategy; Pilots and Innovation; Infrastructure and Integration; and Learning, Evaluation and Impact. These tracks are supported by three enabling layers, Guiding Principles, Governance Enablers, and Digital Foundations and Enabling Systems, which provide the conditions for safe operation, integration, continuity, and accountable scaling. The practical implication is that “progress” is not defined by deployment alone, but by whether governance, delivery integration, and decision-relevant learning are advancing in parallel on top of adequate foundations, capability, and public safeguards.
A second practical contribution is auditability. Many frameworks encourage experimentation, but few specify how experiments are budgeted, maintained, governed, and assessed to support consistent scale/stop decisions. Urban experimentation scholarship has documented both the rise of living labs and the difficulty of translating experiments into durable institutional change [
41]. The SURF addresses this by formalising two “evidence moments”: Gate 1 (Validate & Fund) and Gate 2 (Scale/Refine/Retire), each supported by structured evidence packs that document baselines, results, equity impacts, governance safeguards, and lifecycle implications.
This gate logic is not merely rhetorical. Stage-gate approaches are a well-established way to manage innovation under uncertainty by requiring evidence before escalating investment [
16]. In public administration, similar assurance principles are reflected in gateway-style review processes that strengthen governance and build confidence at key decision points [
17]. The SURF can be understood as a resilience-focused adaptation of this discipline: it justifies scaling when evidence is robust, and it legitimises refinement or discontinuation when evidence is insufficient, while keeping learning transparent across cycles.
Finally, the SURF offers a clearer distinction from many existing smart city and resilience frameworks. Whereas many frameworks primarily offer domains, indicators, or principles, the SURF adds an operational decision spine (phases, gates, and evidence packs) that links strategy to delivery discipline and makes “scale/stop” choices explicit and reviewable.
5.3. Implications for Developing Tier 2 Capacity, Shared Services, and Scalable Delivery
The SURF explicitly addresses tier distinctions by recognising that Tier 2 local governments, which are often smaller or less well-resourced than Tier 1 counterparts, commonly face tighter budgets, smaller specialist teams, and limited in-house expertise in cybersecurity, data governance, analytics, and evaluation. OECD guidance also notes that resourcing constraints are a recurring barrier to scaling smart initiatives [
9]. In this context, the SURF’s early phases matter because they treat “readiness” as a capacity-building roadmap rather than a pre-pilot formality.
Two implications follow. First, the early stages (Phase 0 and Phase 1/Year 0) shift attention from buying tools to establishing the minimum conditions for safe and effective deployment: baseline logic, data readiness and stewardship, privacy and ethics requirements, partner roles, and a viable funding pathway. For Tier 2 councils, this is where expensive dead-ends can be avoided, because initiatives are stress-tested for governability and measurability before pilots are funded.
Second, the SURF provides a governance framework that makes shared services and pooled capabilities more feasible and less risky. Shared services can reduce duplication and enable smaller local governments to access expertise that is difficult to sustain in-house (e.g., cybersecurity assurance, data stewardship, platform procurement, analytics, and evaluation). State guidance also frames shared services as a mechanism for sharing skills and improving outcomes [
42]. The SURF strengthens this approach by offering a common decision structure, standardised gate evidence packs, minimum assurance thresholds, and shared monitoring logic, thereby making scaling across local governments more consistent and defensible.
5.4. Limitations and Research Agenda
5.4.1. Limitations
The primary limitation is contextual. The SURF is derived from evidence from Australian local government, including a high-capacity case used to identify mechanisms for scaling and institutionalisation. While the framework is intended to be transferable, governance structures, privacy and data regimes, and funding arrangements vary across jurisdictions, shaping how gates operate and how responsibilities are distributed. A second limitation concerns the interview evidence. The interview sample was small (
n = 10) and intentionally focused on strategic and advisory participants across Tier 1 and Tier 2 contexts, which strengthened insight into governance, implementation, and institutionalisation, but did not capture frontline operational staff, residents, or the full diversity of implementation actors. A third limitation concerns political and institutional continuity. Although the SURF is designed to align with budgeting and policy cycles, implementation may still be disrupted by electoral turnover, leadership change, shifting strategic priorities, and inter-agency tensions characteristic of multi-level governance. As a result, no framework can fully neutralise the political and organisational conditions that influence sponsorship, risk tolerance, and long-term follow-through [
15]. Fourth, measurement remains challenging because resilience is multidimensional; indicators risk oversimplification if treated as definitive rather than interpreted with local knowledge and context [
3].
5.4.2. Research Agenda
Four near-term directions would strengthen the evidence base and improve portability:
Test the SURF prospectively in real decision settings over 12–24 months across multiple councils (Tier 1 and Tier 2), documenting how Gate 1 and Gate 2 operate in practice, including the prevalence and conditions of “refine” and “retire” outcomes.
Develop a lightweight “indicator core” for IOOI that councils can adopt with minimal burden, while allowing optional context-specific measures to support cross-council learning without enforcing uniformity.
Compare governance conditions across jurisdictions (e.g., strong state coordination versus local autonomy) to identify which SURF elements require adaptation and which remain stable; multi-level governance typologies can guide interpretation [
15].
Evaluate how equity audits and people-centred safeguards influence investment choices and resource distribution over time, and whether embedding these checks at the gate points improves legitimacy and outcomes [
40].
Together, these directions position the SURF not only as a framework for citation but also as an agenda for learning how digital capability can be governed, evaluated, and institutionalised to reliably improve resilience outcomes.