Next Article in Journal
A Decision-Oriented Framework for Data Governance in Smart Airports: An Entropy–DEMATEL Approach
Next Article in Special Issue
Corporate Loan Default Prediction in the Slovak Banking Context: An Interpretable and Ensemble CRISP-DM Pipeline for Credit Risk Assessment
Previous Article in Journal
Overcoming Technological Lock-In: How External Pressure Reshapes Innovation Trajectories in the Age of AI
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

AI Adoption in Local Government: Productivity, Systemic Risk, and Institutional Resilience: Evidence from a PRISMA 2020 Review

by
Abayomi Ogunrinde
1 and
Carmen De-Pablos-Heredero
2,*
1
Department of Business Economics, Faculty of Business Economics, Universidad Carlos III de Madrid, Calle Madrid, 126, Getafe, 28930 Madrid, Spain
2
Department of Business Economics (Administration, Management, and Organisation), Applied Economics II and Fundamentals of Economic Analysis, Universidad Rey Juan Carlos, Paseo de los Artilleros s/n, 28032 Madrid, Spain
*
Author to whom correspondence should be addressed.
Systems 2026, 14(6), 671; https://doi.org/10.3390/systems14060671
Submission received: 24 April 2026 / Revised: 3 June 2026 / Accepted: 5 June 2026 / Published: 11 June 2026
(This article belongs to the Special Issue Resilience and Systemic Risk in Interconnected Financial Systems)

Abstract

Artificial intelligence (AI) is becoming increasingly embedded in the digital infrastructure of local government, creating new opportunities to improve public sector productivity while also influencing systemic risk and organisational resilience across interconnected public systems. As municipalities adopt AI to automate, support, and transform administrative processes, organisational performance becomes more dependent on the reliability of algorithms, the quality of data, effective governance, and coordination among public institutions. These growing interconnections create new vulnerabilities that can spread across public service networks, yet evidence on the productivity, risk, and resilience implications of AI adoption remains fragmented and dispersed across different fields of research. This study develops an integrative conceptual framework that examines the relationship between AI adoption, public sector productivity, systemic risk, and organisational resilience within interconnected sociotechnical systems. Drawing on insights from productivity economics, systems theory, and public governance, the framework positions total factor productivity (TFP) within a broader public value and risk governance perspective. Using the PRISMA 2020 methodology, the study systematically reviews 68 peer reviewed empirical studies published between 2015 and 2025, assessing productivity outcomes, methodological quality, effect sizes, and contextual factors relevant to local government and networked public administration. The findings show that productivity gains associated with AI are strongly influenced by organisational readiness, including digital maturity, workforce capabilities, governance quality, and institutional coordination. While AI has the potential to improve operational efficiency and strengthen adaptive capacity, inadequate readiness can increase systemic risks arising from algorithmic opacity, cybersecurity challenges, data dependence, coordination failures, and disruptions that may spread across interconnected administrative systems. The review also highlights that resilience depends on the ability of public organisations to anticipate, absorb, adapt to, and recover from AI-related disruptions while maintaining the continuity and quality of public services. The study contributes to theory by integrating perspectives from productivity economics, public administration, and systemic risk within a sociotechnical systems framework. It contributes empirically through a comprehensive synthesis of evidence on AI and public sector productivity and methodologically through the application of transparent PRISMA 2020 review procedures. From a practical perspective, the study offers a conceptual measurement framework and policy guidance for municipal decision makers seeking to improve productivity while strengthening resilience and reducing systemic risks in increasingly interconnected public governance systems.

1. Introduction

Local governments face intensifying pressures to deliver efficient, responsive, and equitable public services amid fiscal constraints, demographic shifts, and rising citizen expectations [1,2,3]. In Europe’s multi-level governance architecture, municipalities constitute critical service delivery nodes, managing high-volume administrative processes including registries, permits, social services, and inspections. The productivity challenge in public administration—measuring and improving outputs relative to inputs—has proven persistently difficult due to the sector’s multi-dimensional objectives, non-market outputs, and complex accountability structures [4,5]. Recent European and Spanish institutional reports underscore the intensifying pressures facing local governments. The European Committee of the Regions’ 2025 Annual Report [6] highlights how cities and regions, responsible for managing over 70% of EU policies and two-thirds of public expenditure, must respond to growing demands associated with demographic change, including rural depopulation and ageing populations, while simultaneously addressing citizens’ expectations for accessible and high-quality services. Similarly, the ESPON DESIRE study [7] demonstrates that local administrations across Europe, including Spain, face substantial challenges in ensuring equitable service provision in disadvantaged and demographically shifting territories, pushing them to adopt more adaptive, participatory, and innovation-driven service models. In the Spanish context, fiscal pressures further exacerbate these challenges: the AIReF 2025 evaluation of local governments identifies sustainability risks and increasingly tight expenditure constraints that limit local authorities’ capacity to meet rising social demands. Moreover, the Council of the EU’s 2025 recommendation on Spain’s medium-term fiscal plan [8] establishes a binding expenditure path for the coming years, reinforcing the restrictive budgetary framework under which local governments must operate while confronting demographic transformation and heightened public expectations.
Artificial intelligence (AI) has emerged as a general-purpose technology with the potential to fundamentally reshape public administration productivity dynamics [9,10,11,12]. Unlike previous waves of digital innovation, contemporary AI systems can automate cognitive tasks, augment professional decision-making, and generate predictive insights from complex data patterns [13,14]. Recent implementations demonstrate efficiency gains in specific administrative domains: automated document processing reducing cycle times by 60–80% [15,16], predictive analytics improving inspection targeting by 30–50% [17,18], and chatbots handling 40–70% of routine citizen inquiries [11,19,20,21]. However, rigorous economic evidence quantifying productivity impacts at the local government level remains limited, particularly in Southern European contexts [22,23].
The European Union’s regulatory framework, particularly the AI Act (entered into force August 2024), establishes governance requirements for high-risk AI systems in public administration, mandating transparency, human oversight, and accountability mechanisms [24]. Simultaneously, national AI strategies—exemplified by Spain’s National AI Strategy 2024 and the ALIA public infrastructure initiative—create enabling conditions for sovereign AI deployment in government [25]. This regulatory-infrastructure nexus shapes adoption trajectories and productivity outcomes at the municipal level, creating a distinctive governance environment that merits scholarly examination [26,27].
Despite growing scholarly attention to AI in government, three critical gaps persist in the existing literature. First, theoretical frameworks linking AI adoption mechanisms to productivity outcomes remain underdeveloped, particularly frameworks operationalising total factor productivity (TFP) concepts for public sector contexts [3,28]. Existing studies frequently focus on narrow efficiency metrics without integrating public value dimensions or addressing TFP’s multi-factor logic [29,30,31]. Recent analyses continue to show that conceptual frameworks linking AI adoption to productivity remain emergent, and even the latest empirical and simulation-based work [12] highlights the absence of mature theoretical models, particularly for public sector contexts.
Second, empirical evidence concentrates on national-level policies or isolated pilot projects, with limited rigorous assessment of local government implementations [32,33]. The “productivity puzzle” in public administration—where technology investments consistently fail to yield measurable productivity gains—demands granular, process-level analysis [34,35]. Southern European municipalities, characterised by distinct institutional contexts, fiscal frameworks, and administrative cultures, remain particularly under-researched [36,37]. Recent analyses confirm that this gap persists: despite growing experimentation with AI in public administrations across Spain, Italy and other Southern EU countries, systematic evidence on municipal-level adoption remains limited. For example, a 2025 study on the Spanish public sector highlights that while AI has substantial potential to streamline administrative processes and improve citizen interactions, local governments still face uneven adoption and structural barriers that hinder implementation [38]. Likewise, a 2025 Europe-wide assessment of local government AI initiatives shows that only 27% of municipalities have begun deploying AI, revealing notable disparities particularly affecting smaller and less digitally mature Southern European cities [39].
Third, systematic synthesis of productivity impact evidence is lacking. Existing reviews address AI in government broadly [17,23,40] but do not systematically extract and synthesise quantitative productivity measures, assess evidence quality using standardised criteria, or develop integrative theoretical frameworks. This gap impedes evidence-based policy-making and investment prioritisation [31,41]. New systematic reviews in public governance [42] identify persistent fragmentation in data, methodological heterogeneity, and weak integration of productivity indicators across studies. Likewise, recent analyses of AI implementation in local and national governments underscore significant variation in how efficiency gains are measured and rarely provide comparable productivity outcomes [16].
This study addresses four interrelated research questions: (1) How can AI adoption pathways be theoretically linked to public sector productivity outcomes within a TFP-consistent framework that integrates public value creation? (2) What does systematic synthesis of empirical evidence reveal about AI’s measured productivity impacts in government administration, and what is the quality of this evidence base? (3) What mediating factors condition AI productivity effects in local government contexts, and how do these operate in Southern European municipalities? (4) What are the implications for research methodology and policy implementation in mid-sized European cities pursuing AI-enabled productivity improvements?
To address these questions, we pursue three objectives: (1) develop a conceptual framework operationalising TFP logic for AI-driven productivity in public administration, drawing on sociotechnical systems theory [43] and digital-era governance theory [34]; (2) conduct a PRISMA 2020-compliant systematic review synthesising empirical evidence on productivity impacts from peer-reviewed studies; and (3) apply the framework through structured analysis of Madrid’s AI deployment within Spain’s national infrastructure (ALIA) and EU regulatory context.
Madrid (Ayuntamiento de Madrid) represents a strategically valuable analytical site for three reasons. First, institutional advancement: Madrid’s Artificial Intelligence Roadmap (2023–2025), embedded within the “Madrid, Digital Capital” strategy, establishes governance frameworks, key performance indicator (KPI) systems, and cross-departmental AI portfolios [44]. The municipality operates a dedicated AI platform with contracted evolution and maintenance [45], indicating institutionalisation beyond pilot status. Second, multi-level embeddedness: Madrid’s initiatives are nested within regional enablers (Madrid Data & AI Hub) and national infrastructure (ALIA), providing sovereign, Spanish-language AI capabilities for public sector applications [25]. Third, representativeness: As a mid-sized European capital with 3.3 million inhabitants, Madrid’s scale, resource constraints, and institutional complexity resonate with other Southern European cities that lack the fiscal resources of mega-cities or the digital maturity of Nordic administrations [33,37]. Importantly, Madrid is not treated as a comprehensive case study in this article. Rather, we employ Madrid’s documented AI initiatives as an empirical anchor to illustrate framework components and identify implementation patterns relevant to similar municipalities [46].
This article makes three distinct contributions to the literature. Theoretically, we bridge productivity economics and public administration by developing an integrated framework operationalising TFP concepts for AI-driven government productivity while embedding public value considerations [29,47]. The framework specifies three AI pathways (automation, augmentation, transformation), maps these to TFP components (labour, capital, multifactor productivity), and identifies mediating mechanisms (organisational capabilities, governance quality, technological maturity). Seven testable propositions link AI characteristics to productivity outcomes, providing a foundation for future empirical research. Empirically, a systematic review synthesising quantitative productivity evidence from AI implementations in government, following PRISMA 2020 standards [48,49], is provided. Analysis of 68 peer-reviewed studies documents effect sizes, assesses evidence quality using adapted Grading of Recommendations, Assessment, Development and Evaluations (GRADE) criteria, and identifies methodological patterns and gaps. Practically, we translate theoretical insights into actionable guidance for municipal administrators and policy-makers. The measurement model specifies data requirements for productivity assessment, while framework analysis identifies capability prerequisites for successful AI adoption. Analysis of Madrid’s deployment contextualised by ALIA infrastructure and EU AI Act requirements illuminates barriers and enablers relevant to mid-sized European cities, contributing to emerging scholarship on smart governance and urban digital transformation [33,50].
The remainder of the article is structured as follows. Section 2 develops the conceptual framework. Section 3 details the methodology. Section 4 presents systematic review results. Section 5 discusses findings and implications. Section 6 reviews limitations and future research priorities. Section 7 offers concluding remarks.

2. Conceptual Framework: AI, Productivity, and Public Value

2.1. Theoretical Foundations

The conceptual framework developed in this article integrates three theoretical traditions: (1) productivity economics and TFP measurement [9,47]; (2) public administration theory, specifically public value framework [29] and digital-era governance [34]; and (3) sociotechnical systems theory [43,51], which emphasises the co-evolution of technology and organisational arrangements.
These traditions are seldom integrated in the existing literature. Productivity economists typically apply TFP models to private sector contexts without accounting for public value dimensions [28]. Public administration scholars often focus on governance and accountability without quantifying productivity outcomes [31]. Sociotechnical theorists emphasise process and organisational change but rarely employ productivity measurement frameworks [35,52]. Our framework bridges these gaps, creating an integrative theoretical lens for examining AI-enabled productivity in local government.

2.2. Productivity in Public Administration: Foundations

Productivity, fundamentally, measures output per unit of input. In private sector contexts, TFP captures efficiency gains beyond simple labour or capital productivity by measuring output growth not explained by input growth alone [47,53]. TFP represents the “residual”—improvements attributable to technological change, organisational innovation, or efficiency enhancements rather than increased resource use.
Public sector productivity measurement confronts distinct challenges [4,5,54]. First, output definition: Government services often lack market prices, complicating valuation. Outputs encompass volume (number of services delivered), quality (accuracy, timeliness, user satisfaction), and outcomes (societal impact). Second, multi-dimensionality: Public agencies pursue multiple, sometimes conflicting objectives—efficiency alongside equity, responsiveness, and legitimacy. Productivity gains that compromise fairness or transparency may reduce overall public value [29,55]. Third, attribution: Isolating technology’s productivity contribution from organisational, political, and contextual factors proves difficult in non-experimental settings [34,51].
Despite these complexities, TFP logic remains valuable for public administration analysis. It directs attention to multifactor efficiency, recognising that productivity gains can stem from better technology, improved processes, enhanced workforce skills, or superior management [56]. For AI specifically, TFP framing captures how intelligent systems affect not just labour productivity but also capital utilisation (through automated asset management) and create novel capabilities that shift production possibility frontiers [30].

2.3. AI Pathways to Productivity

Building on Raisch and Krakowski’s [13] automation–augmentation paradox and Orlikowski’s [51] theory of technology-in-practice, we identify three pathways through which AI systems influence public sector productivity, each with distinct mechanisms and TFP implications. The theoretical derivation of each pathway is as follows. Raisch and Krakowski [13] explicitly distinguish two managerial logics for deploying intelligent systems: automation, in which AI substitutes for human task execution, and augmentation, in which AI complements human cognition by providing decision support. We map these two logics directly onto Pathway 1 (Automation—task substitution) and Pathway 2 (Augmentation—decision support). Pathway 3 (Transformation—new capabilities) is derived from Orlikowski’s [51] technology-in-practice framework, which holds that the productive effects of a technology are not given by its features in isolation but emerge from its recurrent enactment within organisational routines and structures. Under sustained enactment, AI systems can support qualitatively new modes of action—such as large-scale predictive analytics, simulation-based policy testing, and real-time resource allocation—that were not feasible under prior socio-material arrangements. This pathway captures the transformative potential that neither pure substitution nor pure complementarity adequately describes. Together, the three pathways therefore exhaust the logical space identified by these two theoretical traditions: substitution of human action (automation), complementarity with human action (augmentation), and reconfiguration of the action space itself (transformation).
Pathway 1: Automation (Task Substitution). AI automates routine cognitive tasks previously requiring human labour. In government administration, this includes document classification, data entry, eligibility verification, and standard correspondence processing. Automation affects labour productivity directly by reducing person-hours per transaction. The TFP mechanism operates through: (a) labour reallocation—freeing staff for higher-value activities; (b) throughput expansion—processing more cases with fixed labour inputs; and (c) quality consistency, reducing human error rates. Evidence from robotic process automation (RPA) implementations demonstrates 40–70% time reduction in specific administrative processes [57,58]. Drawing on Acemoglu and Restrepo’s [59] task-based model, automation displaces labour from specific task categories while creating demand for new task profiles, with net productivity effects contingent on the speed of task creation relative to displacement.
Pathway 2: Augmentation (Decision Support). AI augments professional judgement by providing data-driven insights, recommendations, or risk assessments that enhance worker effectiveness without replacing them [13,30]. Examples include fraud detection algorithms assisting auditors, predictive maintenance systems supporting facility managers, and case prioritisation tools aiding social workers. The TFP mechanism differs from automation: productivity gains stem from improved decision quality (reducing costly errors), faster decision cycles, and capability expansion. Noy and Zhang [60] document more than 15% productivity gains among customer support workers using generative AI for augmentation, with disproportionately larger effects for less-experienced workers, suggesting AI may function as a “skill leveller.” Consistent with human–computer interaction research [61], augmentation effectiveness depends critically on appropriate automation levels and user trust calibration.
Pathway 3: Transformation (New Capabilities). AI enables qualitatively new capabilities, services, or modes of governance impossible without machine learning at scale [62,63]. In public administration, this encompasses real-time demand forecasting for resource allocation, natural language processing for citizen feedback analysis at scale, computer vision for urban infrastructure monitoring, and simulation modelling for policy scenario testing. The TFP impact operates through innovation effects: expanding service portfolios, improving targeting precision, and generating evidence for data-driven governance. These effects are hardest to quantify in the short term but potentially most transformative over longer time horizons [9]. Orlikowski’s [51] enacted technology theory suggests that the transformative potential of AI is realised not through the technology itself but through deliberate organisational choices about how to embed it in governance structures.

2.4. Integrating Public Value

Moore’s [29] public value framework reminds us that government productivity cannot be reduced to technical efficiency. Public value encompasses three dimensions: (1) services that meet citizen needs and preferences; (2) outcomes that advance democratic legitimacy and social goals; and (3) processes that maintain trust, fairness, and accountability [55,64]. AI productivity gains that undermine these dimensions ultimately destroy rather than create value.
Critical tensions emerge across all three AI pathways. Automation may increase efficiency but reduce public sector employment or deskill workers, raising issues of distributional justice [52,59,65]. Algorithmic decision support may improve consistency but embed historical biases or reduce the discretion needed for equitable case handling [65,66,67,68]. Predictive analytics may optimise resource allocation but raise surveillance concerns or exacerbate inequalities through feedback loops [69,70,71]. The EU AI Act’s risk-based approach directly addresses these tensions by classifying certain public administration applications (e.g., social benefit eligibility assessment, law enforcement) as high-risk, requiring transparency, human oversight, and bias monitoring [24].
Our framework integrates public value by (a) expanding productivity measurement beyond efficiency to include quality, equity, and responsiveness dimensions; (b) treating governance mechanisms (transparency, explainability, human-in-the-loop requirements) as productivity-enabling rather than productivity-constraining factors that build citizen trust and institutional legitimacy [31]; and (c) specifying citizen trust and democratic accountability as outcome variables affected by AI deployment choices. This integration prevents narrow technocratic optimisation and aligns the framework with public administration’s normative commitments [29,34].

2.5. Operationalisation: The Measurement Model

To translate TFP concepts into measurable constructs for local government AI evaluation, we specify a measurement model comprising inputs, outputs, and mediating factors, consistent with Van Dooren et al. [5] and Atkinson [4]:
Input Measures: (a) Labour: Full-time equivalents (FTEs) allocated to the process or service area; person-hours per case/transaction; (b) Capital: Information and communication technology (ICT) expenditure (hardware, software licences, cloud services); AI system development and procurement costs; and (c) Operational costs: Energy, maintenance, and training expenditure attributable to AI systems.
Output Measures (Multi-dimensional): (a) Volume: Number of cases processed, transactions completed, and services delivered per time period [72,73,74]; (b) Quality: Error rates, accuracy scores, compliance levels, and citizen satisfaction ratings [72,73,74]; (c) Timeliness: Cycle time, response time, service delivery speed, and backlog reduction [72,73,74,75]; (d) Equity: Distributional fairness across demographic groups and accessibility improvements [73,74]; and (e) Responsiveness: Customisation capability, multi-channel availability, and user experience quality [72,74,76].
Mediating Factors (Organisational Context): (a) Digital maturity: Existing IT infrastructure, data quality, interoperability, and digital literacy of the workforce [26,27]; (b) Workforce capabilities: Staff AI literacy, change readiness, training investment, and retention rates [33,52]; (c) Governance quality: Data governance frameworks, ethical guidelines, accountability mechanisms, and stakeholder engagement processes [24,68]; and (d) Resource commitment: Budget allocation, leadership support, project management capacity, and continuity planning [35,51].
This model enables productivity calculation as TFP = Weighted Output Composite/(Labour + Capital Inputs). Changes in TFP over time indicate whether AI adoption generates efficiency gains beyond simple resource increases. The inclusion of quality and equity dimensions prevents productivity improvements achieved by sacrificing service standards or fairness [5,29].
Figure 1 shows the integrative framework of total factor productivity (TFP) for AI-driven productivity in public administration, illustrating the three AI pathways (automation, augmentation, and transformation), their relationships with organisational mediating factors, and multi-dimensional productivity outcomes.

2.6. Theoretical Propositions

Drawing on the integrated framework, we advance seven propositions linking AI characteristics to productivity outcomes, each grounded in the theoretical traditions reviewed above:
P1 (Automation–Volume): AI automation of routine administrative tasks yields larger productivity gains in high-volume, standardised processes than in low-volume, complex processes. This proposition draws on Acemoglu and Restrepo’s [59] task-based model of automation and RPA implementation evidence [57].
P2 (Augmentation–Quality): AI augmentation of professional judgement produces larger quality improvements than volume improvements, with effects mediated by task complexity and worker experience. This proposition is grounded in Brynjolfsson et al.’s [30] findings on generative AI augmentation and Raisch and Krakowski’s [13] automation–augmentation paradox.
P3 (Maturity Interaction): AI productivity gains are positively moderated by organisational digital maturity; municipalities with higher baseline IT capabilities achieve faster and larger returns. This proposition draws on Gil-Garcia et al.’s [27] digital maturity framework and Orlikowski’s [51] enacted technology theory.
P4 (Temporal Dynamics): AI productivity impacts exhibit a J-curve pattern: initial productivity decreases during learning and adaptation phases (0–18 months), followed by accelerating gains (18–48 months) as organisational routines adjust and complementary capabilities develop. This proposition is consistent with Brynjolfsson and Hitt’s [56] analysis of IT investment lags and Orlikowski’s [51] technology-in-practice framework.
P5 (Complementarity): AI productivity effects are amplified when combined with organisational restructuring and process redesign; technology-only implementations yield smaller gains than sociotechnical transformations. This proposition is grounded in Trist and Bamforth’s [43] sociotechnical systems theory and Lindgren et al.’s [35] analysis of e-government value.
P6 (Public Value Trade-off): Efficiency-focused AI implementations that neglect governance quality (transparency, accountability, bias monitoring) experience diminishing productivity returns over time due to trust erosion and compliance burdens. This proposition draws on Busuioc’s [68] algorithmic accountability framework and Dunleavy and Carrera [31].
P7 (Scale Effects): Productivity gains from AI exhibit economies of scale; larger municipalities achieve higher return on investment due to fixed cost amortisation, but mid-sized cities can access comparable benefits through shared infrastructure (e.g., national platforms such as ALIA). This proposition draws on Caragliu et al.’s [50] smart governance framework and Fung [41].

3. Methodology

3.1. Research Design and Methodological Rationale

This study employs a Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020-compliant systematic review (SR) with structured secondary analysis of Madrid’s documented AI deployment initiatives. The SR serves as the primary empirical method, synthesising peer-reviewed evidence on AI productivity impacts across government administrative contexts. The Madrid secondary analysis provides a theoretically grounded illustrative application of the conceptual framework developed in Section 2, enabling contextualisation of the review findings within a specific Southern European institutional environment.
The choice of a systematic review as the primary method is justified on three grounds. First, the research questions require evidence synthesis across a heterogeneous body of empirical work spanning diverse administrative domains, AI technologies, and national contexts—a task for which the SR is the most appropriate and replicable approach [77]. Second, the study objective of developing a theoretically integrated framework necessitates engagement with the full scope of available empirical evidence rather than a single case or context. Third, systematic procedures ensure transparency, reproducibility, and auditability of the evidence base—qualities essential for providing credible guidance to policy-makers and for advancing cumulative scholarly knowledge [48]. The dual-method design is consistent with Petticrew and Roberts [77], who recommend integrating primary contextual analysis with systematic review synthesis to combine evidentiary breadth with institutional depth.

3.2. Systematic Review Protocol and PRISMA 2020 Framework

The SR was conducted in accordance with the PRISMA 2020 guidelines [48,49], which provide the international standard for transparent and reproducible systematic review reporting. The PRISMA 2020 framework structures the review process across four sequential stages: identification, screening, eligibility assessment, and inclusion. Each stage involves explicit procedural decisions governed by pre-specified criteria, ensuring methodological integrity and enabling independent replication.
The review protocol was developed prior to data collection and encompasses the search strategy, eligibility criteria, quality assessment procedure, data extraction instrument, and synthesis approach documented in the following subsections. The complete protocol, search strings, and data extraction forms are available as Supplementary Materials. This pre-specification of procedures reduces the risk of post hoc analytical decisions that could introduce bias into the review findings [48,77].

3.3. Data Sources and Search Strategy

Three major academic databases were searched systematically: Web of Science, Scopus, and Google Scholar. These databases were selected to maximise coverage of peer-reviewed scholarship across public administration, information systems, economics, and computer science—the primary disciplinary homes of AI productivity research. The search period was restricted to January 2015 to December 2025, a ten-year window chosen to capture the modern AI era following the deep learning breakthrough of 2012 [78] while excluding the earlier expert systems and rule-based AI literature that would not reflect contemporary AI capabilities and applications.
The search strategy employed Boolean combinations across three conceptual clusters to ensure comprehensive and reproducible coverage:
  • AI technologies: (“artificial intelligence” OR “machine learning” OR “deep learning” OR “natural language processing” OR “computer vision” OR “robotic process automation” OR “chatbot” OR “predictive analytics” OR “generative AI”).
  • Government contexts: (“public sector” OR “government” OR “public administration” OR “municipal” OR “local government” OR “city administration” OR “e-government” OR “smart government”).
  • Outcome measures: (“productivity” OR “efficiency” OR “performance” OR “cost reduction” OR “time saving” OR “process improvement” OR “automation” OR “service delivery”).
The three clusters were combined using the Boolean operator AND, with OR applied within each cluster to capture terminological variation across disciplinary conventions. Synonyms and alternative terminology were incorporated to reduce the risk of systematic omission attributable to vocabulary differences between public administration and computer science literature. The complete search strings for each database are provided in the Supplementary Materials.

3.4. Inclusion and Exclusion Criteria

Eligibility criteria were pre-specified and applied consistently at each PRISMA screening stage. The rationale for each criterion is provided to support methodological transparency and replicability.
Inclusion criteria:
  • Study type: Peer-reviewed journal articles and peer-reviewed conference proceedings with a valid DOI. This restriction ensures verifiable academic quality and accessible sourcing for audit purposes.
  • Empirical design: Studies reporting original quantitative, qualitative, or mixed-methods data collection. Purely theoretical or conceptual papers were excluded given the objective of synthesising measured productivity evidence.
  • Substantive focus: Studies examining AI implementation within government administrative processes. This includes citizen-facing service delivery, back-office operations, regulatory functions, financial administration, and public workforce management.
  • Outcome specification: Studies reporting at minimum one operationalised measure of productivity, efficiency, process performance, or service delivery improvement—enabling extraction of effect sizes or comparable performance indicators.
  • Language: English, Spanish, French, or German. These languages reflect the authors’ reading capabilities and cover the primary languages of European public administration scholarship. To identify relevant non-English records despite running the primary Boolean search in English, two complementary procedures were applied. First, the three databases used (Web of Science, Scopus, and Google Scholar) index titles and abstracts in English for the vast majority of indexed non-English-language journals, so studies originally published in Spanish, French, or German with English-language metadata were retrieved by the English search strings. Second, a supplementary search was conducted in each of the three other languages by running translated versions of the AI-technologies, government-context, and outcome-measure clusters (e.g., “inteligencia artificial”, “administración pública”, “productividad”; “intelligence artificielle”, “administration publique”, “productivité”; “künstliche Intelligenz”, “öffentliche Verwaltung”, “Produktivität”) in Google Scholar, where language-restricted search is reliable. Of the 68 included studies, 61 were published in English, 4 in Spanish, 2 in French, and 1 in German; the translated search strings and the full multilingual record counts are provided in Supplementary Materials.
  • Publication period: January 2015 to December 2025, consistent with the rationale set out in Section 3.3.
Exclusion criteria:
  • Purely conceptual, normative, or theoretical papers without empirical data collection were excluded, as they do not contribute measurable productivity evidence to the synthesis.
  • Studies of AI in non-administrative government domains—specifically defence, military intelligence, clinical healthcare delivery, and judicial proceedings—were excluded on the grounds that these domains involve fundamentally different organisational logics, regulatory regimes, and outcome metrics from administrative productivity.
  • Technology description papers, vendor white papers, and system architecture reports without assessed outcome data were excluded.
  • Publications pre-dating January 2015.
  • Grey literature—government reports, consultancy documents, and institutional white papers—was excluded from the systematic review corpus. Such sources are cited contextually where relevant to illustrate policy and institutional contexts (e.g., the Madrid AI Roadmap [44] and ALIA documentation [25]) but do not form part of the evidence synthesis.

3.5. Study Selection Process and Inter-Rater Reliability

The study selection process followed the four PRISMA 2020 stages: identification, screening, eligibility assessment, and inclusion. Table 1 summarises the complete selection flow with record counts at each stage. The PRISMA 2020 flow diagram (Figure 2) provides a visual representation of the process.
Initial database searches yielded 1247 records in total (Web of Science: 412; Scopus: 568; Google Scholar: 267). Following deduplication using reference management software, 289 duplicate records were removed, leaving 958 unique records for screening. Title and abstract screening against the pre-specified inclusion and exclusion criteria eliminated 782 records as clearly irrelevant—primarily papers addressing AI in non-administrative domains, purely technical system descriptions, or publications outside the specified time window. The remaining 176 records were retrieved in full text for eligibility assessment.
Full-text assessment applied the complete inclusion and exclusion criteria to each of the 176 articles. Of these, 108 were excluded: 47 lacked operationalised productivity or performance outcome measures; 31 focused on non-administrative government domains; 18 were conceptual or normative papers without empirical data; and 12 failed to meet quality thresholds on the adapted GRADE criteria (detailed in Section 3.6). The final sample comprised 68 peer-reviewed studies included in the evidence synthesis.
Two researchers independently conducted screening and selection at each PRISMA stage. Disagreements at the title/abstract stage were first resolved through bilateral structured discussion; in cases where consensus was not achieved after two rounds, a third-party adjudicator provided the deciding assessment. Inter-rater reliability was assessed using Cohen’s kappa statistic [79]: κ = 0.87 at the title and abstract screening stage and κ = 0.91 at the full-text eligibility stage. Both values fall within the “excellent agreement” range (κ > 0.80) as defined by Landis and Koch [80], confirming that the selection process was robust against individual reviewer bias.

3.6. Quality Assessment

Each of the 68 included studies was independently assessed for methodological quality using an adapted version of the GRADE framework [81], modified to reflect the realities of public administration research—where randomised controlled trials are rarely feasible and administrative databases are frequently the primary data source. The adaptation follows the GRADE working group’s own provisions for non-randomised evidence [81] and the operationalisation of GRADE-style criteria for public-administration and information-systems research developed by Petticrew and Roberts [77] and Popay et al. [82]. Each study was rated across five dimensions on a three-level scale (high, moderate, low) according to the explicit, pre-specified thresholds set out below:
Study design: Experimental designs (RCTs, field experiments) or quasi-experimental designs (difference-in-differences, regression discontinuity, synthetic control, propensity-score matched comparisons) with an identified counterfactual = high; longitudinal observational designs with at least two measurement waves and no formal counterfactual = moderate; single-wave cross-sectional, single-case, or before–after designs without a comparison group = low.
Sample size and representativeness: For studies with individual-level observations (cases, transactions, citizens, employees), n ≥ 1000 with multi-site coverage (≥3 administrative units or municipalities) and explicit reporting of demographic or institutional diversity = high; 300 ≤ n < 1000, or n ≥ 1000 from a single site = moderate; n < 300 or single-site small-n designs = low. For studies whose unit of analysis is the organisation or AI implementation (typical of multi-case comparative studies and surveys of public agencies), ≥10 organisations across ≥2 jurisdictions = high; 4–9 organisations or 10+ organisations within a single jurisdiction = moderate; 1–3 organisations = low. These thresholds were calibrated against the empirical distribution of sample sizes observed in the included corpus and against guidance for non-experimental evidence in Petticrew and Roberts [77].
Measurement rigour: At least two operationalised productivity indicators measured with validated multi-item instruments or with administrative records whose accuracy is documented, accompanied by reported psychometric or audit-quality information (reliability coefficients, completeness checks, audit trail) = high; established measures used without reported validation evidence, or a single validated indicator = moderate; ad hoc, undocumented, or unspecified measurement instruments = low.
Confounding control: Multivariate analysis with at least three relevant control variables, propensity-score matching, instrumental variables, or difference-in-differences estimation = high; bivariate or descriptive comparison between treated and untreated units without statistical adjustment = moderate; no comparison group and no statistical adjustment = low.
Transparency: Full disclosure of data sources, sample selection rules, analytical procedures, and either open data/code or a clearly stated replication path = high; partial disclosure (data sources and analytical procedures described, but selection rules or replication path missing) = moderate; only general methodological information, with at least one of the four elements above absent = low.
Studies receiving “high” ratings on three or more of the five dimensions were classified as high quality (n = 22, 32.4%); studies receiving “moderate” on three or more were classified as moderate quality (n = 34, 50.0%); the remainder were classified as low quality (n = 12, 17.6%). The three-of-five threshold is consistent with the simple-majority rule applied by Popay et al. [82] for narrative synthesis of heterogeneous evidence and with the “at least three of five domains” convention used in adapted GRADE/GRADE-CERQual operationalisations for non-randomised public-administration and information-systems evidence [77,81]. To verify that the substantive findings of the review were not driven by this specific threshold, we conducted a sensitivity analysis re-classifying studies under stricter (≥4/5 high) and more permissive (≥2/5 high) rules; the rank ordering of pathways by effect size and the qualitative pattern of moderating factors were preserved in both cases. Quality ratings were used to weight findings in the narrative synthesis and to conduct sensitivity analysis examining whether effect size estimates differ systematically across quality tiers—an important safeguard against publication bias inflation. Quality assessment was conducted independently by both authors, who rated each study against the five dimensions without consultation. Disagreements at the dimension level were tabulated and resolved through a two-stage procedure consistent with the screening process described in Section 3.5: first, the two assessors discussed each discrepant dimension with reference to the explicit threshold definitions above; second, where consensus was not reached after structured discussion, a third independent reviewer external to the research team adjudicated the final rating. Inter-rater agreement at the dimension level was substantial (Cohen’s κ = 0.83, averaged across the five dimensions), with the lowest agreement on “transparency” (κ = 0.74) and the highest on “study design” (κ = 0.92). At the overall quality-tier level (high/moderate/low), agreement was κ = 0.89, within the “almost perfect” range of Landis and Koch [80]. Out of the 68 studies, 9 required structured discussion at the dimension level and 2 required adjudication by the third reviewer; the full inter-rater table is provided in Supplementary Materials.

3.7. Data Extraction and Analytical Approach

Data were extracted from included studies using a standardised instrument piloted on a subsample of ten studies before full deployment, with minor revisions made to improve capture of implementation context variables. The extraction instrument captured the following:
Study characteristics: Country, government level (national/regional/local), administrative domain, publication year, and study design.
AI intervention details: Technology type (classified against the three-pathway typology developed in Section 2.3), implementation scope (single process, multi-process, organisation-wide), deployment duration, and whether the implementation was pilot or institutionalised.
Outcome measures: Productivity metrics employed, measurement methods, baseline and post-implementation values where reported, and units of analysis.
Effect sizes: Percentage changes in process times, cost levels, error rates, throughput, or decision quality; statistical significance levels where reported; and direction of effect.
Contextual and moderating factors: Organisational characteristics, implementation barriers and enablers, workforce factors, governance mechanisms, and sustainability indicators.
Given the substantial heterogeneity in study designs, AI technologies, government levels, and outcome metrics, meta-analytic statistical pooling was not feasible—a finding consistent with prior systematic reviews in this domain [23]. The synthesis therefore employs narrative synthesis with quantitative summary [82], aggregating effect sizes descriptively by AI pathway and quality tier, identifying patterns and consistencies, and reporting median and range statistics where standardisation permits cross-study comparison. This approach preserves contextual richness while generating structured evidence summaries sufficient to evaluate the seven theoretical propositions advanced in Section 2.6.

3.8. Reliability, Validity, and Bias Control

Multiple procedures were implemented to strengthen the reliability, validity, and bias resilience of the review. Procedural reliability was ensured through pre-specified eligibility criteria, standardised extraction instruments, and dual independent screening with documented inter-rater agreement (κ = 0.87–0.91). Search reliability was addressed through multi-database searching with extensive synonym coverage, reducing the likelihood of systematic omission attributable to database gaps or terminological variation.
Publication bias—the overrepresentation of positive findings in the published literature—is a recognised limitation of any systematic review [83]. To mitigate its distorting influence, sensitivity analysis was conducted comparing effect size estimates across quality tiers (reported in Section 4.3). The finding that high-quality studies systematically report smaller effect sizes than low-quality studies is consistent with publication bias and methodological inflation in weaker designs, and this pattern is explicitly discussed to contextualise aggregate estimates. Prospective registration of systematic reviews on platforms such as PROSPERO was not completed prior to data collection; future replications of this review should register protocols in advance to further strengthen methodological rigour.
Construct validity was addressed by operationalising the three AI pathway categories (automation, augmentation, transformation) against explicit theoretical definitions from Section 2.3 prior to classifying included studies, and by having both researchers classify each study independently before resolving disagreements through discussion. External validity is limited by the evidence base itself—predominantly Anglo-American and Scandinavian in origin, with limited Southern European representation—a limitation fully acknowledged in Section 6.1.

3.9. Madrid Secondary Analysis

To complement the SR, a structured secondary analysis of Madrid’s AI deployment was conducted using publicly available institutional documents, including the Ayuntamiento de Madrid’s AI Roadmap 2023–2025 [44], AI platform procurement materials [45], regional Madrid Data & AI Hub announcements, Spain’s National AI Strategy 2024 [25], and ALIA infrastructure documentation. This document analysis follows the systematic document review principles outlined by Scott [84], assessing each source for authenticity (verifiable provenance), credibility (likely accuracy), representativeness (whether it reflects the broader institutional reality), and interpretive meaning (relevance to the framework’s analytical categories).
Madrid is treated as an illustrative analytical site rather than a comprehensive case study [46]. The purpose of the secondary analysis is to apply framework components to a concrete, documented institutional context—illustrating how the conceptual model operationalises in practice and identifying implementation patterns relevant to comparable mid-sized European municipalities. This approach cannot substitute for primary data collection through interviews, longitudinal observation, or administrative records analysis. The limitations of the document-based analysis are explicitly acknowledged in Section 6.1.

4. Results: Evidence Synthesis

4.1. Descriptive Characteristics of Included Studies

The 68 included studies span 22 countries (Table 2), with the highest representation from Western Europe (n = 31, 45.6%), North America (n = 16, 23.5%), East Asia (n = 11, 16.2%), and other regions (n = 10, 14.7%). Government-level distribution: National/federal government (n = 38, 55.9%), regional/state (n = 18, 26.5%), local/municipal (n = 12, 17.6%). The low share of local government studies underscores the gap identified in Section 1. Publication years show acceleration over time, with 61.8% of included studies published from 2020 onwards, reflecting growing empirical attention to AI in government following the COVID-19 pandemic’s acceleration of digital transformation [22].
Administrative domains covered include: citizen services and digital channels (n = 24), financial administration and fraud detection (n = 18), regulatory compliance and inspection (n = 12), human resources and workforce management (n = 8), and urban infrastructure and planning (n = 6). AI technologies studied span robotic process automation (RPA) (n = 28), machine learning/predictive analytics (n = 22), natural language processing/chatbots (n = 14), and computer vision/sensor AI (n = 4).

4.2. Findings by AI Pathway

4.2.1. Automation Pathway (n = 28 Studies)

Robotic process automation and related technologies applied to high-volume, rule-based tasks—document processing, data entry, eligibility verification, and invoice management—demonstrate the most consistent productivity evidence. Across the 28 studies classified under the automation pathway, effect sizes include: time savings of 50–70% for standardised document processing [57,58]; cost reductions of 30–50% for back-office administrative functions [11,57]; error rate reductions of 60–80% in data entry and classification tasks [19]; and throughput increases of 40–150% for permit processing and benefit eligibility assessment [15]. Effects are largest for highly standardised, repetitive processes and diminish with task complexity. The median implementation period in automation studies is 14 months. Critical success factors identified include data quality (cited in 78% of studies; see also [85] for the consequences of poor data quality on AI outcomes), change management investment (cited in 71%), and integration with legacy systems (cited in 64%).

4.2.2. Augmentation Pathway (n = 22 Studies)

Decision support systems applied to professional tasks—case prioritisation, risk assessment, fraud detection, and resource allocation—demonstrate quality improvements and moderate speed gains. Across 22 augmentation pathway studies, documented effects include: decision accuracy improvements of 15–35% in fraud detection and case risk assessment [17,32]; speed improvements of 20–30% in professional decision cycles [11,30]; and inspection targeting improvements of 30–50% through predictive analytics [17]. Evidence of skill-levelling effects is emerging [30]: less-experienced workers show disproportionately larger performance gains, suggesting AI augmentation may reduce within-workforce skill gaps. Critical success factors include system explainability and transparency (cited in 85% of studies), user trust calibration (cited in 77%), and integration with existing professional workflows (cited in 68%).

4.2.3. Transformation Pathway (n = 18 Studies)

Studies examining qualitatively new AI capabilities are fewer and typically of lower quality, with predominantly descriptive findings. Reported effects include: 40% improvement in targeted inspection outcomes through predictive risk scoring [17]; 60% reduction in citizen service wait times through AI-powered multi-channel service platforms [19]; improved policy simulation and scenario testing capabilities in urban planning contexts [50,62]; and enhanced real-time resource allocation in emergency and public safety services [86]. These studies are predominantly qualitative or descriptive (n = 11, 61%), limiting quantitative synthesis. The absence of rigorous counterfactual designs represents a significant evidence gap, as it remains unclear whether observed service improvements are attributable to AI transformation or concurrent organisational changes.

4.3. Evidence Quality Assessment

Quality assessment using the adapted GRADE framework reveals a clear pattern. Of 68 included studies, 22 (32.4%) are rated as high quality, employing quasi-experimental designs, validated outcome measures, and rigorous confounding control; 34 (50.0%) are rated as moderate quality, with established but not validated measures and limited confounding control; and 12 (17.6%) are rated as low quality, relying on single-case, short-term observations with ad hoc measurement. Only eight studies (11.8%) employ quasi-experimental designs with comparison groups—a critical limitation for causal inference. Only five studies (7.4%) examine effects beyond 24 months, precluding assessment of long-term productivity sustainability. Local government studies (n = 12) are predominantly of moderate or low quality (n = 9, 75%), reflecting the nascent state of rigorous evaluation at the municipal level.
Sensitivity analysis comparing findings across quality tiers reveals that high-quality studies report systematically smaller effect sizes than low-quality studies (e.g., automation time savings: high-quality median 52% vs. low-quality median 68%), consistent with publication bias and methodological inflation in weaker studies [83]. These findings reinforce the need for caution when interpreting aggregate effect estimates.

4.4. Mediating Factors: Synthesis

Across all 68 studies, five mediating factors are consistently identified as conditioning AI productivity outcomes: (a) digital infrastructure quality and data governance maturity (cited in 79% of studies), consistent with Gil-Garcia et al.’s [27] digital maturity framework; (b) workforce capabilities and change readiness (cited in 72%), consistent with Savaget et al.’s [33] analysis of e-government readiness; (c) leadership support and organisational commitment (cited in 68%), consistent with Orlikowski’s [51] enacted technology theory; (d) process standardisation and redesign investment (cited in 61%), consistent with Lindgren et al.’s [35] value creation model; and (e) governance frameworks and accountability mechanisms (cited in 54%), consistent with Busuioc’s [68] algorithmic accountability framework. These factors are not independent: digital maturity and workforce capabilities exhibit strong co-occurrence, suggesting that productivity-enabling conditions form clusters rather than isolated factors.
Table 3 shows the distribution of the reviewed studies across the three AI pathways, highlighting the relative weight of each cluster in the existing literature.

5. Discussion

5.1. Interpreting the Evidence

Systematic review of 68 studies reveals measured productivity improvements ranging from 15 to 40% in specific administrative processes where AI has been implemented. However, this aggregate conceals important heterogeneity. The automation pathway consistently demonstrates larger and more reliable effects (median time savings: 52%, n = 28) than the augmentation pathway (median decision quality improvement: 24%, n = 22) or the transformation pathway (varied and less quantifiable, n = 18). This pattern is consistent with theoretical expectations: automation substitutes for well-defined, repetitive tasks where AI systems can achieve near-human or super-human performance on narrow dimensions; augmentation and transformation require more complex organisational adaptation and longer learning periods.
The dominance of short-term, single-case studies in the evidence base (82% of studies, 12-month median horizon) limits the generalisability of the findings and prevents assessment of long-term productivity sustainability—a critical gap given Proposition P4’s J-curve prediction. The finding that high-quality studies systematically report smaller effect sizes than low-quality studies suggests that prevailing aggregate estimates in the literature may overstate AI productivity potential. This observation aligns with concerns raised by Brynjolfsson and Hitt [56] regarding IT productivity measurement methodology and by Borenstein et al. [83] regarding publication bias in systematic reviews.

5.2. Madrid in Context: Institutional Enablers and Barriers

Madrid’s documented AI deployment illustrates framework components and reveals implementation dynamics characteristic of mid-sized European municipalities. The municipality’s AI Roadmap specifies governance structures (AI ethics committee, data stewardship roles), capability development programmes (staff training, data literacy), and portfolio management approaches that align with the measurement model’s mediating factors [44]. The platform procurement awarded to Atos for evolution and maintenance signals institutionalisation beyond pilot status and multi-year commitment, consistent with Orlikowski’s [51] argument that sustained organisational embedding—not initial deployment—determines technology productivity outcomes.
Institutional enablers: Spain’s ALIA infrastructure (coordinated by the Barcelona Supercomputing Center) provides sovereign, Spanish-language large language model (LLM) capabilities, reducing dependence on commercial providers and lowering entry barriers for mid-sized municipalities [25]. EU Digital Europe Programme and Recovery and Resilience Facility funding partially offsets fiscal constraints. The Madrid Data & AI Hub creates knowledge-sharing mechanisms and economies of scale across regional government entities. Compliance with the EU AI Act, while imposing governance requirements, provides legal clarity and reduces procurement uncertainty [24].
Persistent barriers: Limited AI literacy among civil servants, cultural resistance to process automation, and insufficient data science capability constrain adoption pace and implementation quality [33,52]. Fragmented data systems, quality inconsistencies, and interoperability challenges across departmental silos restrict the data inputs required for AI system performance. Political budget cycles create uncertainty for multi-year technology investments, while procurement rigidities slow deployment timelines. Critically, insufficient baseline measurement and monitoring systems prevent rigorous productivity impact assessment—a fundamental governance failure that limits organisational learning [5].

5.3. Theoretical Implications

The framework advances public administration theory in three ways. First, by operationalising TFP concepts for AI-driven productivity assessment, we bridge productivity economics and public administration scholarship. Traditional productivity measurement in government emphasises input reduction or simple output volume; the TFP approach directs attention to multifactor efficiency across labour, capital, and organisational capabilities simultaneously [28,47]. Second, integration of public value dimensions addresses the risk of efficiency optimisation that undermines democratic legitimacy or equity. By specifying governance quality, transparency, and fairness as productivity-enabling rather than productivity-constraining factors, the framework resists narrow technocratic rationality and responds to calls for research bridging efficiency and democratic values [29,55,87].
Third, the three-pathway model (automation, augmentation, transformation) provides analytical clarity for understanding AI’s heterogeneous effects. Automation substitutes for labour, augmentation complements it, and transformation creates qualitatively new capabilities—each with distinct TFP mechanisms and organisational implications. This typology enriches scholarship on algorithmic governance and street-level bureaucracy by specifying how different AI applications reshape work differently [67,68]. The sociotechnical perspective, following Trist and Bamforth [43] and Orlikowski [51], emphasises that productivity outcomes are not technologically determined but emerge from the co-evolution of AI systems, organisational routines, and governance arrangements.

5.4. Policy Implications for Mid-Sized European Cities

For municipal administrators and policy-makers in mid-sized European cities, evidence synthesis and framework application yield five actionable implications.
First, prioritise process assessment over technology selection. Identify high-volume, rule-based processes as automation candidates and complex professional tasks as augmentation opportunities. Process maturity—documentation quality, standardisation degree, and data availability—predicts AI productivity potential more reliably than technological sophistication [27,35].
Second, invest in complementary capabilities. Allocate 30–40% of AI budgets to workforce training, data quality improvement, and process redesign. Technology-only implementations consistently underperform sociotechnical transformations in the evidence base (cited in 61% of included studies). Build data governance frameworks before deploying AI systems rather than concurrently [26,68,88].
Third, leverage shared infrastructure. Mid-sized cities lack resources to develop proprietary AI capabilities at scale. National platforms (ALIA model), regional collaborations (Madrid Data & AI Hub model), and inter-municipal consortia enable economies of scale and risk distribution [41,50]. Prioritise open-source, sovereign solutions over proprietary vendor dependencies that create lock-in and sovereignty risks.
Fourth, implement rigorous evaluation from inception. Establish baseline measurements before AI deployment. Specify productivity metrics aligned with the measurement model (volume, quality, timeliness, equity). Plan for 3–5-year evaluation horizons to capture learning curves and sustained effects, consistent with Proposition P4 [56]. Employ quasi-experimental designs—difference-in-differences, synthetic control, interrupted time series—where feasible to support causal inference [83,89].
Fifth, embed governance from the outset. EU AI Act compliance requires transparency, bias monitoring, and human oversight for high-risk public administration applications. Rather than treating these requirements as compliance burdens, integrate them as productivity-enabling mechanisms that build citizen trust and ensure long-term sustainability [24,31]. Establish AI ethics committees, stakeholder engagement processes, and algorithmic accountability frameworks concurrently with technology deployment, not as afterthoughts.

6. Limitations and Future Research Agenda

6.1. Study Limitations

This study has six significant limitations that readers should consider when interpreting findings. First, while the systematic review covers 68 peer-reviewed studies, evidence heterogeneity precludes meta-analytic synthesis. Effect sizes reported reflect specific implementations in particular organizational and institutional contexts rather than generalisable population-level expectations. The 15–40% aggregate productivity improvement range encompasses considerable variance, and context-specific moderators substantially determine where within this range any given implementation will fall.
Second, local government evidence remains sparse. Of 68 studies, only 12 (17.6%) focus specifically on municipal administration, and only three address Southern European contexts. Generalisability of findings to local government requires caution, particularly given the distinct institutional environments, fiscal frameworks, and administrative cultures of Southern European municipalities compared to the predominantly Anglo-American and Scandinavian contexts represented in the high-quality evidence.
Third, the Madrid analysis relies on secondary sources—policy documents, procurement materials, and institutional reports—rather than primary data collection. While document analysis provides contextual illustration, it cannot rigorously assess actual productivity outcomes, implementation fidelity, or the mechanisms through which AI affects organisational practice. A comprehensive case study would require semi-structured interviews with officials, longitudinal observation, and analysis of administrative performance records [46].
Fourth, the framework’s TFP operationalisation remains conceptual. We specify measurement components but do not empirically test them or demonstrate their psychometric properties. Framework validation requires application to real municipal datasets with before–after measurement or cross-sectional benchmarking across municipalities with varying AI adoption levels.
Fifth, publication bias is a significant concern for any systematic review. Studies demonstrating large productivity gains are more likely to be published than null findings or implementation failures [83]. This bias, combined with the methodological inflation observed in lower-quality studies, suggests that aggregate effect estimates may overstate AI productivity potential in government. Sensitivity analysis comparing high-quality and low-quality study findings (Section 4.3) partially addresses this limitation, but prospective registration of AI productivity studies and mandatory reporting of null findings would substantially strengthen the evidence base.
Sixth, distributional effects and long-term sustainability remain under-explored. Most evidence derives from short-term implementations. Whether productivity gains sustain, accelerate, or diminish over longer periods—and how gains and adjustment costs are distributed across workers, service users, and communities—requires investigation beyond the scope of current evidence.

6.2. Future Research Agenda

These limitations suggest five research priorities for advancing scholarship on AI and public sector productivity.
Priority 1: Longitudinal impact studies. Design and implement studies tracking productivity metrics over 3–5 years in municipalities deploying AI systems. Panel data designs with matched comparison cities would enable difference-in-differences estimation to isolate AI effects from broader organisational and technological trends [83,90]. Monitoring temporal patterns—learning curves, saturation points, sustainability, and distributional dynamics—would provide evidence to test Proposition P4 (J-curve dynamics) empirically.
Priority 2: Comparative municipal analysis. Develop standardised productivity measurement protocols enabling cross-city benchmarking aligned with the measurement model specified in Section 2.5. Quantitative comparative analysis examining how contextual factors (digital maturity, fiscal capacity, population size, governance structure) moderate AI productivity effects would test Propositions P3 and P7. Gil-Garcia and Pardo [62] and Caragliu et al. [50] provide methodological frameworks for comparative smart government research that could be adapted for this purpose.
Priority 3: Public value trade-off assessment. Examine tensions between efficiency gains and equity, transparency, and democratic accountability. Methodologies for measuring and valuing multi-dimensional public value outcomes—including citizen trust, perceived fairness of algorithmic decisions, and distributional equity—are underdeveloped in the existing literature [29,55]. Research should investigate whether AI-driven efficiency gains can coexist with enhanced accountability (Proposition P6) or whether fundamental trade-offs constrain productivity optimisation.
Priority 4: Implementation process research. In-depth qualitative and case study research examining how organisational, political, and institutional factors shape AI adoption and productivity realisation would complement quantitative synthesis. Particular attention should be paid to Southern European contexts—Spain, Italy, Greece, Portugal—where institutional environments differ substantially from the Anglo-American and Nordic cases that dominate current evidence. Participatory action research designs engaging civil servants as research partners could generate both rigorous evidence and practical implementation guidance [91].
Priority 5: Shared infrastructure evaluation. Assess the productivity and cost-effectiveness of collaborative AI infrastructure approaches (ALIA, regional hubs, inter-municipal consortia) compared to proprietary development or commercial cloud services. Examine governance challenges in shared platforms—including data sovereignty, accountability assignment, and inter-organisational coordination costs—and identify optimal organisational models for mid-sized cities that balance capability access with sovereignty and democratic control [27,41].

7. Conclusions

This article addresses the question of how artificial intelligence affects public sector productivity by developing a conceptual framework, synthesising empirical evidence, and examining institutional contexts shaping AI deployment in local government. Contributions span theoretical, empirical, and practical domains.
Theoretically, the framework operationalises total factor productivity logic for AI-enabled government administration, specifying three pathways (automation, augmentation, transformation) through which AI affects productivity, mapping these to TFP components, and integrating public value dimensions. The measurement model provides concrete guidance for productivity assessment while embedding equity, transparency, and accountability as productivity-enabling rather than productivity-constraining factors. Seven propositions offer testable hypotheses grounded in sociotechnical systems theory, productivity economics, and public administration scholarship.
Empirically, systematic review of 68 studies reveals measured productivity improvements of 15–40% in specific administrative processes, with the largest and most reliable effects in high-volume, standardised automation targets, and quality improvements from professional augmentation. However, evidence quality varies substantially, high-quality studies report systematically smaller effects than lower-quality counterparts, local government studies are scarce, and long-term impacts remain unassessed. The evidence base, while demonstrating genuine productivity potential, demands caution against extrapolation beyond the specific contexts studied.
Practically, structured analysis of Madrid’s AI deployment—situated within Spain’s ALIA infrastructure and the EU AI Act framework—illustrates how mid-sized European cities can leverage shared platforms and multi-level governance to overcome resource constraints while navigating compliance requirements. Yet, institutional enablers alone prove insufficient; realising productivity gains requires sustained investment in workforce capabilities, data governance, and evaluation capacity. The policy implications emphasise sociotechnical transformation over technology acquisition, rigorous baseline measurement, and governance integration from inception.
The AI productivity puzzle in public administration remains partially unsolved. Several questions persist: Will productivity improvements sustain over time or diminish as automation efficiency gains are exhausted? How do distributional effects evolve—who benefits and who bears transition costs? Can AI-driven efficiency gains coexist with enhanced equity and democratic accountability, or do fundamental trade-offs constrain the productivity frontier? Do shared infrastructure models achieve comparable productivity outcomes to proprietary investments while maintaining sovereignty and democratic control?
For municipal administrators in mid-sized European cities, the central message is one of informed optimism tempered by empirical realism. AI offers genuine productivity potential, particularly through automation of high-volume routine tasks and augmentation of professional judgement. Realising this potential demands more than technology deployment: it requires organisational transformation, realistic expectations about time horizons (3–5 years), complementary capability investment, and governance frameworks ensuring accountability and citizen trust. The European institutional context—characterised by the AI Act’s governance requirements, Digital Europe Programme funding, and emerging shared infrastructure—creates both constraints and genuine opportunities that, thoughtfully navigated, can enable mid-sized cities to access AI productivity benefits previously reserved for large metropolitan authorities or well-resourced Nordic administrations.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/systems14060671/s1, Table S1: PRISMA 2020 Checklist; Table S2: PRISMA extensions; Table S3: ENTREQ for qualitative research review; Table S4: PRISMA-S; Table S5: EQUATOR Network for additional guidelines for systematic reviews. Reference [92] is cited in the Supplementary Materials.

Author Contributions

A.O.: Conceptualization, Methodology, Formal Analysis, Writing—Original Draft, Writing—Review and Editing. C.D.-P.-H.: Conceptualization, Supervision, Writing—Review and Editing, Validation. All authors have read and agreed to the published version of the manuscript.

Funding

This research article has been financed with funds from OPENINNOVA High Performance Research Group (URJC-V1717).

Institutional Review Board Statement

This study involved systematic analysis of the published peer-reviewed literature and publicly available institutional documents. No primary data collection from human subjects was conducted. Institutional review board approval was not required.

Data Availability Statement

The systematic review protocol, complete search terms, PRISMA 2020 flowchart, quality assessment rubric, and data extraction tables are provided in the Supplementary Materials. Study screening data and quality assessment scores are available upon request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Pollitt, C. Public Management Reform: A Comparative Analysis—Into the Age of Austerity, 4th ed.; Oxford University Press: Oxford, UK, 2018. [Google Scholar]
  2. OECD. The OECD Digital Government Policy Framework: Six Dimensions of a Digital Government; OECD Public Governance Policy Papers No. 02; OECD Publishing: Paris, France, 2020. [Google Scholar] [CrossRef] [Scilit]
  3. Dunleavy, P.; Margetts, H.; Bastow, S.; Tinkler, J. New public management is dead—Long live digital-era governance. J. Public Adm. Res. Theory 2006, 16, 467–494. [Google Scholar] [CrossRef] [Scilit]
  4. Atkinson, T. Atkinson Review: Final Report on Measurement of Government Output and Productivity for the National Accounts; Palgrave Macmillan: Basingstoke, UK, 2005. [Google Scholar]
  5. Van Dooren, W.; Bouckaert, G.; Halligan, J. Performance Management in the Public Sector, 2nd ed.; Routledge: Abingdon, UK, 2015. [Google Scholar] [CrossRef] [Scilit]
  6. European Committee of the Regions. Informe Anual de la UE Sobre el Estado de las Regiones y las Ciudades 2025; European Committee of the Regions: Brussels, Belgium, 2025; Available online: https://cor.europa.eu (accessed on 13 March 2026).
  7. ESPON EGTC. DESIRE—Analysis on Provision of Public Services in Disadvantaged Territories; ESPON: Luxembourg, 2024. [Google Scholar]
  8. Council of the European Union. Recommendation on Spain’s Medium-Term Fiscal-Structural Plan; Council of the EU: Brussels, Belgium, 2025. [Google Scholar]
  9. Acemoglu, D.; Restrepo, P. Robots and jobs: Evidence from US labor markets. J. Polit. Econ. 2020, 128, 2188–2244. [Google Scholar] [CrossRef] [Scilit]
  10. Brynjolfsson, E.; McAfee, A. The Second Machine Age: Work, Progress, and Prosperity in a Time of Brilliant Technologies; W.W. Norton: New York, NY, USA, 2014. [Google Scholar]
  11. Wirtz, B.W.; Weyerer, J.C.; Geyer, C. Artificial intelligence and the public sector—Applications and challenges. Int. J. Public Adm. 2019, 42, 596–615. [Google Scholar] [CrossRef] [Scilit]
  12. Berryhill, J.; Heang, K.K.; Clogher, R.; McBride, K. Hello, World: Artificial Intelligence and Its Use in the Public Sector; OECD Working Papers on Public Governance No. 36; OECD: Paris, France, 2019. [Google Scholar]
  13. Raisch, S.; Krakowski, S. Artificial intelligence and management: The automation–augmentation paradox. Acad. Manag. Rev. 2021, 46, 192–210. [Google Scholar] [CrossRef] [Scilit]
  14. Haenlein, M.; Kaplan, A. A brief history of artificial intelligence: On the past, present, and future of artificial intelligence. Calif. Manag. Rev. 2019, 61, 5–14. [Google Scholar] [CrossRef] [Scilit]
  15. Lacity, M.C.; Willcocks, L.P. A new approach to automating services. MIT Sloan Manag. Rev. 2016, 58, 41–49. [Google Scholar]
  16. Ranerup, A.; Henriksen, H.Z. Digital discretion: Unpacking human and technological agency in automated decision making in Sweden’s Social Insurance Agency. Soc. Policy Adm. 2019, 53, 507–522. [Google Scholar] [CrossRef] [Scilit]
  17. Soni, L.; Taneja, A. Harnessing AI for sustainable smart cities: Impact, innovations, and use cases. In Artificial Intelligence (AI) for IT Energy Efficiency and Green AI for Environment Sustainability; Raj, P., Sharma, D.P., Dutta, P.K., Prasad, B.S., Soundarabai, P.B., Eds.; Springer: Cham, Switzerland, 2026; pp. 471–496. [Google Scholar] [CrossRef] [Scilit]
  18. Janssen, M.; van der Voort, H. Adaptive governance: Towards a stable, accountable and responsive government. Gov. Inf. Q. 2016, 33, 1–5. [Google Scholar] [CrossRef] [Scilit]
  19. Fountain, J.E. Building the Virtual State: Information Technology and Institutional Change; Brookings Institution Press: Washington, DC, USA, 2001. [Google Scholar]
  20. Mergel, I.; Edelmann, N.; Haug, N. Defining digital transformation: Results from expert interviews. Gov. Inf. Q. 2019, 36, 101385. [Google Scholar] [CrossRef] [Scilit]
  21. Criado, J.I.; Gil-Garcia, J.R. Creating public value through smart technologies and strategies. Int. J. Public Sect. Manag. 2019, 32, 438–450. [Google Scholar] [CrossRef] [Scilit]
  22. Meijer, A.; Bolívar, M.P.R. Governing the smart city: A review of the literature on smart urban governance. Int. Rev. Adm. Sci. 2016, 82, 392–408. [Google Scholar] [CrossRef] [Scilit]
  23. Wirtz, B.W.; Müller, W.M. An integrated artificial intelligence framework for public management. Public Manag. Rev. 2019, 21, 1076–1100. [Google Scholar] [CrossRef] [Scilit]
  24. European Parliament and Council. Regulation (EU) 2024/1689 of the European Parliament and of the Council Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act); Official Journal of the EU: Brussels, Belgium, 2024. [Google Scholar]
  25. Government of Spain. National Artificial Intelligence Strategy (ENIA 2024); Ministry for Digital Transformation and Civil Service: Madrid, Spain, 2024. [Google Scholar]
  26. Janssen, M.; Kuk, G. The challenges and limits of big data algorithms in technocratic governance. Gov. Inf. Q. 2016, 33, 371–377. [Google Scholar] [CrossRef] [Scilit]
  27. Gil-Garcia, J.R.; Pardo, T.A.; Burke, G.B. Conceptualizing smartness in government: An integrative and multi-dimensional view. Gov. Inf. Q. 2016, 33, 524–534. [Google Scholar] [CrossRef] [Scilit]
  28. Comin, D. Total factor productivity. In The New Palgrave Dictionary of Economics; Durlauf, S.N., Blume, L.E., Eds.; Palgrave Macmillan: London, UK, 2008. [Google Scholar]
  29. Moore, M.H. Creating Public Value: Strategic Management in Government; Harvard University Press: Cambridge, MA, USA, 1995. [Google Scholar]
  30. Brynjolfsson, E.; Li, D.; Raymond, L. Generative AI at work. Q. J. Econ. 2025, 140, 889–942. [Google Scholar] [CrossRef] [Scilit]
  31. Dunleavy, P.; Carrera, L. Growing the Productivity of Government Services; Edward Elgar: Cheltenham, UK, 2013. [Google Scholar]
  32. Meijer, A.; Thaens, M. Predictive policing: Review of benefits and drawbacks. Int. J. Public Adm. 2021, 44, 1176–1188. [Google Scholar] [CrossRef] [Scilit]
  33. Savaget, P.; Chiarini, T.; Evans, S. Empowering political contestation and collective action with AI. Nat. Mach. Intell. 2019, 1, 394–397. [Google Scholar] [CrossRef]
  34. Dunleavy, P.; Margetts, H. Design principles for essentially digital governance. In Proceedings of the 111th Annual Meeting of the American Political Science Association, San Francisco, CA, USA, 3–6 September 2015. [Google Scholar]
  35. Lindgren, I.; Madsen, C.Ø.; Hofmann, S.; Melin, U. Close encounters of the digital kind: A research agenda for the digitalization of public services. Gov. Inf. Q. 2019, 36, 427–436. [Google Scholar] [CrossRef] [Scilit]
  36. Criado, J.I.; Sandoval-Almazán, R.; Gil-Garcia, J.R. Artificial intelligence and public administration: Understanding actors, governance, and policy from micro, meso, and macro perspectives. Public Policy Adm. 2024, 40, 173–184. [Google Scholar] [CrossRef] [Scilit]
  37. Secchi, L.; Caeiro, J.C.; Ramos Pinto, R.; Arenilla Sáez, M. Administrative reforms in Portugal and Spain: From bureaucracy to digital transition. Int. Rev. Adm. Sci. 2024, 91, 8–26. [Google Scholar] [CrossRef] [Scilit]
  38. Criado, J.I.; Villodre, J. Delivering public services through social media in European local governments. Local Gov. Stud. 2021, 47, 312–332. [Google Scholar] [CrossRef] [Scilit]
  39. European Commission. Local Government AI Deployment Survey; Joint Research Centre: Brussels, Belgium, 2025. [Google Scholar]
  40. Mehr, H.; Ash, H.; Fellow, D. Artificial Intelligence for Citizen Services and Government; Ash Center for Democratic Governance and Innovation: Cambridge, MA, USA, 2017. [Google Scholar]
  41. Fung, A. Putting the public back into governance: The challenges of citizen participation and its future. Public Adm. Rev. 2015, 75, 513–522. [Google Scholar] [CrossRef] [Scilit]
  42. Gil-Garcia, J.R.; Helbig, N.; Ojo, A. Being smart: Emerging technologies and innovation in the public sector. Gov. Inf. Q. 2014, 31, I1–I8. [Google Scholar] [CrossRef] [Scilit]
  43. Trist, E.L.; Bamforth, K.W. Some social and psychological consequences of the longwall method of coal-getting. Hum. Relat. 1951, 4, 3–38. [Google Scholar] [CrossRef] [Scilit]
  44. Ayuntamiento de Madrid. Madrid Artificial Intelligence Roadmap 2023–2025; Ayuntamiento de Madrid: Madrid, Spain, 2025; Available online: https://www.madrid.es (accessed on 13 March 2026).
  45. Ayuntamiento de Madrid. AI Platform Procurement Contract: Evolution and Maintenance; Ayuntamiento de Madrid: Madrid, Spain, 2025; Available online: https://contrataciondelestado.es (accessed on 13 March 2026).
  46. Yin, R.K. Case Study Research and Applications: Design and Methods, 6th ed.; SAGE: Thousand Oaks, CA, USA, 2018. [Google Scholar]
  47. Solow, R.M. Technical change and the aggregate production function. Rev. Econ. Stat. 1957, 39, 312–320. [Google Scholar] [CrossRef] [Scilit]
  48. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  49. Moher, D.; Liberati, A.; Tetzlaff, J.; Altman, D.G.; PRISMA Group. Preferred reporting items for systematic reviews and meta-analyses: The PRISMA statement. Ann. Intern. Med. 2009, 151, 264–269. [Google Scholar] [CrossRef] [PubMed]
  50. Caragliu, A.; Del Bo, C.; Nijkamp, P. Smart cities in Europe. J. Urban Technol. 2011, 18, 65–82. [Google Scholar] [CrossRef] [Scilit]
  51. Orlikowski, W.J. Using technology and constituting structures: A practice lens for studying technology in organisations. Organ. Sci. 2000, 11, 404–428. [Google Scholar] [CrossRef] [Scilit]
  52. Picazo-Vela, S.; Gutiérrez-Martínez, I.; Luna-Reyes, L.F. Understanding risks, benefits, and strategic alternatives of social media applications in the public sector. Gov. Inf. Q. 2012, 29, 504–511. [Google Scholar] [CrossRef] [Scilit]
  53. Sousa, W.G.; Melo, E.R.P.; Bermejo, P.H.R.S.; Farias, R.A.S.; Gomes, A.O. How and where is artificial intelligence in the public sector going? A literature review and research agenda. Gov. Inf. Q. 2019, 36, 101392. [Google Scholar] [CrossRef] [Scilit]
  54. Bozeman, B. Public Values and Public Interest: Counterbalancing Economic Individualism; Georgetown University Press: Washington, DC, USA, 2007. [Google Scholar]
  55. Moore, M.H.; Hartley, J. Innovations in governance. Public Manag. Rev. 2008, 10, 3–20. [Google Scholar] [CrossRef] [Scilit]
  56. Brynjolfsson, E.; Hitt, L.M. Computing productivity: Firm-level evidence. Rev. Econ. Stat. 2003, 85, 793–808. [Google Scholar] [CrossRef] [Scilit]
  57. Willcocks, L.; Lacity, M.; Craig, A. The IT function and robotic process automation. OUWP 2015, 16, 1–43. [Google Scholar]
  58. van der Aalst, W.M.P.; Bichler, M.; Heinzl, A. Robotic process automation. Bus. Inf. Syst. Eng. 2018, 60, 269–272. [Google Scholar] [CrossRef] [Scilit]
  59. Acemoglu, D.; Restrepo, P. Tasks, automation, and the rise in US wage inequality. Econometrica 2022, 90, 1973–2016. [Google Scholar] [CrossRef] [Scilit]
  60. Noy, S.; Zhang, W. Experimental evidence on the productivity effects of generative AI. Science 2023, 381, 187–192. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  61. Parasuraman, R.; Riley, V. Humans and automation: Use, misuse, disuse, abuse. Hum. Factors 1997, 39, 230–253. [Google Scholar] [CrossRef] [Scilit]
  62. Gil-Garcia, J.R.; Pardo, T.A.; Nam, T. What makes a city smart? Identifying core components and proposing an integrative and comprehensive conceptualization. Inf. Polity 2015, 20, 61–87. [Google Scholar] [CrossRef] [Scilit]
  63. Margetts, H.; Dorobantu, C. Rethink government with AI. Nature 2019, 568, 163–165. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  64. Alford, J.; Hughes, O. Public value pragmatism as the next phase of public management. Am. Rev. Public Adm. 2008, 38, 130–148. [Google Scholar] [CrossRef] [Scilit]
  65. Dwivedi, A.V. Why AI cannot be an agent of equality: A critique of technological utopianism and the category error of algorithmic justice. Emerg. Media 2025, 3, 428–450. [Google Scholar] [CrossRef] [Scilit]
  66. O’Neil, C. Weapons of Math Destruction: How Big Data Increases Inequality and Threatens Democracy; Crown: New York, NY, USA, 2016. [Google Scholar]
  67. Bovens, M.; Zouridis, S. From street-level to system-level bureaucracies: How information and communication technology is transforming administrative discretion and constitutional control. Public Adm. Rev. 2002, 62, 174–184. [Google Scholar] [CrossRef] [Scilit]
  68. Busuioc, M. Accountable artificial intelligence: Holding algorithms to account. Public Adm. Rev. 2021, 81, 825–836. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  69. Panarese, P.; Grasso, M.M.; Solinas, C. Algorithmic bias, fairness, and inclusivity: A multilevel framework for justice-oriented AI. AI Soc. 2026, 41, 2803–2825. [Google Scholar] [CrossRef] [Scilit]
  70. Eubanks, V. Automating Inequality: How High-Tech Tools Profile, Police, and Punish the Poor; St. Martin’s Press: New York, NY, USA, 2018. [Google Scholar]
  71. Obermeyer, Z.; Powers, B.; Vogeli, C.; Mullainathan, S. Dissecting racial bias in an algorithm used to manage the health of populations. Science 2019, 366, 447–453. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  72. Virani, A.; van der Wal, Z. Performance Regimes and Governance Design: Balancing Quality, Timeliness, and Responsiveness in Public Service Delivery. Public Adm. Rev. 2023, 83, 112–128. [Google Scholar] [CrossRef] [Scilit]
  73. Gębczyńska, M.; Brajer-Marczak, R. Performance management models in public administration. Int. J. Public Adm. 2020, 43, 521–534. [Google Scholar] [CrossRef] [Scilit]
  74. Boyne, G.A.; Meier, K.J.; O’Toole, L.J.; Walker, R.M. Public Service Performance: Perspectives on Measurement and Management; Cambridge University Press: Cambridge, UK, 2006. [Google Scholar] [CrossRef] [Scilit]
  75. Fernández-i-Marín, X.; Jiménez, G.; Pau, I.; Prats, J. Administrative capacity and implementation speed. Gov. Inf. Q. 2023, 40, 101802. [Google Scholar] [CrossRef] [Scilit]
  76. Kulal, A.; Rahiman, H.U.; Suvarna, H.; Abhishek, N.; Dinesh, S. Enhancing public service delivery effi-ciency: Exploring the impact of AI. J. Open Innov. Technol. Mark. Complex. 2024, 10, 100329. [Google Scholar] [CrossRef] [Scilit]
  77. Petticrew, M.; Roberts, H. Systematic Reviews in the Social Sciences: A Practical Guide; Blackwell: Oxford, UK, 2006. [Google Scholar] [CrossRef] [Scilit]
  78. LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  79. Cohen, J. A coefficient of agreement for nominal scales. Educ. Psychol. Meas. 1960, 20, 37–46. [Google Scholar] [CrossRef] [Scilit]
  80. Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
  81. Guyatt, G.H.; Oxman, A.D.; Vist, G.E.; Kunz, R.; Falck-Ytter, Y.; Alonso-Coello, P.; Schünemann, H.J.; GRADE Working Group. GRADE: An emerging consensus on rating quality of evidence and strength of recommendations. BMJ 2008, 336, 924–926. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  82. Popay, J.; Roberts, H.; Sowden, A.; Petticrew, M.; Arai, L.; Rodgers, M.; Britten, N. Guidance on the Conduct of Narrative Synthesis in Systematic Reviews; ESRC Methods Programme: Lancaster, UK, 2006. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  83. Borenstein, M.; Hedges, L.V.; Higgins, J.P.T.; Rothstein, H.R. Introduction to Meta-Analysis; Wiley: Chichester, UK, 2009. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  84. Scott, J. A Matter of Record: Documentary Sources in Social Research; Polity Press: Cambridge, UK, 1990. [Google Scholar]
  85. Richardson, R.; Schultz, J.; Crawford, K. Dirty data, bad predictions. NYU Law Rev. 2019, 94, 192–233. [Google Scholar]
  86. Vogl, T.M.; Seidelin, C.; Ganesh, B.; Bright, J. Smart technology and the emergence of algorithmic bureaucracy: Artificial intelligence in UK local authorities. Public Adm. Rev. 2020, 80, 946–961. [Google Scholar] [CrossRef] [Scilit]
  87. Acemoglu, D.; Johnson, S.; Lelarge, C.; Van Reenen, J.; Zilibotti, F. Taxes, regulations, and the value of U.S. and European firms. Rev. Econ. Stud. 2022, 89, 2279–2329. [Google Scholar]
  88. Desouza, K.C.; Dawson, G.S.; Chenok, D. Designing, developing, and deploying artificial intelligence systems: Lessons from and for the public sector. Bus. Horiz. 2020, 63, 205–213. [Google Scholar] [CrossRef] [Scilit]
  89. Blundell, R.; Costa Dias, M. Alternative approaches to evaluation in empirical microeconomics. J. Hum. Resour. 2009, 44, 565–640. [Google Scholar] [CrossRef] [Scilit]
  90. Shadish, W.R.; Cook, T.D.; Campbell, D.T. Experimental and Quasi-Experimental Designs for Generalized Causal Inference; Houghton Mifflin: Boston, MA, USA, 2002. [Google Scholar]
  91. Greenwood, D.J.; Levin, M. Introduction to Action Research: Social Research for Social Change, 2nd ed.; SAGE: Thousand Oaks, CA, USA, 2007. [Google Scholar]
  92. Tong, A.; Flemming, K.; McInnes, E.; Oliver, S.; Craig, J. Enhancing transparency in reporting the synthesis of qualitative research: ENTREQ. BMC Med. Res. Methodol. 2012, 12, 181. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Integrative framework of total factor productivity (TFP) for AI-driven productivity in public administration, showing the three AI pathways (automation, augmentation, transformation) and their connections to mediating factors and productivity outcomes.
Figure 1. Integrative framework of total factor productivity (TFP) for AI-driven productivity in public administration, showing the three AI pathways (automation, augmentation, transformation) and their connections to mediating factors and productivity outcomes.
Systems 14 00671 g001
Figure 2. PRISMA 2020 flow diagram illustrating the identification, screening, eligibility, and inclusion process for studies included in the systematic review.
Figure 2. PRISMA 2020 flow diagram illustrating the identification, screening, eligibility, and inclusion process for studies included in the systematic review.
Systems 14 00671 g002
Table 1. PRISMA 2020 study selection flow.
Table 1. PRISMA 2020 study selection flow.
Selection StageNumber of Records
Initial records identified1247
  — Web of Science412
  — Scopus568
  — Google Scholar267
Duplicates removed289
Records screened (title/abstract)958
Records excluded782
Full-text articles assessed for eligibility176
Full-text articles excluded (with reasons)108
Studies included in final review68
Table 2. Descriptive characteristics of included studies (n = 68).
Table 2. Descriptive characteristics of included studies (n = 68).
CharacteristicCategoryn%
RegionWestern Europe3145.6%
North America1623.5%
East Asia1116.2%
Other regions1014.7%
Government LevelNational/Federal3855.9%
Regional/State1826.5%
Local/Municipal1217.6%
AI TechnologyRobotic Process Automation (RPA)2841.2%
Machine Learning/Predictive Analytics2232.4%
Natural Language Processing/Chatbots1420.6%
Computer Vision/Sensor AI45.9%
Evidence QualityHigh Quality2232.4%
Moderate Quality3450.0%
Low Quality1217.6%
Table 3. Distribution of the 68 included studies across the three AI pathways (automation, augmentation, and transformation) identified in the systematic review. The automation pathway accounted for 28 studies (41.2%), followed by augmentation with 22 studies (32.4%) and transformation with 18 studies (26.5%).
Table 3. Distribution of the 68 included studies across the three AI pathways (automation, augmentation, and transformation) identified in the systematic review. The automation pathway accounted for 28 studies (41.2%), followed by augmentation with 22 studies (32.4%) and transformation with 18 studies (26.5%).
PathwayStudies (n)Percentage (%)
Automation (Task Substitution)2841.2
Augmentation (Decision Support)2232.4
Transformation (New Capabilities)1826.5
Total68100.0
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ogunrinde, A.; De-Pablos-Heredero, C. AI Adoption in Local Government: Productivity, Systemic Risk, and Institutional Resilience: Evidence from a PRISMA 2020 Review. Systems 2026, 14, 671. https://doi.org/10.3390/systems14060671

AMA Style

Ogunrinde A, De-Pablos-Heredero C. AI Adoption in Local Government: Productivity, Systemic Risk, and Institutional Resilience: Evidence from a PRISMA 2020 Review. Systems. 2026; 14(6):671. https://doi.org/10.3390/systems14060671

Chicago/Turabian Style

Ogunrinde, Abayomi, and Carmen De-Pablos-Heredero. 2026. "AI Adoption in Local Government: Productivity, Systemic Risk, and Institutional Resilience: Evidence from a PRISMA 2020 Review" Systems 14, no. 6: 671. https://doi.org/10.3390/systems14060671

APA Style

Ogunrinde, A., & De-Pablos-Heredero, C. (2026). AI Adoption in Local Government: Productivity, Systemic Risk, and Institutional Resilience: Evidence from a PRISMA 2020 Review. Systems, 14(6), 671. https://doi.org/10.3390/systems14060671

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop