Next Article in Journal
Machine Learning-Based Crisis Detection Framework for Banking Systems: A Case Study of Nigeria
Next Article in Special Issue
Exploring the AHP-AgileITS-ArchDesign: An AHP Model and Tool for Evaluating IT Service Architectural Agile Designs in SMBs
Previous Article in Journal
Behavioral Biases and Retail Investment Decisions in India: The Moderating Role of Financial Literacy and Financial Awareness
Previous Article in Special Issue
Visualising Machine Learning Model Outputs in Data Analytics: A Systematic Review
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Co-Evolution of Artificial Intelligence and Green Technological Innovation: A Computational Mapping and Diagnostic Framework

1
SUSS Academy, Singapore University of Social Sciences, 463 Clementi Road, Singapore 599494, Singapore
2
School of Business, Singapore University of Social Sciences, 463 Clementi Road, Singapore 599494, Singapore
*
Author to whom correspondence should be addressed.
Analytics 2026, 5(3), 27; https://doi.org/10.3390/analytics5030027
Submission received: 11 June 2026 / Revised: 9 July 2026 / Accepted: 28 July 2026 / Published: 3 August 2026
(This article belongs to the Special Issue Reviews on Data Analytics and Its Applications)

Abstract

Artificial intelligence (AI) is increasingly recognised for its transformative implications for sustainability transitions. Yet little is known about how AI co-evolves with green technological innovation systems and whether institutional adaptation keeps pace with technological diffusion. This study maps 3357 peer-reviewed publications between 2003 and 2025 using transformer-based topic modelling and cross-model triangulation to characterise structural evolution across enabling technologies, sectoral applications, and governance domains. Results reveal a reproducible triadic configuration consistent with innovation-system layering, alongside pronounced asymmetry in growth trajectories. While AI-enabled application domains (e.g., energy systems, waste and circularity, agriculture, and urban mobility) exhibit sustained expansion and increasing specialisation, governance and institutional strands demonstrate thinner density, greater model sensitivity, and delayed acceleration. These patterns are consistent with an asynchronous relationship between technological capability and institutional oversight, echoing the innovation-regulation lag observed in other technology-diffusion processes. The findings contribute to technological change literature by providing systematic evidence on the thematic structure and evolution of AI-enabled sustainability research and by proposing a co-evolution diagnostic framework for monitoring alignment between technological expansion and governance attention.

1. Introduction

The unprecedented scale of global environmental deterioration, encompassing phenomena like climate change, resource depletion, and biodiversity loss, calls for the development and implementation of innovative approaches beyond conventional paradigms [1,2]. Artificial intelligence (AI) has emerged as a transformative enabler, offering significant potential across energy, agriculture, mobility, and industrial systems to accelerate the transition toward sustainability at scale [3,4,5,6]. The intersection of AI and green technological innovation represents a structural reconfiguration of sustainability-oriented innovation systems [7].
The existing literature has extensively documented AI applications in multiple domains, ranging from renewable energy forecasting [8,9] and smart grid optimisation [10,11], to climate modelling [4], biodiversity monitoring [3,12], and sustainable supply chain management [13,14].
However, existing reviews largely remain sector-specific, and a field-level view that integrates enabling technologies, sectoral application pathways, and governance/ethics remains incomplete [3,15,16,17,18,19,20]. At the same time, a countervailing discourse has drawn attention to the AI sustainability paradox–the possibility that environmental gains achieved via AI-enabled efficiency or substitution are offset by the energy and material burdens of the AI pipeline itself, alongside risks associated with opacity, bias, and governance gaps [4,21,22]. Therefore, there remains limited visibility into (i) how technical capabilities map to environmental performance across heterogeneous contexts; (ii) where application clusters are consolidating (maturing) versus intensifying rapidly from low baselines (emerging); and (iii) whether governance activity is co-evolving sufficiently to ensure lifecycle-positive outcomes.
This paper responds to these gaps by advancing an integrative, system-oriented review of the use of AI in green technology innovations that is explicitly designed to couple technical, sectoral, and institutional perspectives. Drawing on a corpus of 3357 peer-reviewed publications (2003–2025; data cutoff on 17 July 2025), we apply large-scale computational mapping to identify structural patterns, temporal trajectories, and cross-layer asymmetries.
The study conceptualises the AI-Sustainability Pathways (AISP) framework as a layered system comprising: (i) Layer 1. Enabling AI capabilities (e.g., predictive analytics, optimisation, autonomous control, computer vision, natural-language processing); (ii) Layer 2. Sectoral Application Pathways (e.g., energy and grids, waste and circularity, agriculture and land use, urban mobility, industrial/process monitoring); and (iii) Layer 3. Governance & Institutional Domains (e.g., lifecycle assessment and carbon accounting, data governance and transparency, standards, policy instruments).
Grounded in systems and transitions theory, AISP provides an interpretive scaffold for examining synchronous growth, specialisation, and consolidation [23,24]. In this study, AISP serves a dual role: (i) as a theory-informed analytical device for classifying computationally derived topics into enabling, sectoral, and institutional layers; and (ii) as a diagnostic framework for assessing whether these layers evolve in synchrony or exhibit persistent asymmetry over time. It advances a governance-ready operating model comprising reporting artefacts and evaluation principles that practitioners and policymakers can adopt to reduce risk displacement and support lifecycle-positive impact [4,21,25]. In doing so, it contributes to understanding technological co-evolution and institutional lag in complex socio-technical systems. AISP framework operationalises a co-evolutionary perspective on sustainability transitions. Specifically, it translates abstract innovation-system concepts, such as technological capability development, sectoral embedding, and institutional adaptation, into observable layers that can be supported by the literature. By linking topic-level evidence to these system layers, the framework enables systematic diagnosis of synchronisation or lag across technological, sectoral, and governance domains. In this way, AISP extends the socio-technical transition theory by providing an empirically supportable representation of how digital technologies integrate into sustainability-oriented innovation systems.
The remainder of the paper proceeds as follows. Section 2 elaborates the theoretical background, synthesises prior reviews, identifies field gaps, and formalises the AISP framework. Section 3 details corpus construction, representation learning, clustering, and labelling protocols. Section 4 reports descriptive innovation trends, topic taxonomies, cross-model integration, and temporal dynamics. Section 5 interprets findings and examines layer alignment. Section 6 sets out academic, practical, and policy implications alongside an illustrative unifying AISP field guide. Section 7 concludes with a forward agenda and reflections on directions for future research on AI-enabled green technology innovation efforts.

2. Background and Conceptual Framework

2.1. Evolution of Artificial Intelligence in Green Technology Innovation

Research at the intersection of AI and green technology innovation has evolved distinct phases characterised not only by technical advancement but also by changing configurations of sectoral application and institutional integration. Rather than a linear sequence of algorithmic improvements, this trajectory reflects the progressive embedding of AI capabilities within broader sustainability-oriented innovation systems.
Early implementations in the 1990s focused on deterministic rule-based systems designed to support environmental assessment and process control [26]. These systems operated largely as decision-support tools within existing environmental management regimes. Although limited in scalability and contextual adaptability, they marked the initial integration of computational reasoning into environmental governance and industrial process optimisation.
The maturation of machine learning in the mid-2000s marked a second phase characterised by enhanced adaptive capacity and cross-sector diffusion. Supervised and unsupervised learning techniques were increasingly applied to renewable energy, emission prediction, and process-safety monitoring. Neural-network applications for photovoltaic power forecasting [8] and deep-learning architectures [9] improved predictive accuracy under variable environmental conditions. Importantly, this period marked a transition from isolated decision-support tools toward performance-enhancing subsystems embedded within energy and industrial infrastructures. AI capabilities began to interact more tightly with sectoral innovation pathways, contributing to incremental efficiency gains and operational resilience.
Parallel developments in smart grids, distributed energy systems, and predictive maintenance further deepened this integration [10,11]. These efforts demonstrated that AI functioned not merely as an optimisation layer but as a coordinating mechanism linking data flows, infrastructure management, and system reliability. Such applications illustrate a shift from localised technical augmentation to more systemic reconfiguration, wherein digital intelligence became intertwined with energy-system innovation trajectories.
In the late 2010s, AI diffusion extended beyond energy infrastructures into agriculture, biodiversity monitoring, climate modelling, and circular economy systems. Computer vision and sensor fusion enabled site-specific agricultural interventions that reduced input intensity [27,28]. Ecological monitoring systems leveraged machine-learning classifiers to enhance species recognition and environmental surveillance [12]. Similar progress occurred in climate science, where AI-driven downscaling and extreme-event prediction enhanced spatial and temporal precision in climate models [3]. This phase reflects cross-domain recombination, where AI capabilities became increasingly modular and transferable across sustainability sectors.
More recently, AI applications in circular economy and waste-management systems have demonstrated growing interconnection between digital capabilities and material-flow infrastructures. Image-recognition robotics, predictive routing, and optimised logistics and recycling-stream analytics [29] illustrate how AI now mediates flows of energy, materials, and information simultaneously. These advances reflect a shift to system-level integration, wherein AI functions as a connective infrastructure across entire green innovation ecosystems.
These phases suggest that the evolution of AI in green technology innovation is characterised by increasing embeddedness within sectoral systems and expanding interaction with governance and institutional arrangements. As AI capabilities diffuse and recombine across domains, they reshape not only technical performance but also coordination mechanisms, regulatory requirements, and evaluation standards. Understanding this trajectory therefore requires a co-evolutionary perspective that considers technological capabilities, sectoral application pathways, and institutional adaptation as interdependent components of sustainability-oriented innovation systems.

2.2. Cross-Domain Integrative Reviews and Emerging Debates

A number of influential reviews have sought to examine the intersection of AI and sustainability from sectoral, interdisciplinary, and policy perspectives. Rolnick et al. provided a seminal interdisciplinary synthesis outlining how machine learning can address climate-change mitigation, adaptation, and resilience across energy, agriculture, and materials domains [3]. Vinuesa et al. extended the discussion to a global policy dimension, mapping AI’s contributions to the United Nations Sustainable Development Goals (SDGs) [4]. These syntheses demonstrate that AI capabilities are increasingly embedded within diverse sustainability domains, enabling optimisation, forecasting, monitoring, and resource coordination at unprecedented scales.
From an innovation-system perspective, however, the diffusion of AI across sustainability sectors raises questions not only about technical performance but also about institutional capacity and governance alignment. As AI capabilities become more deeply integrated into sectoral infrastructures, they alter decision architectures, accountability structures, and performance metrics. In doing so, they generate new complementarities but also new coordination requirements across policy, regulatory, and organisational domains. Such dynamics are consistent with co-evolutionary accounts of technological change, in which innovation diffusion interacts recursively with institutional structures and policy arrangements [16,30,31].
Parallel research has begun to investigate the environmental externalities of AI itself, giving rise to the notion of the sustainability paradox of AI. Strubell et al. quantified the carbon emissions associated with training large-scale language models, revealing energy costs comparable to those of industrial manufacturing processes [22]. Schwartz et al. advanced the concept of Green AI, urging the research community to evaluate algorithmic performance beyond accuracy, in terms of computational efficiency and lifecycle energy intensity [21]. Calls for “Green AI” emphasise the need to evaluate algorithmic systems not solely on predictive accuracy but also on lifecycle energy efficiency and carbon intensity, signalling recognition that performance trade-offs and environmental accounting regimes must co-evolve with technological capability.
Similarly, work on responsible AI and sustainability governance highlights the importance of transparency, accountability, and institutional oversight in ensuring that AI-driven innovation yields net-positive environmental outcomes. Cowls et al. emphasised transparency, accountability, and environmental justice as critical principles for ensuring that AI applications do not exacerbate existing inequalities or introduce new risks [25]. Di Vaio et al. similarly argued for the integration of responsible AI principles into sustainable-business models, highlighting the need for managerial and regulatory frameworks that embed ethical oversight within green technological innovation [15]. These studies suggest that the environmental efficacy of AI is contingent on institutional design and alignment between digital capability, sectoral application, and governance structure.
While existing reviews catalogue applications and articulate governance principles, they provide limited insight into how AI-driven sustainability innovation is structurally organised across domains or whether institutional adaptation proceeds proportionately with technological expansion. Addressing these questions requires moving toward a system-level perspective capable of examining structural patterns and temporal dynamics across domains. A co-evolutionary lens is particularly relevant in this context. Technological capabilities, sectoral innovation pathways, and governance arrangements interact recursively, shaping one another over time.
A structured comparison of these and other prior reviews and science-mapping studies is provided in Table A1 (Appendix A), summarising each along its type and method, data source, coverage, domain, principal contribution, and key limitation. The comparison indicates that existing work is predominantly sector specific. They are mostly confined to individual domains such as solar energy [8,9], energy systems [10,11], agriculture [27,28], or waste management [29], while the few cross-domain efforts are either non-systematic expert syntheses [3,4,16] or single-model and descriptive bibliometric mappings [15,32,33]. None combines a unified cross-sector corpus, cross-model transformer triangulation, and an explicit enabling-sectoral-governance layering, which is the gap the present study addresses (positioned in the final row of Table A1).

2.3. Gaps and Fragmentation in the Existing Literature

Despite significant progress, the literature remains conceptually fragmented and unevenly developed across technological, sectoral, and institutional dimensions. From an innovation-system perspective, this unevenness may reflect differentiated rates of development across system components.
First, diffusion has proceeded primarily through sector-specific pathways (Appendix A, Table A1). Research in renewable energy, agriculture, urban systems, and circular economy tends to evolve within relatively bounded knowledge communities, and the scarcity of integrative analyses limits visibility into cross-sector recombination. For example, optimisation methods developed for energy scheduling might inform circular-economy logistics or climate-adaptation planning (Rolnick et al. [3]). Because technological transformation often depends on complementarities across sectors and institutions, such compartmentalisation leaves opportunities for systemic integration underdeveloped [23,30,31].
Second, methodological approaches have largely focused on application performance or normative governance principles, with limited examination of structural patterns across domains. Existing reviews rely on traditional bibliometric or qualitative content-analysis methods (e.g., Di Vaio et al. [15]; Vinuesa et al. [4]) which, while informative, lack the granularity to trace evolving interconnections across disciplines or to establish whether technological diffusion and institutional adaptation develop proportionally. This is a synchronisation that is critical in complex socio-technical transitions [23].
Third, governance and institutional dimensions appear comparatively thinner relative to technological and application-focused research. Governance-focused analyses remain comparatively limited, with the exception of recent commentaries like Strubell et al. [22] or Cowls et al. [25] examining the environmental costs or ethical implications of AI applications, indicating that such issues remain peripheral. From a co-evolutionary standpoint, such asymmetry may signal institutional lag, or a condition in which oversight mechanisms, evaluation standards, and policy instruments expand more slowly than the technological domains they seek to govern.
These observations suggest that AI-enabled sustainability innovation exhibits differentiated rates of development across technological, sectoral, and institutional domains, motivating a co-evolutionary perspective capable of examining whether these system components evolve in a coordinated or asynchronous manner. The gaps identified above directly inform the analytical design of this study. The fragmentation of existing research into largely sector-specific literatures motivates the AI–Sustainability Pathways (AISP) framework, which conceptualises AI-enabled sustainability research as a layered innovation system and, together with cross-model computational mapping over a unified corpus, enables field-level analysis beyond individual application domains. The limited understanding of structural organisation and temporal evolution motivates the use of transformer-based topic modelling and longitudinal trajectory analysis to characterise thematic architecture and developmental patterns across the literature. Finally, the comparatively limited treatment of governance and institutional dimensions motivates the explicit inclusion of a governance layer within the AISP framework and the examination of inter-layer temporal alignment, thereby providing an empirical basis for assessing the extent to which institutional development keeps pace with technological and sectoral expansion.

2.4. The AI-Sustainability Pathways Framework (AISP)

To analyse the structural organisation and developmental dynamics of AI-enabled green technological innovation, this study conceptualises the AISP framework as a layered innovation-system architecture. Grounded in systems theory and sustainability transitions literature [23,24], the AISP model delineates how AI functions within an ecosystem of technological, sectoral, and institutional interactions. It posits that such sustainability innovation outcomes emerge from the alignment of three analytically distinct but interdependent system layers. While systems theory and the multi-level perspective explain how socio-technical transitions unfold through interactions among technologies, institutions, and regimes, they do not by themselves specify how a general-purpose digital capability such as AI diffuses simultaneously across multiple sectoral innovation pathways. AISP addresses this gap by representing AI as a cross-cutting capability layer that recombines across energy, waste, agriculture, mobility, and industrial systems while generating corresponding governance requirements. In this sense, AISP does not replace existing socio-technical frameworks; rather, it operationalises them for the specific case of AI-enabled green technological innovation.
Positioned against existing frameworks, AISP’s contribution is twofold. Established socio-technical accounts—the multi-level perspective on transitions and the systems-of-innovation tradition [23,24,30,31]—characterise how individual technologies emerge within, and eventually reconfigure, sector-bound niches and regimes; they are oriented toward technologies that are themselves sectoral and offer limited purchase on a general-purpose capability that diffuses horizontally across many sectors at once. AISP’s first departure is therefore to model AI as a crosscutting enabling layer whose recombination across energy, waste, agriculture, mobility, and industrial pathways simultaneously generates distinct governance requirements, with the enabling-application-governance stratification making this horizontal diffusion analytically visible. Its second and more consequential departure is diagnostic. As the three layers are rendered as analytically distinct but directly comparable strata, the framework converts the abstract proposition of technology-institution co-evolution into an observable, measurable question, whether the layers expand in synchrony or whether governance activity lags behind. This is operationalised here through cross-model structural recurrence (Table 1) and annual inter-layer co-movement (Table 2 and Table 3). It is this measurable treatment of co-evolution and institutional lag that constitutes AISP’s principal advance over the frameworks on which it builds.
The framework is used both as an interpretive scaffold and as a diagnostic device. Topic discovery is conducted inductively through unsupervised clustering, after which the resulting themes are mapped to the AISP layers to evaluate structural correspondence and relative alignment across capability, sectoral, and governance domains. This approach preserves the exploratory nature of computational mapping while allowing theoretical interpretation of emergent patterns.
Layer 1 Enabling Capabilities. This layer encompasses the computational infrastructures and algorithmic mechanisms, including machine learning, deep reinforcement learning, optimisation algorithms, computer vision, and natural-language processing, that expand predictive, diagnostic, and coordination capacities. These capabilities constitute the knowledge and technological base of the system. Their evolution influences the scope and intensity of application across sustainability domains.
Layer 2 Sectoral Application Pathways. This layer captures the embedding of AI capabilities within specific green innovation trajectories, such as smart-grid management, waste-stream optimisation, precision agriculture, urban mobility, and industrial-process monitoring. These pathways reflect sectoral innovation subsystems where digital capabilities interact with infrastructure, firms, and domain-specific regulatory conditions. Their expansion signals diffusion and consolidation of AI-enabled sustainability applications.
Layer 3 Governance and Institutional Structures. This layer encompasses regulatory frameworks, lifecycle protocols, standards and mechanisms ensuring fairness, transparency, and accountability, and policy instruments that shape the direction and legitimacy of AI-enabled innovation. Institutional arrangements define boundary conditions under which technological deployment translates into socially and environmentally acceptable outcomes. Feedback loops between this layer and the lower layers determine whether AI advances reinforce or undermine long-term sustainability objectives. In operational terms, this layer can be observed through concrete instruments such as life-cycle assessment protocols, carbon-accounting and disclosure requirements, model documentation standards, audit trails, procurement criteria, and sector-specific oversight rules.
Crucially, these layers are not independent. Innovation-system dynamics emerge through their interaction. Technological capabilities enable new sectoral applications; expanding applications generate new governance requirements; institutional arrangements, in turn, influence incentives for technological development and sectoral diffusion. Co-evolution therefore depends on the growth of individual layers and their relative alignment.

2.5. Co-Evolution and Propositions

The AISP framework (as illustrated in Figure 1) provides a conceptual scaffold for the empirical analyses presented in this study. It captures how enabling technologies interface with domain-specific applications under varying governance regimes across three analytically distinct but interdependent layers: enabling AI capabilities, sectoral application pathways, and governance/institutional structures. The framework is not intended as a strictly linear sequence; rather, it depicts reciprocal influence and feedback, whereby technological advances expand the scope of feasible applications, expanding applications generate new regulatory, accountability, and coordination requirements, and institutional responses, in turn, shape the direction and pace of subsequent technological development.
Mapping research activity across enabling, application, and governance layers allows assessment of whether system components exhibit coordinated development or asymmetric expansion. This multi-layered structure thus serves both as an interpretive lens and a diagnostic representation of innovation-system dynamics. Persistent asymmetries across layers may indicate institutional lag, misaligned incentives, or coordination deficits within the broader sustainability-oriented innovation system.
To assess whether AI-enabled sustainability research exhibits the layered dynamics anticipated in innovation-system and co-evolution theory, the analysis formulates three analytical propositions concerning structural differentiation, temporal diffusion, and institutional synchronisation within the literature. In this study, “maturity” is used in a bibliometric and thematic sense; not in the form of technology-readiness. Specifically, mature domains are inferred from sustained publication volume, recurrence across independent representation models, and persistent temporal visibility, whereas emerging domains are inferred from lower baselines followed by accelerated growth.
  • P1 (Structural differentiation). Thematic structures extracted from large-scale text embeddings should display stable differentiation between capability-oriented, application-oriented, and institutional knowledge domains across independent representation models.
  • P2 (Uneven diffusion). Temporal trajectories of application-oriented research themes should exhibit heterogeneous growth patterns consistent with uneven diffusion across sectoral innovation subsystems.
  • P3 (Institutional asynchrony). Institutional and governance-oriented themes should display structural asynchronous development relative to capability and application domains.
These propositions function as theory-derived analytical expectations that structure the descriptive mapping; they are not formal hypotheses subjected to confirmatory testing. Consistent with the study’s exploration, thematic-mapping design, each proposition is examined through pattern-level correspondence: structural recurrence across representation models (P1), the shape of temporal trajectories (P2), and the relative mass and timing of governance-related themes (P3). The simple Spearman rank correlations reported in Section 4 (Table 3; the governance-trajectory association) are bivariate association measures included for transparency; given the short annual series, they are interpreted as indicative of temporal co-movement and do not constitute formal tests of the propositions or evidence of causal or co-evolutionary dynamics. Subsequent references to evidence “supporting” or “consistent with” a proposition accordingly denote descriptive correspondence between observed and theoretically anticipated patterns; they do not denote confirmatory or causal validation. Formal temporal-inference and model-based testing are identified as directions for future work.

3. Data and Methodology

This section outlines the corpus construction, computational mapping strategy, and validation procedures used to examine the propositions derived in Section 2. The empirical design seeks to assess structural organisation, temporal differentiation, and cross-layer alignment within AI-enabled green technological innovation.

3.1. Corpus Construction and Retrieval Protocol

The empirical analysis is based on a corpus of 3357 peer-reviewed publications from the WoS Core Collection to map research at the intersection of AI and sustainability-oriented green technologies. The search strategy employed two controlled keyword groups combined using the Boolean operator AND, while the terms within each group were combined using the Boolean operator OR in the title, abstract, and keyword fields:
  • AI-related terms: “artificial intelligence” OR “AI”.
  • Sustainability/green-related terms: “green technology” OR “green innovation” OR “green sustainability” OR “ESG” OR “Environmental, Social, and Governance” OR “environmental sustainability”.
To ensure conceptual focus while minimising noise from purely technical machine-learning research unrelated to sustainability contexts, the search string centred on the umbrella descriptors “artificial intelligence” and “AI”. During query development, pilot tests incorporating narrower algorithmic terms (e.g., “machine learning”, “deep learning”, “neural networks”) substantially expanded the retrieval set with engineering studies lacking explicit sustainability framing, thereby reducing topical precision. As sustainability-oriented research typically situates such methods under the broader AI descriptor, the final query retained these umbrella terms to preserve conceptual relevance. It is noted that studies framing methods purely in algorithmic terms without explicit AI terminology may be underrepresented.
As the retrieval strategy targets the intersection of AI and sustainability-oriented innovation at the title/abstract/keyword level, the corpus may include a limited number of sustainability-adjacent records whose connection to green technological innovation is indirect. This is an expected feature of broad bibliometric retrieval in interdisciplinary domains. To guard against over-interpretation of such records, the study assigns interpretive priority to topics that recur across independent representation models and persist across adjacent clustering resolutions.
Retrieval and filtering were conducted using the built-in filtering functions of the WoS Core Collection interface, without the use of AI-based or manual content screening tools, as all filters applied were structural (metadata-based) rather than judgement-based. Records were restricted to English-language journal articles and conference papers published between 2003 and the data cutoff of 17 July 2025; restricting document type to journal articles and conference papers inherently excluded review articles, editorials, notes, corrections, and letters. This yielded an initial set of 3359 records. Exact-title and DOI de-duplication were then applied, removing 2 duplicate records and resulting in a final corpus of 3357 publications (2850 journal articles and 507 conference papers).

3.2. Text Representation and Structural Mapping

Abstracts were adopted as concise, author-curated statements of contribution suitable for large-scale thematic mapping in review studies [34,35]. To explore research scope and topics in this field, the study utilised three transformer-based models to extract meaningful summaries in abstracts known for their effectiveness in representing scientific content. Extracting topics from abstracts has been widely used and proven effective in understanding the development of emerging technologies across various fields, such as blockchain [36], generative AI [37], generative AI in finance [38], ESG and AI [39], etc. Most prior studies employ either Latent Dirichlet Allocation (LDA) or transformer-based encoders such as Bidirectional Encoder Representations from Transformers (BERT) for abstract-level topic discovery. In comparative settings, LDA is typically outperformed by contextual embeddings because it treats words as exchangeable and ignores intra-sentential context, whereas BERT models encode context-dependent semantics that improve thematic discrimination [40]. However, the original BERT is pre-trained on general-domain corpora (Wikipedia and BooksCorpus), which may limit transfer to scholarly prose with domain-specific terminology and citation-driven semantics. Accordingly, this study adopts BERTopic [41] as the baseline neural topic model and further incorporates two science-oriented representation models for performance comparison and triangulation, including SciBERT, which is pre-trained on scientific text to better capture technical vocabulary [42], and SPECTER, which learns document-level, citation-informed embeddings tailored to scholarly similarity [43]. The pre-training data and representation characteristics of these three models are summarised in Table 4.
The topic-extraction workflow comprised three stages:
1. Embedding representation. Sentence/document embeddings were generated separately for each model, yielding three feature spaces of differing dimensionality and granularity. For BERTopic, the default all-MiniLM-L6-v2 encoder was used to obtain 384-dimensional sentence-level embeddings. For SciBERT (scibert_scivocab_uncased), 768-dimensional embeddings were computed via mean token pooling to provide stable sentence/document representations suitable for downstream clustering [44]. In contrast, SPECTER (allenai/specter) provides a 768-dimensional document-level embedding based on the CLS token and is trained to encode citation-informed similarity among scientific papers [43]. The three representations therefore differ not only in dimensionality but also in granularity of sentence-level and document-level; cross-model comparison is consequently conducted at the level of higher-order thematic correspondence rather than through one-to-one geometric alignment of the underlying vectors.
2. Document grouping. Abstracts were partitioned by applying K-means to each model’s embedding space. Although density-based clustering (e.g., HDBSCAN) is usually used to pair embeddings, K-means was selected after comparative evaluation. We applied HDBSCAN to all three embedding spaces under multiple parameter configurations using the same evaluation metric and clustering profile evaluation protocol. The results were unsatisfactory across models. K-means allows fixed values of K to be applied consistently across representation regimes, enabling controlled granularity, reproducible partitions, and direct comparison for triangulation and structural alignment analysis. The detailed results and comparison are shown in the modelling and result section.
A grid search over K (3–25) was conducted to identify analytically useful resolutions. Cluster quality was assessed using the Silhouette coefficient (range −1 to 1), which captures relative cohesion and separation [45]. Because coherence indices (e.g., c_v) require token-based topics rather than continuous embedding clusters, Silhouette provides a model-agnostic internal metric suitable for embedding-space partitions.
3. Keyword induction and topic labelling. For each cluster, KeyBERT [46] extracted the top 10 representative keyphrases. Candidates were refined in a post-processing step (lemmatisation, semantic de-duplication, synonym consolidation, fuzzy matching), after which topics were labelled using the keyphrases together with sampled abstracts from the corresponding cluster. Topic labels were assigned through a consensus-coding procedure by two independent researchers, who reviewed the extracted keyphrases alongside a random sample of abstracts for each cluster; disagreements were resolved through discussion until convergence on a final label. Initial coder agreement exceeded 85% across clusters prior to discussion, indicating substantial convergence in topic interpretation before consensus reconciliation.
The clustering and labelling pipeline produced the topic taxonomy reported in Section 4. To support interpretation, abstract embeddings were reduced to 60 dimensions using Principal Component Analysis (PCA) and projected into two dimensions using t-Distributed Stochastic Neighbor Embedding (t-SNE) with PCA initialization for visualisation. This two-stage reduction improves projection stability, although the results are interpreted heuristically due to t-SNE’s distortion of global distances and sensitivity to hyperparameters [47]. PCA was used only for visualisation and not for clustering. Topic prevalence over time was analysed by mapping publication years to cluster assignments and plotting normalised frequencies (2003–2025) to illustrate topic evolution.
4. Structural mapping to AISP layers. Each labelled topic was mapped to one of three layers in the AISP framework: Layer 1 (Enabling AI Capabilities), Layer 2 (Sectoral Application Pathways), or Layer 3 (Governance and Institutional Domains). Assignments were made independently by three researchers on the basis of each topic’s keyphrase profile, representative abstracts, and alignment with the layer definitions. Initial percentage agreement was 85% prior to discussion, with a pooled Fleiss’ kappa = 0.932 (BERTopic: 1.0; SciBERT: 0.880; SPECTER: 0.928), indicating almost perfect inter-rater reliability after accounting for chance agreement [48]. Disagreements were resolved through discussion until consensus. The resulting cross-model layer mapping is reported in Table 1.

3.3. Methodological Validity and Limitations

The empirical analysis is based on a corpus of 3357 peer-reviewed publications from the WoS Core Collection to map research at the intersection of AI and sustainability-oriented green technologies. Two controlled keyword groups were specified and combined using Boolean operators at the title/abstract/keyword fields. The study delineates bound, method-consistent limitations and implements proportionate mitigations that preserve the validity of its claims.
First, on construct validity, it is recognised that abstract-level analysis may not recover all methodological or results-level detail available in full texts. Nevertheless, working at the title/abstract level is established practice in large-scale thematic mapping and bibliometric text mining. Consequently, inferences are deliberately confined to field-level themes and trajectories [25,34,49]. Any residual summarisation bias is contained through cross-model triangulation, manual label checks, and claim discipline (no unit-level causal assertions).
Second, on external validity, restricting retrieval to WoS and English-language records may introduce coverage bias (e.g., under-representation of regional venues). This design choice privileges-controlled indexing and metadata quality, thereby prioritising precision and replicability over exhaustiveness. Conclusions are explicitly framed as WoS/English-bounded patterns, with emphasis on the internal coherence of the derived thematic taxonomy.
Third, topic induction entails interpretive subjectivity. In particular, boundary clusters that emerge in only one representation regime are interpreted cautiously and are not used on their own to support the paper’s structural claims. To minimise arbitrariness, a pre-specified protocol is applied: (i) cross-model convergence (BERTopic, SciBERT, SPECTER) to distinguish replicated from model-sensitive topics; (ii) parameter sensitivity (adjacent values of K) to ensure headline themes persist beyond a single partition; (iii) transparent labelling via two-coder consensus coding with discrepancy resolution; and (iv) fixed random seeds with complete software/model disclosure for reproducibility. These controls align with recognised standards for robustness and auditability in thematic mapping [25,34,45].
It is useful to note that this study does not attempt formal manifold analysis or geometric optimisation of embedding spaces for representation-learning optimisation or supervised classification, as its objective is computational thematic mapping, structural interpretation, and theory-integrated review. Accordingly, its validity is not assessed through supervised accuracy or formal model specification, but through structural reproducibility, cross-model convergence, and interpretive coherence.

4. Modelling and Results

4.1. Descriptive Trends

The annual distribution of publications (2003–2025; data cutoff 17 July 2025) exhibits a prolonged, low-volume tail through 2019, followed by a marked inflection after 2020 (Figure 2). Activity peaks in 2022–2024; the 2025 figure reflects partial-year coverage at the data-freeze date. Notably, the partial-year total for 2025 has already exceeded the full-year count for 2024, suggesting that the growth trajectory is not only sustained but is further intensifying. Journal outputs dominate the corpus, while 507 conference papers indicate a sustained methods pipeline characteristic of rapidly evolving AI research domains. Although causal attribution is not claimed, this post-2020 acceleration is congruent with heightened institutional attention to digital enablers of net-zero and circular-economy agendas and the broader integration of AI within sustainability policy discourse [3,4]. This pattern also aligns with global decarbonisation and climate-adaptation initiatives under multilateral processes (e.g., COP26–COP28). It further aligns with national strategies that frame AI as a lever for emission mitigation and resource efficiency, such as the EU Green Deal/Digital Europe, U.S. AI for Climate, and China’s 14th Five-Year Plan.

4.2. Internal Validity and Model Selection for Analysis

4.2.1. Clustering Method and K Selection

Validation measures such as topic coherence and topic diversity serve as proxies for inherently subjective evaluations. In this study, we used two complementary internal validation indices-Silhouette score and Davies–Bouldin (DB) index-to assess the quality of the clustering results (Table 5). The Silhouette score measures combined cluster cohesion and separation (higher is better), while the DB index captures the ratio of within-cluster scatter to between-cluster distance (lower is better) [50]. Both metrics consistently favour low values of K across all models: Silhouette scores peak at K = 3 (BERTopic: 0.0469; SciBERT: 0.0468; SPECTER: 0.0728), and DB values are generally lowest at coarser resolutions. Although these values are modest, this is typical for high-dimensional semantic embeddings derived from interdisciplinary corpora, where topical boundaries are inherently fuzzy and documents often span multiple conceptual domains.
To balance separation and thematic resolution, subsequent analyses emphasise BERTopic K = 8 (Silhouette: 0.0349; DB: 4.0031), SciBERT K = 9 (Silhouette: 0.0370; DB: 3.4982), and SPECTER K = 15 (Silhouette: 0.0609; DB: 2.7270). In the Silhouette curve, BERTopic peaks at K = 8 relative to adjacent values (K = 7: 0.0283; K = 9: 0.0300), SciBERT peaks at K = 9 (K = 8: 0.0355; K = 10: 0.0307), and SPECTER peaks at K = 15 within the K = 10–20 range (K = 10: 0.0538; K = 20: 0.0597). The DB index corroborates these selections: BERTopic achieves a local DB minimum at K = 8 (K = 7: 3.9465; K = 9: 3.8251), SciBERT at K = 9 (K = 8: 3.5392; K = 10: 3.6176), and SPECTER at K = 15 (K = 10: 2.9000; K = 20: 2.7936). The convergence of both metrics on the same local optima strengthens confidence in the chosen partitions.
Although the global Silhouette maximum for all three models occurs at K = 3, this resolution was judged too coarse for meaningful thematic differentiation, collapsing distinct sub-themes evident in the finer-grained KeyBERT keyphrase sets into overly broad categories. The selected K values were determined based on two criteria: strong cluster separation and qualitative validation through keyphrase inspection and consensus coding. This ensured that the resulting clusters were thematically distinct, non-redundant, and balanced in terms of topic granularity and coherence [47]. Note that these values are not treated as directly equivalent partitions across models; rather, they represent model-specific operating points chosen to preserve interpretable granularity within each embedding regime. Cross-model comparison is therefore conducted at the level of higher-order thematic correspondence and recurrent structural patterns, not through one-to-one comparison of individual clusters.
The preceding analysis assumes K-means as the base clustering algorithm. To justify this choice, we also applied HDBSCAN to all three embedding spaces under multiple parameter configurations. For SciBERT, even with a minimum cluster size as low as 10, HDBSCAN failed to separate the abstracts into meaningful groups and returned only a single cluster, with the remaining documents assigned to the noise category. For SPECTER, HDBSCAN identified 3 clusters but classified 71.8% of documents (2412 of 3357) as noise, with the remaining assignments highly imbalanced (915, 15, and 15 documents), although the resulting Silhouette score was comparable to K-means. For BERTopic embeddings, HDBSCAN identified only 2 clusters with 38.0% of documents (1275 of 3357) labelled as noise and one cluster containing 2070 documents while the other contained just 12, indicating a degenerate partition despite a higher Silhouette score (0.1596). Across all three models, a considerable proportion of documents were consistently assigned to the noise category, reducing effective corpus coverage and undermining controlled cross-model comparison. These outcomes are consistent with the dense, semantically overlapping geometry of interdisciplinary scholarly embeddings, where density-based methods struggle to identify well-separated regions of high local density. In contrast to these limitations, K-means allows fixed values of K to be applied consistently for each representation regime, enabling controlled granularity, reproducible partitions, and structural alignment analysis. In such contexts, the comparative pattern across K provides a stable internal benchmark for selecting analytically interpretable resolutions.

4.2.2. Robustness Assessment

To assess the sampling stability of the three selected solutions, we further evaluated the sampling stability of the three headline solutions via bootstrap resampling. For each model, the corpus of 3357 abstract embeddings was resampled with replacement at the original sample size, and K-means was refit at the model-specific K on each resample. The Adjusted Rand Index between each bootstrap clustering and the reference clustering obtained from the full corpus was computed. Following standard bootstrap practice, B = 500 resamples were used as the primary specification, with a B = 200 sensitivity analysis under an independent seed providing a Monte Carlo convergence check [51].
Across B = 500 resamples, mean ARI values were 0.643 (SD 0.104; 95% percentile interval [0.439, 0.845]) for BERTopic (K = 8), 0.552 (SD 0.087; [0.405, 0.744]) for SciBERT (K = 9), and 0.621 (SD 0.061; [0.516, 0.749]) for SPECTER (K = 15). The B = 200 sensitivity analysis yielded mean ARIs of 0.676, 0.538, and 0.613 respectively, indicating absolute deviations of 0.033, 0.014, and 0.008 from the primary estimates, consistent with Monte Carlo convergence and low simulation error. SPECTER additionally exhibits the tightest sampling distribution (lowest standard deviation), indicating comparatively higher stability under resampling. These values fall within the moderate stability regime typically observed for K-means clustering in high-dimensional embedding spaces for interdisciplinary corpora [52].
These results support the robustness of the identified clustering structures under resampling variation, while the primary structural claim of the paper rests on the cross-model recurrence of the triadic enabling–application–governance configuration reported in Table 1. In this context, moderate stability is consistent with expectations for unsupervised clustering of high-dimensional interdisciplinary embeddings, where thematic boundaries are not discrete, but inherently overlapping.
The comparatively stronger separation and stability observed for SPECTER (Table 5; bootstrap SD = 0.061) are consistent with its document-level, citation-informed pooling. SPECTER achieves the lowest Davies–Bouldin values across all K values tested (e.g., 2.7270 at K = 15 versus 3.8251 and 3.4982 for BERTopic and SciBERT at their respective selected K), corroborating the Silhouette-based finding with a centroid-distance measure that penalises within-cluster scatter relative to between-cluster separation. Aggregating an abstract into a single CLS-based vector suppresses sentence-level noise arising from heterogeneous rhetorical structure (e.g., background, method, and result sentences). The citation-network fine-tuning signal further injects a domain-relevant similarity axis that raw semantic embeddings do not capture. BERTopic and SciBERT, which rely on sentence-level embeddings, are correspondingly more sensitive to this within-abstract heterogeneity, which likely contributes to their lower and less stable Silhouette scores and higher DB values. These differences should be interpreted as indicating relative confidence in cluster boundaries rather than fundamental deficiencies in topic extraction. The modest Silhouette scores and elevated DB values observed across all three models are consistent with the expected geometric fuzziness of high-dimensional interdisciplinary embeddings, where documents routinely span multiple thematic domains. Even at these levels, KeyBERT-derived keyphrases remained coherent and interpretable under manual review, suggesting that topic-extraction accuracy is more robust to embedding-space fuzziness than cluster-geometric separation alone would imply.

4.2.3. Visualisation of the Clustering Results

Figure 3 further presents two-dimensional t-SNE projections of the clustering results for the four model–granularity combinations examined in this study. For each model, abstract embeddings were first reduced to a 60-dimensional subspace using PCA and then projected to two dimensions via t-SNE (initialised with PCA for projection stability). Each point represents an individual publication abstract and is coloured by its K-means cluster assignment; centroid label boxes mark the position of each cluster. Thematic labels correspond to the topics reported in Table 6, Table 7 and Table 8.
The four panels offer complementary views of the corpus structure. The BERTopic K = 3 projection (Figure 3a, left) reveals a coarse tri-partite organisation, with three broad thematic regions clearly visible across the projected space: enabling capabilities, governance/ethics, and sectoral applications. Increasing the resolution to K = 8 (Figure 3a, right) preserves this overall layout while revealing finer sub-themes within the larger regions, but at the cost of greater local overlap compared to K = 3. The SciBERT K = 9 projection (Figure 3b, left) emphasises science-domain semantics, with denser thematic transitions reflecting the interdisciplinary overlap typical of scholarly embeddings. It yields finer science-domain sub-topics with noticeable local overlap typical of interdisciplinary semantics. The SPECTER K = 15 projection (Figure 3b, right), based on citation-informed document embeddings, displays the most clearly delineated neighbourhood boundaries, with distinguishable regions corresponding to the finer-grained topics identified in Table 8.
These projections preserve local neighbourhood structure but distort global distances, relative cluster sizes, and the magnitude of gaps between regions. Spatial proximity therefore indicates relative semantic similarity, not absolute distance, and the projections should be interpreted as qualitative neighbourhood maps rather than quantitative inter-cluster distance plots. Boundary overlap between adjacent regions is expected and reflects the inherently graded nature of interdisciplinary thematic content, where individual abstracts often span multiple conceptual domains simultaneously.

4.3. Topic Taxonomy and Cross-Model Integration

Across models, K = 3 solutions yield relatively higher separation compared to higher K values, whereas K = 8 to 15 capture richer sub-topics at modest costs to separation, which is an expected granularity-separation trade-off in interdisciplinary corpora. Post hoc keyphrase normalisation and manual checks confirm intuitive topical centroids. Topic clusters are interpreted as thematic approximations, reflecting the probabilistic nature of embedding-based semantic partitioning.
At coarse granularity (BERTopic, K = 3), the corpus decomposes into three macroscopic themes (Table 6), namely Green AI Infrastructure, Governance, Ethics and Sustainability, and Sustainable Healthcare Systems, which align naturally with the enabling, governance/ethics, and sectoral application lenses of the AISP framework. Increasing the resolution to K = 8 yields sector-anchored subtopics, namely Smart Urban Systems, AI in Green Agriculture, Sustainable Cities, and Emissions and Environmental Analytics, indicating stronger domain localisation while preserving environmental endpoints. The presence of healthcare-related clusters reflects the broad interpretation of sustainability within portions of the literature, where health-system resilience and environmental health outcomes are sometimes framed within sustainability discourse. This cluster is treated as a boundary case rather than a core sustainability theme; consistent with the interpretive caution applied to boundary clusters (Section 3.3), it is not used on its own to support the paper’s central structural claims. The central cross-model architecture remains driven by recurrent themes in energy, waste/circularity, agriculture/land use, urban systems, and governance.
SciBERT at K = 9 emphasises application-level semantics, surfacing Sustainable Energy and BioSystems, Cybersecurity and Smart Infrastructure, Sustainable Consumption and Behaviour, AI Governance and Ethics, and Digital Sustainability and Health (Table 7). The relative de-emphasis of the token “AI” as a surface keyword reflects discipline-specific abstract writing and SciBERT’s science-domain pretraining, which captures technical vocabulary and research framings beyond generic terms [42].
SPECTER at K = 15 produces finer-grained clusters, including Energy and Smart Grids, Waste Management and Biofuels, Agriculture and Food Systems, Urban Planning and Transport, Blockchain and Governance, Sustainability Education, and Forest and Land Use, with comparatively clearer boundaries among adjacent subfields (Table 8). This behaviour is consistent with SPECTER’s citation-aware objective that encodes document-level topical proximity in scholarly networks [43].
Evidence for P1. A cross-model synthesis (Table 1) organises these clusters under the AISP layers: Enabling Technologies (e.g., Green AI Infrastructure), Application Pathways (e.g., Energy/Grids; Waste/Biofuels; Agriculture/Food; Urban/Transport; Industry/Processes), and Governance & Ethics (e.g., Ethics & Social Change; Blockchain & Governance). The corresponding AISP layer labels are added in Figure 3. The recurrence of this triadic configuration across independent representation regimes provides convergent structural evidence consistent with P1, that the field’s thematic organisation mirrors enabling-application-governance layers grounded in systems and transition theory [16,17]. While cluster boundaries differ across embedding models, the repeated alignment of enabling, application, and governance strata suggests a stable field architecture. Reproducibility of the three-layered structure is assessed through cross-model convergence, adjacent-K persistence, and independent coder agreement in topic interpretation. The recurrence of enabling, application, and governance strata under these conditions provides empirical support for claiming structural robustness.
Interestingly, it turns out that the models provide a complementary rather than substitutive effect. BERTopic covers macro-partitioning, SciBERT engenders science-domain framings, and SPECTER scopes niche research areas with citation-informed precision. This triangulation logic is important for interpretation. Single-model clusters are not treated as sufficient evidence for field-level structure unless supported by cross-model recurrence or by stable alignment with the broader thematic architecture. The cross-model mapping in Table 6 functions as the primary basis for inference, with the reliability of the underlying layer assignments confirmed by independent coder agreement.
The replication of these sectoral anchors across models and K-values indicates thematic stability at the application-pathway layer. By contrast, governance-oriented topics appear consistently but with smaller mass and greater model sensitivity, signalling a more diffuse or emergent discourse. This pattern coheres with the gap analysis in Section 2 and reinforces the need to couple technical deployments with environmental and governance dimensions such as lifecycle assessment, carbon accounting, and transparency mechanisms to advance sustainability impact [4,25]. Keeping in view the sustainability paradox in Section 2.2, this pattern suggests that governance-related knowledge remains less consolidated than capability and application domains, thereby creating thematic conditions under which lifecycle burdens and accountability gaps may remain under-addressed.
To provide a preliminary quantitative illustration of inter-layer evolutionary relationships within the AISP framework, let:
L i ( t ) = n i ( t ) N ( t )
where n i ( t ) denotes the number of publications assigned to AISP layer i at time t , and N ( t ) denotes the total number of publications published in year t . Here, i { E , S , G } corresponds respectively to enabling capability, sectoral application, and governance and institutional layers. Normalising by annual publication volume removes the influence of overall corpus growth, allowing the analysis to examine changes in the relative thematic prominence of each AISP layer over time.
Cross-layer evolutionary association may then be represented as:
ρ i j = c o r r ( L i ( t ) , L j ( t ) )
where ρ i j denotes the Spearman correlation coefficient between the annual prevalence trajectories of layers i and j . Under this formulation, stronger positive associations indicate greater synchronisation in thematic development across layers, whereas weaker associations may reflect asynchronous expansion or institutional lag within the broader AI-sustainability innovation system.
A correlation analysis was conducted on the annual layer share trajectories associated with the three macro-level AISP layers derived from the BERTopic K = 3 solution (Table 7). The publication years from the WoS export were merged with the clustering assignments, producing annual layer trajectories (2003–2025; data cutoff on 17 July 2025). Spearman correlation coefficients were then computed in order to examine the extent of temporal co-movement across enabling capability, sectoral application, and governance and institutional domains.
The findings in Table 8 provide preliminary evidence of thematic synchronisation across the AISP layers, indicating that technological capability, sectoral diffusion, and governance-oriented research domains have expanded concurrently over time. As these patterns are correlational, they are interpreted as evidence of temporal co-movement and differential growth. At the same time, the comparatively weaker association involving governance themes is consistent with possible institutional lag, suggesting that governance adaptation may evolve less synchronously relative to technological and application-oriented expansion.
As an inferential extension, a Spearman trend analysis was conducted on the annual trajectory of governance-oriented themes within the BERTopic macro-structure. The analysis yielded a strong positive temporal association (ρ = 0.81, p < 0.001). This serves as an indicative corroboration of the sustained post-2020 expansion in governance and institutional research activity; it does not constitute a formal test of temporal or co-evolutionary dynamics.

4.4. Research Trends

Evidence for P2. This was assessed using bibliometric maturity signals, namely publication volume over time and recurrence across model outputs. Figure 4 presents the distribution of publications per year for six representative topics identified in the corpus. Data for 2025 are excluded as only the first half of the year was available at the time of data collection. Including a partial year in a normalised-share analysis would risk distorting compositional proportions, as topics with uneven seasonal publication patterns would be systematically over- or under-represented.
The curves report normalised shares and therefore reflect the joint effect of (i) topic-specific activity and (ii) the overall post-2020 expansion of the corpus. The trajectories reveal three consistent patterns:
  • Cross-cutting infrastructure peaking. Green AI Infrastructure rose rapidly from roughly 7% before 2020 to nearly 40% of annual output by 2023, before retreating to around 25% in 2024. This trajectory is consistent with its role as an umbrella category whose foundational research progressively diffused into more specialised application and governance topics, a pattern indicative of maturity and wide absorption across adjacent subfields.
  • Accelerating governance and strategy themes. Sustainability education climbed from approximately 16% before 2020 to 26% by 2024, consistent with growing attention to climate literacy and workforce upskilling. Sustainability innovation strategy rose from roughly 4% to 20%, reflecting intensified interest in the policy, managerial, and market mechanisms that organise technology adoption. Waste management and biofuels increased from 5% to approximately 14%, consistent with the operationalisation of circular-economy practices and alternative fuel systems. These steepening profiles suggest emerging domains; however, given overall corpus expansion after 2020, the signals should be interpreted as indicative rather than causal evidence of thematic emergence [4,25].
  • Diverging application pathways. Before 2020, empirically grounded domains dominated: livestock/biotechnology and forest/land use together accounted for over two-thirds of all publications, reflecting the early concentration of AI-for-sustainability research in sectors with established data infrastructures. Both shares declined steadily as the field diversified, falling to approximately 5% and 11% respectively by 2024. By contrast, waste management and biofuels moved in the opposite direction, rising from roughly 5% to 14% over the same period, consistent with the operationalisation of circular-economy practices and alternative fuel systems. This divergence suggests that research attention is shifting away from legacy empirical domains toward application areas with stronger policy and industrial momentum.

5. Discussion

The discussion that follows interprets the results at the level of cross-model recurrent domains, triadic layer alignment, and topic-specific temporal trajectories.

5.1. Addressing P1: Interpreting the Field Structure Through AISP

Across modelling regimes, the corpus exhibits a reproducible tri-layered structural configuration corresponding to enabling capabilities, sectoral application pathways, and governance domains. This triadic organisation persists across BERTopic, SciBERT, and SPECTER at analytically defensible granularities.
The result is theoretically coherent, where sustainable outcomes arise when technological capabilities are embedded within sectoral contexts under appropriate institutional arrangements [23,24]. From this perspective, enabling AI functions (forecasting, optimisation, anomaly detection, autonomous control) constitute capability levers; application domains (energy, waste/circularity, agriculture/land, urban mobility, industry/processes) provide the operational domains; governance and ethics (lifecycle accounting, policy, standards, equity) set the high level governing and boundary conditions that mediate risk transfer and ensure net environmental benefit [4,25]. The reproducibility of this structure across models strengthens construct validity and offers a robust scaffold for interpreting topical dynamics.

5.2. Addressing P2: Maturity and Emergence from Temporal Trajectories

Interpreting Figure 4 as raw volume trajectories, the results indicate two complementary dynamics aligned with P2. First, mature application spaces, such as Waste Management & Biofuels and Forest & Land Use exhibit sustained or steady growth from substantial bases and recur across modelling regimes. These attributes are consistent with knowledge diffusion in operationally grounded areas where data availability and decision cycles favour the uptake of data-driven optimisation and control [3]. Second, emerging governance/policy fronts, including Sustainability Education and Green Innovation Strategy accelerate from low baselines after 2020. Their steepening trajectories suggest consolidation around institutional capacity, literacy, and strategic orchestration that enable responsible scaling of AI systems [4,25]. Therefore, the corpus reflects a field moving from predominantly efficiency-oriented deployments toward alignment-oriented infrastructures that tie technical work to organisational and policy mechanisms.
The practical implication is twofold. In mature pathways, the priority is impact verification and scaling across contexts (e.g., evaluating lifecycle performance and process-safety outcomes). In emerging pathways, the priority is design-oriented, longitudinal work that integrates technical, behavioural, and governance dimensions to accelerate knowledge consolidation and policy uptake. Interpretations are bound to observational maturity/emergence signals inferred from topic density, replication, and temporal evolution.

5.3. Addressing P3: Interpreting the Field Structure Through AISP

Section 5.3 interprets observed thematic asymmetries in light of the paradox literature; it does not estimate causal effects. Across modelling regimes, the corpus exhibits a reproducible tri-layered structural configuration corresponding to enabling capabilities, sectoral application pathways, and governance domains. This triadic organisation persists across BERTopic, SciBERT, and SPECTER at analytically defensible granularities. Across clustering solutions, governance-related clusters account for a smaller proportion of total documents compared to major application pathways such as energy, waste, and agriculture, reinforcing the interpretation of relative asymmetry rather than absence. Extending the paradox discussion in Section 2.2, three corpus-level patterns indicate asymmetric growth across layers and help explain why governance-related safeguards may lag behind technical diffusion:
  • Relative scale. Governance-ethics strata (e.g., Ethics/Justice, Blockchain & Governance, policy/legislation terms) occupy a comparatively smaller proportion of thematic mass relative to technological and sectoral domains across BERTopic, SciBERT, and SPECTER solutions (Table 1, Table 6, Table 7 and Table 8).
  • Model sensitivity. Institutional and governance strata exhibit greater boundary instability across representation regimes (labels and boundaries shift more across K and representation regimes) than do application anchors (e.g., energy, waste, agriculture, urban), which replicate across models.
  • Temporal lag. Governance- and strategy-related topics (e.g., Sustainability Education, sustainability Innovation Strategy) show post-2020 acceleration from low baselines (Figure 4). In normalised terms, both topics gained substantial share between 2020 and 2024, yet they remain lower in absolute volume than major application domains such as Green AI Infrastructure. This pattern suggests that governance alignment is catching up rather than co-evolving at pace with technical diffusion, narrowing the compositional gap but not yet closing it.
Within enabling-technology layer, co-occurring terms such as Green AI, carbontracker, and emissions appear, indicating awareness of computational energy and lifecycle concerns. However, this signal is embedded inside technical strata more than it is anchored in dedicated governance clusters. These patterns are consistent with relative layer asymmetry, in which technical and sectoral application research expands more rapidly and with greater structural stability than governance-oriented strands. Technical/application work is denser and more stable, whereas governance-ethics activity is thinner, more volatile, and lagged. The existing literature documents non-trivial energy and carbon costs associated with construction of AI pipelines (training and inference), raising the prospect that efficiency gains in targeted domains may be partially offset by the footprint of the AI systems themselves if lifecycle considerations are not integrated [21,22]. In parallel, sustainability reviews emphasise that institutional design, transparency, and accountability mediate whether AI contributions translate into net-positive environmental [4,25]. The corpus patterns of relatively smaller scale, greater cross-model sensitivity, and later acceleration in governance-oriented topics are consistent with a scenario in which technical diffusion advances more rapidly than institutional consolidation.
Here, the study would like to emphasize that it does not infer causality between layer misalignment and realised environmental impacts. Rather, it reports an observation of emphases in the literature, where higher activities in technical/application topics co-exist alongside comparatively thin, lagging governance-ethics works. This aligns with theoretical expectations that uncoordinated diffusion of enablers and applications increases the risk of unmeasured or unmanaged externalities. While the study does not evaluate governance effectiveness directly, the observed asymmetry in thematic emphasis motivates several governance considerations:
  • Lifecycle visibility. Routine reporting of training and inference energy/carbon and adoption of carbon-accounting protocols alongside performance metrics [4,21].
  • Design-time integration. Incorporate lifecycle assessment and energy-aware model selection within MLOps and procurement (linking technical choices to environmental budgets).
  • Data and auditability. Strengthen data governance, reproducibility, and model documentation (e.g., consistent disclosure templates) to enable verification of claimed benefits and costs [25].
  • Sector coupling. Couple application rollouts (energy, waste, agriculture, mobility) with context-specific oversight, including sectoral standards, boundary conditions, and fail-safes tailored to process safety and environmental protection.
For industrial contexts, predictive maintenance and anomaly detection remain critical. However, without transparency on computational overheads and governance of data and model updates, risk can be transferred (from physical to digital/energy domains) rather than reduced. Aligning layers ensures that efficiency gains are not offset by hidden lifecycle burdens.
The observed layer asymmetry and lagged governance are consistent with the persistence of the AI sustainability paradox as characterised in the literature. Addressing this requires synchronised evolution of the AISP layers (technical enablers, sectoral applications, and governance), so that environmental accounting and accountability co-proceed with innovation rather than follow it.

6. Implications

The implications below are derived from the combination of empirical clustering patterns, cross-layer asymmetries, and the supporting literature on lifecycle governance, Green AI, and responsible innovation.

6.1. Academic Implications

The AISP triadic organisation identified in this review (i.e., enablers-applications-governance) demonstrates that AI-enabled greentech innovation is best understood as the interaction of technological enablers, application contexts, and governance mechanisms. This perspective has three implications for future research.
First, construct refinement and measurable alignment. The evidence supports formalising layer alignment as a construct capturing the co-evolution of technical deployments, sectoral integration, and governance artefacts within a domain. Alignment can be operationalised using observable signals such as topical persistence, cross-model replication, and co-occurring governance markers (e.g., disclosure, lifecycle assessment terms) within application clusters. This provides testable pathways for P1 (structural organisation) and P3 (misalignment).
Second, mechanism testing beyond accuracy. Research, for instance in process-safety and environmental systems, may treat AI/ML model predictive uncertainty and model drift as first-order mechanisms with consequences for hazard identification, risk prioritisation, and environmental variability. Designs that integrate uncertainty quantification and drift diagnostics into safety cases enable falsifiable tests of whether AI capability translates into dependable operational benefit [21].
Third, lifecycle-aware evaluation. A recurring gap is the weak coupling between technical benefits and lifecycle burdens (e.g., energy and carbon). Comparative and quasi-experimental studies that link deployments to measured outcomes, including emissions, diversion rates, grid losses, incident rates are needed to adjudicate net environmental effect and to assess whether governance instruments (e.g., transparency, auditability, documentation) mediate that effect [4]. Open, cross-domain datasets on compute and energy disclosure, model documentation, and deployment metadata would materially advance research in this area.
These directions can help establish a basis for theory-building on how digital enablers condition sustainable transitions.

6.2. Practical Innovation Management Implications

Results from this study can help translate into implementable controls that reduce incident risk and ensure net environmental benefit across mature and emerging deployment contexts. These include:
  • Safety-by-design for industrial and infrastructure operations. In mature enabler-application combinations, such as forecasting for grid stability, anomaly detection for rotating equipment, routing and scheduling for waste logistics, AI can reduce incident probabilities and improve reliability. To avoid risk displacement from the physical to the digital domain (e.g., energy burden, privacy exposure, model drift), engineering practice may require, for instance: (i) performance and uncertainty reporting of AI/ML models mapped to hazards and safety limits; (ii) drift monitoring (data and concept drift) with escalation, rollback, and fail-safe modes; (iii) lifecycle assessment and carbon accounting as acceptance criteria alongside accuracy and latency [21]; (iv) reproducible pipelines and provenance controls (versioning, lineage, audit trails) to enable traceability; (v) context-specific validation, reflecting duty cycles, operating envelopes, and failure modes of target assets.
  • Environmental programme execution. As the use of AI diffuses across waste/circularity, land use, and energy, deployments should be paired with sectoral safeguards. These include data-quality and provenance checks, reproducible workflows for regulatory review, and accountability among developers, operators, and regulators. To ensure comparability across sites, one may couple rollouts to standards for emissions accounting, explainability thresholds, and model documentation [4].
  • Carbon-aware MLOps. Consider establishing energy metering for training and inference, track compute-energy budgets, schedule workloads to low-carbon windows where feasible, and prefer frugal architectures when performance is equivalent. These practices narrow the gap identified in P3 between rapid technical diffusion and slower governance uptake.
Application rollouts should be coupled to sectoral standards (e.g., for emissions accounting, explainability thresholds, and model documentation) to ensure traceability and comparability across deployments.

6.3. Policy and Innovation Systems Implications

The review indicates that governance has developed more slowly than technical and application-oriented research, creating a persistent imbalance across the AISP layers. From an innovation systems perspective, governance, capability development, and institutional learning must evolve together if AI is to deliver sustainable environmental outcomes. The following priorities can help strengthen this alignment.
  • Embedding governance upstream, for instance, requiring life-cycle assessment and carbon accounting for training and inference at procurement and design stage, rather than post-deployment; adopt disclosure of compute-energy budgets and standardised model documentation. These instruments are observable and actionable mechanisms through which institutional adaptation becomes measurable in organisational and policy practice.
  • Prioritise evaluability in mature domains. In energy, waste, and land sectors, funding implementations with before–after or counterfactual designs linked to verifiable environmental indicators, coupled with data-sharing arrangements can help in independent audit efforts.
  • Accelerate consolidation in emergent domains. For sustainability education and innovation strategy, supporting integrative longitudinal, multi-stakeholder programmes that integrate behavioural, institutional, and technical elements can help inform holistic practice.
  • Strengthen data and model governance. Establish shared taxonomies, metadata standards, and audit trails across the data lifecycle to ensure interoperability, compliance, and replication across jurisdictions.
  • Carbon-aware MLOps. Incorporate carbon budgets as acceptance criteria for public procurements and regulated deployments of AI systems with environmental objectives.
These implications provide an aligned agenda across AISP layers, in particular, robust technical deployment anchored in domain operations, evaluable pathways to environmental benefit, and governance instruments that proceed in step with innovation. This alignment can help ensure that AI delivers net-positive, lifecycle-aware contributions to AI-for-green technology research and deployment.

6.4. AISP Field Guide: From Evidence to Action

This study takes a further step towards providing a single, unifying reference table that brings forth the study’s findings into an illustrative actionable schema spanning the three AISP layers (Appendix B). It is designed for academics (in terms of study designs and constructs), industry-engineering (in terms of controls and procedures), and governance-policy (in terms of instruments and standards). Cells emphasise what to look for, how to measure it and key risks.
A brief illustration may clarify its use. In the Energy & Smart Grids pathway, forecasting and optimisation capabilities are already structurally mature in the corpus. However, Table A2 shows that effective deployment depends on governance complements such as model-update disclosure, explainability thresholds, lifecycle-emissions reporting, and rollback procedures. In AISP terms, the framework therefore helps diagnose whether strong capability and application development is matched by sufficiently explicit oversight. Where such instruments are absent, the likely gap is not technical feasibility but institutional readiness.

7. Conclusions

This review set out to clarify how AI is being mobilised for green technological innovations, and to provide a defensible foundation for future inquiry and deployment. Three contributions emerge. First, it provides a reproducible computational mapping of the AI-sustainability research landscape using transformer-based embeddings and cross-model triangulation, establishing a defensible field-level taxonomy and providing support for P1. This structure is reproduced across independent representation regimes and granularities, strengthening construct validity. Second, it introduces the AISP framework as an empirically tractable representation of how technological capabilities, sectoral applications, and governance domains interact within sustainability-oriented innovation systems. Temporal volume patterns distinguish maturing domains (sustained activity from a substantial base) from emerging domains (steep post-onset growth), providing evidence for P2. Third, by tracing thematic trajectories across these layers, the study identifies asymmetries in the evolution of technological and institutional strands, offering a diagnostic perspective on the co-evolution of digital technologies and sustainability governance. The relative thinness and lag of governance strands, compared with denser, more stable technical and application work, suggests asymmetry in thematic emphasis, thereby supporting P3.
Through the AISP lens, effective green innovation emerges when: (i) enabling capabilities are designed with uncertainty and lifecycle costs in mind; (ii) applications are validated within their operational and institutional contexts; and (iii) governance mechanisms foster transparency, accountability, and evaluability. The corpus indicates that application diffusion is advancing fastest, whereas governance is catching up. As such, governance should be conceived not as a post hoc compliance layer, but as an integral design input and scaling enabler within the innovation system.
The study’s claims are explicitly bounded to mapping and synthesis using abstracts from a WoS/English-language corpus and topic models at multiple granularities [49]. Triangulation across models, sensitivity checks across K , consensus labelling, and reproducibility controls mitigate, though do not eliminate, well-known limitations of thematic mapping approaches. Within these bounds, the findings furnish a governance-ready operating model, aligning enabling technologies, sectoral applications, and institutional instruments so that AI’s technical gains translate into measurable, durable contributions to sustainable innovation and environmental outcomes.

Limitations and Future Work

Several further limitations of the present design warrant explicit acknowledgement and motivate a structured agenda for future research.
First, both the cross-layer alignment construct and temporal trajectories are not supported by formal quantitative specifications. Asymmetry is inferred from differences in topical mass, model sensitivity, and temporal acceleration, while temporal patterns are characterised descriptively and with bivariate rank-correlation measures. A key extension may be to formalise these constructs through explicit measures, such as prevalence-normalised co-movement indices, lag-adjusted alignment scores, and weighted governance indicators. Complementary inferential approaches, including non-parametric trend tests, segmented regression, and structural break detection, could enable rigorous testing of temporal dynamics. At the modelling level, vector autoregression, lead–lag estimation, and coupled dynamical systems could further transform descriptive diagnostics into testable representations of capability-application-governance interactions, including formal tests of co-evolutionary dynamics among AISP layers.
Second, while the study establishes internal, construct, interpretive, and sampling-based validation, it does not incorporate external validation against independently labelled data, nor does it directly link observed research activity to realised environmental outcomes. A natural extension may be a stratified, multi-coder annotation protocol (e.g., 200–300 abstracts across layers), with clustering–label correspondence evaluated using the Adjusted Rand Index and benchmarked against external evidence sources such as policy databases, patent records, and sustainability disclosures. In parallel, quasi-experimental designs (e.g., difference-in-differences and synthetic controls) could help establish causal links between AI deployment and outcomes such as emissions, resource intensity, grid efficiency, and operational risk. We intend to pursue this external validation protocol as a direct extension of the present study, beginning with the stratified multi-coder annotation exercise described above.
Third, cross-model triangulation remains constrained by differences in representational granularity. Bootstrap stability analysis (B = 500) yields a mean Adjusted Rand Index of approximately 0.64 (SD ≈ 0.10), indicating moderate intra-representational clustering stability under resampling. However, this assessment is inherently within-representation and does not support direct comparison across heterogeneous embedding spaces. Developing alignment-aware approaches, such as normalised mutual information on shared assignments or co-clustering frameworks, are helpful for formal cross-representation evaluation.
Finally, the analysis primarily captures research activity, leaving net lifecycle benefit empirically open. From an applied perspective, future work could evaluate drift detection, uncertainty quantification, and fail-safe policies as first-order design variables. Their effects on process-safety metrics (e.g., alarm fidelity, false-negative rates, and near-miss frequency) could be explicitly quantified to ensure that claimed environmental and safety gains remain auditable, reproducible, and lifecycle-positive.
The field AI-enabled green technological innovation has gained momentum in applications; the most reliable path to meaningful, lifecycle-positive green innovation is to synchronise capability, context, and governance, and evidencing that coherence in practice. From a technological-forecasting perspective, the observed stratification implies that future system performance will depend less on incremental gains in algorithmic accuracy and more on institutional synchronisation. If governance mechanisms, such as lifecycle disclosure standards, carbon-aware MLOps, and sector-specific oversight protocols, scale proportionately with technical capability, AI-enabled sustainability systems may converge toward lifecycle-positive equilibria. Conversely, persistent asymmetry may entrench rebound effects and externality displacement. Monitoring cross-layer co-movement can become a useful diagnostic of socio-technical system resilience.

Author Contributions

Conceptualization, C.G.; data curation, C.G.; Methodology, C.G., J.R. and T.L.; Formal analysis, C.G., J.R. and T.L.; Investigation, C.G., J.R. and T.L.; Writing—original draft, C.G., J.R. and T.L.; Writing—review and editing, C.G., J.R. and T.L.; Visualisation, J.R. and T.L.; Project administration, C.G., J.R. and T.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

Restrictions apply to the availability of the original bibliographic data. The data were obtained from Web of Science (Clarivate) through an institutional subscription and are subject to Clarivate’s licencing terms. Derived datasets generated during this study are available from the corresponding author upon reasonable request, subject to the applicable licencing terms.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AISPAI-Sustainability Pathways
SDGSustainable Development Goal
ESGEnvironmental, Social, and Governance
LDALatent Dirichlet Allocation
BERTBidirectional Encoder Representations from Transformers
BERTopicBERT-based Topic Modelling
SciBERTScientific BERT
SPECTERScientific Paper Embeddings using Citation-informed Transformers
HDBSCAN Hierarchical Density-Based Spatial Clustering of Applications with Noise
PCAPrincipal Component Analysis
t-SNEt-Distributed Stochastic Neighbor Embedding
CO2Carbon dioxide
kgCO2eKilograms of carbon dioxide equivalent
kWh Kilowatt-hour
MLOpsMachine Learning Operations
LCALife Cycle Assessment
SAIDISystem Average Interruption Duration Index
SAIFISystem Average Interruption Frequency Index
EROIEnergy Return on Investment
MRVMeasurement, Reporting, and Verification
PAYTPay-As-You-Throw
EVElectric Vehicle
VARVolt-ampere reactive
MEPMinimum Evidence Package
FPGAField-Programmable Gate Array
WSNWireless Sensor Network

Appendix A

Table A1. Comparative summary of selected prior reviews and science-mapping studies on AI and sustainability, and the positioning of the present study.
Table A1. Comparative summary of selected prior reviews and science-mapping studies on AI and sustainability, and the positioning of the present study.
Study [Ref.]Type & MethodDataDomain/ScopeKey Focus & ContributionKey Limitation (Relative to an Integrated Field View)
A. Integrative/cross-domain syntheses of AI and sustainability
Rolnick et al. (2022) [3]Narrative expert synthesis (non-systematic)No systematic corpus; curated expert selection of ML workML and climate action across many sectors (e.g., energy, agriculture, materials, etc.)Wide agenda of ML applications for climate mitigation, adaptation and resilience; high-impact framingNot corpus-based or reproducible; no thematic-structure extraction; does not separate enabling/sectoral/governance layers or trace temporal dynamics
Vinuesa et al. (2020) [4]Expert consensus assessment (non-bibliometric)Structured expert evaluation of AI against the 17 SDGs/169 targetsAI and all SDGs (environmental, social, economic)Maps enabling vs. inhibiting effects of AI on the SDGs; widely cited positioningQualitative; no literature corpus or topic modelling; no innovation-system layering or co-movement analysis
Nishant et al. (2020) [16]Conceptual review and research agendaSelective literature; conceptual synthesis (corpus not systematic)AI for environmental sustainability (cross-sector)Challenges/opportunities framing; proposes an AI-for-sustainability research agendaConceptual, not empirical/computational; no structural mapping of the field
Di Vaio et al. (2020) [15]Systematic literature review with bibliometric analysis73 publications, 1990–2019AI and sustainable business models (SDG perspective; esp. SDG#12)Structures the AI-SBM literature; identifies a knowledge-management-systems research gapSingle thematic domain (business models); no cross-sector or governance layering; coverage ends pre-2020
B. Sector-specific reviews and surveys
Mellit & Kalogirou (2008) [8]Technical reviewNo systematic corpus; curated narrative review of AI-technique studiesPhotovoltaic/solar energyAI techniques for PV sizing, forecasting and controlSingle sector; technique-focused; pre-deep-learning; no sustainability-system or governance view
Voyant et al. (2017) [9]Methodological reviewNo systematic corpus; curated review of ML forecasting-method studiesSolar radiation forecasting/energyComparative review of ML methods for solar forecastingSingle task/sector; no institutional or temporal-structure analysis
Ahmad et al. (2021) [10]Status quo/challenges reviewNarrative synthesisSustainable energy industryStatus, challenges and opportunities of AI in energySingle sector; narrative; no computational field mapping
Mosavi et al. (2019) [11]Systematic reviewML-model studies in energy systems (70 publications, 2000–2018)Energy systemsState-of-the-art catalogue of ML models across energy applicationsSingle sector; model-catalogue focus; no governance layer or co-evolution lens
Wäldchen & Mäder (2018) [12]ReviewMethods & datasets survey (5 publications: 2016–2018)Biodiversity monitoringML for image-based species identificationNarrow task/sector; no sustainability-system or governance dimension
Naz et al. (2022) [14]Systematic review, bibliometric study and research propositionsLiterature synthesis (353 publications: 2000–2021)Sustainable supply chain/operationsApplications and future propositions for AI in sustainable SCMSingle sector; propositional, not empirical field mapping
Kamilaris & Prenafeta-Boldú (2018) [27]Survey40 publicationsAgricultureDeep-learning techniques, data and performance in agricultureSingle sector; technique survey; no governance or temporal-structure analysis
Liakos et al. (2018) [28]ReviewML-application studies (40 publications, 2004–2018)AgricultureTaxonomy of ML across crop, livestock, water and soil tasksSingle sector; no cross-layer or co-evolution perspective
Fotovvatikhah et al. (2025) [29]Systematic reviewAI waste-classification studies (97 publications, 2020–2025)Waste management/circular economyAI techniques for automated waste classificationSingle application; technique-focused; no field-level structural mapping
C. Bibliometric/topic-modelling studies of AI and sustainability
Raman et al. (2024) [32]Bibliometric analysis using linkage mapping, topic prominence analysis, and flow vergence gradient network analysis2410 publications, 2017–2022Sustainable development, focusing on SDG 12 (Responsible Consumption and Production) and its interlinkages with other SDGsDevelops an integrated bibliometric framework for SDG linkage mapping, topic mining, and policy recommendationsFocuses on SDG 12 using scientometric and citation-network analyses. Unlike this study, it does not model AI-enabled green innovation as a layered innovation system, employ cross-model transformer triangulation, or diagnose temporal governance lag.
Lampropoulos et al. (2024) [33]Bibliometric review; scientific mapping9182 publications, 1989–2022AI, IoT, AIoT for sustainability and SDGsMaps AIoT evolution, themes, collaborations, and SDG-related research directionsTechnology-centric; lacks innovation-system layering, governance dynamics, cross-model triangulation, inter-layer temporal diagnostics.
D. Topic-modelling/science-mapping in adjacent domains
Sharma et al. (2022) [36]Topic-modelling review (LDA)LDA on blockchain corpus (933 publications, 2000–2021)Blockchain technologyTrends and research patterns via LDALDA is bag-of-words/non-contextual; single technology domain; single model
Gupta et al. (2024) [37]Systematic review via topic modelling (BERTopic)1319 publications, 1985–2023Generative AISeven topic clusters mapping the generative-AI research landscapeSingle model (BERTopic); single domain; no cross-model triangulation or governance-layer diagnosis
Lee et al. (2024) [38]Systematic review; BERTopic neural topic modelling90 publications, 2018–2024Generative AI applications in financeBERTopic reveals GAI themes, risks, finance LLMs, synthetic data research agendaFinance-only scope; limited corpus; lacks innovation-system layering, cross-domain integration, governance-lag diagnostics
Lim (2024) [39]Systematic literature mapping, topic modelling and network analysis370 publications, 2008–2023ESG & AI in financeTheme structure and AI-technique evolution in ESG-finance researchSingle domain; single embedding/topic approach; no cross-model triangulation or enabling/sectoral/governance layering
Present study
Present study (this paper)Computational science mapping with cross-model triangulation (BERTopic, SciBERT, SPECTER); bootstrap stability; Spearman inter-layer co-movement3357 publications, 2003–2025AI and green technological innovation, spanning enabling capabilities, sectoral applications and governance/institutional domainsIntegrative, field-level AISP framework; a reproducible triadic structure recovered across three independent embedding regimes; diagnostic of thematic synchronisation and governance lagExtends single-model topic-mapping of green/sustainable AI and descriptive AI-SDG science mapping via three-model transformer triangulation on a green-technology-specific corpus; adds the AISP enabling–sectoral–governance layering and an inter-layer co-movement/governance-lag diagnostic absent from prior mappings. Remains exploratory/correlational (formal co-evolution testing flagged as future work)

Appendix B

Table A2. Evidence-to-action matrix with a cross-layer alignment.
Table A2. Evidence-to-action matrix with a cross-layer alignment.
AISP LayerWhat the Evidence Shows (This Study)Readiness Signal (Maturity/Emergence)Key Outcomes & Metrics to ReportStudy Designs & Data (Research)Operational Controls & Checklists (Industry/Engineering)Policy & Oversight Instruments (Governance)Common Pitfalls/Red FlagsIllustrative Use-Cases
Enabling AI Capabilities-Coherent clusters around forecasting, optimisation, anomaly detection
-Rising attention to “Green AI”, compute/energy and emissions terms
-Mature when persistent volume and cross-model recurrence
-Emergent when rapid post-2020 growth from low baseline
-Predictive error and uncertainty bounds
-False-negative rate (safety)
-Training-inference kWh & kgCO2e
-Model/version provenance
-Time-to-drift
-Uncertainty-integrated hazard/risk models
-Drift detection logging
-Paired performance-energy reporting datasets
-Pre-set benchmarks
Safety-by-design:
-Performance and uncertainty mapped to hazards
-Data/concept-drift monitors with escalation/rollback
-Carbon-aware MLOps (e.g., budget, meter, schedule)
-Reproducible pipelines (e.g., data lineage, model cards)
-Procurement clauses for compute-energy/CO2 disclosure
-Acceptance gates that combine accuracy and lifecycle metrics
-Minimum documentation standards
-Optimising purely for accuracy-latency
-Opaque pipelines
-No drift monitoring; missing energy/CO2 disclosures
-AI models changed without change control
-Load/solar/wind forecasting
-Predictive maintenance for rotating equipment
-Computer vision for materials classification
Sectoral Application Pathway: Energy & Smart Grids-Stable, replicated topic
-Strong growth
-Clear sub-themes (micro/macro-grids, storage, electrification)
Mature-SAIDI/SAIFI
-Loss reduction (%)
-Renewable curtailment avoided
-Demand forecast error (with CIs)
-Lifecycle emissions avoided
-Quasi-experimental before-after on feeders
-Synthetic controls for grid sections
-Telemetry and weather fused datasets
-Model governance board
-Envelope-aware validation
-Rollback manuals
-Human-in-the-loop for critical switching
-Event post-mortems tied to model decisions
-Inter-connection rules with explainability thresholds
-Disclosure of model updates to regulators
-Grid-code annexes for AI
-Silent model drift
-Data-quality shocks
-Optimisation that increases hidden curtailment elsewhere
-Peak-shaving optimisation
-Outage prediction
-Voltage/VAR control
Sectoral Application Pathway: Waste Mgmt & Biofuels-Sustained increase across years
-Circular economy logistics and conversion pathways visible
Mature-Diversion/recycling rate
-Route efficiency (km/ton)
-Contamination rate
-Biofuel yield & EROI
-Lifecycle GHG per ton
-Matched-pair depot studies
-Route-level A/B tests
-LCA integrated with operational telemetry
-Traceable routing (telematics and audit)
-Contamination classifiers with confidence thresholds
-Labelling & operator feedback loops
-Municipal data-sharing mandates
-PAYT and AI audit rules
-LCA reporting for public contracts
-Shifting burdens to downstream processors
-No contamination verification
-“Black-box” routing
-Dynamic routing
-Robotic sorting
-Feedstock blending for biogas/biochar
Sectoral Application Pathway: Agriculture & Food Systems-Strong presence
-Finer splits (crop, livestock, forestry) in citation-aware model
-Maturing (crop/forest)/Slower (livestock/biotech)-Yield/ha
-Water & nutrient intensity
-Pest/disease detection precision/recall
-Deforestation avoided
-Leakage checks
-Field-cluster stepped-wedge trials
-Satellite and IoT fusion
-Counter-factual land-use modelling
-Sensor calibration SOPs
-Label drift reviews by agronomists
-Boundary checks for land-use leakage
-MRV (measurement-reporting-verification) for land-use
-Sustainability certification tie-ins
-Overfitting to season/site
-Neglecting smallholder constraints
-Leakage to adjacent lands
-Variable-rate irrigation
-Early blight detection
-Deforestation alerts
Sectoral Application Pathway: Urban Planning & Transport-Clear, replicated clusters
-Integration with mobility & planning
Maturing-Travel time reliability
-Modal share; emissions per p-km
-Equity of service distribution
-Interrupted time-series around policy changes
-Agent-based sims calibrated to observed flows
-Bias/impact assessment in siting
-Resilience checks under disruptions
-Citizen-facing transparency
-Open mobility data standards
-Algorithmic impact assessments
-Auditable procurement
-Optimising for average user only
-Induced demand
-Surveillance creep
-Bus priority optimisation
-EV charging siting
-Curb management
Governance & Institutional Domains: Sustainability EducationPost-2020 acceleration from low baseEmergent (catch-up)-Workforce competency indices
-Curricula coverage
-Training hours
-Uptake by professional bodies
-Longitudinal training-outcome studies
-Competency frameworks linked to safety/env. KPIs
-Mandatory operator training
-Competency matrices for AI-enabled roles
-Accreditation standards including AI & LCA competencies-Training not tied to SOPs or KPIs-Grid operator training
-Waste-facility upskilling
Governance & Institutional Domains: Innovation Strategy/Blockchain & Governance-Fast growth (strategy)
-Smaller but rising governance clusters
-Higher model sensitivity
Emergent-Presence/quality of LCA
-Transparency & audit scores
-Adoption of disclosure templates
-Time-to-policy after pilot
-Mixed-method case studies of policy uptake
-Document analysis of model cards & LCAs
-Release gates requiring LCA
-Contract clauses for audit trails
-Chain-of-custody for data
-Standardised model cards
-Compute/energy disclosure
-Regulatory sandboxes with pre-agreed outcomes
-Performative compliance
-Carbon accounting bolted on post hoc
-Unverifiable claims
-Carbon-labelled AI services
-Provenance on environmental data
Cross-Layer Alignment-Technical/application growth outpaces governance volume
-Governance topics more volatile/lagged
Misalignment risk-Alignment index (co-movement across layers)
-Share of deployments with LCA and audit
-Lag (months) between tech release and governance artifacts
-Normalised trend and citation-coupling analyses
-Mediation tests, including governance as mechanism from AI to outcomes
-Minimum Evidence Package (MEP): performance and uncertainty
-Drift controls
-Lifecycle footprint
-Provenance
-Context validation
-Phased compliance: disclosure-audit performance-linked incentives
-Carbon budgets in acceptance criteria
-Impact claims without lifecycle net effect
-No independent audit
-Frequent untracked model changes
-Portfolio reviews for alignment
-Governance milestones in programme charters

References

  1. Parmesan, C.; Morecroft, M.D.; Trisurat, Y. Climate Change 2022: Impacts, Adaptation and Vulnerability; GIEC: Melbourne, Australia, 2022. [Google Scholar]
  2. UNEP. Emissions Gap Report 2022; United Nations Environment Programme: Nairobi, Kenya, 2022. [Google Scholar]
  3. Rolnick, D.; Donti, P.L.; Kaack, L.H.; Kochanski, K.; Lacoste, A.; Sankaran, K.; Ross, A.S.; Milojevic-Dupont, N.; Jaques, N.; Waldman-Brown, A.; et al. Tackling climate change with machine learning. ACM Comput. Surv. (CSUR) 2022, 55, 1–96. [Google Scholar] [CrossRef] [Scilit]
  4. Vinuesa, R.; Azizpour, H.; Leite, I.; Balaam, M.; Dignum, V.; Domisch, S.; Felländer, A.; Langhans, S.D.; Tegmark, M.; Nerini, F.F. The role of artificial intelligence in achieving the Sustainable Development Goals. Nat. Commun. 2020, 11, 233. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Behera, B.; Behera, P.; Pata, U.K.; Sethi, L.; Sethi, N. Artificial intelligence-driven green innovation for sustainable development: Empirical insights from India’s renewable energy transition. J. Environ. Manag. 2025, 389, 126285. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Raman, R.; Pattnaik, D.; Lathabai, H.H.; Kumar, C.; Govindan, K.; Nedungadi, P. Green and sustainable AI research: An integrated thematic and topic modeling analysis. J. Big Data 2024, 11, 55. [Google Scholar] [CrossRef] [Scilit]
  7. Chotia, V.; Cheng, Y.; Agarwal, R.; Vishnoi, S.K. AI-enabled Green Business Strategy: Path to carbon neutrality via environmental performance and green process innovation. Technol. Forecast. Soc. Change 2024, 202, 123315. [Google Scholar] [CrossRef] [Scilit]
  8. Mellit, A.; Kalogirou, S.A. Artificial intelligence techniques for photovoltaic applications: A review. Prog. Energy Combust. Sci. 2008, 34, 574–632. [Google Scholar] [CrossRef] [Scilit]
  9. Voyant, C.; Notton, G.; Kalogirou, S.; Nivet, M.-L.; Paoli, C.; Motte, F.; Fouilloy, A. Machine learning methods for solar radiation forecasting: A review. Renew. Energy 2017, 105, 569–582. [Google Scholar] [CrossRef] [Scilit]
  10. Ahmad, T.; Zhang, D.; Huang, C.; Zhang, H.; Dai, N.; Song, Y.; Chen, H. Artificial intelligence in sustainable energy industry: Status Quo, challenges and opportunities. J. Clean. Prod. 2021, 289, 125834. [Google Scholar] [CrossRef] [Scilit]
  11. Mosavi, A.; Salimi, M.; Ardabili, S.F.; Rabczuk, T.; Shamshirband, S.; Varkonyi-Koczy, A.R. State of the art of machine learning models in energy systems, a systematic review. Energies 2019, 12, 1301. [Google Scholar] [CrossRef] [Scilit]
  12. Wäldchen, J.; Mäder, P. Machine learning for image based species identification. Methods Ecol. Evol. 2018, 9, 2216–2225. [Google Scholar] [CrossRef] [Scilit]
  13. Bag, S.; Telukdarie, A.; Pretorius, J.H.C.; Gupta, S. Industry 4.0 and supply chain sustainability: Framework and future research directions. Benchmarking Int. J. 2021, 28, 1410–1450. [Google Scholar]
  14. Naz, F.; Agrawal, R.; Kumar, A.; Gunasekaran, A.; Majumdar, A.; Luthra, S. Reviewing the applications of artificial intelligence in sustainable supply chains: Exploring research propositions for future directions. Bus. Strategy Environ. 2022, 31, 2400–2423. [Google Scholar] [CrossRef] [Scilit]
  15. Di Vaio, A.; Palladino, R.; Hassan, R.; Escobar, O. Artificial intelligence and business models in the sustainable development goals perspective: A systematic literature review. J. Bus. Res. 2020, 121, 283–314. [Google Scholar] [CrossRef] [Scilit]
  16. Nishant, R.; Kennedy, M.; Corbett, J. AArtificial intelligence for sustainability: Challenges, opportunities, and a research agenda. Int. J. Inf. Manag. 2020, 53, 102104. [Google Scholar] [CrossRef] [Scilit]
  17. Cicerone, G.; Faggian, A.; Montresor, S.; Rentocchini, F. Regional artificial intelligence and the geography of environmental technologies: Does local AI knowledge help regional green-tech specialization? Reg. Stud. 2023, 57, 330–343. [Google Scholar]
  18. Mansour, M.; Zobi, M.A.; Alomair, M. Artificial Intelligence, ESG Governance, and Green Innovation Efficiency in Emerging Economies. Economies 2026, 14, 11. [Google Scholar]
  19. Iqbal, A.; Zhang, W.; Jahangir, S. Building a Sustainable Future: The Nexus Between Artificial Intelligence, Renewable Energy, Green Human Capital, Geopolitical Risk, and Carbon Emissions Through the Moderating Role of Institutional Quality. Sustainability 2025, 17, 990. [Google Scholar] [CrossRef] [Scilit]
  20. Feng, B.; Chen, X.; Tang, H. AI-driven green governance: Assessing the impact of artificial intelligence on corporate sustainability performance. J. Innov. Knowl. 2026, 11, 100869. [Google Scholar] [CrossRef] [Scilit]
  21. Schwartz, R.; Dodge, J.; Smith, N.A.; Etzioni, O. Green AI. Commun. ACM 2020, 63, 54–63. [Google Scholar] [CrossRef] [Scilit]
  22. Strubell, E.; Ganesh, A.; McCallum, A. Energy and policy considerations for modern deep learning research. Proc. AAAI Conf. Artif. Intell. 2020, 34, 13693–13696. [Google Scholar] [CrossRef] [Scilit]
  23. Geels, F.W. Technological transitions as evolutionary reconfiguration processes: A multi-level perspective and a case-study. Res. Policy 2002, 31, 1257–1274. [Google Scholar] [CrossRef] [Scilit]
  24. Von Bertalanffy, L. An outline of general system theory. Br. J. Philos. Sci. 1950, 1, 134–165. [Google Scholar] [CrossRef] [Scilit]
  25. Cowls, J.; Tsamados, A.; Taddeo, M.; Floridi, L. The AI gambit: Leveraging artificial intelligence to combat climate change—Opportunities, challenges, and recommendations. AI Soc. 2023, 38, 283–307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Geraghty, P. Environmental assessment and the application of expert systems: An overview. J. Environ. Manag. 1993, 39, 27–38. [Google Scholar] [CrossRef] [Scilit]
  27. Kamilaris, A.; Prenafeta-Boldú, F.X. Deep learning in agriculture: A survey. Comput. Electron. Agric. 2018, 147, 70–90. [Google Scholar] [CrossRef] [Scilit]
  28. Liakos, K.G.; Busato, P.; Moshou, D.; Pearson, S.; Bochtis, D. Machine learning in agriculture: A review. Sensors 2018, 18, 2674. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Fotovvatikhah, F.; Ahmedy, I.; Noor, R.M.; Munir, M.U. A Systematic Review of AI-Based Techniques for Automated Waste Classification. Sensors 2025, 25, 3181. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Lundvall, B.A. National Systems of Innovation: Towards a Theory of Innovation and Interactive Learning; Pinter Publishers: London, UK, 1992. [Google Scholar]
  31. Freeman, C. Technology Policy and Economic Performance: Lessons from Japan; Pinter Publishers: New York, NY, USA, 1987. [Google Scholar]
  32. Raman, R.; Lathabai, H.H.; Nedungadi, P. Sustainable development goal 12 and its synergies with other SDGs: Identification of key research contributions and policy insights. Discov. Sustain. 2024, 5, 150. [Google Scholar] [CrossRef] [Scilit]
  33. Lampropoulos, G.; Garzón, J.; Misra, S.; Siakas, K. The Role of Artificial Intelligence of Things in Achieving Sustainable Development Goals: State of the Art. Sensors 2024, 24, 1091. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Zupic, I.; Čater, T. Bibliometric Methods in Management and Organization. Organ. Res. Methods 2014, 18, 429–472. [Google Scholar] [CrossRef] [Scilit]
  35. van Eck, N.J.; Waltman, L. Visualizing Bibliometric Networks, in Measuring Scholarly Impact; Ding, Y., Rousseau, R., Wolfram, D., Eds.; Springer: Cham, Switzerland, 2014. [Google Scholar]
  36. Sharma, C.; Sharma, S.; Sakshi. Latent DIRICHLET allocation (LDA) based information modelling on BLOCKCHAIN technology: A review of trends and research patterns used in integration. Multimed. Tools Appl. 2022, 81, 36805–36831. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Gupta, P.; Ding, B.; Guan, C.; Ding, D. Generative AI: A systematic review using topic modelling techniques. Data Inf. Manag. 2024, 8, 100066. [Google Scholar] [CrossRef] [Scilit]
  38. Lee, D.K.C.; Guan, C.; Yu, Y.; Ding, Q. A comprehensive review of generative AI in finance. FinTech 2024, 3, 460–478. [Google Scholar] [CrossRef] [Scilit]
  39. Lim, T. Environmental, social, and governance (ESG) and artificial intelligence in finance: State-of-the-art and research takeaways. Artif. Intell. Rev. 2024, 57, 76. [Google Scholar] [CrossRef] [Scilit]
  40. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minnesota, MI, USA, 2–7 June 2019; Volume 1. long and short papers. [Google Scholar]
  41. Grootendorst, M. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv 2022, arXiv:2203.05794. [Google Scholar]
  42. Beltagy, I.; Lo, K.; Cohan, A. SciBERT: A pretrained language model for scientific text. arXiv 2019, arXiv:1903.10676. [Google Scholar]
  43. Cohan, A.; Feldman, S.; Beltagy, I.; Downey, D.; Weld, D.S. Specter: Document-level representation learning using citation-informed transformers. arXiv 2020, arXiv:2004.07180. [Google Scholar]
  44. Reimers, N.; Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv 2019, arXiv:1908.10084. [Google Scholar]
  45. Rousseeuw, P.J. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J. Comput. Appl. Math. 1987, 20, 53–65. [Google Scholar] [CrossRef] [Scilit]
  46. Grootendorst, M. KeyBERT: Minimal Keyword Extraction with BERT; Zenodo: Geneva, Switzerland, 2020. [Google Scholar]
  47. van der Maaten, L.; Hinton, G. Visualizing Data using t-SNE. J. Mach. Learn. Res. 2008, 9, 2579–2605. [Google Scholar]
  48. Fleiss, J.L. Measuring nominal scale agreement among many raters. Psychol. Bull. 1971, 76, 378. [Google Scholar] [CrossRef] [Scilit]
  49. Parida, R.; Dash, M.K.; Kumar, A.; Zavadskas, E.K.; Luthra, S.; Mulat-Weldemeskel, E. Evolution of supply chain finance: A comprehensive review and proposed research directions with network clustering analysis. Sustain. Dev. 2022, 30, 1343–1369. [Google Scholar] [CrossRef] [Scilit]
  50. Halkidi, M.; Batistakis, Y.; Vazirgiannis, M. On clustering validation techniques. J. Intell. Inf. Syst. 2001, 17, 107–145. [Google Scholar] [CrossRef] [Scilit]
  51. Hall, P. The Bootstrap and Edgeworth Expansion; Springer Science & Business Media: Berlin/Heidelberg, Germany, 2023. [Google Scholar]
  52. Hennig, C. Cluster-wise assessment of cluster stability. Comput. Stat. Data Anal. 2007, 52, 258–271271. [Google Scholar] [CrossRef] [Scilit]
Figure 1. The AI-Sustainability Pathways Framework (AISP).
Figure 1. The AI-Sustainability Pathways Framework (AISP).
Analytics 05 00027 g001
Figure 2. Number of publications on the Web of Science (2003–2025).
Figure 2. Number of publications on the Web of Science (2003–2025).
Analytics 05 00027 g002
Figure 3. (a) Two-dimensional t-SNE projections of BERTopic clustering at K = 3 (left) and K = 8 (right). (b) Two-dimensional t-SNE projections of SciBERT clustering at K = 9 (left) and SPECTER clustering at K = 15 (right).
Figure 3. (a) Two-dimensional t-SNE projections of BERTopic clustering at K = 3 (left) and K = 8 (right). (b) Two-dimensional t-SNE projections of SciBERT clustering at K = 9 (left) and SPECTER clustering at K = 15 (right).
Analytics 05 00027 g003
Figure 4. Topic trends by publication year (share percentage per year).
Figure 4. Topic trends by publication year (share percentage per year).
Analytics 05 00027 g004
Table 1. Cross-model mapping onto AISP layers (enablers, applications, governance).
Table 1. Cross-model mapping onto AISP layers (enablers, applications, governance).
General TopicsBERTopicSciBERTSPECTER
Enabling AI CapabilitiesGreen AI InfrastructureGreen AI InfrastructureGreen AI Infrastructure
Green Technology SystemSustainability Assessment and Lifecycle Systems;
Sustainable Industrial Systems
Sectoral Application PathwaysSmart Urban Systems;
Sustainable Cities
Urban Planning and Transport
AI in Sustainable HealthcareDigital Sustainability and Health;
Agroecology and Natural Systems
Sustainable Healthcare Systems
AI in Green Agriculture Agriculture and Food Systems;
Livestock and Biotechnology;
Forest and Land Use
Industrial Sustainability SystemsSustainable Consumption and BehaviourCircular Economy and Lifecycle;
Sustainable Business;
Sustainability Innovation Strategy
Emissions & Environmental AnalyticsSustainable Energy and BiosystemsEnergy and Smart Grids;
Waste Management and biofuels
Cybersecurity and Smart Infrastructure
Sustainability Education
Governance & Institutional DomainsAI Governance & Environmental EthicsAI Governance and Ethics;
Sustainability Ethics and Society
Blockchain and Governance
Table 2. Mapping of BERTopic (K = 3) clusters to the three AISP layers.
Table 2. Mapping of BERTopic (K = 3) clusters to the three AISP layers.
BERTopic ClusterAISP Layer
Cluster 0Enabling capability
Cluster 1Governance & institutional
Cluster 2Sectoral application
Table 3. Spearman correlation between annual AISP layer shares.
Table 3. Spearman correlation between annual AISP layer shares.
Layer PairSpearman ρp-Value
Enabling-Sectoral0.858<0.001
Enabling-Governance0.837<0.001
Sectoral-Governance0.767<0.001
Table 4. Transformer-based representation models for topic discovery and triangulation.
Table 4. Transformer-based representation models for topic discovery and triangulation.
ModelsPre Training Data & SignalTypical Representation
BERTopicGeneral documentsSentence-level embeddings
SciBERTScientific text
(Semantic scholar corpus)
Sentence-level embeddings
SPECTERScientific papers, and fine-tuned using citation networks to reflect document similarityDocument-level embeddings
Table 5. Internal validation across clustering granularities (K = 3–25).
Table 5. Internal validation across clustering granularities (K = 3–25).
Silhouette Score
ModelsK = 3K = 5K = 6K = 7K = 8K = 9K = 10K = 15K = 20K = 25
BERTopic 0.04690.03140.02980.02830.03490.03000.03010.03410.03240.0317
SciBERT 0.04680.03990.03810.03750.03550.03700.03070.03010.02660.0246
SPECTER0.07280.05340.05290.05450.04740.04700.05380.06090.05970.0584
Davies–Bouldin (DB) Index
ModelsK = 3K = 5K = 6K = 7K = 8K = 9K = 10K = 15K = 20K = 25
BERTopic 3.96124.11164.10363.94654.00313.82513.76913.46943.46963.4083
SciBERT 3.77433.39353.54013.41363.53923.49823.61763.47673.42193.4382
SPECTER3.38243.14493.14343.00323.01792.97322.90002.72702.79362.6977
Table 6. BERTopic topic solutions at coarse and fine granularities (K = 3 and K = 8).
Table 6. BERTopic topic solutions at coarse and fine granularities (K = 3 and K = 8).
BERTopicsKeywordsSize
K = 3
Green AI Infrastructureai, tensorflow, programmable, fpgas, supercomputer, environmentally, greenrunner, intel, greenai, fpga, supercomputing, sustainable, eco2ai, ais, efficiencies, machines, computing, cloud1326
Governance, Ethics and Sustainabilitysustainability, sustainable, environmentally, ai, ecoinvent, ethicality, renewable, ethics, eco, econeighborhoods, greenability, technoscientific, intelligentization, intellectual, ecologies, ecoresponsibility, ecovillages1459
Sustainable Healthcare Systemsappointment, patient, consultation, inpatient, consultation, calendar, scheduled, outpatient, interprofessional, reminder, clinician, telemedicine, assistant, hospitalizations, timetable572
K = 8
Smart Urban Systemsurbanization, urbanizing, ai, smartdevops, urban, sustainability, greening, macroenvironment, autonomous, automation381
Industrial Sustainability Systemssustainability, ai, ecoinvent, industrial, intelligentization, industrialized, creation, innovation, automation, technological353
AI in Sustainable Healthcarehealthcare, ai, bioethics, ethics, biomedicine, health, sustainability, environmental, healthdata, wellness262
AI in Green Agricultureagriculture, farm, ai, crops, agronomic, agri, farmer, agrifood, agribusiness, foodbioprocesses471
Green AI Infrastructureai, sustainability, tensorflow, environmental, fpgas, programmable, greenrunner, renewable, greenai, intel374
AI Governance and Environmental Ethicssustainability, environmental, ai, ethics, renewable, eco, technoscientific, intellectual, ecologies, ecovillages513
Sustainable Citiessustainability, ai, environmental, green, eco, renewable, ais, emissions, econeighborhoods, econ458
Emissions and Environmental Analyticsai, ecoinvent, greenability, sustainability, environmental, eco, emissions, greenhouse, renewable, economized545
Table 7. SciBERT topic solution at K = 9 (science-domain representation).
Table 7. SciBERT topic solution at K = 9 (science-domain representation).
SciBERT TopicsKeywordsSize
Sustainable Energy and BioSystemstelehealth, powertrains, biosensing, sustainability, bioenergy, analytics, biofuels, geoscience, biogas, nanoflowers529
Agroecology and Natural Systemsbiogas, ejaculate, transhumant, evapotranspiration, inventories, remanufacturing, biodiesel, cambisol, botanical, agroecologies203
Cybersecurity and Smart Infrastructurecyberattacks, interoperability, extensibility, datacenters, wearables, crypto, smartness, programmability, wsn, meshwork256
Digital Sustainability and Healthmhealth, emancipation, citizenship, contextualism, pedagogy, humanity, sustainability, telemedicine, policymaking, deliberation387
AI Governance and Ethicsmorality, sustainability, telehealth, humanity, transhumanism, sociomateriality, individuality, epistemology, futurism400
Green Technology Systemunsustainability, blockchains, ecoinvent, geopolymer, biofuels, metaheuristics, energyplus, metaheuristic, audiovisual, bioenergy377
Green AI Infrastructureworkflows, roadmaps, softmax, sustainability, kernel, verification, metaheuristic, footprint, tracking, hvac382
Sustainability Ethics and Societyhumanity, bioethics, discourse, sustainability, feminism, pedagogy, posthumanism, dynamism, telehealth, overconsumption561
Sustainable Consumption and Behaviourentrepreneurship, personalization, overconsumption, remanufacturing, lifestyles, sustainability, meaningfulness, behavior, personalizing, decisionmaking262
Table 8. SPECTER topic solution at K = 15 (citation-aware representation).
Table 8. SPECTER topic solution at K = 15 (citation-aware representation).
SPECTER TopicsKeywordsSize
Urban Planning and Transportroadmap, city, transportation, urban, cityscape, urbanization, municipality, placelessness, geoai211
Agriculture and Food Systemsagro, farmland, farm, agrifood, agriculture, bioenergy, agroecology, cropland, crop, agritourism296
Sustainability Assessment and Lifecycle Systemssustainability, green, ecologies, criticality, lifecycle, vitality, empowers, greenwash, renewable, restructuring372
Energy and Smart Gridssmartgrid, powertrains, renewables, microgrids, bioenergy, biofuel, electrification, sustainable, battery, hydropower217
Sustainability Educationpedagogy, curriculum, empowers, sustainable, curricular, classroom, empowering, didactical, green, lifecycle309
Blockchain and Governanceblockchain, blockchainops, legislation, protections, governance, policymaking, roadmap, lifecycle, ethics, legally180
Sustainable Industrial Systemssmarter, industrial, building, industry, occupants, green, woodbased, sustainable, technology, architects, automation202
Sustainable Businessgreen, retailers, consumers, sourcing, retailing, business, brand, businessmen, sustainable, innovation299
Circular Economy and Lifecyclelifecycle, industrial, business, sustainable, ecofriendly, enterprise, manufacturing, restructuring, remanufacturing, company175
Livestock and Biotechnologycowpea, slaughtering, breeders, broiler, biotech, feedstuffs, straws, meat, breading, pests67
Waste Management and biofuelswaste, biofuel, recycled, bioenergy, wastewater, biochar, biogas, landfill, effluent, biowaste203
Sustainability Innovation Strategygreenovation, leaning, strategy, innovativeness, innovation, entrepreneurial, enterprise, business, microenterprise, strategize253
Sustainable Healthcare Systemshealthcare, carers, ethics, care, bioethics, beneficence, hospitals, practice, nursing, medicolegal229
Green AI Infrastructuregreenauto, greenrunner, industry, carbontracker, big, greenai, green, greentsf, greenness, technology295
Forest and Land Useforestland, afforestation, deforestation, blueforest, reforestation, landholders, agroforestry, peatland, landowners, grassland178
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Guan, C.; Ren, J.; Lim, T. Co-Evolution of Artificial Intelligence and Green Technological Innovation: A Computational Mapping and Diagnostic Framework. Analytics 2026, 5, 27. https://doi.org/10.3390/analytics5030027

AMA Style

Guan C, Ren J, Lim T. Co-Evolution of Artificial Intelligence and Green Technological Innovation: A Computational Mapping and Diagnostic Framework. Analytics. 2026; 5(3):27. https://doi.org/10.3390/analytics5030027

Chicago/Turabian Style

Guan, Chong, Jing Ren, and Tristan Lim. 2026. "Co-Evolution of Artificial Intelligence and Green Technological Innovation: A Computational Mapping and Diagnostic Framework" Analytics 5, no. 3: 27. https://doi.org/10.3390/analytics5030027

APA Style

Guan, C., Ren, J., & Lim, T. (2026). Co-Evolution of Artificial Intelligence and Green Technological Innovation: A Computational Mapping and Diagnostic Framework. Analytics, 5(3), 27. https://doi.org/10.3390/analytics5030027

Article Metrics

Back to TopTop