AI-Enhanced Evolutionary Game Theory for Intelligent Coordination and Adaptive Optimization in Low-Carbon Energy Systems: A Multi-Scale Review from Smart Grids to Carbon Markets
Abstract
1. Introduction
1.1. Research Background and Problem Posing
1.2. Research Status and Literature Review
1.3. Research Objectives and Overview Framework
1.4. Review Methodology and Literature Selection Protocol
| Algorithm 1. Variance-Gated Regime-Switching Replicator Dynamics (VG-RSRD) | |
| Input: populations P = {S, B, D, G} with strategy shares xp; coupled payoffs πp from Equations (6)–(9); competitor baselines φp; renewable scenario set Ω; stable step band [αlo, αhi]; exploratory step αex; window W; oscillation tolerance ε; hysteresis margins δin < δout; gains κ, ρ; tolerance τ; horizon Tmax | |
| Output: stabilized strategy profile x*; oscillation trace O; regime-switch log | |
| 1 | xp ← xp0 for all p ∈ P; regime ← EXPLORE; α ← αex; O ← ∅ |
| 2 | for t = 1 to Tmax do |
| 3 | draw scenario Ω ~ Ω ▷ inject intermittency into selection |
| 4 | for each p ∈ P: ← E[πp(x, Ω)]; μp ← Σq xq πq |
| 5 | Δxp ← xp ( − μp) for all p ▷ raw replicator drift |
| 6 | push ‖Δx‖2 to O; if |O| > W pop oldest; σ ← std(O) |
| 7 | if σ > ε then α ← max(αlo, α/(1 + κσ)) ▷ variance gate: contract |
| 8 | else α ← min(αhi, α (1 + ρ)) ▷ relax toward upper edge |
| 9 | d ← ‖xt − xt−W‖ ▷ windowed drift |
| 10 | if regime = EXPLORE and d < δin then regime ← LOCK-IN; α ← αlo |
| 11 | if regime = LOCK-IN and d > δout then regime ← EXPLORE; α ← αex |
| 12 | xp ← Πsimplex(xp + α Δxp) for all p ▷ project onto probability simplex |
| 13 | if maxp ‖Δxp‖ < τ and regime = LOCK-IN then break |
| 14 | end for |
| 15 | return x*, O, regime-switch log |
| Algorithm 2. Price-Band Triggered Adaptive Quota Co-Evolution (PBT-AQE) | |
| Input: emitter population with strategy shares yi ∈ {BUY, INVEST}; marginal abatement cost MACi; firm investment-trigger price pi*; regulator price band [plo, phi]; reserve pool Qres; cap Qcap; CCER offset ratio λ; shortfall penalty πpen; replicator rate η; band-adjustment gain γ; emission target E*; horizon T | |
| Output: equilibrium price path; firm strategy distribution {yi*}; reserve-pool trajectory; realized emissions | |
| 1 | allocate quotas by benchmark/historical rule; yi ← yi0; p0 ← market clear |
| 2 | for t = 1 to T do |
| 3 | aggregate quota demand from {yi}; clear market → price pt |
| 4 | // regulator regime switch on price band |
| 5 | if pt > phi then release ΔQ from Qres; mode ← RELEASE ▷ loosen supply |
| 6 | elif pt < plo then repurchase ΔQ into Qres; mode ← REPURCHASE ▷ firm up price |
| 7 | else hold supply; mode ← NEUTRAL |
| 8 | // firm regime switch on investment trigger |
| 9 | for each emitter i do |
| 10 | if pt ≥ pi* then bias yi → INVEST (abatement/CCUS) |
| 11 | else bias yi → BUY (quota + CCER offset at ratio λ) |
| 12 | ui ← −(quota cost) or −MACi·ai, minus πpen if short ▷ period payoff |
| 13 | yi ← replicator_update(yi, ui, η) for all i |
| 14 | Et ← realized emissions; [plo, phi] ← band ± γ(Et − E*) ▷ adapt band |
| 15 | settle compliance; update Qres and emission ledger |
| 16 | if pt ∈ [plo, phi] and {yi} stable then break |
| 17 | end for |
| 18 | return price path, {yi*}, Qres trace, emissions |
| Algorithm 3. Federated Payoff Co-Learning with Differential-Privacy Masking (FPC-DP) | |
| Input: K clients with private datasets Dk (cost/emission/output records); local payoff model fθ; global rounds R; local epochs E; clip bound C; DP noise scale σ; privacy budget (ε, δ); local strategy shares xk; replicator rate η | |
| Output: shared payoff model θ*; per-client strategy profiles {xk*}; consumed privacy budget | |
| 1 | server init θ0; broadcast to all clients; spent ← 0 |
| 2 | for r = 1 to R do |
| 3 | for each client k in parallel do |
| 4 | fit fθ on Dk for E epochs ▷ local empirical fitness |
| 5 | Δθk ← parameter update; clip ‖Δθk‖ ≤ C ▷ bound per-client sensitivity |
| 6 | Δ ← Δθk + N(0, σ2C2I) ▷ differential-privacy mask |
| 7 | secret-share Δ for secure aggregation ▷ no raw update revealed |
| 8 | θr ← θr−1 + (1/K) Σk Δ ▷ aggregate masked updates only |
| 9 | broadcast θr; each client sets (·) ← fθr(·) |
| 10 | xk ← replicator_update(xk, , η) ▷ game step stays local |
| 11 | spent ← spent + per-round privacy cost |
| 12 | if spent ≥ (ε, δ) then break ▷ stop at budget ceiling |
| 13 | end for |
| 14 | return θ*, {xk*}, spent |
2. The Theoretical Basis of EGT in Cleaner Production
2.1. The Core Theory of EGT and Its Applicability in Cleaner Production
2.2. Game Theory Modeling Framework of Cleaner Production System
2.3. Evolutionary Game Mechanism in Complex Network Environment
3. Evolutionary Game Model and Algorithm of Cleaner Production System
3.1. Evolutionary Game Model of Industrial Symbiosis Network
3.2. Evolutionary Game Algorithm for Renewable Energy Configuration
3.3. Mechanism Design and Evolution Analysis of Carbon Trading Market
4. Artificial Intelligence Enhanced Evolutionary Game Method
4.1. Fusion Innovation of Machine Learning and Evolutionary Game
4.2. Application of Multi-Agent Reinforcement Learning in Clean Energy System
4.3. Decentralized Evolutionary Game Mechanism Supported by Blockchain Technology
5. Application of EGT in Clean Energy System
5.1. Evolutionary Game Mechanism of Smart Grid Demand Response
5.2. Coordinated Optimization Game of Distributed Renewable Energy
5.3. Evolutionary Game Optimization of Electric Vehicle Charging Network
5.4. Extension to Broader Cleaner-Production Domains: Waste Heat, Water–Energy Coupling, Hydrogen, and Circular Configurations
6. The Application of EGT in Environmental Policy Design
6.1. Evolutionary Game Analysis of Environmental Tax Policy
6.2. Evolutionary Game Mechanism of Regional Environmental Cooperation
6.3. Game Optimization of Green Development Incentive Mechanism
7. Illustrative Case Analyses: Multi-Scale Applications of EGT in Cleaner Production Systems
7.1. Case Study I: Evolutionary Dynamics of Industrial Symbiosis in Regional Eco-Industrial Parks
- (1)
- Research Motivation and Objectives
- (2)
- Methodological Framework
- (3)
- Core Parameter Configuration and Methodological Settings
- (4)
- Quantitative Results Analysis
7.2. Case Study II: Multi-Agent Evolutionary Game Optimization in Integrated Smart Energy Systems
- (1)
- Research Motivation and Objectives
- (2)
- Methodological Framework
- (3)
- Core Parameter Configuration and Methodological Settings
- (4)
- Quantitative Results Analysis
7.3. Synthesis and Policy Implications: Toward Integrated Evolutionary Governance of Cleaner Production Systems Cross-Case Synthesis
- (1)
- Policy Implications for Cleaner Production Governance
- (2)
- Methodological Contributions and Future Directions
8. Summary and Prospect
8.1. Main Research Conclusions and Theoretical Contributions
8.2. Analysis of Existing Problems and Challenges
8.3. Future Development Direction and Research Prospect
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
References
- Giannetti, B.F.; Agostinho, F.; Eras, J.J.C.; Yang, Z.; Almeida, C.M.V.B. Cleaner production for achieving the sustainable development goals. J. Clean. Prod. 2020, 271, 122127. [Google Scholar] [CrossRef] [Scilit]
- Mudhee, K.H.; Hilal, M.M.; Alyami, M.; Rendal, E.; Algburi, S.; Sameen, A.Z.; Khurramov, A.; Abboud, N.G.; Barakat, M. Assessing climate strategies of major energy corporations and examining projections in relation to Paris Agreement objectives within the framework of sustainable energy. Unconv. Resour. 2025, 5, 100127. [Google Scholar] [CrossRef] [Scilit]
- Woon, K.S.; Phuang, Z.X.; Taler, J.; Varbanov, P.S.; Chong, C.T.; Klemeš, J.J.; Lee, C.T. Recent advances in urban green energy development towards carbon emissions neutrality. Energy 2023, 267, 126502. [Google Scholar] [CrossRef] [Scilit]
- Chung, S.Y.; Ives, M.C.; Allen, M.R.; Doorga, J.R.S.; Xu, Y. Accelerating carbon neutrality in China: Sensitive intervention points for the energy and transport sectors in Beijing and Hong Kong. J. Clean. Prod. 2024, 450, 141681. [Google Scholar] [CrossRef] [Scilit]
- Vieira, L.C.; Amaral, F.G. Barriers and strategies applying Cleaner Production: A systematic review. J. Clean. Prod. 2016, 113, 5–16. [Google Scholar] [CrossRef] [Scilit]
- Li, F.; Cao, X.; Sheng, P. Impact of pollution-related punitive measures on the adoption of cleaner production technology: Simulation based on an evolutionary game model. J. Clean. Prod. 2022, 339, 130703. [Google Scholar] [CrossRef] [Scilit]
- Han, W.; Zhang, Z.; Zhu, Y.; Xia, C. Co-evolutionary dynamics in optimal multi-agent game with environment feedback. Neurocomputing 2024, 581, 127510. [Google Scholar] [CrossRef] [Scilit]
- Křivan, V.; Cressman, R. The asymmetric Hawk-Dove game with costs measured as time lost. J. Theor. Biol. 2022, 547, 111162. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Breed, M.D. Evolutionarily Stable Strategies. In Ecology; Oxford University Press: Oxford, UK, 2019. [Google Scholar] [CrossRef] [Scilit]
- Janssen, M.A.; Walker, B.H.; Langridge, J.; Abel, N. An adaptive agent model for analysing co-evolution of management and policies in a complex rangeland system. Ecol. Model. 2000, 131, 249–268. [Google Scholar] [CrossRef] [Scilit]
- Bai, H.; Chen, Y.; Bai, H.; Liu, M.; Fan, Y. Energy optimization and efficiency improvement model for enterprise production process based on deep learning under the background of carbon peak and carbon neutrality. Int. J. Comput. Intell. Syst. 2025, 18, 169. [Google Scholar] [CrossRef] [Scilit]
- Tian, Y.; Zhu, Z.; Zhao, X.; Chen, X.; Huang, W.; Zhang, X. A Dynamic and Heterogeneous Representation for Topology Optimization Using Evolutionary Algorithms [Research Frontier]. IEEE Comput. Intell. Mag. 2025, 20, 71–82. [Google Scholar] [CrossRef] [Scilit]
- Cheng, L.; Li, M.; Tan, C.; Huang, P.; Zhang, M.; Sun, R. Computational Game-Theoretic Models for Adaptive Urban Energy Systems: A Comprehensive Review of Algorithms, Strategies, and Engineering Applications. Arch. Comput. Methods Eng. 2026, 33, 2037–2114. [Google Scholar] [CrossRef] [Scilit]
- Wang, K.; Cheng, L.; Yin, M.; Zhang, K.; Wang, R.; Zhang, M.; Sun, R. Evolutionary Game Theory in Energy Storage Systems: A Systematic Review of Collaborative Decision-Making, Operational Strategies, and Coordination Mechanisms for Renewable Energy Integration. Sustainability 2025, 17, 7400. [Google Scholar] [CrossRef] [Scilit]
- Cheng, L.; Wei, X.; Li, M.; Tan, C.; Yin, M.; Shen, T.; Zou, T. Integrating Evolutionary Game-Theoretical Methods and Deep Reinforcement Learning for Adaptive Strategy Optimization in User-Side Electricity Markets: A Comprehensive Review. Mathematics 2024, 12, 3241. [Google Scholar] [CrossRef] [Scilit]
- Durlauf, S.N.; Blume, L.E. Learning and Evolution in Games: ESS. In Game Theory; Palgrave Macmillan: London, UK, 2010; pp. 199–206. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mabrok, M. Passivity Analysis of Replicator Dynamics and its Variations (Version 1). arXiv 2018. [Google Scholar] [CrossRef] [Scilit]
- Simon, H.A. Rationality, Bounded. In The New Palgrave Dictionary of Economics; Palgrave Macmillan: London, UK, 2008; pp. 1–4. [Google Scholar] [CrossRef]
- Gunarathne, A.N.; Lee, K.H. Environmental and managerial information for cleaner production strategies: An environmental management development perspective. J. Clean. Prod. 2019, 237, 117849. [Google Scholar] [CrossRef] [Scilit]
- Newman, M.E.J.; Watts, D.J. Scaling and percolation in the small-world network model. Phys. Rev. E 1999, 60, 7332–7342. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Barabási, A.-L. Scale-Free Networks: A Decade and Beyond. Science 2009, 325, 412–413. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Johnston, J.; Andersen, T. Random Processes with High Variance Produce Scale Free Networks. arXiv 2022. [Google Scholar] [CrossRef] [Scilit]
- Rodrigues, F.A. Network centrality: An introduction (Version 1). arXiv 2019. [Google Scholar] [CrossRef] [Scilit]
- Gentner, M.; Heinrich, I.; Jäger, S.; Rautenbach, D. Large Values of the Clustering Coefficient (Version 1). arXiv 2016. [Google Scholar] [CrossRef] [Scilit]
- Furutani, S.; Shibahara, T.; Akiyama, M.; Aida, M. Analysis of Homophily Effects on Information Diffusion on Social Networks. IEEE Access 2023, 11, 79974–79983. [Google Scholar] [CrossRef] [Scilit]
- Tozlu, B.; Akgunduz, A.; Zeng, Y. Unbiased criteria identification for two-sided matching: An environment-based design approach. Expert Syst. Appl. 2025, 277, 127233. [Google Scholar] [CrossRef] [Scilit]
- Ruini, A.; Sporchia, F.; Niccolucci, V.; Pulselli, F.M.; Bastianoni, S. Rethinking environmental benefit allocation in industrial symbiosis. Sci. Total Environ. 2025, 992, 179932. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- De Poorter, E.; Latré, B.; Moerman, I.; Demeester, P. Symbiotic Networks: Towards a New Level of Cooperation Between Wireless Networks. Wirel. Pers. Commun. 2008, 45, 479–495. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Zhang, Q.; Zhang, G.; Wang, D.; Liu, C. Can industrial symbiosis policies be effective? Evidence from the nationwide industrial symbiosis system in China. J. Environ. Manag. 2023, 331, 117346. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Das, P.; Mathur, J.; Bhakar, R.; Kanudia, A. Implications of short-term renewable energy resource intermittency in long-term power system planning. Energy Strategy Rev. 2018, 22, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Lv, Y.; Qin, R.; Sun, H.; Guo, Z.; Fang, F.; Niu, Y. Research on energy storage allocation strategy considering smoothing the fluctuation of renewable energy. Front. Energy Res. 2023, 11, 1094970. [Google Scholar] [CrossRef] [Scilit]
- Jiang, T.; Ju, P.; Lin, Z.; Wu, Q.; Lin, Z.; Chung, C.Y. Competitive Incentive Mechanism for Multi-Agents in Demand Response via a Hierarchical Game Considering Joint Uncertainties. IEEE Trans. Smart Grid 2025, 16, 2913–2925. [Google Scholar] [CrossRef] [Scilit]
- Zitzler, E.; Thiele, L. Multiobjective evolutionary algorithms: A comparative case study and the strength Pareto approach. IEEE Trans. Evol. Comput. 1999, 3, 257–271. [Google Scholar] [CrossRef] [Scilit]
- Mallipeddi, R.; Suganthan, P.N. Ensemble of Constraint Handling Techniques. IEEE Trans. Evol. Comput. 2010, 14, 561–579. [Google Scholar] [CrossRef] [Scilit]
- Qi, X.; Han, Y. Research on the evolutionary strategy of carbon market under “dual carbon” goal: From the perspective of dynamic quota allocation. Energy 2023, 274, 127265. [Google Scholar] [CrossRef] [Scilit]
- Cheng, Z.; Yu, X. China’s carbon emissions trading system and energy directed technical change. Environ. Impact Assess. Rev. 2024, 105, 107417. [Google Scholar] [CrossRef] [Scilit]
- Liang, Z.; Mu, L. Multi-agent low-carbon optimal dispatch of regional integrated energy system based on mixed game theory. Energy 2024, 295, 130953. [Google Scholar] [CrossRef] [Scilit]
- Cheng, L.; Zhang, M.; Wang, K.; Yuan, M.; Liu, Z.; Wang, J.; Zhang, K.; Huang, P. Evolutionary smart contracts for virtual power plant trading: Integrating prospect theory and multi-stage negotiation in cross-regional energy markets. Int. J. Electr. Power Energy Syst. 2025, 173, 111453. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Zhang, C.; Liu, Y.; Liang, X.; Yang, C.; Gui, W. Advancing Industrial Process Control with Deep Learning-Enhanced Model Predictive Control for Nonlinear Time-Delay Systems. IEEE Trans. Ind. Inform. 2025, 21, 6823–6833. [Google Scholar] [CrossRef] [Scilit]
- Wang, K.; Zhong, H. Simulation of Evolutionary Game Decision Model Based on Reinforcement Learning Algorithm. In Proceedings of the 2023 International Conference on Telecommunications, Electronics and Informatics (ICTEI), Lisbon, Portugal, 11–13 September 2023; pp. 550–555. [Google Scholar] [CrossRef] [Scilit]
- Rafi, T.H.; Noor, F.A.; Hussain, T.; Chae, D.-K. Fairness and privacy preserving in federated learning: A survey. Inf. Fusion 2024, 105, 102198. [Google Scholar] [CrossRef] [Scilit]
- Xu, Y.; Qiu, X.; Zhang, F.; Hao, J. Crowdsourced Federated Learning Architecture with Personalized Privacy Preservation. Intell. Converg. Netw. 2024, 5, 192–206. [Google Scholar] [CrossRef] [Scilit]
- Shen, J.; Zhao, Y.; Huang, S.; Ren, Y. Secure and flexible privacy-preserving federated learning based on multi-key fully homomorphic encryption. Electronics 2024, 13, 4478. [Google Scholar] [CrossRef] [Scilit]
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial networks. Commun. ACM 2020, 63, 139–144. [Google Scholar] [CrossRef] [Scilit]
- Franci, B.; Grammatico, S. A game–theoretic approach for generative adversarial networks. In Proceedings of the 2020 59th IEEE Conference on Decision and Control (CDC), Jeju, Republic of Korea, 14–18 December 2020; IEEE: New York, NY, USA, 2020; pp. 1646–1651. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Bollig, B.; Függer, M.; Nowak, T.; Dréau, V.L. Centralized Permutation Equivariant Policy for Cooperative Multi-Agent Reinforcement Learning (Version 1). arXiv 2025. [Google Scholar] [CrossRef] [Scilit]
- Cao, X.; Başar, T.; Diggavi, S.; Eldar, Y.C.; Letaief, K.B.; Poor, H.V.; Zhang, J. Communication-efficient distributed learning: An overview. IEEE J. Sel. Areas Commun. 2023, 41, 851–873. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Qiu, T.; Pu, Z.; Yi, J. A cooperation and decision-making framework in dynamic confrontation for multi-agent systems. Comput. Electr. Eng. 2024, 118, 109300. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Tian, X. On Blockchain We Cooperate: An Evolutionary Game Perspective (Version 3). arXiv 2022. [Google Scholar] [CrossRef] [Scilit]
- Taherdoost, H. Smart contracts in blockchain technology: A critical review. Information 2023, 14, 117. [Google Scholar] [CrossRef] [Scilit]
- Golinucci, N.; Tonini, F.; Rocco, M.V.; Colombo, E. Towards BitCO2, an individual consumption-based carbon emission reduction mechanism. Energy Policy 2023, 183, 113851. [Google Scholar] [CrossRef] [Scilit]
- Xiong, H.; Huang, S. A study on a blockchain-based waste classification management model and the effect evaluation of the model based on the entropy matter-element method. J. Mater. Cycles Waste Manag. 2023, 26, 222–238. [Google Scholar] [CrossRef] [Scilit]
- AlKhader, W.; Musamih, A.; Salah, K.; Jayaraman, R.; Omar, M. Promoting sustainability with blockchain incentives in timber-based construction. Environ. Dev. Sustain. 2025, 1–35. [Google Scholar] [CrossRef] [Scilit]
- Yamashiro, H.; Omote, K.; Imakura, A.; Sakurai, T. Toward the Application of Differential Privacy to Data Collaboration. IEEE Access 2024, 12, 63292–63301. [Google Scholar] [CrossRef] [Scilit]
- Majeed, A.; Lee, S. Anonymization Techniques for Privacy Preserving Data Publishing: A Comprehensive Survey. IEEE Access 2021, 9, 8512–8545. [Google Scholar] [CrossRef] [Scilit]
- Hamza, R. Homomorphic encryption for ai-based applications: Challenges and opportunities. In Proceedings of the 2023 15th International Conference on Knowledge and Systems Engineering (KSE), Hanoi, Vietnam, 18–20 October 2023; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Zhao, C.; Zhao, S.; Zhao, M.; Chen, Z.; Gao, C.Z.; Li, H.; Tan, Y.A. Secure multi-party computation: Theory, practice and applications. Inf. Sci. 2019, 476, 357–372. [Google Scholar] [CrossRef] [Scilit]
- Dong, J.; Jiang, Y.; Liu, D.; Dou, X.; Liu, Y.; Peng, S. Promoting dynamic pricing implementation considering policy incentives and electricity retailers’ behaviors: An evolutionary game model based on prospect theory. Energy Policy 2022, 167, 113059. [Google Scholar] [CrossRef] [Scilit]
- Sachdev, R.S.; Singh, O. Consumer’s demand response to dynamic pricing of electricity in a smart grid. In Proceedings of the 2016 International Conference on Control, Computing, Communication and Materials (ICCCCM), Allahbad, India, 21–22 October 2016; IEEE: New York, NY, USA, 2016; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Yang, S.; Gao, H.O.; You, F. Demand flexibility and cost-saving potentials via smart building energy management: Opportunities in residential space heating across the US. Adv. Appl. Energy 2024, 14, 100171. [Google Scholar] [CrossRef] [Scilit]
- Taştan, M. IoT-Based Smart Energy Management and Load Shifting for Residential Consumption Optimization. Celal Bayar Üniversitesi Fen Bilim. Derg. 2025, 21, 166–183. [Google Scholar] [CrossRef] [Scilit]
- Vahid-Ghavidel, M.; Javadi, M.S.; Santos, S.F.; Gough, M.; Shafie-khah, M.; Catalão, J.P.S. Energy storage system impact on the operation of a demand response aggregator. J. Energy Storage 2023, 64, 107222. [Google Scholar] [CrossRef] [Scilit]
- Pinyo, A.; Bangviwat, A. Smart Contracts-Based Demand Response Bidding Mechanism to Enhance the Load Aggregator Model in Thailand. Energies 2023, 16, 3606. [Google Scholar] [CrossRef] [Scilit]
- Mahmoudi, N.; Saha, T.K.; Eghbal, M. Modelling demand response aggregator behavior in wind power offering strategies. Appl. Energy 2014, 133, 347–355. [Google Scholar] [CrossRef] [Scilit]
- Ruggiero, S.; Kangas, H.-L.; Annala, S.; Lazarevic, D. Business model innovation in demand response firms: Beyond the niche-regime dichotomy. Environ. Innov. Soc. Transit. 2021, 39, 1–17. [Google Scholar] [CrossRef] [Scilit]
- Afentoulis, K.D.; Vagropoulos, S.I. Are current demand response baseline designs suitable for electric vehicles? Policy insights from the independent aggregation business model. Appl. Energy 2025, 396, 126281. [Google Scholar] [CrossRef] [Scilit]
- Abapour, S.; Mohammadi-Ivatloo, B.; Tarafdar Hagh, M. Robust bidding strategy for demand response aggregators in electricity market based on game theory. J. Clean. Prod. 2020, 243, 118393. [Google Scholar] [CrossRef] [Scilit]
- Cheng, L.; Huang, P.; Zhang, M.; Wang, K.; Zhang, K.; Zou, T.; Lu, W. Optimizing virtual power plants cooperation via evolutionary game theory: The role of reward–punishment mechanisms. Mathematics 2025, 13, 2428. [Google Scholar] [CrossRef] [Scilit]
- Fang, Y.; Xiong, B.; Qin, K.; Li, Y.; Tang, J.; Wang, Z. Optimal Home Energy Management with Distributed Generation and Energy Storage Systems. In Proceedings of the 2021 31st Australasian Universities Power Engineering Conference (AUPEC), Perth, Australia, 26–30 September 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Kuo, W.-C.; Chen, C.-H.; Wang, C.-C.; Chang, Y.-C. Application of Mobile Energy Storage System in Micro-Grid Management System. In Proceedings of the 2021 International Conference on Electronic Communications, Internet of Things and Big Data (ICEIB), Jiaoxi, Taiwan, 10–12 December 2021; pp. 314–317. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Wu, J.; Zhou, Y.; Cheng, M.; Long, C. Peer-to-Peer energy trading in a Microgrid. Appl. Energy 2018, 220, 1–12. [Google Scholar] [CrossRef] [Scilit]
- Tushar, W.; Yuen, C.; Mohsenian-Rad, H.; Saha, T.; Poor, H.V.; Wood, K.L. Transforming Energy Networks via Peer-to-Peer Energy Trading: The Potential of Game-Theoretic Approaches. IEEE Signal Process. Mag. 2018, 35, 90–111. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Yao, E.; Pan, L. Electric vehicle drivers’ charging behavior analysis considering heterogeneity and satisfaction. J. Clean. Prod. 2021, 286, 124982. [Google Scholar] [CrossRef] [Scilit]
- Jonas, T.; Macht, G.A. Analyzing the urban-rural divide: Understanding geographic variations in charging behavior for a user-centered EVSE infrastructure. J. Transp. Geogr. 2024, 116, 103859. [Google Scholar] [CrossRef] [Scilit]
- Almaghrebi, A.; James, K.; Al Juheshi, F.; Alahmad, M. Insights into Household Electric Vehicle Charging Behavior: Analysis and Predictive Modeling. Energies 2024, 17, 925. [Google Scholar] [CrossRef] [Scilit]
- Guo, Z.; Bian, H.; Zhou, C.; Ren, Q.; Gao, Y. An electric vehicle charging load prediction model for different functional areas based on multithreaded acceleration. J. Energy Storage 2023, 73, 108921. [Google Scholar] [CrossRef] [Scilit]
- Cai, H.; Li, B.; Li, W.; Wang, J. Heterogeneity in electric taxi charging behavior: Association with travel service characteristics. Travel Behav. Soc. 2025, 38, 100917. [Google Scholar] [CrossRef] [Scilit]
- Sprei, F.; Kempton, W. Mental models guide electric vehicle charging. Energy 2024, 292, 130430. [Google Scholar] [CrossRef] [Scilit]
- Shao, Q.; Lyu, Y.; Cao, J. Evolutionary Dynamics and Policy Coordination in the Vehicle–Grid Interaction Market: A Tripartite Evolutionary Game Analysis. Mathematics 2025, 13, 2356. [Google Scholar] [CrossRef] [Scilit]
- Deng, Z.; Yang, Z. Theoretic and Empirical Analysis of Fiscal and Tax Policies Inducing Enterprises’ Technological Innovation. Financ. Trade Econ. 2011, 5, 5–10. [Google Scholar] [CrossRef]
- Wei, S. A sequential game analysis on carbon tax policy choices in open economies: From the perspective of carbon emission responsibilities. J. Clean. Prod. 2021, 283, 124588. [Google Scholar] [CrossRef] [Scilit]
- Que, W.; Zhang, Y.; Liu, S.; Yang, C. The spatial effect of fiscal decentralization and factor market segmentation on environmental pollution. J. Clean. Prod. 2018, 184, 402–413. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Chang, D.; Wang, X. Does the regional atmospheric quality punishment incentive mechanism (AQPI) promote environmental regulation? Subordinate government as an agent of superior environmental policies. J. Clean. Prod. 2023, 414, 137718. [Google Scholar] [CrossRef] [Scilit]
- Takashima, N. International environmental agreements between asymmetric countries: A repeated game analysis. Jpn. World Econ. 2018, 48, 38–44. [Google Scholar] [CrossRef] [Scilit]
- Zhao, L.; Chong, K.M.; Gooi, L.-M.; Yan, L. Research on the impact of government fiscal subsidies and tax incentive mechanism on the output of green patents in enterprises. Financ. Res. Lett. 2024, 61, 104997. [Google Scholar] [CrossRef] [Scilit]
- Tan, R.; Zhu, W.; Xu, M.; Zhang, Z. From voluntary to mandatory implementation: The impact of green credit policy on de-zombification in China. Energy Econ. 2025, 141, 108045. [Google Scholar] [CrossRef] [Scilit]
- Han, X.; Cai, Q. Environmental regulation, green credit, and corporate environmental investment. Innov. Green Dev. 2024, 3, 100135. [Google Scholar] [CrossRef] [Scilit]
- Dong, H.; Zhang, L.; Zheng, H. Green bonds: Fueling green innovation or just a fad? Energy Econ. 2024, 135, 107660. [Google Scholar] [CrossRef] [Scilit]
- Lei, G.; Zhong, C.; Zhang, J.; Zheng, Y. Environmental reward–punishment policy and collaborative green innovation in China: Lessons for global environmental management. J. Environ. Manag. 2025, 395, 127681. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, S.; Liu, C.; Zhou, Z. Government-enterprise green collaborative governance and urban carbon emission reduction: Empirical evidence from green PPP programs. Environ. Res. 2024, 257, 119335. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kumar, P.; Date, A.; Shabani, B. Techno-economic analysis of an integrated desalination-renewable-hydrogen system for zero-emission freshwater and electricity production. Energy Convers. Manag. 2026, 353, 121231. [Google Scholar] [CrossRef] [Scilit]












| Prior Review | Scope Covered | Treatment of AI Integration | Methodological Status | Not Covered/Left Open |
|---|---|---|---|---|
| Vieira and Amaral [5] | Barriers to and strategies for cleaner production adoption | None; predates the AI-integration literature | Systematic; explicit protocol | No game-theoretic dynamics; no energy-system or carbon-market layer |
| Cheng et al. [13] | Game-theoretic models for adaptive urban energy systems; algorithms and engineering applications | Algorithmic emphasis; learning treated as a solution method | Narrative, broad coverage | Urban energy focus; industrial symbiosis and carbon-market governance outside scope |
| Wang et al. [14] | EGT in energy storage systems; collaborative decision-making and coordination for renewable integration | Limited; coordination mechanisms rather than learning architectures | Systematic; storage-specific | Single technology domain; no cross-scale comparison, no policy-instrument analysis |
| Cheng et al. [15] | EGT combined with deep reinforcement learning in user-side electricity markets | Central; EGT–DRL coupling examined in depth | Narrative; market-side focus | One market layer only; no enterprise-scale symbiosis, no carbon-market mechanism design |
| This review | Three scales in one frame: enterprise symbiosis, system-level smart energy, market-level carbon governance | Five functional roles of AI on the evolutionary object, including federated and blockchain execution | Systematic; protocol and selection flow reported in Section 1.4 | Contribution: cross-scale mechanism invariance, privacy-preserving payoff estimation, and instrument-to-dynamics mapping |
| Dimension | Included If | Excluded If, and Why |
|---|---|---|
| Method | An evolutionary game formulation is constructed, solved, or applied; or a foundational contribution to that formalism | Classical one-shot or purely cooperative game models with no population dynamics—outside the dynamic selection mechanism this review examines |
| Domain | Cleaner production, industrial symbiosis, smart grid and demand response, renewable coordination, electric vehicles, or carbon market and environmental policy | Applications in unrelated domains such as epidemiology or general social evolution—payoff structures are not transferable to production and energy settings |
| AI integration | Present (retained and tagged AI-integrated) or absent (retained as the comparison strand) | Not an exclusion criterion; used only to partition the corpus, since the review’s central claim requires both strands |
| Publication type | Peer-reviewed journal articles and archival conference papers | Editorials, abstracts, theses, preprints without peer review, agency and grey literature—methodological detail insufficient for the critical appraisal undertaken in this review |
| Reporting quality | Model assumptions, payoff structure, and solution method are recoverable from the text | Results reported without a recoverable model—the study cannot be placed on the modeling dimensions against which the retained corpus is compared |
| Language and window | English; 1973–2025 | Other languages, and records outside the window—a stated limit of the protocol, acknowledged in Section 1.4 and again in Section 8.1 |
| Ref. | Year of Publication | Research Field | Core Content | Key Methods/Models | Main Conclusions |
|---|---|---|---|---|---|
| [20] | 1999 | Small-World Network | Modified Watts–Strogatz model, studied scaling properties and site percolation | Modified small-world model, numerical simulations, series expansion, Pade’ approximants | Length scale governs distance scaling and diverges with decreasing shortcut density; effective dimension is scale-dependent; percolation properties close to random graphs |
| [21] | 2009 | Network Science, Scale-Free Networks | Review scale-free networks’ universality, formation mechanisms, and impacts | Growth-Preferential Attachment Model, Network Topology Analysis | Real networks converge to scale-free architectures; topology shapes dynamics like epidemic spread |
| [22] | 2022 | Network Science, Scale-Free Networks | Propose alternative model for scale-free networks without preferential attachment | Randomly Stopped Linking, Generalized Central Limit Theorem | High variance in linking parameters drives power-law distributions; preferential attachment is not mandatory |
| [23] | 2019 | Network Science, Centrality Measures | Review main centrality metrics and their impacts on dynamical processes | Degree, k-Core, Betweenness, Eigenvector Centrality | Centrality choice depends on network type; influences epidemic spreading and synchronization |
| [24] | 2016 | Graph Theory, Clustering Coefficient | Determine maximum clustering coefficient for connected regular/subcubic graphs | Graph Construction (G(k, ℓ)), Extremal Graph Analysis | Characterize extremal graphs; adding a single edge can drastically increase clustering coefficient |
| [25] | 2023 | Social Networks, Information Diffusion | Analyze homophily’s impact on information diffusion in modular networks | Generalized Contagion Model (GCM), Message-Passing Approach | Homophily facilitates local diffusion but inhibits global diffusion in strongly modular networks |
| Network Class | Structural Signature | Consequence for Strategy Evolution | Applicable Cleaner- Production Setting | Anchors |
|---|---|---|---|---|
| Watts–Strogatz small-world | High clustering with short average path | Cooperative clusters consolidate locally and reach the whole network quickly; diffusion fast and comparatively even | Technology diffusion among geographically or supply-chain proximate enterprises | [20]; Section 2.3 |
| Scale-free | Few hubs, many low-degree nodes; power laws need not arise from preferential attachment | Diffusion routed through hubs; robust to random failure, fragile to hub loss; hub strategies weigh disproportionately in imitation | Markets and platforms with dominant firms; anchor-enterprise targeting | [21,22]; Section 3.1 and Section 7.1 |
| Modular with homophily | Dense within-module ties, sparse bridges | Fast diffusion inside modules, inhibited between them; risk of local lock-in on distinct conventions | Eco-industrial parks and regional clusters; cross-park cooperation | [25]; Section 6.2 |
| Centrality-heterogeneous | Node importance ranking depends on the index chosen | Predicted diffusion path and best intervention target change with the centrality measure | Selecting which enterprise a subsidy or audit should reach first | [23]; Section 7.1 |
| High-clustering graph families | Clustering statistic sensitive to single edges | Calibrated local closure can misstate the ease of cooperation; sensitivity analysis required | Calibrating park network models from sparse relational data | [24] |
| Functional Role | What It Does to the Evolutionary Game | Acts at | AI Techniques/ Refs in This Review |
|---|---|---|---|
| Bounded-rationality modeling | Replaces the assumed behavioral rule with an adaptive rule learned from interaction, so the selection pressure driving the dynamics is inferred rather than fixed | Replicator Equation (4); user-group replicator Equation (17) | Reinforcement learning/Q-learning [40]; DQN-EGT coupling [15] |
| Replicator acceleration | Solves high-dimensional and coupled replicator systems that fixed-step derivation handles poorly, via approximation and adaptive step control | Coupled-population dynamics, Algorithm 1; multi-agent training [46,47] | Deep learning function approximation [15,39]; centralized-training/distributed-execution MARL [46] |
| Strategy forecasting | Conditions an agent’s play on predicted opponent behavior, so response anticipates rather than only reacts | Game step within Algorithm 1; DRL learning rate, Section 7.2 | Deep reinforcement learning agents [15,38,40]; heterogeneous-agent coordination [48] |
| Payoff approximation | Fits the fitness surface where an analytic payoff matrix cannot express severe heterogeneity or nonlinearity | Payoff surrogate in Algorithm 3; deep nonlinear modeling | Deep neural networks [39]; GAN-based equilibrium approximation [44,45] |
| Decentralized execution | Enforces rules, settles outcomes, and distributes incentives without a trusted central coordinator | Carbon-market co-evolution, Algorithm 2; blockchain game mechanism | Smart contracts/consensus [49,50]; token-based settlement [51] |
| Ref. | Year of Publication | Research Field | Core Content | Key Methods/Models | Main Conclusions |
|---|---|---|---|---|---|
| [15] | 2024 | User-Side Electricity Market | Integrate EGT with DQN for adaptive strategy optimization | EGT, Deep Q-Network (DQN) | EGT ensures long-term stability; DQN enhances real-time adaptability to supply-demand fluctuations |
| [39] | 2025 | Industrial Process Control | Propose DNNs-MPC for nonlinear time-delay systems | DNNs-MPC, TCM Network, Adaptive Gradient Descent | Enhances control performance, stability and response accuracy; outperforms traditional methods |
| [40] | 2023 | Evolutionary Game, Reinforcement Learning | Combine Q-Learning with game model for dynamic decision-making | Q-Learning, Prisoner’s Dilemma Model | Multi-agent model converges faster; distinguishes enterprise reputation via income |
| [41] | 2024 | Federated Learning (FL) | Survey privacy and fairness in FL, and their trade-offs | Differential Privacy (DP), Homomorphic Encryption (HE), Secure Multiparty Computation (SMC) | Need balanced privacy, fairness and accuracy; existing studies lack integrated solutions |
| [42] | 2024 | Crowdsourced Federated Learning | Propose personalized privacy preservation architecture | Two-stage Stackelberg Game, Weight Priority Perturbation (PMWP) | Achieves better model performance with same privacy budget; reaches unique Nash equilibrium |
| [43] | 2024 | Privacy-Preserving Federated Learning | Propose mMFHE for FL privacy protection | Multi-Key Fully Homomorphic Encryption (mMFHE) | Resists N-1 user-server collusion; supports homomorphic addition/multiplication |
| [44] | 2014 | Generative Modeling | Propose Generative Adversarial Networks (GANs) | Generator-Discriminator Game, Minimax/Non-Saturating GAN | Generates high-quality realistic samples; relies on game theory for unsupervised learning |
| [45] | 2020 | Generative Adversarial Networks (GANs) | Propose SRFB algorithm for GAN training via stochastic Nash equilibrium | Stochastic Relaxed Forward-Backward (SRFB), Variational Inequality | Converges to exact solution (large samples) or its neighborhood (finite samples); low computational cost |
| Dimension | AI-Enhanced Evolutionary Game (Section 4.1) | Multi-Agent Reinforcement Learning (Section 4.2) |
|---|---|---|
| Object of analysis | Population strategy shares and their stability | Joint policy of interacting learners |
| What carries the dynamics | Evolutionary update (replication, imitation, best response); learning estimates its inputs | Learned value functions and policy gradients are the dynamics |
| Solution concept | Evolutionarily stable strategy; asymptotically stable states of replicator systems | Equilibrium (in practice, approximate) of a stochastic game |
| Status of parameters | Behavioral: imitation intensity, payoff coefficients; interpretable term by term | Algorithmic: discount factor, exploration schedule, network capacity; no direct behavioral reading |
| Role of data | Calibration of payoffs and heterogeneity, federated where records are fragmented | Experience for policy improvement through interaction |
| Representative anchors in this review | [38,39,40,41,42,43,44,45]; Algorithm 3 | [46,47,48] |
| Where deployed | Constructions of Section 3, Section 5 and Section 6 | Application settings surveyed in this subsection |
| Ref. | Year of Publication | Research Field | Core Content | Key Methods/Models | Main Conclusions |
|---|---|---|---|---|---|
| [49] | 2023 | Blockchain Consensus | Model blockchain consensus as evolutionary game with bounded rationality | Evolutionary Stable Strategy (ESS), Imitative Learning, Assortative Matching | Three stable equilibria; honest equilibrium optimal for safety/liveness/welfare |
| [50] | 2023 | Blockchain Smart Contracts | Critical review of smart contracts (2012–2022) | Literature Review (252 papers) | Highlights applications (healthcare/supply chain); identifies challenges (security/scalability); suggests AI/data science integration |
| [51] | 2023 | Carbon Emission Reduction | Propose BitCO2 mechanism to incentivize BEV adoption for emission reduction | System Dynamics, Life Cycle Assessment (LCA) | Cumulative 973 ktonCO2eq reduction over 20 years; boosts BEV registrations |
| [52] | 2024 | Waste Classification Management | Build blockchain-based waste classification model with software/hardware integration | Blockchain, Entropy Matter-Element Method, RFID | Participation rate rises to 95.9%; overflow rate drops 41.67%; high violation traceability (82.4%) |
| [53] | 2025 | Timber Construction Circular Supply Chain | Develop blockchain tokenization framework to incentivize circular economy practices | Solidity Smart Contracts, TOPSIS, ERC-20 Token | Technically feasible; ensures transparency/accountability; aligns with SDGs |
| [54] | 2024 | Privacy-Preserving Data Collaboration | Apply Differential Privacy to Data Collaboration via PCA dimension reduction | Differential Privacy (DP), PCA, Gaussian Mechanism | DP Data Collaboration performs comparably to DP Federated Learning; minimal utility loss |
| [55] | 2021 | Privacy-Preserving Data Publishing | Survey anonymization techniques for tabular and social network data | k-anonymity, ℓ-diversity, t-closeness, Differential Privacy (DP) | Provides systematic coverage of PPDP techniques; identifies challenges and future research directions |
| [56] | 2023 | Homomorphic Encryption (HE) for AI | Survey HE’s application in AI, software engineering aspects, and library comparisons | Fully HE (FHE), Somewhat HE (SWHE), Multi-key HE (MKHE) | HE enables privacy-preserving AI but faces challenges in performance, noise management, and scalability |
| [57] | 2019 | Secure Multi-Party Computation (SMPC) | Survey SMPC theory, cloud-assisted protocols, and application-oriented solutions | Garbled Circuits, Oblivious Transfer, Secret Sharing, Homomorphic Encryption | SMPC is mature in theory; cloud assistance and application-specific protocols improve practicality |
| Ref. | Year of Publication | Research Field | Core Content | Key Methods/Models | Main Conclusions |
|---|---|---|---|---|---|
| [58] | 2022 | Dynamic Electricity Pricing | Establish tripartite evolutionary game model for dynamic pricing promotion | Evolutionary Game, Prospect Theory, CVaR, Price Elasticity Matrix | Regulator subsidies, retailer promotion costs, and consumer risk aversion influence pricing adoption |
| [59] | 2016 | Smart Grid, Dynamic Pricing | Propose real-time pricing algorithm for consumer utility and grid efficiency | Quadratic Utility Function, Lagrange Optimization, MATLAB | RTP maximizes consumer satisfaction; cuts peak load vs. fixed pricing |
| [60] | 2024 | Residential Heating, Demand Flexibility | Assess PCM thermal storage + smart control for DR across US metros | MPC, Active/Passive PCM, 5 min Simulation | 98.5% peak load shifting, 338.3% cost reduction; viable in 50% of areas |
| [62] | 2023 | DR, Energy Storage | Analyze ESS impact on DR aggregator in short-term markets | Robust Optimization, ESS, Rooftop PV, TOU/Reward-based DR | ESS boosts profit (20% at Γ = 12) and flexibility; larger ESS benefits worst-case scenarios |
| [63] | 2023 | DR, Smart Grid, Blockchain | Propose smart contract bidding mechanism for Thailand’s LA model | Smart Contracts (BSC), Reverse Auction, Guaranteed Fund System | Enhances transparency; 60% of expected compensation as guaranteed fund suffices |
| [64] | 2014 | Wind Power, DR | Integrate DR into wind power offering via aggregator collaboration | Bilevel Programming, MPEC, CVaR, Stochastic Programming | Risk-neutral producers favor day-ahead market; DR reduces uncertainty |
| [65] | 2021 | BMI, DR | Explore BMI drivers/behaviors of Finnish DR firms | Semi-structured Interviews, Morphological Box Model | BMI drivers vary by firm type; hybrid niche/incumbent behaviors; multi-sector interaction matters |
| [66] | 2025 | EV Aggregation, DR Baselines | Evaluate DR baselines for independent EV aggregators | Mixed-Integer Linear Programming, 4 Baseline Designs | Some baselines prone to manipulation; dynamic pricing boosts savings |
| [67] | 2020 | DR Aggregator Bidding | Propose robust bidding strategy via game theory | Game Theory (Nash Equilibrium), Robust Optimization | Medium bidding (k = 1.2) is Nash equilibrium; RO mitigates price uncertainty |
| Ref. | Year of Publication | Research Field | Core Content | Key Methods/Models | Main Conclusions |
|---|---|---|---|---|---|
| [73] | 2021 | EV Charging Behavior | Analyze EV drivers’ charging choice with satisfaction and heterogeneity | Binary Logit Model, Latent Class Model | Two user classes (service/pragmatic concerned); satisfaction and SOC are key factors |
| [74] | 2024 | EV Charging, Urban-Rural Divide | Investigate geographic/urbanity differences in public Level 2 charging | Mann–Whitney U Test, Kruskal–Wallis Test | Clear urban-rural charging differences; weekly repetitive patterns exist |
| [75] | 2024 | Household EV Charging Prediction | Predict charging session parameters (duration, demand, next session time) | Random Forest, XGBoost, Artificial Neural Network | RF performs best (R2 0.40–0.48); time-of-day and historical data are key |
| [76] | 2023 | EV Charging Load Prediction | Propose multithreaded acceleration-based prediction model for functional areas | Improved Floyd Algorithm, Two-stage Charging Power Model | Functional area load differs significantly; four-thread speedup ratio > 2.5 |
| [77] | 2025 | Electric Taxi Charging Behavior | Explore charging behavior heterogeneity and service characteristic association | Covariate-enhanced Latent Profile Analysis, Multinomial Logistic Regression | 5 charging patterns identified; dual-shift/single-shift drivers differ significantly |
| [78] | 2024 | EV Charging Mental Models | Analyze mental models influencing charging strategies (novice vs. experienced users) | In-depth Interviews, Qualitative Analysis | 3 core models; event-triggered model reduces anxiety and improves convenience |
| [79] | 2025 | Vehicle–Grid Interaction (VGI) | Construct tripartite evolutionary game model for EV aggregators, governments, users | Tripartite Evolutionary Game, Replicator Dynamics | VGI evolves through V0G, V1G, V2G; peak-valley price difference and subsidies drive transition |
| Instrument | Model Element It Moves | Evolutionary Reading | Anchor |
|---|---|---|---|
| Fiscal subsidy | Subsidy term of Equations (24) and (25); economic weights of Equations (32) and (33) | Lowers the compliance threshold; strongest where public-good character starves market provision | [85] |
| Tax preference | Effective tax burden in Equations (24) and (25) | Raises the after-tax return to compliance; most binding for front-loaded technologies | — |
| Green credit | Financing constraint on entry into the compliant strategy set | Voluntary form shifts payoffs; mandatory form acts as a forcing constraint on participation | [86,87] |
| Green securities | External-financing constraint on innovation capacity | Expands the reachable strategy set by relieving constraints through the market channel | [88] |
| Reward–punishment policy | Fine and inspection terms of Equation (26) | Raises the expected cost of evasion; stabilizes the compliant state | [89] |
| Green public–private partnership | Risk and benefit-sharing coefficients of the government–private game | Aligns private strategy with policy signals under shared risk | [90] |
| Ref. | Year of Publication | Research Field | Core Content | Key Methods/ Models | Main Conclusions |
|---|---|---|---|---|---|
| [85] | 2024 | Fiscal Subsidies, Tax Incentives Green Patents | Compare impacts of fiscal subsidies and tax incentives on firms’ green patents | Baseline Regression, Heterogeneity Test | Fiscal subsidies have stronger incentive effects; clean energy firms are more sensitive |
| [86] | 2025 | Green Credit Policy Firm De-zombification | Compare voluntary vs. mandatory green credit policy’s impact | DID Model, Heterogeneity Test | Mandatory policy (green credit in bank assessment) effectively promotes de-zombification |
| [87] | 2024 | Environmental Regulation Corporate Environmental Investment | Explore environmental regulation’s impact and green credit’s moderating role | Fixed Effect Model, Robustness Test | Environmental regulation promotes environmental investment; green credit positively moderates |
| [88] | 2024 | Green Bonds Green Innovation | Investigate green bond issuance’s impact on green innovation | Time-varying DID, IV Approach | Green bonds promote green innovation via alleviating financial constraints and increasing R&D investment |
| [89] | 2025 | Environmental Reward-Punishment Policy Collaborative Green Innovation | Analyze policy’s impact on corporate collaborative green innovation | Staggered DID, Mechanism Analysis | Policy enhances high-quality joint green invention patents; works via data disclosure, university-industry collaboration, risk reduction |
| [90] | 2024 | Government-Enterprise Green Collaborative Governance Carbon Emission Reduction | Explore green PPP projects’ role in urban carbon emission reduction | Generalized DID, ChatGPT for Green PPP Identification | Collaborative governance reduces urban carbon emissions via structural, technological, co-investment effects |
| Transaction Cost | σ = 0.00 | σ = 0.07 | σ = 0.14 | σ = 0.21 | σ = 0.29 | σ = 0.36 | σ = 0.43 | σ = 0.50 |
|---|---|---|---|---|---|---|---|---|
| c = 0.10 | 0.15 | 0.12 | 0.10 | 0.08 | 0.06 | 0.05 | 0.04 | 0.03 |
| c = 0.17 | 0.22 | 0.18 | 0.15 | 0.12 | 0.10 | 0.08 | 0.06 | 0.05 |
| c = 0.24 | 0.30 | 0.25 | 0.21 | 0.17 | 0.14 | 0.11 | 0.09 | 0.07 |
| c = 0.31 | 0.38 | 0.32 | 0.27 | 0.22 | 0.18 | 0.15 | 0.12 | 0.10 |
| c = 0.39 | 0.47 | 0.40 | 0.34 | 0.28 | 0.23 | 0.19 | 0.16 | 0.13 |
| c = 0.46 | 0.56 | 0.48 | 0.41 | 0.35 | 0.29 | 0.24 | 0.20 | 0.17 |
| c = 0.53 | 0.65 | 0.57 | 0.49 | 0.42 | 0.35 | 0.30 | 0.25 | 0.21 |
| c = 0.60 | 0.75 | 0.66 | 0.57 | 0.49 | 0.42 | 0.36 | 0.30 | 0.25 |
| Adoption Strategy | Low Subsidy (σ = 0.05) | Medium Subsidy (σ = 0.15) | High Subsidy (σ = 0.30) | Strategy Mean | Improvement (Low → High) |
|---|---|---|---|---|---|
| Central First | 0.83 ± 0.08 | 0.89 ± 0.06 | 0.94 ± 0.04 | 0.887 | +13.3% |
| Peripheral First | 0.81 ± 0.09 | 0.88 ± 0.07 | 0.93 ± 0.05 | 0.873 | +14.8% |
| Random Sequence | 0.82 ± 0.10 | 0.89 ± 0.07 | 0.95 ± 0.04 | 0.887 | +15.9% |
| Anchor First | 0.84 ± 0.07 | 0.90 ± 0.05 | 0.96 ± 0.03 | 0.900 | +14.3% |
| Policy Mean | 0.825 | 0.890 | 0.945 | — | — |
| Effect Size (η2) | — | — | — | — | 0.72 |
| Performance Metric | Baseline | High Subsidy | Low Cost | Anchor Focus | Combined | Best-Baseline Δ |
|---|---|---|---|---|---|---|
| Final Coop. Rate | 0.45 (5) | 0.82 (3) | 0.75 (4) | 0.88 (2) | 0.92 (1) | +0.47 |
| Convergence Time | 0.55 (5) | 0.70 (4) | 0.80 (3) | 0.85 (2) | 0.90 (1) | +0.35 |
| Threshold Value (lower is better) | 0.60 (5) | 0.35 (3) | 0.42 (4) | 0.30 (2) | 0.25 (1) | −0.35 |
| Network Robustness | 0.50 (5) | 0.65 (3) | 0.72 (2) | 0.58 (4) | 0.78 (1) | +0.28 |
| Policy Efficiency | 0.40 (5) | 0.60 (4) | 0.75 (2) | 0.70 (3) | 0.85 (1) | +0.45 |
| Aggregate Score † | 0.46 | 0.68 | 0.72 | 0.74 | 0.84 | +0.38 |
| Rank | 5 | 4 | 3 | 2 | 1 | — |
| Learning Type | Price Signal (ρ) | Final Coordination () | Convergence Time (Iterations) | Stability Index | Oscillation Amplitude | Improvement vs. Baseline | Coordination Efficiency | Social Welfare |
|---|---|---|---|---|---|---|---|---|
| SIMPLE | 0.5 | 0.345 | 85 | 0.82 | 0.018 | +12.3% | 0.68 | 0.72 |
| SIMPLE | 1.0 | 0.425 | 62 | 0.88 | 0.015 | +38.5% | 0.78 | 0.81 |
| SIMPLE | 2.0 | 0.382 | 48 | 0.91 | 0.012 | +24.5% | 0.85 | 0.86 |
| DRL | 0.5 | 0.318 | 78 | 0.75 | 0.025 | +3.6% | 0.62 | 0.68 |
| DRL | 1.0 | 0.358 | 42 | 0.82 | 0.022 | +16.7% | 0.75 | 0.78 |
| DRL | 2.0 | 0.378 | 28 | 0.85 | 0.019 | +23.2% | 0.82 | 0.84 |
| Baseline | — | 0.307 | 120 | 0.70 | 0.035 | — | 0.55 | 0.62 |
| Combined Optimal | 2.0 | 0.425 | 35 | 0.92 | 0.010 | +38.4% | 0.88 | 0.91 |
| Theoretical Max | — | 1.000 | — | 1.00 | 0.000 | +225.7% | 1.00 | 1.00 |
| Market Average | — | 0.365 | 58 | 0.83 | 0.018 | +18.9% | 0.74 | 0.78 |
| Agent Type | Mean Payoff ($/MWh) | Std. Dev. ($/MWh) | Median ($/MWh) | 25th Percentile | 75th Percentile | Max Payoff | Min Payoff | Coefficient of Variation (CV) | Skewness | Coordination Benefit |
|---|---|---|---|---|---|---|---|---|---|---|
| Renewable Suppliers | 28.0 | 12.5 | 26.5 | 18.2 | 38.5 | 52.0 | 8.5 | 0.446 | 0.35 | +45.2% |
| Storage Operators | 14.6 | 5.8 | 13.2 | 10.5 | 18.2 | 28.5 | 5.2 | 0.397 | 0.62 | +38.5% |
| Demand Response | 15.8 | 6.2 | 15.0 | 11.2 | 19.8 | 32.0 | 6.0 | 0.392 | 0.28 | +42.1% |
| Grid Operator | 22.5 | 8.5 | 21.0 | 16.0 | 28.5 | 42.0 | 10.0 | 0.378 | 0.45 | +35.8% |
| System Average | 20.2 | 8.3 | 19.0 | 14.0 | 26.0 | 42.0 | 7.4 | 0.411 | 0.43 | +40.4% |
| High Coordination | 25.8 | 7.2 | 24.5 | 20.2 | 30.5 | 45.0 | 12.0 | 0.279 | 0.25 | +58.2% |
| Low Coordination | 15.2 | 9.8 | 14.0 | 8.5 | 21.0 | 38.0 | 3.5 | 0.645 | 0.55 | +15.6% |
| Baseline (No Coord.) | 12.8 | 11.2 | 11.5 | 5.2 | 18.5 | 42.0 | 1.5 | 0.875 | 0.72 | — |
| Theoretical Maximum | 35.0 | 0.0 | 35.0 | 35.0 | 35.0 | 35.0 | 35.0 | 0.000 | 0.00 | +173.4% |
| Market Equilibrium | 18.5 | 7.5 | 17.5 | 12.5 | 23.5 | 38.0 | 6.5 | 0.405 | 0.38 | +32.5% |
| Market Mechanism | Coordination Level | Convergence Speed | Stability | Equity | Efficiency | Aggregate Score | Rank | Implementation Cost | Robustness Index | Scalability |
|---|---|---|---|---|---|---|---|---|---|---|
| Dynamic Pricing | 0.85 | 0.80 | 0.70 | 0.60 | 0.82 | 0.754 | 3 | Medium | 0.72 | High |
| Capacity Remuneration | 0.72 | 0.65 | 0.85 | 0.78 | 0.70 | 0.740 | 4 | High | 0.85 | Medium |
| Information Provision | 0.78 | 0.75 | 0.80 | 0.85 | 0.75 | 0.786 | 2 | Low | 0.82 | High |
| Gradual Liberalization | 0.82 | 0.70 | 0.88 | 0.72 | 0.78 | 0.780 | 3 | Medium | 0.88 | Medium |
| Combined Approach | 0.92 | 0.88 | 0.85 | 0.80 | 0.90 | 0.870 | 1 | High | 0.90 | High |
| Baseline (No Mechanism) | 0.45 | 0.35 | 0.55 | 0.50 | 0.42 | 0.454 | 6 | None | 0.52 | — |
| Theoretical Optimum | 1.00 | 1.00 | 1.00 | 1.00 | 1.00 | 1.000 | — | — | 1.00 | — |
| Industry Average | 0.68 | 0.62 | 0.72 | 0.65 | 0.68 | 0.670 | — | Medium | 0.70 | Medium |
| Emerging Markets | 0.55 | 0.50 | 0.60 | 0.58 | 0.55 | 0.556 | — | Low | 0.58 | Low |
| Mature Markets | 0.82 | 0.78 | 0.85 | 0.75 | 0.82 | 0.804 | — | High | 0.85 | High |
| Validation Modality | Representative Studies (This Review) | What Is Actually Tested |
|---|---|---|
| Pure numerical simulation (MATLAB/Python) | DQN–EGT user-side market [15]; Q-learning game [40]; real-time pricing in MATLAB [59]; DRA robust bidding [67]; both case studies of Section 7 | Convergence, thresholds, and sensitivity of the modeled dynamics under synthetic or historical scenarios |
| Simulation with empirical/field-data calibration | PCM demand flexibility across US metros [60]; ESS-augmented DR in short-term markets [62]; enterprise deep-learning energy optimization [91] | Model behavior under parameters fitted to measured field or market data |
| Simulation with real-data natural experiment (policy) | Green-credit corporate investment [86]; green-bond innovation effect [88]; reward–punishment collaborative innovation [89] | Policy effect identified from observational data, not a controlled model test |
| Real-world hardware-in-the-loop (HIL) | None identified among the surveyed AI–EGT studies | The validation gap this review reports |
| System Scale | Core AI Technique | EGT Model Type | Objective Function | Key Limitation Identified |
|---|---|---|---|---|
| Enterprise/industrial symbiosis [12,15] | Deep Q-network; deep learning payoff fitting | Multi-population replicator on symbiosis network; bilateral matching | Maximize cooperative surplus/cooperation share subject to subsidy and transaction cost | Homogeneous agent simplification; threshold sensitivity to initial conditions; numerical validation only |
| Microgrid/distributed energy [46,47,48] | Multi-agent RL (centralized-training/distributed-execution) | Heterogeneous-agent coordination game; coupled-population replicator | Coordinated dispatch approximating social optimum under uncertainty | Communication constraints; convergence stable only within bounded learning rate; no HIL test |
| Smart grid demand response [17,40,58] | Q-learning; DRL; dynamic-pricing learning | Tripartite/multi-group user–grid replicator, Equations (14)–(17) | Maximize user and grid payoff/load smoothing via price incentives | Bounded-rationality rule still stylized; user heterogeneity partially captured; privacy of load data |
| Smart grid EV charging [66,67] | MILP-assisted learning; game-theoretic robust bidding | Nash/evolutionary charging-strategy game | Minimize charging cost/baseline manipulation; vehicle–grid interaction | Baseline gaming; sensitivity to price-signal design; scenario-limited validation |
| Carbon market–trading mechanism [35,51] | Blockchain smart contracts; token settlement; price-band control (Algorithm 2) | Emitter buy/invest replicator with regulator price band | Meet emission target at minimum abatement cost; price stability | Regime-switching missed by smooth analysis; data tracking and contract-security risk |
| Carbon market—tax and penalty policy [80,81] | Deep-learning scenario analysis; multi-group learning | Large-firm/SME/government tripartite replicator, Equations (24)–(29) | Government net regulatory benefit; firm compliance under tax, subsidy, penalty | Firm heterogeneity coarse; time-consistency of commitment; calibration data scarce |
| Carbon market—green incentives [85,86,88] | Data-driven policy evaluation; federated calibration (Algorithm 3) | Multi-level multi-objective incentive game, Equations (32) and (33) | Weighted economic–environmental–social objective under budget | Fiscal-burden trade-off; identification from observational data; privacy of firm ledgers |
| Study Group | Behavioral Assumption | Payoff Construction | Equilibrium Concept | Solution Technique | Cost/Validation |
|---|---|---|---|---|---|
| Industrial symbiosis [12] | Bounded rationality; imitation of higher payoff | Linear surplus net of transaction cost and subsidy | Evolutionary stable strategy of a two-population system | Analytic stability plus numerical integration | Low; closed-form thresholds. Validated by simulation only |
| Network games [20,21] | Imitation restricted to graph neighbors | Local payoff aggregated over degree | Stability conditional on topology; no unique ESS | Agent-based simulation on generated graphs | Grows with node count; results sensitive to generative model |
| Demand response [58] and Equations (14)–(17) | Heterogeneous user groups, each boundedly rational | Price-difference revenue less response cost, group-specific | Multi-group replicator fixed point | Coupled ODE integration | Moderate; scales with group count, not user count |
| EGT–DRL coupling [15,40] | Learned policy replaces fixed imitation rule | Approximated from interaction data, not specified ex ante | Convergence to coordinated equilibrium, not ESS in the strict sense | Deep Q-learning; policy gradient | High; training dominates. Stability bounded by learning rate |
| Federated payoff learning [41,42,43] | Bounded rationality plus private information | Surrogate fitted locally, aggregated under noise | Fixed point of the estimated, not the true, dynamics | Federated averaging with differential privacy | High and communication-bound; accuracy traded for privacy |
| Blockchain mechanisms [49,50,51] | Strategic agents under enforceable rules | Payoff realized through contract settlement | Mechanism-induced equilibrium; enforcement assumed exact | Smart-contract execution; consensus-dependent | Consensus overhead; on-chain cost rarely reported |
| Environmental tax [80]; Equations (24)–(29) | Firms and regulator both adaptive | Revenue less abatement cost, tax, penalty, subsidy | Tripartite replicator equilibrium | Jacobian stability analysis | Low; but equilibrium multiplicity often unexamined |
| Gap | Statement | Established in | What Remains Open | Answered by (Section 8.3) |
|---|---|---|---|---|
| G1 | Network-game conclusions are topology-conditioned and do not transfer across structures | Section 2.3, Table 4 | Evolutionary stability results derived from empirically measured symbiosis and market networks, with topology reported as a scope condition | Direction 1 |
| G2 | AI-enhanced models lack interpretability; parameters mix behavioral and algorithmic meanings and are weakly identifiable | Section 4.1 | Surrogates whose parameters retain behavioral readings, with identifiability reported per parameter | Direction 2 |
| G3 | Agent heterogeneity is compressed into two or three homogeneous populations | Section 1.2 and Section 8.2; Equations (14)–(29) | Multi-population replicator systems with empirically grounded type distributions | Direction 1 |
| G4 | Calibration data are fragmented among competing holders; empirical validation is thin | Section 4.1 and Section 8.2 | Federated calibration on real enterprise ledgers; out-of-sample tests against observed policy episodes | Direction 4 |
| G5 | Policy is modeled as exogenous although it co-evolves with enterprise strategy | Section 6.1, Section 6.3 and Section 8.2; Equation (29) | Endogenous policy as an evolving player, with commitment and credibility constraints | Directions 3, 5 |
| G6 | Simulation claims lack verification protocols and reproducibility standards | Section 7 and Section 8.2 | Pre-registered parameterizations; audits of threshold predictions against realized adoption | Directions 4, 6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, G.; Zhong, L.; Zeng, Y. AI-Enhanced Evolutionary Game Theory for Intelligent Coordination and Adaptive Optimization in Low-Carbon Energy Systems: A Multi-Scale Review from Smart Grids to Carbon Markets. Processes 2026, 14, 2568. https://doi.org/10.3390/pr14162568
Wang G, Zhong L, Zeng Y. AI-Enhanced Evolutionary Game Theory for Intelligent Coordination and Adaptive Optimization in Low-Carbon Energy Systems: A Multi-Scale Review from Smart Grids to Carbon Markets. Processes. 2026; 14(16):2568. https://doi.org/10.3390/pr14162568
Chicago/Turabian StyleWang, Guorui, Liang Zhong, and Yixuan Zeng. 2026. "AI-Enhanced Evolutionary Game Theory for Intelligent Coordination and Adaptive Optimization in Low-Carbon Energy Systems: A Multi-Scale Review from Smart Grids to Carbon Markets" Processes 14, no. 16: 2568. https://doi.org/10.3390/pr14162568
APA StyleWang, G., Zhong, L., & Zeng, Y. (2026). AI-Enhanced Evolutionary Game Theory for Intelligent Coordination and Adaptive Optimization in Low-Carbon Energy Systems: A Multi-Scale Review from Smart Grids to Carbon Markets. Processes, 14(16), 2568. https://doi.org/10.3390/pr14162568
