A Value-Driven Multi-Agent Reinforcement Learning Framework for Decentralized Adaptive Energy Management in Prosumer Smart Grids
Abstract
1. Introduction
2. Literature Review
3. Methodology
| Algorithm 1: VA-HMARL Training Procedure |
| Input: N agents, T timesteps/episode, E = 300 episodes, J_min = 0.90, ε_DP = 1.0, δ_DP = 10−5 |
| Output: converged policy parameters {θI*}. |
| (1) Initialize actor networks πI(θI) and critic networks VI(φI) with orthogonal initialisation for all i ∈ ; set λ(0) = 0.01. |
| (2) For episode e = 1,…, E: |
| (3) For timestep t = 0,…, T: each agent i observes oI(t) and selects action aI(t) ~ πI(|oI(t)) from the stochastic Gaussian policy; store transition (oI(t), aI(t), rI(t), oI(t + 1)) in the on-policy rollout buffer. |
| (4) Compute (r1:n(t)) via Equation (6); update λ(t) via Equation (7). |
| (5) At the end of the episode: sample minibatch (size 256); update actor/critic via MAPPO clipped objective with ε_clip = 0.20, entropy coefficient 0.01, GAE λ = 0.95, γ_disc = 0.99. |
| (6) Every 10 episodes: federated aggregation via Equation (4) with DP noise σ ≈ 3.93; discard gradients with cosine similarity <0.1 to the global mean (free-rider detection). |
| (7) Declare convergence when 50-episode rolling mean reward changes <0.5% over 20 consecutive evaluation windows. |
| (8) Return {θI*}. |
4. Proposed Application
4.1. System Architecture
4.2. Experimental Results and Tests
4.3. Sensitivity Analysis
4.4. Ablation Study
4.5. Discussions
5. Conclusions and Future Research Directions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Kumar, P.; Singh, O. Thermodynamic Analysis of Solid Oxide Fuel Cell-Gas Turbine-Organic Rankine Cycle Combined System. Int. J. Mater. Sci. Mech. Eng. 2019, 6, 98–102. [Google Scholar]
- Kumar, P.; Abhijit, D.; Bahman, S. Techno-economic analysis of an integrated desalination-renewable-hydrogen system for zero-emission freshwater and electricity production. Energy Convers. Manag. 2026, 353, 121231. [Google Scholar] [CrossRef] [Scilit]
- Dragomir, O.E.; Dragomir, F. Application of Scheduling Techniques for Load-Shifting in Smart Homes with Renewable-Energy-Sources Integration. Buildings 2023, 13, 134. [Google Scholar] [CrossRef] [Scilit]
- Dragomir, O.E.; Dragomir, F.; Gurgu, V.; Păun, M.; Duca, O.; Drăgoi, I.-C. Multi-agent System for Smart Grids with Produced Energy from Photovoltaic Energy Sources. In Proceedings of the 9th International Conference on Electronics, Computers and Artificial Intelligence (ECAI), Ploiesti, Romania, 30 June–1 July 2022; pp. 1–6. [Google Scholar]
- Vázquez-Canteli, J.R.; Nagy, Z. Reinforcement learning for demand response: A review of algorithms and modeling techniques. Appl. Energy 2019, 235, 1072–1089. [Google Scholar] [CrossRef] [Scilit]
- Perera, A.T.D.; Kamalaruban, P. Applications of reinforcement learning in energy systems. Renew. Sustain. Energy Rev. 2021, 137, 110618. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Badia, A.P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; Kavukcuoglu, K. Asynchronous Methods for Deep Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA, 20–22 June 2016; PMLR 48. pp. 1928–1937. [Google Scholar]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 6380–6391. [Google Scholar]
- Rashid, T.; Samvelyan, M.; Schroeder de Witt, C.; Farquhar, G.; Foerster, J.; Whiteson, S. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In Proceedings of the 35th International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018; PMLR 80. pp. 4295–4304. [Google Scholar]
- Jendoubi, I.; Bouffard, F. Multi-agent hierarchical reinforcement learning for energy management. Appl. Energy 2023, 332, 120500. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Zhao, T.; Yan, H.; Liu, M.; Liu, N. Hierarchical Hybrid Multi-Agent Deep Reinforcement Learning for Peer-to-Peer Energy Trading Among Multiple Heterogeneous Microgrids. IEEE Trans. Smart Grid 2023, 14, 4649–4665. [Google Scholar] [CrossRef] [Scilit]
- Tushar, W.; Saha, T.K.; Yuen, C.; Smith, D.; Poor, H.V. Peer-to-Peer Trading in Electricity Networks: An Overview. IEEE Trans. Smart Grid 2020, 11, 3185–3200. [Google Scholar] [CrossRef] [Scilit]
- Mengelkamp, E.; Notheisen, B.; Beer, C.; Dauer, D.; Weinhardt, C. A blockchain-based smart grid: Towards sustainable local energy markets. Comput. Sci.—Res. Dev. 2018, 33, 207–214. [Google Scholar] [CrossRef] [Scilit]
- Andoni, M.; Robu, V.; Flynn, D.; Abram, S.; Geach, D.; Jenkins, D.; McCallum, P.; Peacock, A. Blockchain technology in the energy sector: A systematic review of challenges and opportunities. Renew. Sustain. Energy Rev. 2019, 100, 143–174. [Google Scholar] [CrossRef] [Scilit]
- Qiu, D.; Ye, Y.; Papadaskalopoulos, D.; Strbac, G. Scalable Coordinated Management of Peer-to-Peer Energy Trading: A Multi-Cluster Deep Reinforcement Learning Approach. Appl. Energy 2021, 292, 116940. [Google Scholar] [CrossRef] [Scilit]
- Qiu, D.; Xue, J.; Zhang, T.; Wang, J.; Sun, M. Federated Reinforcement Learning for Smart Building Joint Peer-to-Peer Energy and Carbon Allowance Trading. Appl. Energy 2023, 333, 120526. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Xu, X.; Hu, D. Multiobjective Reinforcement Learning: A Comprehensive Overview. IEEE Trans. Syst. Man Cybern. Syst. 2015, 45, 385–398. [Google Scholar] [CrossRef] [Scilit]
- Roijers, D.M.; Vamplew, P.; Whiteson, S.; Dazeley, R. A Survey of Multi-Objective Sequential Decision-Making. J. Artif. Intell. Res. 2013, 48, 67–113. [Google Scholar] [CrossRef] [Scilit]
- Tessler, C.; Mankowitz, D.J.; Mannor, S. Reward Constrained Policy Optimization. In Proceedings of the 7th International Conference on Learning Representations (ICLR), New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Konečný, J.; McMahan, H.B.; Ramage, D.; Richtárik, P. Federated Optimization: Distributed Machine Learning for Mobile Devices. arXiv 2016, arXiv:1610.02527. [Google Scholar]
- Geyer, R.C.; Klein, T.; Nabi, M. Differentially Private Federated Learning: A Client Level Perspective. arXiv 2017, arXiv:1712.07557. [Google Scholar]
- Bonawitz, K.; Ivanov, V.; Kreuter, B.; Marcedone, A.; McMahan, H.B.; Patel, S.; Ramage, D.; Segal, A.; Seth, K. Practical Secure Aggregation for Privacy-Preserving Machine Learning. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS), Dallas, TX, USA, 30 October–3 November 2017; pp. 1175–1191. [Google Scholar] [CrossRef] [Scilit]
- Abadi, M.; Chu, A.; Goodfellow, I.; McMahan, H.B.; Mironov, I.; Talwar, K.; Zhang, L. Deep Learning with Differential Privacy. In Proceedings of the ACM SIGSAC Conference on Computer and Communications Security (CCS), Vienna, Austria, 24–28 October 2016; pp. 308–318. [Google Scholar] [CrossRef] [Scilit]
- McMahan, H.B.; Moore, E.; Ramage, D.; Hampson, S.; Agüera y Arcas, B. Communication-Efficient Learning of Deep Networks from Decentralized Data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS), Fort Lauderdale, FL, USA, 20–22 April 2017; PMLR 54. pp. 1273–1282. [Google Scholar]
- Lee, S.; Choi, D.-H. Federated Reinforcement Learning for Energy Management of Multiple Smart Homes with Distributed Energy Resources. IEEE Trans. Ind. Inform. 2022, 18, 488–497. [Google Scholar] [CrossRef] [Scilit]
- Matlab Software, Mathworks. Available online: https://mathworks.com/products/matlab-online.html (accessed on 26 January 2026).






| Method | Cost [€] | Jain’s J | Gini G | CO2 [kg] | Self-Suf. | J ≥ 0.90 |
|---|---|---|---|---|---|---|
| VA-HMARL | 31.87 ± 11.27 | 0.735 ± 0.051 | 0.338 ± 0.046 | 122.5 ± 20.0 | 0.448 ± 0.002 | ✓ YES |
| MADDPG | 33.06 ± 11.62 | 0.781 ± 0.013 | 0.225 ± 0.006 | 144.6 ± 23.6 | 0.673 ± 0.017 | ✗ NO |
| MAPPO | 32.71 ± 11.79 | 0.823 ± 0.017 | 0.180 ± 0.009 | 134.8 ± 22.0 | 0.720 ± 0.021 | ✗ NO |
| MPC | 31.61 ± 11.24 | 0.883 ± 0.016 | 0.110 ± 0.009 | 128.7 ± 21.0 | 0.764 ± 0.015 | ✗ NO |
| Rule-Based | 33.97 ± 12.32 | 0.738 ± 0.019 | 0.268 ± 0.006 | 149.5 ± 24.4 | 0.638 ± 0.022 | ✗ NO |
| VA-HMARL | 31.87 ± 11.27 | 0.735 ± 0.051 | 0.338 ± 0.046 | 122.5 ± 20.0 | 0.448 ± 0.002 | ✓ YES |
| Comparison | p-Value | Sig. | Interpretation |
|---|---|---|---|
| VA-HMARL vs. MADDPG | p = 0.0002 | *** | Very strong evidence of cost improvement |
| VA-HMARL vs. MAPPO | p = 0.0353 | * | Moderate evidence of cost improvement |
| VA-HMARL vs. MPC | p = 0.2772 | n.s. | Comparable cost; MPC violates privacy |
| VA-HMARL vs. Rule-Based | p = 0.0004 | *** | Very strong evidence of cost improvement |
| VA-HMARL vs. MADDPG | p = 0.0002 | *** | Very strong evidence of cost improvement |
| Variant | Cost [€/ep] | Jain’s J (Conv.) | CO2 [kg] | Self-Suf. | Component Removed |
|---|---|---|---|---|---|
| VA-HMARL (full) | 31.87 ± 11.27 | 0.912 ± 0.031 | 122.5 ± 20.0 | 0.448 ± 0.002 | —(full framework) |
| No VAM (economic only) | 30.94 ± 10.81 | 0.768 ± 0.042 | 141.2 ± 22.6 | 0.441 ± 0.009 | VAM reward shaping removed |
| No DP-FedAvg | 30.71 ± 10.65 | 0.921 ± 0.028 | 120.8 ± 19.4 | 0.452 ± 0.008 | Differential privacy removed |
| No dual ascent (λ = 0) | 31.10 ± 11.02 | 0.811 ± 0.059 | 124.3 ± 20.8 | 0.445 ± 0.010 | Fairness enforcement removed |
| No attention coordination | 32.88 ± 12.04 | 0.893 ± 0.041 | 126.7 ± 21.5 | 0.438 ± 0.012 | Attention coordination removed |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Dragomir, O.E.; Dragomir, F. A Value-Driven Multi-Agent Reinforcement Learning Framework for Decentralized Adaptive Energy Management in Prosumer Smart Grids. Buildings 2026, 16, 1974. https://doi.org/10.3390/buildings16101974
Dragomir OE, Dragomir F. A Value-Driven Multi-Agent Reinforcement Learning Framework for Decentralized Adaptive Energy Management in Prosumer Smart Grids. Buildings. 2026; 16(10):1974. https://doi.org/10.3390/buildings16101974
Chicago/Turabian StyleDragomir, Otilia Elena, and Florin Dragomir. 2026. "A Value-Driven Multi-Agent Reinforcement Learning Framework for Decentralized Adaptive Energy Management in Prosumer Smart Grids" Buildings 16, no. 10: 1974. https://doi.org/10.3390/buildings16101974
APA StyleDragomir, O. E., & Dragomir, F. (2026). A Value-Driven Multi-Agent Reinforcement Learning Framework for Decentralized Adaptive Energy Management in Prosumer Smart Grids. Buildings, 16(10), 1974. https://doi.org/10.3390/buildings16101974

