An Enhanced MADDPG–A2C Framework for Optimized Resource Allocation in High-Speed Vehicular Networks
Abstract
1. Introduction
- (1)
- Construct a high-speed dynamic model of the vehicle network environment and formulate the RA problem as an MADRL problem in a continuous action space. Establish an objective optimization problem and a reward function under multiple demand constraints, with the primary aim of maximizing V2I throughput and ensuring V2V link transmission stability, thereby fulfilling diverse QoS requirements.
- (2)
- Develop an optimized algorithm framework to address the convergence issues that traditional reinforcement learning algorithms encounter in high-speed dynamic environments. The algorithm combines the MADDPG algorithm with the A2C architecture. By integrating the A2C framework to optimize gradient estimation, the advantage function is embedded into the loss functions of both the actor and critic and participating in gradient calculation. This approach reduces the mean squared error and enhances the algorithm’s convergence and stability.
- (3)
- Design a multidimensional reward function that encompasses not only V2I and V2V link rates but also includes regularization penalty terms. Set specific constraints to ensure the reward function adapts effectively to environmental changes, enhancing learning efficiency to fulfil the requirements of real-world V2V communication environments.
2. Related Work
3. System Model and Problem Formulation
3.1. System Model
3.2. Communication Model
3.2.1. Interference Analysis
3.2.2. Throughput Analysis
3.3. Problem Formulation
4. Proposed Resource Allocation Optimization Strategy Based on MADDPG and A2C
4.1. State Space Design
4.2. Action Space Design
4.3. Reward Function Design
4.4. Joint Optimization Algorithm Based on MADDPG and A2C
| Algorithm 1. Proposed MADDPG-A2C Algorithm. |
![]() |
5. Experimental Analysis and Discussion
5.1. Simulation Environment Setup
5.2. Ablation Experiment
5.3. Comparative Experimental Analysis
5.4. Statistical Analysis
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Xiao, Y.; Zhao, J.; Zhang, Q.; Huang, Y.; Quan, H.; Fan, L. MEC-enabled resource allocation in Internet of Vehicles. Phys. Commun. 2024, 65, 102368. [Google Scholar] [CrossRef]
- Hua, M.; Chen, D.; Qi, X.; Jiang, K.; Liu, Z.E.; Sun, H.; Zhou, Q.; Xu, H. Multi-Agent Reinforcement Learning for Connected and Automated Vehicles Control: Recent Advancements and Future Prospects. IEEE Trans. Autom. Sci. Eng. 2025, 22, 16266–16286. [Google Scholar] [CrossRef]
- Liu, W.; Hua, M.; Deng, Z.; Meng, Z.; Huang, Y.; Hu, C.; Song, S.; Gao, L.; Liu, C.; Shuai, B.; et al. A systematic survey of control techniques and applications in connected and automated vehicles. IEEE Internet Things J. 2023, 10, 21892–21916. [Google Scholar] [CrossRef]
- Jin, H.; Seo, J.; Park, J.; Kim, S.C. A Deep Reinforcement Learning-Based Two-Dimensional Resource Allocation Technique for V2I Communications. IEEE Access 2023, 11, 78867–78878. [Google Scholar] [CrossRef]
- Ji, B.; Huang, J.; Wang, Y.; Song, K.; Li, C.; Han, C.; Wen, H. Multi-Relay cognitive network with anti-fragile relay communication for intelligent transportation system under aggregated interference. IEEE Trans. Intell. Transp. Syst. 2023, 24, 7736–7745. [Google Scholar] [CrossRef]
- Chandrasekharan, P.; Jaekel, A. Transmission Power Based Congestion Control Using Q-Learning Algorithm in Vehicular Ad Hoc Networks (VANET). In Innovations in Mechatronics Engineering III; Springer: Cham, Switzerland, 2024. [Google Scholar]
- Zhang, X.; Peng, M.; Yan, S.; Sun, Y. Deep reinforcement learning based mode selection and resource allocation for cellularV2Xcommunications. IEEE Internet Things J. 2020, 7, 6380–6391. [Google Scholar]
- Xu, X.; Wu, Q.; Fan, P.; Wang, K.; Cheng, N.; Chen, W.; Letaief, K.B. Enhanced velocity-adaptive scheme: Joint fair access and age of information optimization in vehicular networks. IEEE Trans. Mob. Comput. 2025, 25, 3488–3505. [Google Scholar] [CrossRef]
- Ji, Y.; Wang, Y.; Zhao, H.; Gui, G.; Gacanin, H.; Sari, H.; Adachi, F. Multi-Agent Reinforcement Learning Resources Allocation Method Using Dueling Double Deep Q-Network in Vehicular Networks. IEEE Trans. Veh. Technol. 2023, 72, 13447–13460. [Google Scholar] [CrossRef]
- Lingayya, S.; Jodumutt, S.B.; Pawar, S.R.; Vylala, A.; Chandrasekaran, S. Correction: Dynamic task offloading for resource allocation and privacy-preserving framework in Kubeedgebased edge computing using machine learning. Clust. Comput. 2024, 27, 9433. [Google Scholar] [CrossRef]
- Gyawali, S.; Qian, Y.; Hu, R.Q. Deep Reinforcement Learning Based Dynamic Reputation Policy in 5G Based Vehicular Communication Networks. IEEE Trans. Veh. Technol. 2021, 70, 6136–6146. [Google Scholar] [CrossRef]
- Upadhyay, P.; Marriboina, V.; Goyal, S.J.; Kumar, S.; El-Kenawy, E.-S.M.; Ibrahim, A.; Alhussan, A.A.; Khafaga, D.S. An improved deep reinforcement learning routing technique for collision-free VANET. Rep. Sci. 2023, 13, 21796. [Google Scholar] [CrossRef] [PubMed]
- Ergün, S. Resource allocation optimization for effective vehicle network communications using multi-agent deep reinforcement learning. J. Dyn. Games 2025, 12, 134–156. [Google Scholar] [CrossRef]
- Fang, W.W.; Wang, Y.P.; Zhang, H. Optimization of Communication Resource Allocation in Vehicular Networks Based on Multi-Agent Deep Reinforcement Learning. J. Beijing Jiaotong Univ. 2022, 46, 64–72. [Google Scholar]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-agent actor-critic for mixed cooperative-competitive environments. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS’17); Curran Associates Inc.: New York, NY, USA, 2017; pp. 6382–6393. [Google Scholar]
- Mnih, V.; Badia, A.P.; Mirza, M. Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning (ICML’16); JMLR.org: Cambridge, MA, USA, 2016; Volume 48, pp. 1928–1937. [Google Scholar]
- Tian, J.; Liu, Q.; Zhang, H.; Wu, D. Multiagent Deep-Reinforcement-Learning-Based Resource Allocation for Heterogeneous QoS Guarantees for Vehicular Networks. IEEE Internet Things J. 2022, 9, 1683–1695. [Google Scholar] [CrossRef]
- Xu, L.; Chen, W.; Liu, X.; Chen, Y.Y. MADDPG: Multi-agent Deep Deterministic Policy Gradient Algorithm for Formation Elliptical Encirclement and Collision Avoidance. In Proceedings of 2021 5th Chinese Conference on Swarm Intelligence and Cooperative Control; Lecture Notes in Electrical Engineering; Springer: Singapore, 2023; Volume 934, pp. 978–981. [Google Scholar]
- Wan, K.; Wu, D.; Li, B.; Gao, X.; Hu, Z.; Chen, D. ME-MADDPG: An efficient learning-based motionplanning method for multiple agents in complex environments. Int. J. Intell. Syst. 2022, 37, 2393–2427. [Google Scholar] [CrossRef]
- Zhang, H.; Li, J.; Yuan, Q. Edge Service Migration for Vehicular Networks Based onMulti-agent Deep Reinforcement Learning. In Internet of Vehicles, Proceedings of the Technologies and Services Toward Smart Cities: 6th International Conference, IOV 2019 Proceedings, Kaohsiung, Taiwan, 18–21 November 2019; Springer: Berlin/Heidelberg, Germany, 2020. [Google Scholar]
- Kumar, A.S.; Zhao, L.; Fernando, X. Task Offloading and Resource Allocation in Vehicular Networks: A Lyapunov-Based Deep Reinforcement Learning Approach. IEEE Trans. Veh. Technol. 2023, 72, 13360–13373. [Google Scholar] [CrossRef]
- Kang, J.; Chen, J.; Xu, M.; Xiong, Z.; Jiao, Y.; Han, L.; Niyato, D.; Tong, Y.; Xie, S. UAV-Assisted Dynamic Avatar Task Migration for Vehicular Metaverse Services: A Multi-Agent Deep Reinforcement Learning Approach. IEEE/CAA J. Autom. Sin. 2024, 11, 430–445. [Google Scholar] [CrossRef]
- Peng, H.X.; Shen, X.M. Deep reinforcement learning based resource management for multi-access edge computing in vehicular networks. IEEE Trans. Netw. Sci. Eng. 2020, 7, 2416–2428. [Google Scholar] [CrossRef]
- Tang, T.; Li, C.; Liu, F. Collaborative cloud-edge-end task offloading with task dependency based on deep reinforcement learning. Comput. Commun. 2023, 209, 78–90. [Google Scholar] [CrossRef]
- Li, Z.; Xu, C.; Wu, Z.R. Deep reinforcement learning based trajectory design and resource allocation for task-aware multi-UAV enabled MEC networks. Comput. Commun. 2024, 213, 88–98. [Google Scholar] [CrossRef]
- Fan, C.; Xu, H.; Wang, Q. Multi-agent deep reinforcement learning for trajectory planning in UAVs-assisted mobile edge computing with heterogeneous requirements. Comput. Netw. 2024, 248, 110469. [Google Scholar] [CrossRef]
- Nayak, B.P.; Hota, L.; Kumar, A.; Turuk, A.K.; Chong, P.H.J. Autonomous Vehicles: Resource Allocation, Security, and Data Privacy. IEEE Trans. Green Commun. Netw. 2022, 6, 117–131. [Google Scholar] [CrossRef]
- Ahmed, M.; Liu, J.S.; Muhammad, A.M.M.; Wali, U.K.; Fahd, N.; Al-Wesabi, F.N. MARL based resource allocation scheme leveraging vehicular cloudlet in automotive-industry 5.0. Comput. Inf. Sci. 2023, 35, 101420. [Google Scholar] [CrossRef]
- Liu, Q.; Ma, Y. Communication resource allocation method in vehicular networks based on federated multi-agent deep reinforcement learning. Sci. Rep. 2025, 15, 30866. [Google Scholar] [CrossRef]
- Cao, T.; Huang, D.; Zhao, J. Application of Improved A2C Algorithm in Traffic Signal Control. Comput. Eng. Des. 2024, 45, 1713–1719. [Google Scholar]
- Wang, Q.; Li, Y.; Zhang, K. Multi-agent Asynchronous Advantage Actor-Critic for Dynamic Task Offloading in Mobile Edge Computing. IEEE Trans. Mob. Comput. 2024, 23, 5121–5135. [Google Scholar]
- Wiseman, Y. Adapting the H. 264 Standard to the Internet of Vehicles. Technologies 2023, 11, 103. [Google Scholar] [CrossRef]
- Chen, J.; Wang, Q.; Liu, Y. AoI-Aware Resource Allocation for C-V2X Networks via Multi-Agent Reinforcement Learning with Attention. In Proceedings of the 2024 IEEE 100th Vehicular Technology Conference (VTC2024-Fall), Washington, DC, USA, 7–10 October 2024; pp. 1–7. [Google Scholar]
- Ji, B.; Dong, B.; Li, D.; Wang, Y.; Yang, L.; Tsimenidis, C.; Menon, V.G. Optimization of Resource Allocation for V2X Security Communication Based on Multi-Agent Reinforcement Learning. IEEE Trans. Veh. Technol. 2025, 74, 1849–1861. [Google Scholar] [CrossRef]
- 3rd Generation Partnership Project; Technical Specification Group Radio Access Network. Evolved Universal Terrestrial Radio Access (E-UTRA); User Equipment (UE) Radio Transmission and Reception (Release 18): 3GPP TS 36. 101 Version 18.5.0; 3rd Generation Partnership Project: Sophia, France, 2024. [Google Scholar]
- Song, X.Q.; Zhang, W.J.; Lei, L.; Song, T.; Zhao, L. Dynamic resource allocation algorithm for multi-objective joint optimization in Internet of vehicle. J. Southeast Univ./Dongnan Daxue Xuebao 2025, 55, 266–274. [Google Scholar]















| Algorithm | Reference | Limitations | Environment | Convergence | Scalability |
|---|---|---|---|---|---|
| DQN | [6], 2024; [9], 2023; [12], 2023 | Hard to handle continuous actions; unstable in high-dynamic vehicular networks | Single-agent; discrete spectrum allocation | Slow; unstable in complex scenarios | Limited scalability |
| DDPG | [5], 2023; [18], 2023; [23], 2020 | Sensitive to hyperparameters; overestimation bias | Continuous power control; single-agent | Better than DQN but unstable under dynamic channels | Limited scalability to multi-agent |
| PPO | [4], 2023; [24], 2023; [25], 2024 | High sampling cost; limited fairness; weak generalization in large-scale IoV | Single-agent; continuous and discrete hybrid resource allocation | Stable convergence, better than DDPG/DQN | Poor scalability to multi-agent |
| MAPPO | [2], 2025; [29], 2025; [33], 2024 | High computational cost for centralized critic; bottleneck in large networks | Multi-agent; CTDE framework | Stable; faster convergence than MADDPG | Better than PPO but still limited in very large-scale IoV |
| MADDPG | [13], 2025 [15], 2017; [19], 2022; [20], 2020 | Training instability due to non-stationarity; high variance | Multi-agent; CTDE with deterministic policies | Converges faster than DDPG but is less stable than MAPPO | Limited scalability when the number of agents increases |
| MADQN | [9], 2023; [17], 2022; [34], 2025 | Limited to discrete action spaces; poor adaptability to mixed continuous/discrete tasks | Multi-agent; value-based | Converges more slowly than policy gradient methods | Moderate scalability but struggles in continuous IoV tasks |
| A2C | [16], 2016; [30], 2024; [31], 2024 | Sample inefficiency; sensitive to learning rates | Single-agent; synchronous actor–critic | Faster convergence than vanilla policy gradient; less stable than PPO | Limited scalability; mostly used in small IoV environments |
| Vehicular Environment Parameters | Value |
|---|---|
| Carrier frequency | 2 GHz |
| Number of V2I links | 4 |
| Number of V2V links | 4 |
| Bandwidth | 4 MHz |
| Noise power | −114 dBm |
| Base station antenna height | 25 m |
| Base station single-antenna gain | 8 dBi |
| Base station noise figure | 5 dB |
| Vehicle noise figure | 9 dB |
| Vehicle speed | 5–50m/s |
| Time constraint | 100 ms |
| Max transmission power of the V2I | 0.199 w |
| Max transmission power of the V2V | 0.199 w |
| Transmission power of the V2I | [−10, 5, 15, 23] dBm |
| Vehicle antenna height | 1.5 m |
| V2V link path loss model | |
| V2I link path loss model | |
| Fast-fading update | 1 ms |
| Index | Methods | Reward | V2V Links’ Success Probability | V2I Links’ Throughput (Mbps) | MFLOPS (M) | |||
|---|---|---|---|---|---|---|---|---|
| DDPG | A2C | MADDPG | Reward* | |||||
| 1 | √ | 64.47 | 0.688 | 27.43 | 0.103 | |||
| 2 | √ | 85.97 | 0.8217 | 28.61 | 0.108 | |||
| 3 | √ | 87.59 | 0.837 | 29.8 | 0.43 | |||
| 4 | √ | √ | 87.72 | 0.84 | 29.75 | 0.516 | ||
| 5 | √ | √ | √ | 88.75 | 0.851 | 30.84 | 0.516 | |
| Methods | Reward | V2I Links’ Throughput (Mbps) | V2V Links’ Success Probability | MFLOPS (M) |
|---|---|---|---|---|
| RANDOM | 85.09 ± 2.65 | 29.24 ± 0.87 | 0.797 ± 0.023 | — |
| DDPG | 64.47 ± 1.63 | 27.43 ± 0.42 | 0.688 ± 0.015 | 0.103 |
| MADQN | 69.65 ± 3.40 | 27.86 ± 0.62 | 0.735 ± 0.031 | 0.207 |
| MADDPG | 87.59 ± 1.54 | 29.80 ± 0.46 | 0.837 ± 0.021 | 0.43 |
| MADDPG-A2C | 88.75 ± 1.52 | 30.84 ± 0.40 | 0.851 ± 0.014 | 0.516 |
| ANOVA p | 1.15 × 10−216 | 1.19 × 10−288 | 8.62 × 10−198 | |
| ANOVA F | 1297.14 | 285.27 | 1294.17 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hu, L.; Zha, W.; Xue, P.; Xie, S.; Guo, B.; Wang, W. An Enhanced MADDPG–A2C Framework for Optimized Resource Allocation in High-Speed Vehicular Networks. Electronics 2026, 15, 1214. https://doi.org/10.3390/electronics15061214
Hu L, Zha W, Xue P, Xie S, Guo B, Wang W. An Enhanced MADDPG–A2C Framework for Optimized Resource Allocation in High-Speed Vehicular Networks. Electronics. 2026; 15(6):1214. https://doi.org/10.3390/electronics15061214
Chicago/Turabian StyleHu, Linna, Weixian Zha, Penghao Xue, Shuhao Xie, Bin Guo, and Wei Wang. 2026. "An Enhanced MADDPG–A2C Framework for Optimized Resource Allocation in High-Speed Vehicular Networks" Electronics 15, no. 6: 1214. https://doi.org/10.3390/electronics15061214
APA StyleHu, L., Zha, W., Xue, P., Xie, S., Guo, B., & Wang, W. (2026). An Enhanced MADDPG–A2C Framework for Optimized Resource Allocation in High-Speed Vehicular Networks. Electronics, 15(6), 1214. https://doi.org/10.3390/electronics15061214


