AI-Driven Traffic Control Method and Reliability Analysis for Digital City Local Narrow-Road, Dense-Network
Abstract
1. Introduction
- A discrete action embedding table and a CVAE [7] are employed to construct a latent representation space for discrete and continuous actions, while the PPO algorithm is applied for latent policy learning. Additionally, a nearest-neighbor search algorithm and a CVAE decoder conditioned on states and discrete action embeddings are introduced to reconstruct the discrete and continuous actions corresponding to the latent policy, enabling bidirectional conversion between original actions and their latent representations.
- Targeting the collective intelligence control requirements for continuous intersections, a globally shared latent action representation space and a locally shared state mechanism are designed. The state space incorporates the queue lengths of both upstream and downstream intersections. Guided by adjacent collaboration rewards, each agent achieves global cooperation while making independent decisions. This approach facilitates latent information interaction between agents and ensures consistent decision logic, thereby improving cooperative efficiency.
- To enhance the CVAE’s capability in learning the representation space of latent actions that dynamically capture the specificity of narrow-road, dense-network environments and improve the model’s generalization, a regularization loss function based on state residual prediction is incorporated. This assigns environmental impact-aware semantic information to each latent action, enabling agents to achieve a more in-depth comprehension of the higher-dimensional state space, thereby guiding the parameter updates of the hybrid action representation (HyAR).
- To improve model convergence efficiency and minimize training redundancy, a three-stage training process is implemented: HyAR pretraining, RL policy pre-training, and policy training. During the RL policy pre-training stage, a supervised learning loss function is incorporated to guide the preliminary updating of actor network parameters, thereby mitigating the risk of failing to achieve global optimization due to the high computational complexity inherent in high-dimensional HyAR spaces.
2. Related Work
3. HyAR-PPO-Based Traffic Signal Control Framework
3.1. Problem Description and Environment Modeling
3.1.1. Hybrid Action Space
3.1.2. Consistent Design of State Space and Reward Function
3.2. HyAR Reinforcement Learning Model
| Algorithm 1 AI-Driven Traffic Signal Control |
| Input initial traffic signal agents; environment state space; hyperparameters |
| 1: Initialize traffic signal agents |
| 2: for each agent i do |
| 3: Generate random state |
| 4: Select signal phase and green time |
| 5: Execute action and obtain state residual |
| 6: Train HyAR module to learn latent representation |
| 7: end for |
| 8: for each agent i do |
| 9: Obtain latent actions and from HyAR |
| 10: Train PPO agent to learn actions in latent space |
| 11: end for |
| 12: while training not complete do |
| 13: for each agent i do |
| 14: Observe state |
| 15: Select latent actions and from HyAR |
| 16: Decode to original actions (,) |
| 17: Execute action and collect reward ri,t |
| 18: Store experience in replay buffer |
| 19: Calculate HyAR loss LHyAR and update module |
| 20: Calculate PPO loss LPPO and update policy |
| 21: Update state residual = − |
| 22: Compute state residual loss to guide learning |
| 23: end for |
| 24: end while |
| 25: Assess traffic efficiency and control stability |
| Output optimized traffic signal control policy |
3.2.1. HyAR Module
3.2.2. Reinforcement Learning Policy Module
3.2.3. Reliable Training Process
4. Experimental Results and Discussion
4.1. Experimental Scenario Selection and Initialization
- 1.
- Fixed-time control
- 2.
- PPO-Continuous [11]
- 3.
- PPO-Discrete [11]
- 4.
- HPPO [21]
4.2. Reliable Analysis for Experimental Results
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| CVAE | Conditional Variational Autoencoder |
| PPO | Proximal Policy Optimization |
| ITS | Intelligent Transportation Systems |
| RL | Reinforcement Learning |
| DRL | Deep Reinforcement Learning |
| MARL | Multi-agent Reinforcement Learning |
| HPPO | Hybrid Proximal Policy Optimization |
| HyAR | Hybrid Action Representation |
| DQN | Deep Q-Network |
| DDQN | Double Deep Q-Network |
| A2C | Advantage Actor–Critic |
| SARSA | State–Action–Reward–State–Action |
| P-DQN | Parameterized Deep Q-Network |
References
- Saeidizand, P.; Savieri, P.; Boussauw, K. Car Dependency Contributors in Global Metropolitan Areas over Time. J. Transp. Geogr. 2025, 123, 104152. [Google Scholar] [CrossRef] [Scilit]
- Chen, W.; Wang, H.; Lu, J.; Xiao, H.; He, D.; Wang, P.; Ding, X.; Ding, W. Coupled Impacts of Urban Development Patterns and Policy Interventions on Motor Vehicle Ownership Based on Multi-Source Big Data. Sustainability 2026, 18, 3449. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Wang, F.Y.; Wang, K.; Lin, W.-H.; Xu, X.; Chen, C. Data-Driven Intelligent Transportation Systems: A Survey. IEEE Trans. Intell. Transp. Syst. 2011, 12, 1624–1639. [Google Scholar] [CrossRef] [Scilit]
- Ren, F.; Dong, W.; Zhao, X.; Zhang, F.; Kong, Y.; Yang, Q. Two-Layer Coordinated Reinforcement Learning for Traffic Signal Control in Traffic Network. Expert Syst. Appl. 2024, 235, 121111. [Google Scholar] [CrossRef] [Scilit]
- Rasheed, F.; Yau, K.L.A.; Noor, R.M.; Wu, C.; Low, Y.C. Deep Reinforcement Learning for Traffic Signal Control: A Review. IEEE Access 2020, 8, 208016–208044. [Google Scholar] [CrossRef] [Scilit]
- Haydari, A.; Yilmaz, Y. Deep Reinforcement Learning for Intelligent Transportation Systems: A Survey. IEEE Trans. Intell. Transp. Syst. 2020, 23, 11–32. [Google Scholar] [CrossRef] [Scilit]
- Sadeghi, M.; Leglaive, S.; Alameda-Pineda, X.; Girin, L.; Horaud, R. Audio-Visual Speech Enhancement Using Conditional Variational Auto-Encoders. IEEE/ACM Trans. Audio Speech Lang. Process. 2020, 28, 1788–1800. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-Level Control Through Deep Reinforcement Learning. Nature 2014, 518, 529–533. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Schaul, T.; Hessel, M.; Hasselt, H.; Lanctot, M.; Freitas, N. Dueling Network Architectures for Deep Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning (ICML), New York, NY, USA, 20–22 June 2016. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Fang, Z.; Chen, Y.; Dai, H.; Jiang, Q.; Zeng, X. Traffic Signal Optimization Control Method Based on Attention Mechanism Updated Weights Double Deep Q Network. Complex Intell. Syst. 2025, 11, 217. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Chu, T.; Wang, J.; Codeca, L.; Li, Z. Multi-Agent Deep Reinforcement Learning for Large-Scale Traffic Signal Control. IEEE Trans. Intell. Transp. Syst. 2020, 21, 1086–1095. [Google Scholar] [CrossRef] [Scilit]
- Fujimoto, S.; Meger, D.; Precup, D. Off-Policy Deep Reinforcement Learning without Exploration. In Proceedings of the 36th International Conference on Machine Learning (ICML), Long Beach, CA, USA, 9–15 June 2019. [Google Scholar] [CrossRef] [Scilit]
- Zhao, D.; Wang, H.; Shao, K.; Zhu, Y. Deep Reinforcement Learning with Experience Replay Based on SARSA. In Proceedings of the 2016 IEEE Symposium Series on Computational Intelligence (SSCI), Athens, Greece, 6–9 December 2016. [Google Scholar] [CrossRef] [Scilit]
- Yen, C.-C.; Ghosal, D.; Zhang, M.; Chuah, C.-N. A Deep On-Policy Learning Agent for Traffic Signal Control of Multiple Intersections. In Proceedings of the 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC), Rhodes, Greece, 20–23 September 2020. [Google Scholar] [CrossRef] [Scilit]
- Zheng, G.; Xiong, Y.; Zang, X.; Feng, J.; Wei, H.; Zhang, H.; Li, Y.; Xu, K.; Li, Z. Learning Phase Competition for Traffic Signal Control. arXiv 2019, arXiv:1905.04722. [Google Scholar] [CrossRef] [Scilit]
- Zhou, B.; Zhou, Q.; Hu, S.; Ma, D.; Jin, S.; Lee, D.-H. Cooperative Traffic Signal Control Using a Distributed Agent-Based Deep Reinforcement Learning With Incentive Communication. IEEE Trans. Intell. Transp. Syst. 2024, 25, 10147–10160. [Google Scholar] [CrossRef] [Scilit]
- Zeng, J.; Xin, J.; Cong, Y.; Zhu, J.; Zhang, Y.; Jiang, W.; Pu, S. HALight: Hierarchical Deep Reinforcement Learning for Cooperative Arterial Traffic Signal Control with Cycle Strategy. In Proceedings of the 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC), Macau, China, 8–12 October 2022. [Google Scholar] [CrossRef] [Scilit]
- Liang, X.; Du, X.; Wang, G.; Zhu, H. A Deep Reinforcement Learning Network for Traffic Light Cycle Control. IEEE Trans. Veh. Technol. 2019, 68, 1243–1253. [Google Scholar] [CrossRef] [Scilit]
- Aslani, M.; Mesgari, M.S.; Wiering, M. Adaptive Traffic Signal Control with Actor-Critic Methods in a Real World Traffic Network with Different Traffic Disruption Events. Transp. Res. Part C Emerg. Technol. 2017, 85, 732–752. [Google Scholar] [CrossRef] [Scilit]
- Luo, H.; Bie, Y.; Jin, S. Reinforcement Learning for Traffic Signal Control in Hybrid Action Space. IEEE Trans. Intell. Transp. Syst. 2024, 25, 5225–5241. [Google Scholar] [CrossRef] [Scilit]
- Bouktif, S.; Cheniki, A.; Ouni, A. Traffic Signal Control Using Hybrid Action Space Deep Reinforcement Learning. Sensors 2021, 21, 2302. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Genders, W.; Razavi, S. Using a Deep Reinforcement Learning Agent for Traffic Signal Control. arXiv 2016, arXiv:1611.01142. [Google Scholar] [CrossRef] [Scilit]
- Li, L.; Lv, Y.; Wang, F.Y. Traffic Signal Timing via Deep Reinforcement Learning. IEEE/CAA J. Autom. Sin. 2016, 3, 247–254. [Google Scholar] [CrossRef] [Scilit]
- Tan, T.; Bao, F.; Deng, Y.; Jin, A.; Dai, Q.; Wang, J. Cooperative Deep Reinforcement Learning for Large-Scale Traffic Grid Signal Control. IEEE Trans. Cybern. 2020, 50, 2687–2700. [Google Scholar] [CrossRef] [Scilit]
- Le, T.; Kovacs, P.; Walton, N.; Vu, H.L.; Andrew, L.L.; Hoogendoorn, S.S. Decentralized Signal Control for Urban Road Networks. Transp. Res. Part C Emerg. Technol. 2015, 58, 431–450. [Google Scholar] [CrossRef] [Scilit]
- Masson, W.; Ranchod, P.; Konidaris, G. Reinforcement Learning with Parameterized Actions. In Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016. [Google Scholar] [CrossRef] [Scilit]
- Zheng, G.; Zang, X.; Xu, N.; Wei, H.; Yu, Z.; Gayah, V.; Xu, K.; Li, Z. Diagnosing Reinforcement Learning for Traffic Signal Control. arXiv 2019, arXiv:1905.04716. [Google Scholar] [CrossRef] [Scilit]
- Bouktif, S.; Cheniki, A.; Ouni, A.; El-Sayed, H. Deep Reinforcement Learning for Traffic Signal Control with Consistent State and Reward Design Approach. Knowl.-Based Syst. 2023, 267, 110440. [Google Scholar] [CrossRef] [Scilit]
- Zhu, H.; Nakamura, H.; Alhajyseen, W.; Iryo-Asano, M. Modeling Traffic Flows on Urban Arterials Considering the Downstream Influence. Transp. Res. Rec. 2020, 2674, 475–485. [Google Scholar] [CrossRef] [Scilit]
- Wu, T.; Zhou, P.; Liu, K.; Yuan, Y.; Wang, X.; Huang, H.; Wu, D.O. Multi-Agent Deep Reinforcement Learning for Urban Traffic Light Control in Vehicular Networks. IEEE Trans. Veh. Technol. 2020, 69, 8243–8256. [Google Scholar] [CrossRef] [Scilit]
- Li, B.; Tang, H.; Yan, Z.; Hao, J.; Li, P.; Wang, Z.; Meng, Z.; Wang, L. HyAR: Addressing Discrete-Continuous Action Reinforcement Learning via Hybrid Action Representation. arXiv 2021, arXiv:2109.05490. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.; Abbeel, P. High-Dimensional Continuous Control Using Generalized Advantage Estimation. arXiv 2015, arXiv:1506.02438. [Google Scholar] [CrossRef] [Scilit]
- Bing Maps. Microsoft Corporation, NavInfo, OpenStreetMap; Map Review Number: GS(2025)3133. 2026. Available online: https://cn.bing.com/maps?FORM=Z9LH2&cp=qnmkk2tph6bv&lvl=12.8&style=r (accessed on 22 April 2026).











| Algorithm | Action Space | Coordination Mechanism | Evaluation |
|---|---|---|---|
| DDQN-Attention [10] | Discrete | Single-agent | Fixed discrete phases only; no continuous optimization |
| PPO-Discrete [11] | Discrete | Single-agent | Fixed interval; poor adaptation to dynamic flow |
| PPO-Continuous [11] | Continuous | Single-agent | Fixed phase sequence; uneven demand handling |
| SARSA [14,15] | Discrete | Single-agent/Multi-agent | On-policy; sample inefficient |
| HPPO [21] | Hybrid (independent) | Single-agent | Independent actions; ignores dependency |
| P-DQN [22] | Hybrid | Single-agent | Error propagation; limited scalability |
| Dueling DQN [9] | Discrete | Multi-agent | Weak coupling coordination |
| A2C [12] | Discrete | Multi-agent (centralized) | Centralized; high complexity; poor scalability |
| Centralized MARL [25] | Discrete | Centralized | Poor scalability |
| Distributed MARL [26] | Discrete | Decentralized | Weak coordination |
| HyAR-PPO | Hybrid (latent representation) | Decentralized with global latent sharing | Solves action dependency, and multi-agent coordination |
| Name of Intersection | Traffic Flow Classification | North Approach Lane | East Approach Lane | South Approach Lane | West Approach Lane | ||||
|---|---|---|---|---|---|---|---|---|---|
| Straight | Left | Straight | Left | Straight | Left | Straight | Left | ||
| Intersection of Shanrong Rd. and Lemin St. | Authentic traffic flow proportion | 0.3036 | 0.3429 | 0.8186 | 0.1002 | 0.3388 | 0.3224 | 0.7933 | 0.0772 |
| Traffic flow proportion under spatial balance | 0.5 | 0.25 | 0.5 | 0.25 | 0.5 | 0.25 | 0.5 | 0.25 | |
| Intersection of Fujia Rd. and Lemin St. | Authentic traffic flow proportion | 0.5413 | 0.2783 | 0.8914 | 0.0407 | 0.5405 | 0.1081 | 0.7844 | 0.1367 |
| Traffic flow proportion under spatial balance | 0.75 | 0.125 | 0.75 | 0.125 | 0.75 | 0.125 | 0.75 | 0.125 | |
| Intersection of North Chongwen Rd. and Lemin St. | Authentic traffic flow proportion | 0.2674 | 0.5116 | 0.9266 | 0.0169 | 0.5962 | 0.1871 | 0.7643 | 0.1067 |
| Traffic flow proportion under spatial balance | 0.625 | 0.1875 | 0.625 | 0.1875 | 0.625 | 0.1875 | 0.625 | 0.1875 | |
| Traffic Flow Classification | Road Classification | Time Period | |||
|---|---|---|---|---|---|
| [0, 15) min | [15, 30) min | [30, 45) min | [45, 60] min | ||
| Off-peak hours | East–west arterial road | 720 | 900 | 1080 | 900 |
| North–south minor roads | 240 | 300 | 360 | 300 | |
| Peak hours | East–west arterial road | 1200 | 1500 | 1800 | 1500 |
| North–south minor roads | 400 | 500 | 600 | 500 | |
| Hybrid Action Representation Module | Reinforcement Learning Policy Module | ||
|---|---|---|---|
| Model Parameter | Set Value | Model Parameter | Set Value |
| Latent Discrete Action Dimension d1 | 6 | Actor Network Learning Rate | 0.0003 |
| Latent Continuous Action Dimension d2 | 6 | Critic Network Learning Rate | 0.001 |
| Learning Rate | 0.0001 | Discount Factor γ | 0.99 |
| Conditional Encoding Hidden Layer | 256 | Actor Network Hidden Layer | (256, 128, 128) |
| Continuous Action Encoding Hidden Layer | 256 | Critic Network Hidden Layer | (256, 128, 64) |
| Latent Action Encoding Hidden Layer | (128, 128) | Batch Size | 256 |
| Continuous Action Decoding Hidden Layer | (256, 258) | Number of Training Episodes per Batch | 20 |
| State Residual Decoding Hidden Layer | 128 | Clipping Hyperparameter | 0.2 |
| Scenario | Evaluation Index | Fixed-Time | PPO- Continuous | PPO- Discrete | HPPO | Hyar-PPO | Optimization |
|---|---|---|---|---|---|---|---|
| Spatially imbalanced traffic scenarios during off-peak hours | Average delay time | 95.03 s | 80.51 s | 38.68 s | 35.98 s | 33.20 s | 7.73% |
| Average queue length | 7.5 veh | 5.7 veh | 2.5 veh | 2.5 veh | 2.2 veh | 6.3% | |
| Spatially imbalanced traffic scenarios during peak hours | Average delay time | 155.32 s | 127.41 s | 80.01 s | 61.37 s | 52.26 s | 14.84% |
| Average queue length | 12.5 veh | 9.8 veh | 5.9 veh | 4.7 veh | 4.3 veh | 9.2% | |
| Temporally imbalanced traffic scenarios during off-peak hours | Average delay time | 45.71 s | 38.62 s | 37.41 s | 35.04 s | 34.14 s | 2.57% |
| Average queue length | 3.6 veh | 2.5 veh | 2.4 veh | 2.3 veh | 2.2 veh | 4.0% | |
| Temporally imbalanced traffic scenarios during peak hours | Average delay time | 85.74 s | 63.28 s | 70.80 s | 59.35 s | 56.17 s | 5.36% |
| Average queue length | 7.9 veh | 5.5 veh | 6.2 veh | 5.2 veh | 4.8 veh | 7.2% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ji, A.; Wang, J.; Deng, H.; Wang, Z.; Zhang, M.; Wang, P. AI-Driven Traffic Control Method and Reliability Analysis for Digital City Local Narrow-Road, Dense-Network. Appl. Sci. 2026, 16, 4430. https://doi.org/10.3390/app16094430
Ji A, Wang J, Deng H, Wang Z, Zhang M, Wang P. AI-Driven Traffic Control Method and Reliability Analysis for Digital City Local Narrow-Road, Dense-Network. Applied Sciences. 2026; 16(9):4430. https://doi.org/10.3390/app16094430
Chicago/Turabian StyleJi, Aixu, Jie Wang, Hui Deng, Zipeng Wang, Mingfang Zhang, and Pangwei Wang. 2026. "AI-Driven Traffic Control Method and Reliability Analysis for Digital City Local Narrow-Road, Dense-Network" Applied Sciences 16, no. 9: 4430. https://doi.org/10.3390/app16094430
APA StyleJi, A., Wang, J., Deng, H., Wang, Z., Zhang, M., & Wang, P. (2026). AI-Driven Traffic Control Method and Reliability Analysis for Digital City Local Narrow-Road, Dense-Network. Applied Sciences, 16(9), 4430. https://doi.org/10.3390/app16094430

