Comparative Analysis of MADQN and QMIX Multi-Agent Reinforcement-Learning Methods for Urban Traffic Signal Control
Featured Application
Abstract
1. Introduction
- The successes of a 16-intersection urban network have been evaluated against a comparative multi-objective metric, run 10 times with random seeds, and mobility, emissions, and surrogate safety indicators have been reported.
- The study evaluates the behavior of multi-agent reinforcement learning methods against each other and against other classical methods. To contextualize this, Max-Pressure was compared with Fixed-Time control as an analytical adaptive baseline.
- The multi-agent approaches MADQN and QMIX have been examined in a simulation environment across different traffic demand densities, using the same network and observation criteria, providing an application-level summary of the advantages of independent and coordinated reinforcement learning.
- The resilience of the algorithms against lighter or heavier congestion scenarios in response to systematic demand scaling has been discussed.
2. Related Work
2.1. Multi-Agent Traffic-Signal Control
2.2. State, Action and Reward Design
2.3. Robustness, CO2 Emissions, Safety, and Study Positioning
3. Methods
3.1. Fixed-Time Control (FT)
3.2. Max-Pressure Control (MP)
3.3. Multi-Agent Deep Q-Network (MADQN) Algorithm
| Algorithm 1. Pseudocode of the MADQN Algorithm. |
| Require: Environment E, 16 signalized intersections, shared online Q-network Qθ and target network Qθ−, shared replay buffer B, discount factor γ, learning rate α, ε-greedy schedule, decision interval Δt, signal constraints Output: Trained shared Q-network weights θ; best model checkpoint 1. Initialization Initialize θ and θ− (copy), empty replay buffer B, ε ← 1.0, global step ← 0 2. Episode loop for episode k = 1 to 400 do Initial local observation oᵢ0 for each intersection i ∈ {1,…,16} Local observation: oᵢ = [vehicle count, halting count, mean speed, occupancy] × 8 lanes + one-hot(active green phase) → dim = 34 3. Decision step at each decision epoch t do 3a. Compute feasible action set Aᵢ(t) per intersection under minimum-green and maximum-green constraints 3b. Select action aᵢ with masked ε-greedy policy: aᵢ ~ Uniform(Aᵢ(t)) with prob. ε; aᵢ ← argmaxₐ∈Aᵢ(t) Qθ(oᵢ, a) otherwise 3c. If phase change requested, insert yellow (3 s) then activate target green; else extend current green 3d. Observe oᵢᵗ+1 and local reward rᵢᵗ = w_q q̄ᵢ + w_d Δw̄ᵢ + w_tl Δτ̄ᵢ + w_epv eᵢ + w_em c̄ᵢ − w_sw zᵢ (Equation (7)) 3e. Push transition (oᵢᵗ, aᵢᵗ, rᵢᵗ, oᵢᵗ+1) to B; ε ← decay(ε) 4. Learning update if |B| ≥ warm-up threshold do Sample mini-batch of 256 transitions from B Compute Double-DQN target: yᵢᵗ = rᵢᵗ + γ · Qθ−(oᵢᵗ+1, argmaxₐ Qθ(oᵢᵗ+1, a)) (Equations (3) and (4)) Minimize L(θ) = E[(yᵢᵗ − Qθ(oᵢᵗ, aᵢᵗ))2]; gradient clipping (norm 10.0); update θ via Adam Every 2000 steps: synchronise target network θ− ← θ 5. Checkpoint if episode metric improves, save best model weights θ* 6. Return trained shared Q-network θ (used for decentralised evaluation) |
3.4. QMIX Multi-Agent Deep Reinforcement Learning Algorithm
| Algorithm 2. Pseudocode of the QMIX Algorithm. |
| Require: The environment E, 16 agents, shared agent utility network Qθ, mixing network fmix with hypernetworks, target counterparts Qθ− and fmix−, joint replay buffer B, discount factor γ, learning rate α, ε-greedy schedule, reward scale ρ, running state normalizer N (mean–variance, clip ± 5), decision interval Δt, signal constraints Output: Trained agent utility network weights θ and mixing network weights; best model checkpoint 1. Initialization Initialize θ, θ−, fmix, fmix− (copies), empty buffer B, running normalizer N, ε ← 1.0, global step ← 0 2. Episode loop for episode k = 1 to 400 do Reset environment; collect local observations oᵢ0 for each intersection i ∈ {1,…,16} (dim = 34 per agent) Form global state s0 = concat(o10, …, o160) (dim = 544); normalize: ŝ0 ← N(s0) 3. Decision step at each decision epoch t do 3a. Compute feasible action set Aᵢ(t) per agent under minimum-green and maximum-green constraints 3b. Forward pass: Qθ(oᵢᵗ, ·) for all i in one batch; select aᵢ with masked ε-greedy policy (same masking rule as Algorithm 1) 3c. Apply joint action with yellow insertion and green-time feasibility (identical to Algorithm 1): If phase change → insert yellow (3 s), then activate target green; otherwise extend current green 3d. Compute local reward ṙᵢᵗ = −queueᵢ − w_em · CO2ᵢ + w_epv · EPVᵢ − w_sw · switchᵢ; team reward r̂ᵗ = ρ · Σᵢ ṙᵢᵗ (Equation (8)) 3e. Collect next observations; form and normalize ŝᵗ+1; push joint transition (ŝᵗ, aᵗ, r̂ᵗ, ŝᵗ+1) to B; ε ← decay(ε) 4. Centralized learning update if |B| ≥ warm-up threshold do Sample mini-batch of 256 joint transitions; reshape to per-agent local states Compute chosen utilities: Qᵢ = Qθ(oᵢ, aᵢ); mix: Qtot = fmix(ŝ, Q1,…,Q16) (Equation (5)) Compute masked Double-Q target: yᵗ = r̂ᵗ + γ · fmix−(ŝᵗ+1, argmaxā Qθ(·, ā)) · (1 − dᵗ) (Equation (6)) Minimize Huber loss (Qtot, y); gradient clipping (norm 10.0); update θ and mixing weights via Adam Every 2000 steps: synchronize θ− ← θ and fmix− ← fmix 5. Checkpoint if episode metric improves, save best agent utility and mixing network weights 6. Return trained θ (local agent policy for decentralized evaluation; mixer discarded at execution) |
4. Experimental Study
4.1. Simulation Environment
4.1.1. State Space
4.1.2. Action Space Definition
4.1.3. Reward Design
4.2. Simulation Parameter Settings
5. Results and Discussion
5.1. Coordination Versus Independence



5.2. Robustness Across Demand Shifts
5.3. Trade-Offs Across Mobility, CO2 Emissions, and Safety
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| DQN | Deep Q-Network |
| EPV | Emergency-Priority Vehicle |
| FDR | False Discovery Rate |
| KPI | Key Performance Indicator |
| MADQN | Multi-Agent Deep Q-Network |
| PET | Post-Encroachment Time |
| QMIX | QMIX value decomposition method |
| RL | Reinforcement Learning |
| SSM | Surrogate Safety Module |
| SUMO | Simulation of Urban Mobility |
| TraCI | Traffic Control Interface |
| TTC | Time to Collision |
| CO2 | Carbon Dioxide |
References
- Webster, F.V. Traffic Signal Settings; Road Research Technical Paper No. 39; Road Research Laboratory: London, UK, 1957. [Google Scholar]
- Tang, D.; Duan, Y. Traffic signal control optimization based on neural network in the framework of model predictive control. Actuators 2024, 13, 251. [Google Scholar] [CrossRef]
- Ouyang, C.; Zhan, Z.; Lv, F. A comparative study of traffic signal control based on reinforcement learning algorithms. World Electr. Veh. J. 2024, 15, 246. [Google Scholar] [CrossRef]
- Rashid, T.; Samvelyan, M.; Schroeder de Witt, C.; Farquhar, G.; Foerster, J.; Whiteson, S. Monotonic value function factorisation for deep multi-agent reinforcement learning. J. Mach. Learn. Res. 2020, 21, 1–52. [Google Scholar]
- Bouktif, S.; Cheniki, A.; Ouni, A.; El-Sayed, H. Deep reinforcement learning for traffic signal control with consistent state and reward design approach. Knowl.-Based Syst. 2023, 267, 110440. [Google Scholar] [CrossRef]
- Tan, X.; Zhou, Y.; Jiao, X. Traffic signal control based on deep reinforcement learning using state fusion and trend reward. Eng. Appl. Artif. Intell. 2025, 159, 111701. [Google Scholar] [CrossRef]
- Zhou, R.; Nousch, T.; Wei, L.; Wang, M. Constrained traffic signal control under competing public transport priority requests via safe reinforcement learning. Expert Syst. Appl. 2025, 284, 127676. [Google Scholar] [CrossRef]
- Michailidis, P.; Michailidis, I.; Lazaridis, C.R.; Kosmatopoulos, E. Traffic Signal Control via Reinforcement Learning: A Review on Applications and Innovations. Infrastructures 2025, 10, 114. [Google Scholar] [CrossRef]
- Graves, R.T.; Nelson, Z.E.; Chakraborty, S. A decentralized intersection management system through collaborative negotiation be-tween smart signals. J. Intell. Transp. Syst. 2023, 27, 272–294. [Google Scholar] [CrossRef]
- Mushtaq, A.; Haq, I.U.; Sarwar, M.A.; Khan, A.; Khalil, W.; Mughal, M.A. Multi-Agent Reinforcement Learning for Traffic Flow Management of Autonomous Vehicles. Sensors 2023, 23, 2373. [Google Scholar] [CrossRef] [PubMed]
- Fang, B.; Zheng, C.; Wang, H.; Yu, T. Two-Stream Fused Fuzzy Deep Neural Network for Multiagent Learning. IEEE Trans. Fuzzy Syst. 2023, 31, 511–520. [Google Scholar] [CrossRef]
- Yang, X.; Yu, Y.; Feng, Y.; Ochieng, W.Y. Improving the Urban Transport System Resilience Through Adaptive Traffic Signal Control Enabled by Decentralised Multiagent Reinforcement Learning. J. Adv. Transp. 2024, 2024, 3035753. [Google Scholar] [CrossRef]
- Han, G.; Liu, X.; Han, Y.; Peng, X.; Wang, H. CycLight: Learning traffic signal cooperation with a cycle-level strategy. Expert Syst. Appl. 2024, 255, 124543. [Google Scholar] [CrossRef]
- Wei, H.; Chen, C.; Zheng, G.; Wu, K.; Gayah, V.; Xu, K.; Li, Z. PressLight: Learning max pressure control to coordinate traffic signals in arterial network. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), Anchorage, AK, USA, 4–8 August 2019; pp. 1290–1298. [Google Scholar] [CrossRef]
- Wei, H.; Xu, N.; Zhang, H.; Zheng, G.; Zang, X.; Chen, C.; Zhang, W.; Zhu, Y.; Xu, K.; Li, Z. CoLight: Learning network-level cooperation for traffic signal control. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM), Beijing, China, 3–7 November 2019; pp. 1913–1922. [Google Scholar] [CrossRef]
- Ault, J.; Sharon, G. Reinforcement learning benchmarks for traffic signal control. In Proceedings of the 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Datasets and Benchmarks Track, virtual, 6–14 December 2021; pp. 1–11. [Google Scholar]
- Wei, H.; Zheng, G.; Gayah, V.; Li, Z. A survey on traffic signal control methods. arXiv 2019, arXiv:1904.08117. [Google Scholar]
- Fan, L.; Yang, Y.; Ji, H.; Xiong, S. Research on cooperative control of traffic signals based on deep reinforcement learning. In Proceedings of the 2023 IEEE 12th Data Driven Control and Learning Systems Conference (DDCLS), Xiangtan, China, 12–14 May 2023; pp. 1608–1612. [Google Scholar] [CrossRef]
- Rasheed, F.; Yau, K.-L.A.; Noor, R.M.; Chong, Y.-W. Deep reinforcement learning for addressing disruptions in traffic light control. Comput. Mater. Contin. 2022, 71, 2225–2247. [Google Scholar] [CrossRef]
- Cao, K.; Yang, S.; Yang, C.; Yu, M.; Geng, J.; Jung, H. Research on intelligent traffic signal control based on multi-agent deep reinforcement learning. Mathematics 2026, 14, 149. [Google Scholar] [CrossRef]
- Bokade, R.; Jin, X.; Amato, C. Multi-agent reinforcement learning based on representational communication for large-scale traffic signal control. IEEE Access 2023, 11, 47646–47658. [Google Scholar] [CrossRef]
- Sattarzadeh, A.R.; Pathirana, P.N. Unification of probabilistic graph model and deep reinforcement learning (UPGMDRL) for multi-intersection traffic signal control. Knowl.-Based Syst. 2024, 305, 112663. [Google Scholar] [CrossRef]
- Wang, L.; Zhang, W.; Yan, Z. Vehicle–infrastructure cooperation framework for vehicle navigation and traffic signal control using deep reinforcement learning. Transp. Res. Rec. 2026, 2680, 568–583. [Google Scholar] [CrossRef]
- Zeynivand, A.; Javadpour, A.; Bolouki, S.; Sangaiah, A.K.; Ja’fari, F.; Pinto, P.; Zhang, W. Traffic flow control using multi-agent reinforcement learning. J. Netw. Comput. Appl. 2022, 207, 103497. [Google Scholar] [CrossRef]
- Rasheed, F.; Yau, K.-L.A.; Low, Y.-C. Deep reinforcement learning for traffic signal control under disturbances: A case study on Sunway City, Malaysia. Future Gener. Comput. Syst. 2020, 109, 431–445. [Google Scholar] [CrossRef]
- Kodama, N.; Harada, T.; Miyazaki, K. Traffic signal control system using deep reinforcement learning with emphasis on reinforcing successful experiences. IEEE Access 2022, 10, 128943–128950. [Google Scholar] [CrossRef]
- Akbar, A.; Ullah, S.S.; Malik, A.; Qaisar, S.M. RT-FedFlow: An efficient framework for real-time traffic signal optimization using federated multi-agent reinforcement learning. Eng. Appl. Artif. Intell. 2025, 161, 112147. [Google Scholar] [CrossRef]
- Hassan, M.A.; Elhadef, M.; Khan, M.U.G. Collaborative Traffic Signal Automation Using Deep Q-Learning. IEEE Access 2023, 11, 136016–136033. [Google Scholar] [CrossRef]
- Guzmán, J.A.; Pizarro, G.; Núñez, F. A Reinforcement Learning-Based Distributed Control Scheme for Cooperative Intersection Traffic Control. IEEE Access 2023, 11, 57038–57046. [Google Scholar] [CrossRef]
- An, Y.; Zhang, J. Traffic signal control method based on modified proximal policy optimization. In Proceedings of the 2022 10th International Conference on Traffic and Logistic Engineering (ICTLE), Macau, China, 12–14 August 2022. [Google Scholar] [CrossRef]
- Meepokgit, T.; Wisayataksin, S. Traffic signal control with state-optimizing deep reinforcement learning and fuzzy logic. Appl. Sci. 2024, 14, 7908. [Google Scholar] [CrossRef]
- Mei, X.; Fukushima, N.; Yang, B.; Wang, Z.; Takata, T.; Nagasawa, H.; Nakano, K. Reinforcement learning-based intelligent traffic signal control considering sensing information of railway. IEEE Sens. J. 2023, 23, 31125–31136. [Google Scholar] [CrossRef]
- Bouktif, S.; Cheniki, A.; Ouni, A. Traffic signal control using hybrid action space deep reinforcement learning. Sensors 2021, 21, 2302. [Google Scholar] [CrossRef]
- Xu, Y.; Wang, Y.; Liu, C. Training a reinforcement learning agent with AutoRL for traffic signal control. In Proceedings of the 2022 Euro-Asia Conference on Frontiers of Computer Science and Information Technology (FCSIT), Beijing, China, 16–18 December 2022. [Google Scholar] [CrossRef]
- Wang, P.; Wu, X.; He, X. Vibration-Theoretic Approach to Vulnerability Analysis of Nonlinear Vehicle Platoons. IEEE Trans. Intell. Transp. Syst. 2023, 24, 11334–11344. [Google Scholar] [CrossRef]
- Xu, D.; Liao, X.; Yu, Z.; Gu, T.; Guo, H. Robustness enhancement of deep reinforcement learning-based traffic signal control model via structure compression. Knowl.-Based Syst. 2025, 310, 113022. [Google Scholar] [CrossRef]
- Jagdish, A.; Liu, T. Impact of reward function selection on DQN-based traffic signal control. In Proceedings of the 2024 IEEE International Conference on Green Energy and Smart Systems (GESS), Long Beach, CA, USA, 11–12 November 2024. [Google Scholar] [CrossRef]







| Weight | Description | Reward Component |
|---|---|---|
| 1.0 | Mean departure delay per vehicle (s) | delay |
| 0.5 | Mean time loss per vehicle (s) | timeLoss |
| 0.2 | Mean halting vehicle count per lane | queue |
| 1.0 | Priority bonus for emergency vehicles | epv_bonus |
| 0.1 | Phase change frequency penalty | switch_penalty |
| 0.0 | CO2 emissions per vehicle | CO2 |
| Parameter | Setting |
|---|---|
| Network topology | Synthetic 4 × 4 urban grid with 16 signalized intersections |
| Road configuration | Bidirectional urban links with two lanes per edge; nominal speed 13.89 m/s |
| Vehicle classes | Passenger car, truck, and emergency-priority vehicle (EPV) |
| Simulation start/end | 0 s/9200 s |
| Simulation step length | 1 s |
| Decision interval | 10 s |
| Signal constraints | Two green actions per intersection; yellow 3 s; minimum green 10 s; maximum green 60 s |
| Safety logging | Deterministic SSM with TTC and PET; thresholds 3 s and 2 s; range 200 m; extra time 5 s |
| Training demand level | scaled300 |
| Evaluation demand levels | scaled025, scaled300, and scaled350 |
| Scenario | Demand Multiplier | Approximate Scheduled Demand (Loading Horizon) | Interpretation in This Study |
|---|---|---|---|
| scaled025 | 2.5× | ≈7204 vehicles (≈7140 cars, ≈58 trucks, 6 EPVs) | Lower-demand evaluation scenario |
| scaled300 | 3.0× | ≈8643 vehicles (≈8568 cars, ≈69 trucks, 6 EPVs) | Nominal scenario used for RL training and evaluation |
| scaled350 | 3.5× | ≈10,083 vehicles (≈9996 cars, ≈81 trucks, 6 EPVs) | Higher-demand evaluation scenario |
| Item | Value Used in This Study | Clarification |
|---|---|---|
| Car-following model | Krauss model | No explicit car-following-model override was defined in the network, route, or controller configurations. |
| Lane-changing model | LC2013 model | No explicit lane-changing-model override was defined. |
| Lateral-resolution setting | Standard lane-based simulation | Because no sub-lane extension was activated, vehicles occupied one lane laterally, and lane changes were represented without an explicit continuous lateral-motion model. |
| Routing assumption | Predefined routes with fixed edge sequences | Demand varied by flow intensity, not by online rerouting. |
| Seed usage | Training seed 42; evaluation seeds 1–10 | Mean ± standard deviation values were computed from 10 seeded evaluation runs per controller and demand scenario. |
| Parameter | MADQN | QMIX |
|---|---|---|
| Training paradigm | Independent multi-agent Double DQN with parameter sharing | Centralized training with decentralized execution (QMIX) |
| Number of agents | 16 | 16 |
| Local state size | 34 | 34 |
| Global state size | Not used | 544 |
| Action size per agent | 2 | 2 |
| Replay capacity | 200,000 | 500,000 |
| Batch size | 256 | 256 |
| Discount factor γ | 0.99 | 0.99 |
| Target update interval | 2000 steps | 2000 steps |
| Exploration schedule | ε: 1.0 → 0.05 over 200,000 steps | ε: 1.0 → 0.05 over 60,000 steps |
| Learning rate | 0.0005 | 0.0001 |
| Agent hidden layers | [256, 256] | [256, 256] |
| Additional architecture | Dueling network; Double DQN | Mixer hidden dim = 32; hypernet layers = [64]; Double-Q + Huber; state norm; reward scale 0.01; replay warm-up 5000 |
| Training episodes | 400 | 400 |
| Random seed | 42 | 42 |
| Demand | Controller | Waiting Time (s) | Travel Time (s) | Speed (m/s) | Throughput (veh/h) | CO2 (×109) | TTC < 1 s (%) | PET < 1 s (%) |
|---|---|---|---|---|---|---|---|---|
| scaled025 | Fixed-Time | 42.62 ± 0.08 | 132.78 ± 0.10 | 7.80 ± 0.01 | 3552.09 ± 0.59 | 2.48 ± 0.00 | 27.78 ± 0.25 | 22.58 ± 0.64 |
| scaled025 | Max-Pressure | 22.48 ± 0.63 | 117.05 ± 0.75 | 8.82 ± 0.05 | 3557.07 ± 5.77 | 2.28 ± 0.02 | 23.76 ± 0.32 | 9.49 ± 0.65 |
| scaled025 | MADQN | 19.24 ± 0.55 | 112.57 ± 0.72 | 9.16 ± 0.06 | 3562.10 ± 3.64 | 2.18 ± 0.02 | 22.09 ± 0.41 | 9.23 ± 0.68 |
| scaled025 | QMIX | 19.86 ± 0.54 | 112.58 ± 0.63 | 9.19 ± 0.04 | 3570.35 ± 4.24 | 2.17 ± 0.01 | 22.37 ± 0.36 | 9.10 ± 0.53 |
| scaled300 | Fixed-Time | 42.65 ± 0.08 | 133.27 ± 0.11 | 7.78 ± 0.01 | 4256.91 ± 3.25 | 3.00 ± 0.00 | 28.66 ± 0.28 | 20.25 ± 1.04 |
| scaled300 | Max-Pressure | 22.53 ± 0.93 | 117.38 ± 1.08 | 8.79 ± 0.08 | 4263.38 ± 5.60 | 2.74 ± 0.03 | 24.67 ± 0.40 | 8.37 ± 0.72 |
| scaled300 | MADQN | 19.56 ± 0.79 | 113.49 ± 0.95 | 9.09 ± 0.07 | 4268.30 ± 7.96 | 2.64 ± 0.02 | 23.45 ± 0.65 | 8.53 ± 0.63 |
| scaled300 | QMIX | 20.11 ± 0.53 | 113.44 ± 0.60 | 9.11 ± 0.04 | 4278.29 ± 6.90 | 2.63 ± 0.02 | 23.54 ± 0.20 | 8.60 ± 0.56 |
| scaled350 | Fixed-Time | 43.45 ± 0.08 | 134.50 ± 0.11 | 7.70 ± 0.01 | 4938.11 ± 18.23 | 3.70 ± 0.00 | 29.00 ± 0.20 | 20.99 ± 0.48 |
| scaled350 | Max-Pressure | 23.04 ± 1.29 | 118.41 ± 1.60 | 8.70 ± 0.11 | 4957.53 ± 7.54 | 3.39 ± 0.05 | 25.13 ± 0.40 | 8.00 ± 0.27 |
| scaled350 | MADQN | 19.38 ± 1.13 | 113.71 ± 1.33 | 9.05 ± 0.09 | 4974.73 ± 7.95 | 3.24 ± 0.04 | 23.90 ± 0.40 | 8.30 ± 0.27 |
| scaled350 | QMIX | 20.21 ± 0.33 | 114.17 ± 0.33 | 9.05 ± 0.02 | 4977.78 ± 11.56 | 3.24 ± 0.01 | 24.06 ± 0.32 | 8.50 ± 0.36 |
| Demand | Metric | MADQN (Mean ± SD) | QMIX (Mean ± SD) | Welch p | FDR-Adjusted p | Bonferroni-Adjusted p | Hedges’ g |
|---|---|---|---|---|---|---|---|
| scaled025 | Waiting time (s) | 19.24 ± 0.55 | 19.86 ± 0.54 | 0.0188 | 0.0939 | 0.2816 | −1.1060 |
| scaled025 | Travel time (s) | 112.57 ± 0.72 | 112.58 ± 0.63 | 0.9839 | 0.9839 | 1.0000 | −0.0090 |
| scaled025 | Throughput (veh/h) | 3562.10 ± 3.64 | 3570.35 ± 4.24 | 0.0002 | 0.0030 | 0.0030 | 2.0040 |
| scaled025 | CO2 (×109) | 2.18 ± 0.02 | 2.17 ± 0.01 | 0.0675 | 0.2024 | 1.0000 | 0.8350 |
| scaled025 | TTC < 1 s (%) | 22.09 ± 0.41 | 22.37 ± 0.36 | 0.1270 | 0.2381 | 1.0000 | −0.6860 |
| scaled300 | Waiting time (s) | 19.56 ± 0.79 | 20.11 ± 0.53 | 0.0859 | 0.2148 | 1.0000 | −0.7850 |
| scaled300 | Travel time (s) | 113.49 ± 0.95 | 113.44 ± 0.60 | 0.8917 | 0.9553 | 1.0000 | 0.0590 |
| scaled300 | Throughput (veh/h) | 4268.30 ± 7.96 | 4278.29 ± 6.90 | 0.0079 | 0.0589 | 0.1178 | 1.2840 |
| scaled300 | CO2 (×109) | 2.64 ± 0.02 | 2.63 ± 0.02 | 0.1160 | 0.2381 | 1.0000 | 0.7120 |
| scaled300 | TTC < 1 s (%) | 23.45 ± 0.65 | 23.54 ± 0.20 | 0.6982 | 0.8728 | 1.0000 | −0.1710 |
| scaled350 | Waiting time (s) | 19.38 ± 1.13 | 20.21 ± 0.33 | 0.0494 | 0.1851 | 0.7403 | −0.9510 |
| scaled350 | Travel time (s) | 113.71 ± 1.33 | 114.17 ± 0.33 | 0.3163 | 0.5032 | 1.0000 | −0.4510 |
| scaled350 | Throughput (veh/h) | 4974.73 ± 7.95 | 4977.78 ± 11.56 | 0.5017 | 0.6841 | 1.0000 | 0.2940 |
| scaled350 | CO2 (×109) | 3.24 ± 0.04 | 3.24 ± 0.01 | 0.8740 | 0.9553 | 1.0000 | −0.0700 |
| scaled350 | TTC < 1 s (%) | 23.90 ± 0.40 | 24.06 ± 0.32 | 0.3355 | 0.5032 | 1.0000 | −0.4240 |
| Demand | Metric | MADQN SD | QMIX SD | QMIX/MADQN SD Ratio | More Stable | Brown–Forsythe p | FDR-Adjusted p |
|---|---|---|---|---|---|---|---|
| scaled025 | Waiting time (s) | 0.5482 | 0.5380 | 0.9813 | QMIX | 0.9505 | 0.9505 |
| scaled025 | Travel time (s) | 0.7247 | 0.6301 | 0.8695 | QMIX | 0.6763 | 0.8413 |
| scaled025 | Throughput (veh/h) | 3.6351 | 4.2358 | 1.1652 | MADQN | 0.7051 | 0.8413 |
| scaled025 | CO2 (×109) | 0.0151 | 0.0126 | 0.8348 | QMIX | 0.5746 | 0.8413 |
| scaled025 | TTC < 1 s (%) | 0.4145 | 0.3625 | 0.8746 | QMIX | 0.7291 | 0.8413 |
| scaled300 | Waiting time (s) | 0.7867 | 0.5262 | 0.6688 | QMIX | 0.3869 | 0.7254 |
| scaled300 | Travel time (s) | 0.9517 | 0.5974 | 0.6277 | QMIX | 0.3289 | 0.7254 |
| scaled300 | Throughput (veh/h) | 7.9563 | 6.9019 | 0.8675 | QMIX | 0.5769 | 0.8413 |
| scaled300 | CO2 (×109) | 0.0224 | 0.0151 | 0.6733 | QMIX | 0.3436 | 0.7254 |
| scaled300 | TTC < 1 s (%) | 0.6497 | 0.1973 | 0.3037 | QMIX | 0.0397 | 0.1488 |
| scaled350 | Waiting time (s) | 1.1285 | 0.3290 | 0.2916 | QMIX | 0.0004 | 0.0018 |
| scaled350 | Travel time (s) | 1.3268 | 0.3337 | 0.2515 | QMIX | 0.0001 | 0.0015 |
| scaled350 | Throughput (veh/h) | 7.9468 | 11.5567 | 1.4543 | MADQN | 0.2064 | 0.6191 |
| scaled350 | CO2 (×109) | 0.0393 | 0.0087 | 0.2210 | QMIX | 0.0002 | 0.0015 |
| scaled350 | TTC < 1 s (%) | 0.4028 | 0.3241 | 0.8045 | QMIX | 0.8195 | 0.8780 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ali, A.O.; Köylü, F. Comparative Analysis of MADQN and QMIX Multi-Agent Reinforcement-Learning Methods for Urban Traffic Signal Control. Appl. Sci. 2026, 16, 5008. https://doi.org/10.3390/app16105008
Ali AO, Köylü F. Comparative Analysis of MADQN and QMIX Multi-Agent Reinforcement-Learning Methods for Urban Traffic Signal Control. Applied Sciences. 2026; 16(10):5008. https://doi.org/10.3390/app16105008
Chicago/Turabian StyleAli, Ahmed Osman, and Fehim Köylü. 2026. "Comparative Analysis of MADQN and QMIX Multi-Agent Reinforcement-Learning Methods for Urban Traffic Signal Control" Applied Sciences 16, no. 10: 5008. https://doi.org/10.3390/app16105008
APA StyleAli, A. O., & Köylü, F. (2026). Comparative Analysis of MADQN and QMIX Multi-Agent Reinforcement-Learning Methods for Urban Traffic Signal Control. Applied Sciences, 16(10), 5008. https://doi.org/10.3390/app16105008

