Deep Reinforcement Learning-Based Adaptive Protocol Optimization for Heterogeneous IoT Networks in 5G-Enabled Smart Cities
Abstract
1. Introduction
- APO-DRL is proposed, a novel framework that jointly addresses communication protocol selection and transmission parameter optimization for heterogeneous IoT networks in 5G-enabled smart cities, employing a Dueling Double DQN architecture augmented with QoS-Aware Prioritized Experience Replay.
- QoS-Aware Prioritized Experience Replay (QA-PER) is introduced, a modification of standard Prioritized Experience Replay (PER) wherein the sampling priority of each transition is augmented by a factor where flags QoS-violating transitions. This mechanism ensures that critical rare events remain high-priority even after their TD-errors decrease, directly addressing a fundamental limitation of standard PER in QoS-sensitive network optimization.
- The adaptive protocol optimization problem is formulated as a multi-objective Markov Decision Process that concurrently optimizes throughput, latency, energy efficiency, and QoS satisfaction across heterogeneous categories of smart city IoT devices.
- A hierarchical state representation is designed, capturing device-level characteristics, network-level conditions, and application-level QoS requirements, enabling context-aware protocol decisions that adapt dynamically to real-time network states.
- Extensive simulations are conducted demonstrating that APO-DRL achieves higher throughput than the static allocation baseline and the other DRL algorithms (DQN, D3QN+PER) across multiple performance metrics, with all methods evaluated at N = 30 devices for the main comparison, and APO-DRL, Static Allocation, and AHP-TOPSIS additionally evaluated at N = 50, 100, and 200 devices to assess scalability (Section 5.5).
2. Related Work
2.1. IoT Communication Protocol Management in 5G Networks
2.2. Deep Reinforcement Learning for Wireless Network Optimization
2.3. Smart City IoT Network Optimization
2.4. Research Gap Analysis
3. System Model and Problem Formulation
3.1. Network Architecture
- Class 1—Environmental Monitoring (mMTC): Low data rate (≤100 kbps), high latency tolerance (≤10 s), ultra-low power consumption, high device density. Examples: air quality sensors, temperature/humidity monitors, noise-level detectors.
- Class 2—Smart Infrastructure (mMTC/eMBB): Moderate data rate (100 kbps–1 Mbps), moderate latency (≤1 s), medium power budget. Examples: smart meters, streetlight controllers, waste management sensors.
- Class 3—Public Safety (URLLC/eMBB): High data rate (1–50 Mbps), low latency (≤10 ms), high reliability (99.99%). Examples: surveillance cameras, emergency communication, gunshot detection.
- Class 4—Intelligent Transportation (URLLC): Variable data rate (1–100 Mbps), ultra-low latency (≤1 ms), ultra-high reliability (99.999%). Examples: traffic signal controllers, connected vehicle V2I communication, and autonomous shuttle coordination.
3.2. Communication Model
3.3. Problem Formulation
- (minimum data rate constraint)
- (maximum latency constraint)
- (maximum power constraint)
- (bandwidth constraint)
3.4. MDP Formulation
- Protocol selection: ∈ {0, 1, 2, 3} corresponding to {NB-IoT, LTE-M, LTE Cat-1, 5G NR}
- MCS index: ∈ {0, 1, …, 15}
- Power level: discrete power levels ∈ {−10, −5, 0, 5, 10, 15, 20, 23} dBm
4. Proposed APO-DRL Framework
4.1. Framework Architecture
4.2. Dueling Double DQN Architecture
- Input layer: 16-dimensional per-device input vector x_n,t ∈ R^16, extracted from the full state s_t = [s_t^dev ||s_t^net ||s_t^env] as follows. For device n selected at step t: x_n,t = [c_n (4), pos_n (2), q_n,t (1), SINR_n,t^p for p in P (4), L_p,t for p in P (4), t/T (1)], where c_n is a one-hot device class indicator, pos_n = (x_n/R, y_n/R) is the normalized 2-D position, q_n,t is the pending queue occupancy, SINR_n,t^p is the measured SINR on each of the |P| = 4 protocols, L_p,t is the load fraction on each protocol cell, and t/T is the episode progress. The mapping extracts 4 + 2 + 1 + 4 + 4 + 1 = 16 features per device from the full N × |P|-dimensional state; all other N-1 devices’ per-device sub-states are present in s_t^dev but are not passed to the network at step t. The input is fixed size regardless of N.
- Shared feature extraction: three fully connected layers [512, 256, 128] with ReLU activation and batch normalization.
- Value stream: 2 FC layers [128, 1].
- Advantage stream: 2 FC layers [128, |A|].
- Output: Q-values for all actions.
4.3. Prioritized Experience Replay
QoS-Aware Prioritized Experience Replay
4.4. Algorithm Description
| Algorithm 1: APO-DRL Training Procedure |
| Initialize: D3QN online network Q(s,a;θ), target network Q(s,a;θ−) Initialize: Prioritized Replay Buffer B with capacity M Initialize: ε ← 1.0, episode ← 0 FOR episode = 1 TO max_episodes DO: Reset smart city environment, observe initial state s0 FOR t = 1 TO T DO: // Select one device for this step Select device i_t~Uniform(1,…,N) // ε-greedy action selection WITH probability ε: select random action a_t OTHERWISE: a_t = argmax_a Q(s_t, a; θ) // Execute action: assign protocol and parameters to devices Execute a_t in environment Observe reward r_t, next state s_{t + 1}, done flag // Store transition with QA-PER priority Compute TD-error: δ_t = |r_t + γ·Q(s_{t + 1}, argmax_a Q(s_{t + 1},a;θ); θ−) − Q(s_t,a_t;θ)| Set v_t = 1 if device i_t violated QoS constraint, else v_t = 0 Compute priority: p_t = (|δ_t| + ε) × (1 + λ·v_t) // Equation (16), with i = t at storage time [λ = 1.5; α_p = 0.6 (PER prioritization exponent, Equation (14)); reward weights α = 0.3, β = 0.3, δ = 0.2, η = 0.2 (Equation (11))] Store (s_t, a_t, r_t, s_{t + 1}, done, v_t) in B with priority p_t // Sample prioritized mini-batch and update Sample mini-batch of size K from B using PER priorities Compute importance-sampling weights w_i Compute loss: L = (1/K) Σ w_i⋯(Y_i − Q(s_i, a_i; θ))2 Update θ via gradient descent Update priorities in B: p_i ← (|δ_i| + ε) × (1 + λ·v_i) // Equation (16) // Periodic target network update Every C steps: θ− ← θ // Decay exploration ε ← max(ε_end, ε − ε_decay) END FOR END FOR |
5. Simulation Results and Analysis
5.1. Simulation Environment
5.2. Baseline Methods
- Static Allocation (SA): Giving a device a fixed protocol based on its class (Class 1 → NB-IoT, Class 2 → LTE-M, Class 3 → LTE Cat-1, Class 4 → 5G NR).
- Random Selection (RS): A random protocol is chosen at each time step.
- AHP-TOPSIS: Multi-criteria network selection using the Analytic Hierarchy Process with load-aware protocol preference ranking. Protocol preference order by device class:
- Class 1–2: [NB-IoT, LTE-M, LTE Cat-1, 5G NR]
- Class 3–4: [5G NR, LTE Cat-1, LTE-M, NB-IoT]
- At each step, the highest-preference protocol with network load < 60% is selected.
- MCS index: 5 (Class 1–2), 10 (Class 3–4). Transmit power: 10 dBm
- 4.
- Standard DQN [25]: A simple Deep Q-Network with a consistent replay.
- 5.
- D3QN + Standard PER: Dueling Double DQN with standard Prioritized Experience Replay (Schaul et al. [35]), serving as the primary ablation baseline for QA-PER evaluation.
- 6.
- QoS-Greedy Heuristic: A stronger deterministic heuristic added in response to reviewer feedback. At each step, 16 candidate actions are sampled uniformly from the 512-action space; the candidate maximizing a local QoS utility estimate combining expected throughput (weight 0.4), transmit-power efficiency (weight 0.3), and a protocol-load penalty is selected. No learning or training is required. This baseline is more sophisticated than Random Selection and avoids the class-fixed limitations of Static Allocation, yet it lacks temporal credit assignment. Evaluated at N = 30 devices, seeds {0, 42, 123}, over 10 evaluation episodes per seed, QoS-Greedy achieved 20.35 ± 1.41 Mbps throughput, 67.66 ± 0.24% QoS satisfaction, 20.22 ± 1.13 ms latency, 0.51 ± 0.00 mJ/step energy, 17.56 ± 0.24% packet loss, and 38.41 ± 0.51% switch rate. These results fall below Static Allocation (25.12 Mbps, 81.30% QoS), confirming that greedy per-step optimization without a learned temporal state representation is insufficient for this task.
5.3. Convergence Analysis
5.4. Performance Comparison
| Metric | SA | RS | AHP-TOPSIS | QoS-Greedy | DQN | D3QN+PER | APO-DRL |
|---|---|---|---|---|---|---|---|
| Avg. Throughput (Mbps) | 25.12 ± 0.16 | 34.99 ± 0.16 | 31.90 ± 0.23 | 20.35 ± 1.41 | 52.52 ± 0.17 | 47.60 ± 11.03 ‡ | 60.00 ± 0.88 |
| Avg. Latency (ms) | 5.60 ± 0.12 | 37.41 ± 0.43 | 21.57 ± 0.52 | 20.22 ± 1.13 | 7.26 ± 0.33 | 13.80 ± 7.27 | 6.04 ± 0.19 |
| Energy (mJ/step) | 0.82 ± 0.00 | 0.93 ± 0.00 | 0.87 ± 0.00 | 0.51 ± 0.00 | 0.91 ± 0.01 | 0.92 ± 0.02 | 0.89 ± 0.01 |
| QoS Satisfaction (%) | 81.30 ± 0.07 | 62.16 ± 0.04 | 55.79 ± 0.10 | 67.66 ± 0.24 | 81.64 ± 0.27 | 75.45 ± 7.20 | 83.38 ± 0.23 |
| Packet Loss (%) | 7.20 ± 0.07 | 18.39 ± 0.01 | 8.73 ± 0.07 | 17.56 ± 0.24 | 10.89 ± 0.13 | 13.81 ± 3.26 | 10.19 ± 0.14 |
| Switch Rate (%) | 0.00 ± 0.00 | 74.66 ± 0.05 | 0.00 ± 0.00 | 38.41 ± 0.51 | 44.79 ± 0.52 | 24.68 ± 4.61 | 28.93 ± 0.52 |
| Inference Time (ms) | <0.1 | <0.1 | 0.4 | <0.1 | 2.2 | 2.2 | 2.2 |
5.4.1. Per-Class QoS Breakdown
5.4.2. Cross-Seed Generalization
5.5. Scalability Analysis
| N (Devices) | Method | Throughput (Mbps) | Latency (ms) | QoS Sat. (%) | Packet Loss (%) | Switch Rate (%) |
|---|---|---|---|---|---|---|
| 30 | APO-DRL | 60.00 ± 0.72 | 6.0 ± 0.2 | 83.4 ± 0.2 | 10.19 ± 0.12 | 28.93 ± 0.43 |
| 30 | Static | 25.12 ± 0.13 | 5.6 ± 0.1 | 81.3 ± 0.1 | 7.20 ± 0.06 | 0.00 ± 0.00 |
| 30 | AHP-TOPSIS | 31.90 ± 0.19 | 21.6 ± 0.4 | 55.8 ± 0.1 | 8.73 ± 0.05 | 0.00 ± 0.00 |
| 50 | APO-DRL | 54.92 ± 8.14 | 10.0 ± 4.0 | 79.1 ± 4.2 | 12.04 ± 1.91 | 35.53 ± 2.44 |
| 50 | Static | 25.01 ± 0.19 | 5.6 ± 0.0 | 81.1 ± 0.1 | 7.25 ± 0.06 | 0.00 ± 0.00 |
| 50 | AHP-TOPSIS | 31.71 ± 0.32 | 22.3 ± 0.5 | 55.8 ± 0.1 | 8.79 ± 0.06 | 0.00 ± 0.00 |
| 100 | APO-DRL | 58.16 ± 9.11 | 8.6 ± 4.5 | 80.8 ± 4.9 | 11.16 ± 2.13 | 49.44 ± 2.79 |
| 100 | Static | 25.09 ± 0.06 | 5.6 ± 0.1 | 81.2 ± 0.1 | 7.20 ± 0.03 | 0.00 ± 0.00 |
| 100 | AHP-TOPSIS | 31.84 ±0.10 | 21.8 ± 0.6 | 55.8 ± 0.0 | 8.75 ± 0.04 | 0.00 ± 0.00 |
| 200 | APO-DRL | 63.38 ± 1.59 | 6.3 ± 0.4 | 83.3 ± 0.4 | 9.89 ± 0.14 | 59.08 ± 0.67 |
| 200 | Static | 25.03 ± 0.09 | 5.7 ± 0.1 | 81.2 ± 0.1 | 7.23 ± 0.06 | 0.00 ± 0.00 |
| 200 | AHP-TOPSIS | 31.78 ± 0.18 | 21.7 ± 0.4 | 55.8 ± 0.1 | 8.76 ± 0.07 | 0.00 ± 0.00 |
5.6. Ablation Study
6. Discussion
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| APO-DRL | Adaptive Protocol Optimization using Deep Reinforcement Learning |
| D3QN | Dueling Double Deep Q-Network |
| DRL | Deep Reinforcement Learning |
| eMBB | Enhanced Mobile Broadband |
| IoT | Internet of Things |
| LTE-M | Long Term Evolution for Machines |
| MDP | Markov Decision Process |
| mMTC | Massive Machine-Type Communication |
| NB-IoT | Narrowband Internet of Things |
| PER | Prioritized Experience Replay |
| QA-PER | QoS-Aware Prioritized Experience Replay |
| QoS | Quality of Service |
| RAT | Radio Access Technology |
| URLLC | Ultra-Reliable Low-Latency Communication |
| 5G NR | 5G New Radio |
References
- Hamza, E.K.; Ibraheem, E.K.; Humaidi, A.J.; Al-Qassar, A.A. Energy-Efficient Mac Protocol and Scalable Communication Systems in WSN-IoT. J. Eng. Sci. Technol. 2025, 20, 2195–2218. [Google Scholar]
- IoT Analytics. State of IoT 2025: Number of Connected IoT Devices Growing 14% to 21.1 Billion Globally; IoT Analytics Research: Hamburg, Germany, 2025. [Google Scholar]
- Shafique, K.; Khawaja, B.A.; Sabir, F.; Qazi, S.; Mustaqim, M. Internet of Things (IoT) for next-generation smart systems: A review of current challenges, future trends and prospects for emerging 5G-IoT scenarios. IEEE Access 2020, 8, 23022–23040. [Google Scholar]
- 3GPP. Service Requirements for the 5G System; 3GPP TS 22.261, Release 18; ETSI: Valbonne, France, 2023. [Google Scholar]
- Rafique, W.; Khan, M.; Yakubu, A.; He, J. A survey on beyond 5G network slicing for smart cities applications. IEEE Commun. Surv. Tutor. 2024, 26, 1904–1942. [Google Scholar]
- Ogbodo, E.U.; Abu-Mahfouz, A.M.; Kurien, A.M. A Survey on 5G and LPWAN-IoT for Improved Smart Cities and Remote Area Applications: From the Aspect of Architecture and Security. Sensors 2022, 22, 6313. [Google Scholar] [PubMed]
- Barakabitze, A.A.; Ahmad, A.; Mijumbi, R.; Hines, A. 5G network slicing using SDN and NFV: A survey of taxonomy, architectures and future challenges. Comput. Netw. 2020, 167, 106984. [Google Scholar] [CrossRef]
- Phyu, H.P.; Ma, M.; Wang, X. Machine learning in network slicing A survey. IEEE Access 2023, 11, 39123–39153. [Google Scholar] [CrossRef]
- Hurtado Sánchez, J.A.; Casilimas, K.; Caicedo Rendon, O.M. Deep reinforcement learning for resource management on network slicing: A survey. Sensors 2022, 22, 3031. [Google Scholar] [CrossRef] [PubMed]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [PubMed]
- Wang, Z.; Schaul, T.; Hessel, M.; Hasselt, H.; Lanctot, M.; de Freitas, N. Dueling network architectures for deep reinforcement learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning, New York, NY, USA, 19–24 June 2016; pp. 1995–2003. [Google Scholar]
- Bendaoud, F.; Abdennebi, M. A machine learning access network selection in a heterogeneous wireless environment. Concurr. Comput. Pract. Exp. 2023, 36, e7989. [Google Scholar] [CrossRef]
- Bendaoud, F. A modified-SAW for network selection in heterogeneous wireless networks. ECTI Trans. Electr. Eng. Electron. Commun. 2017, 15, 8–17. [Google Scholar] [CrossRef]
- Cheng, P.; Chen, Y.; Ding, M.; Chen, Z.; Liu, S.; Chen, Y.P.P. Deep reinforcement learning for online resource allocation in IoT networks: Technology, development, and future challenges. IEEE Commun. Mag. 2023, 61, 111–117. [Google Scholar] [CrossRef]
- Wijethilaka, S.; Liyanage, M. Survey on network slicing for Internet of Things realization in 5G networks. IEEE Commun. Surv. Tutor. 2021, 23, 957–994. [Google Scholar] [CrossRef]
- Afolabi, I.; Taleb, T.; Samdanis, K.; Ksentini, A.; Flinck, H. Network slicing and softwarization: A survey on principles, enabling technologies, and solutions. IEEE Commun. Surv. Tutor. 2018, 20, 2429–2453. [Google Scholar] [CrossRef]
- Ebrahimi, S.; Bouali, F.; Haas, O. Resource management from single-domain 5G to end-to-end 6G network slicing: A survey. IEEE Commun. Surv. Tutor. 2024, 26, 2836–2866. [Google Scholar]
- Li, X.; Zhao, L.; Yu, K.; Aloqaily, M.; Jararweh, Y. A cooperative resource allocation model for IoT applications in mobile edge computing. Comput. Commun. 2021, 173, 183–191. [Google Scholar] [CrossRef]
- Seid, A.M.; Boateng, G.O.; Mareri, B.; Sun, G.; Jiang, W. Multi-agent DRL for task offloading and resource allocation in multi-UAV enabled IoT edge network. IEEE Trans. Netw. Serv. Manag. 2021, 18, 4531–4547. [Google Scholar]
- Li, R.; Wang, C.; Zhao, Z.; Guo, R.; Zhang, H. The LSTM-based advantage actor-critic learning for resource management in network slicing with user mobility. IEEE Commun. Lett. 2020, 24, 2005–2009. [Google Scholar]
- Mai, T.; Yao, H.; Zhang, N.; He, W.; Guo, D.; Guizani, M. Transfer reinforcement learning aided distributed network slicing optimization in industrial IoT. IEEE Trans. Ind. Inform. 2022, 18, 4308–4316. [Google Scholar]
- Dubey, V.; Chinara, S. AI based resource management for 5G network slicing: History, use cases, and research directions. Concurr. Comput. Pract. Exp. 2025, 37, e8327. [Google Scholar]
- Ssengonzi, C.; Kogeda, O.P.; Olwal, T.O. A survey of deep reinforcement learning application in 5G and beyond network slicing and virtualization. Array 2022, 14, 100142. [Google Scholar] [CrossRef]
- Abba Ari, A.A.; Gueroui, A.L.; Titouna, C.; Thiare, O.; Aliouat, Z. IoT-5G and B5G/6G resource allocation and network slicing orchestration using learning algorithms. IET Netw. 2025, 14, 89–118. [Google Scholar]
- Liang, F.; Yu, W.; Liu, X.; Griffith, D.; Golmie, N. Towards deep Q-network based resource allocation in Industrial Internet of Things. IEEE Internet Things J. 2022, 9, 9138–9150. [Google Scholar]
- Nivetha, A.; Preetha, K.S. Efficient joint resource allocation using self-organized map based deep reinforcement learning for cybertwin enabled 6G networks. Sci. Rep. 2025, 15, 19795. [Google Scholar]
- Wei, R.; Qin, T.; Huang, J.; Yang, Y.; Ren, J.; Yang, L. Resource allocation scheduling scheme for task migration and offloading in 6G cybertwin internet of vehicles based on DRL. IET Commun. 2024, 18, 1244–1265. [Google Scholar] [CrossRef]
- Anand, J.; Karthikeyan, B. Adaptive and intelligent customized deep Q-network for energy-efficient task offloading in mobile edge computing environments. Sci. Rep. 2026, 16, 5456. [Google Scholar] [CrossRef] [PubMed]
- Bendaoud, F.; Amraoui, A.; Sehimi, K. A DQN-based model for intelligent network selection in heterogeneous wireless systems. arXiv 2026, arXiv:2601.04978. [Google Scholar] [CrossRef]
- del Rio, A.; Jimenez, D.; Serrano, J. Comparative analysis of A3C and PPO algorithms in reinforcement learning: A survey on general environments. IEEE Access 2024, 12, 146795–146806. [Google Scholar] [CrossRef]
- Jamal, S.; Siddiqui, F.; Alam, M.A.; Ayman-Mursaleen, M.; Zafar, S.; Naaz, S. Deep reinforcement learning for sustainable urban mobility: A bibliometric and empirical review. Sensors 2026, 26, 376. [Google Scholar] [CrossRef] [PubMed]
- Louati, H.; Javed, M.H.; Bechikh Ali, M.; Louati, A.; Lahyani, R.; Elarbi, M. Multi-agent reinforcement learning for cooperative autonomous vehicles. Front. Sustain. Cities 2024, 6, 1449404. [Google Scholar]
- Singh, A.R.; Sujatha, M.S.; Kadu, A.D.; Bajaj, M.; Addis, H.K.; Sarada, K. A deep learning and IoT-driven framework for real-time adaptive resource allocation and grid optimization in smart energy systems. Sci. Rep. 2025, 15, 19309. [Google Scholar] [PubMed]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double Q-learning. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, AAAI 2016, Phoenix, AZ, USA, 12–17 February 2016; pp. 2094–2100. [Google Scholar]
- Schaul, T.; Quan, J.; Antonoglou, I.; Silver, D. Prioritized experience replay. In Proceedings of the ICLR 2016, San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- 3GPP. Study on Channel Model for Frequencies from 0.5 to 100 GHz; 3GPP TR 38.901, Release 18; ETSI: Valbonne, France, 2024. [Google Scholar]
- Salman, A.D.; Zeyad, A.T.; Jumaa, S.S.; Raafat, S.M.; Jasim, F.H.; Humaidi, A.J. Hybrid LLM-assisted fault diagnosis framework for 5G/6G networks using real-world logs. Computers 2025, 14, 551. [Google Scholar]
- Salman, A.D.; Zeyad, A.T.; Al-karkhi, A.A.S.; Raafat, S.M.; Humaidi, A.J. Hybrid CDN Architecture Integrating Edge Caching, MEC Offloading, and Q-Learning-Based Adaptive Routing. Computers 2025, 14, 433. [Google Scholar] [CrossRef]
- Fujimoto, S.; Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; pp. 1587–1596. [Google Scholar]
- Chen, M.; Challita, U.; Saad, W.; Yin, C.; Debbah, M. Artificial neural networks-based machine learning for wireless networks: A tutorial. IEEE Commun. Surv. Tutor. 2019, 21, 3039–3071. [Google Scholar] [CrossRef]
- Salman, A.D.; Khudheer, U.; Abdulsahib, G.M. An adaptive smart street light system for smart city. J. Comput. Theor. Nanosci. 2019, 16, 262–268. [Google Scholar] [CrossRef]
- Abbas, O.K.; Abdullah, F.B.; Radzi, N.A.M.; Salman, A.D.; Kadir, S.J.A. Survey on clustered routing protocols adaptivity for fire incidents: Architecture challenges, data losing, and recommended solutions. IEEE Access 2024, 12, 113518–113552. [Google Scholar] [CrossRef]
- Salman, A.D.; Seitz, J. An approach for QoS-aware routing in mobile ad hoc networks. In Proceedings of the International Symposium on Wireless Communication Systems (ISWCS), Brussels, Belgium, 25–28 August 2015; pp. 626–630. [Google Scholar] [CrossRef]
- Talib, A.A.; Salman, A.D. Design and develop authentication in electronic payment systems based on IoT and biometric. TELKOMNIKA (Telecommun. Comput. Electron. Control) 2022, 20, 1297–1306. [Google Scholar] [CrossRef]
- Korial, A.E.; Gorial, I.I.; Humaidi, A.J. An Improved Ensemble-Based Cardiovascular Disease Detection System with Chi-Square Feature Selection. Computers 2024, 13, 126. [Google Scholar]
- Mansoor, M.I.; Tuama, H.M.; Humaidi, A.J. Application of correlation-based recurrent neural network in porosity prediction for petroleum exploration. Eng. Res. Express 2025, 7, 015241. [Google Scholar] [CrossRef]
- Hady, H.N.; Hadi, R.H.; Hassoon, O.H.; Hasan, A.M.; Humaidi, A.J. Predicting process quality in multi-stage manufacturing using AE-BilA: An autoencoder-BiLSTM with attention mechanism. Eng. Res. Express 2025, 7, 015424. [Google Scholar]
- Samaan, S.S.; Korial, A.E.; Sarra, R.R.; Humaidi, A.J. Multilingual Web Traffic Forecasting for Network Management Using Artificial Intelligence Techniques. Results Eng. 2025, 26, 105262. [Google Scholar] [CrossRef]
- Khudhur, S.D.; Samaan, S.S.; Taher, O.N.M.; Salman, A.D.; Humaidi, A.J. NetGuard: A Hybrid Framework for Intelligent and Scalable Malicious URL Detection. J. Cybersecur. Priv. 2026, 6, 102. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2019, arXiv:1509.02971. [Google Scholar]
- Kibria, M.G.; Nguyen, K.; Villardi, G.P.; Zhao, O.; Ishizu, K.; Kojima, F. Big data analytics, machine learning, and artificial intelligence in next-generation wireless networks. IEEE Access 2018, 6, 32328–32338. [Google Scholar] [CrossRef]
- Luong, N.C.; Hoang, D.T.; Gong, S.; Niyato, D.; Wang, P.; Liang, Y.C.; Kim, D.I. Applications of deep reinforcement learning in communications and networking: A survey. IEEE Commun. Surv. Tutor. 2019, 21, 3133–3174. [Google Scholar] [CrossRef]
- Liang, Z.; Liu, Y.; Lok, T.M.; Huang, K. Multi-cell mobile edge computing: Joint service migration and resource allocation. IEEE Trans. Wirel. Commun. 2021, 20, 5898–5912. [Google Scholar] [CrossRef]
- Zaman, M.; Puryear, N.; Abdelwahed, S.; Zohrabi, N. A review of IoT-based smart city development and management. Smart Cities 2024, 7, 1462–1501. [Google Scholar]





| Reference | Approach | Protocol Selection | Parameter Optimization | Multi-Objective | Scalability |
|---|---|---|---|---|---|
| [12] | ML-Based Selection | ✓ | ✗ | ✗ | ✗ |
| [25] | DQN | ✗ | ✓ | ✗ | ✗ |
| [26] | SOM-DRL | ✗ | ✓ | ✓ | ✗ |
| [27] | D3QN-DDPG | ✗ | ✓ | ✓ | ✗ |
| [28] | AICDQN | ✗ | ✓ | ✓ | ✗ |
| [29] | DQN | ✓ | ✗ | ✗ | ✗ |
| APO-DRL (Ours) | D3QN+PER | ✓ | ✓ | ✓ | ✓ (N = 30–200) |
| Symbol | Domain | Definition |
|---|---|---|
| N | Z+ | Number of IoT devices |
| t | Z+ | Decision step index |
| s_t | ℝd | State vector at step t (channel, queue, QoS) |
| a_t | {0,…,511} | Action: joint (protocol, MCS, Tx power) for selected device |
| r_t | ℝ | Reward signal at step t (Equation (11)) |
| α, β, δ, η | [0, 1] | Reward weights for throughput, latency, energy, QoS (Equation (11)) |
| α_p, β_p | [0, 1] | PER prioritization and IS exponents (distinct from reward weights) |
| λ | ℝ+ | QA-PER QoS boost factor (Equation (16)) |
| ε | [0, 1] | ε-greedy exploration rate (decays from 1.0 to 0.05) |
| θ, θ^− | ℝ^p | Online and target network parameters |
| δ_i | ℝ | TD-error for transition i (Equation (14)) |
| p_i | ℝ+ | QA-PER priority for transition i (Equation (16)) |
| v_i | {0, 1} | QoS-violation flag for transition i |
| w_i | ℝ+ | Importance-sampling (IS) correction weight |
| U_r, U_l, U_e, U_q | [0, 1] | Normalized throughput, latency, energy, QoS utilities (Equations (12)–(15)) |
| Parameter | Value |
|---|---|
| Learning rate | 1 × 10−4 (Adam optimizer) |
| Discount factor (γ) | 0.99 |
| Batch size | 128 |
| Replay buffer size | 100,000 |
| Target network update frequency | Every 500 steps |
| Epsilon (ε) start/end/decay | 1.0/0.05/50,000 steps |
| PER α_p (prioritization exponent) | 0.6 |
| PER β (importance sampling) | 0.4 → 1.0 (linearly annealed) |
| Hidden layer activation | ReLU |
| Training episodes | 3000 (DRL); 50 (baselines) |
| Steps per episode | 5 (DRL); 20 (baselines) |
| Parameter | Value |
|---|---|
| Simulation area | 2 km × 2 km urban grid |
| Base station (gNB) | 1 macro + 4 small cells |
| Number of IoT devices | 30–200 (all methods evaluated) |
| Device distribution | Uniform random + hotspot clusters |
| Class 1 (Environmental) | 40% of devices |
| Class 2 (Infrastructure) | 30% of devices |
| Class 3 (Public Safety) | 20% of devices |
| Class 4 (Transportation) | 10% of devices |
| NB-IoT bandwidth | 200 kHz |
| LTE-M bandwidth | 1.4 MHz |
| LTE Cat-1 bandwidth | 20 MHz |
| 5G NR bandwidth | 100 MHz (3.5 GHz band) |
| Path loss model | 3GPP Urban Macro (TR 38.901) |
| Carrier frequency | 3.5 GHz (5G NR sub-6 GHz band) |
| UE height | = 1.5 m |
| Base station height | Macro: = 25 m; Small cell: = 10 m |
| Noise figure | NF = 9 dB |
| Thermal noise density | N0 = −174 dBm/Hz |
| Antenna configuration | SISO (single-input single-output, no beamforming) |
| Scheduler | Round-robin device selection per step |
| Shadowing | Log-normal, σ = 8 dB |
| Max Tx power (NB-IoT/LTE-M) | 23 dBm |
| Max Tx power (5G NR) | 23 dBm |
| Traffic model | Poisson arrivals, class-specific rates |
| Simulation tool/hardware | Python 3.10, PyTorch 2.6.0, Gymnasium|Apple MacBook M4 (MPS acceleration) |
| Simulation duration | 3000 episodes × 5 steps (DRL methods); 50 episodes × 20 steps (conventional baselines) |
| Variant | Throughput (Mbps) | QoS Sat. (%) | Pkt Loss (%) | Switch Rate (%) |
|---|---|---|---|---|
| D3QN + Standard PER (λ = 0) | 47.60 ± 11.03 ‡ | 75.45 ± 7.20 | 13.81 ± 3.26 | 24.68 ± 4.61 |
| APO-DRL + QA-PER (λ = 1.5) ★ | 60.00 ± 0.88 | 83.38 ± 0.23 | 10.19 ± 0.14 | 28.93 ± 0.52 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Alwane, S.K.; Jumaa, S.S.; Saleh, M.H.; Salman, A.D.; Al-Dujaili, A.Q.; Humaidi, A.J. Deep Reinforcement Learning-Based Adaptive Protocol Optimization for Heterogeneous IoT Networks in 5G-Enabled Smart Cities. IoT 2026, 7, 52. https://doi.org/10.3390/iot7030052
Alwane SK, Jumaa SS, Saleh MH, Salman AD, Al-Dujaili AQ, Humaidi AJ. Deep Reinforcement Learning-Based Adaptive Protocol Optimization for Heterogeneous IoT Networks in 5G-Enabled Smart Cities. IoT. 2026; 7(3):52. https://doi.org/10.3390/iot7030052
Chicago/Turabian StyleAlwane, Saddam K., Shereen S. Jumaa, Muna H. Saleh, Aymen D. Salman, Ayad Q. Al-Dujaili, and Amjad J. Humaidi. 2026. "Deep Reinforcement Learning-Based Adaptive Protocol Optimization for Heterogeneous IoT Networks in 5G-Enabled Smart Cities" IoT 7, no. 3: 52. https://doi.org/10.3390/iot7030052
APA StyleAlwane, S. K., Jumaa, S. S., Saleh, M. H., Salman, A. D., Al-Dujaili, A. Q., & Humaidi, A. J. (2026). Deep Reinforcement Learning-Based Adaptive Protocol Optimization for Heterogeneous IoT Networks in 5G-Enabled Smart Cities. IoT, 7(3), 52. https://doi.org/10.3390/iot7030052

