Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review
Abstract
1. Introduction
Novel Contributions of This Review
2. Technical Background of PPO for Communication Systems
2.1. Applications of PPO in Communication Systems
2.2. Alternatives to PPO
3. Materials and Methods
3.1. Review Design and Research Questions
3.2. Search Strategy
3.3. Study Selection Criteria
3.4. Screening Process (PRISMA Core)
3.5. PRISMA Flow Diagram
3.6. Data Extraction and Analysis
- (i)
- Communication application domain;
- (ii)
- PPO variant and architectural design;
- (iii)
- Primary system benefit;
- (iv)
- Recurring technical limitation or future research direction.
4. Results and Discussion
4.1. Resource Allocation and Management
4.2. Adaptive Sensing and Beamforming Systems
4.3. Communication–Computation Co-Design
- Executed locally;
- Offloaded to nearby edge servers through Vehicle-to-Infrastructure (V2I) communication;
- Delegated to neighboring vehicles through Vehicle-to-Vehicle (V2V) communication.
4.4. Integrated Sensing and Communication (ISAC)
4.5. UAV-Assisted Communications
4.6. Network Slicing, Service Function Chaining, and Orchestration
4.7. Satellite Communications
4.8. Vehicle-to-Everything (V2X) and Cognitive Radio
4.9. Common Trends in PPO-Based Communication Systems
4.10. Comparative Performance
4.11. Synthesis of Research Questions
5. Challenges and Limitations
5.1. Challenges with PPO Frameworks
5.2. Limitations of the Present Review
6. Future Research Directions
7. Implications for Sustainable and Smart Infrastructure
7.1. Sustainability and Practical Implications
7.2. Implications for Smart Infrastructure
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| 5G | Fifth Generation |
| 6G | Sixth Generation |
| B5G | Beyond Fifth Generation |
| A3C | Advantage Actor–Critic |
| ABR | Adaptive Bitrate |
| ACER | Actor–Critic with Experience Replay |
| AI | Artificial Intelligence |
| AoI | Age of Information |
| AoU | Age of Updates |
| CD-PPO | Contribution-based Dual-clip PPO |
| CTDE | Centralized Training with Decentralized Execution |
| CRN | Cognitive Radio Network |
| D2D | Device-to-Device |
| DDPG | Deep Deterministic Policy Gradient |
| DL | Deep Learning |
| DPPO | Distributed Proximal Policy Optimization |
| DRL | Deep Reinforcement Learning |
| DSA | Dynamic Spectrum Access |
| DNN | Deep Neural Network |
| ECN | Explicit Congestion Notification |
| eMBB | Enhanced Mobile Broadband |
| FL | Federated Learning |
| FWA | Fixed Wireless Access |
| GAN | Generative Adversarial Network |
| GAT | Graph Attention Networks |
| GAE | Generalized Advantage Estimation |
| GEO | Geostationary Earth Orbit |
| HTS | High-Throughput Satellite |
| IoT | Internet of Things |
| IoV | Internet of Vehicles |
| IPPO | Independent Proximal Policy Optimization |
| IRS | Intelligent Reflecting Surface |
| ISAC | Integrated Sensing and Communication |
| ISCC | Integrated Sensing, Communication, and Computing |
| KL | Kullback–Leibler |
| LEO | Low Earth Orbit |
| LSTM | Long Short-Term Memory |
| MAPPO | Multi-agent Proximal Policy Optimization |
| MARL | Multi-agent Reinforcement Learning |
| MEC | Multi-Access Edge Computing |
| MDP | Markov Decision Process |
| MIMO | Multiple-Input Multiple-Output |
| MISO | Multiple-Input Single-Output |
| ML | Machine Learning |
| mmWave | Millimeter Wave |
| MOS | Mean Opinion Scores |
| O-RAN | Open Radio Access Network |
| PPO | Proximal Policy Optimization |
| QoE | Quality of Experience |
| QoS | Quality of Service |
| RIS | Reconfigurable Intelligent Surface |
| RL | Reinforcement Learning |
| RSMA | Rate-Splitting Multiple Access |
| SAC | Soft Actor-Critic |
| SCA | Successive Convex Approximation |
| SFC | Service Function Chaining |
| SLA | Service Level Agreement |
| SNN | Spiking Neural Networks |
| STAR-RIS | Simultaneous Transmitting and Reflecting Reconfigurable Intelligent Surface |
| SU | Secondary User |
| SWIPT | Simultaneous Wireless Information and Power Transfer |
| TD3 | Twin Delayed DDPG |
| THz | Terahertz |
| TRPO | Trust Region Policy Optimization |
| UAV | Unmanned Aerial Vehicle |
| URLLC | Ultra-Reliable Low-Latency Communication |
| USD | Unmet System Demand |
| V2X | Vehicle-to-Everything |
| WP-IoT | Wireless Powered Internet of Things |
Appendix A
| Ref. No. | Authors | Year | Communication Domain | PPO Variant/Study Type | Contribution |
|---|---|---|---|---|---|
| [9] | Eskandari et al. | 2024 | RIS-aided MU-MISO Systems | PPO | Proposes a joint beamforming algorithm using statistical CSI and the PPO algorithm to maximize the ergodic sum rate while significantly reducing channel estimation overhead compared to instantaneous CSI methods. |
| [14] | Sheikh et al. | 2026 | Smart City Systems/System-of-Systems Networks | MAPPO/PPO-based MARL Framework | Multi-agent reinforcement learning framework for optimizing smart-city communication coordination and infrastructure management |
| [27] | Shen et al. | 2025 | Tactical Communication Networks | PPO + GRU + Graph Attention | Intelligent path selection |
| [28] | Zhang et al. | 2023 | RIS-Assisted SWIPT Networks with RSMA | PPO | PPO-based DRL framework to jointly optimize transmit beamforming vectors, power splitting (PS) ratios, common message rates, and RIS phase shifts in unison to maximize energy efficiency. |
| [30] | Viana et al. | 2025 | UAV Communications | MAPPO-Transformer | Secure UAV communication and resilience optimization. Integrates MAPPO with Transformer-based detection for resilient UAV links. |
| [31] | Zuo et al. | 2025 | 6G Space-Air-Ground Integrated Network (SAGIN) | MAPPO | Shows that a multi-agent PPO architecture can effectively coordinate communication-network decisions under realistic network constraints |
| [35] | Hong et al. | 2026 | Wireless Channel Access | IPPO (Specifically quantized QIPPO/CA) | Distributed channel-access optimization with quantized communication. Introduces a quantized IPPO framework (QIPPO/CA) for efficient channel access. |
| [36] | Wang et al. | 2025 | Cognitive Radio Networks | MAPPO | Dynamic spectrum access in heterogeneous wireless systems. Employs the HAPPO algorithm for heterogeneous-agent spectrum access. |
| [37] | Wang et al. | 2025 | High-Speed Data Center Networks (DCNs) | IPPO | Automatic ECN tuning scheme based on the IPPO algorithm that uses a decentralized training and execution paradigm to handle in-cast congestion and mixed mice–elephant traffic patterns. |
| [38] | Lin et al. | 2026 | Probabilistic Routing/Traffic Engineering | MAPPO (Specifically JGAT-MAPPO) | Constrained multi-agent DRL approach using JointGAT (Graph Attention Network) arc |
| [40] | Xu et al. | 2023 | Vehicular Networks | Contribution-based Dual-clip PPO | CD-PPO algorithm to jointly optimize subchannel selection and power allocation, maximizing intra-platoon transmission success ratios and V2I Mean Opinion Scores (MOS). |
| [41] | Alsahfi et al. | 2025 | Vehicular Big Data (VBD) Offloading | PPO | Develops a multi-tier offloading framework using PPO to intelligently allocate vehicular tasks across edge, regional, and cloud layers based on real-time feedback like congestion and CPU utilization. |
| [42] | Rehman et al. | 2025 | Cross-layer 6G Systems/6G A2G-TN Systems | MAPPO | Adaptive resource allocation in A2G-TN systems. Introduces CL-MAPPO for cross-layer resource block allocation. |
| [43] | Shang et al. | 2026 | ISAC-assisted V2X Networks | SNN-driven PPO (Spiking Actor-Critic PPO) | Integrates energy-efficient Spiking Neural Networks (SNNs) into a PPO-based Actor-Critic framework to jointly optimize beamforming and power allocation for integrated sensing and communication. |
| [45] | Wang et al. | 2024 | V2X/Integrated Sensing/Resource Optimization | PPO | An optimization framework for energy-harvesting vehicular communication networks, and explicitly evaluates the Age of Information (AoI) performance |
| [47] | Lin et al. | 2025 | MIMO Systems | PPO | Adaptive transmission in nonstationary environments. Develops a PPO-based adaptive mode and modulation selection scheme for dynamic channels. |
| [49] | Mao et al. | 2025 | Maritime Wireless Networks | PPO/MetaRL | Analyzes MAML-PPO as a comparator in a knowledge-embedded resource allocation study |
| [50] | Lin et al. | 2023 | In-Vehicle Heterogeneous Networks (HetNets) | PPO | Implements a PPO-based intelligent resource allocation mechanism to maximize device energy efficiency and satisfy dynamic traffic demands within vehicle cabins. |
| [51] | Ma et al. | 2026 | Secure Mobile Communications | PPO-BiLSTM | Secure communication via enhanced PPO. Proposes PPO-BOP incorporating BiLSTM and off-policy feedback for secrecy rates |
| [52] | Fang et al. | 2026 | IoT Edge Systems/Computing | MAPPO | Collaborative inference optimization. Introduces the MAHPPO multi-agent framework to optimize DNN partitioning and resource scheduling. |
| [54] | Mustafa et al. | 2025 | Vehicular Edge Computing | PPO | PPO-based algorithm using Generalized Advantage Estimation (GAE) and surrogate clipping to optimize offloading decisions, minimizing delays and task drop ratios in dynamic vehicular networks |
| [58] | Wang et al. | 2026 | UAV Communication/Multi-UAV Wireless Networks | MAPPO/DRL-based RSMA Framework | Energy-efficient RSMA optimization in UAV-assisted wireless-powered communication networks. EEMACO algorithm based on MAPPO to maximize weighted sum rates. |
| [64] | Zhao et al. | 2025 | Satellite Communications/LEO Satellite Networks | Hybrid PPO-DQN | Handover and power allocation optimization/Proposes a hierarchical framework using PPO for timing and DQN for location in handovers. |
| [65] | Marzuk et al. | 2025 | V2X/O-RAN | DRL | Proposes an optimized distributed computation offloading (ODCO) framework using PPO to minimize latency and energy consumption in 5G/6G O-RAN-based V2X networks. |
| [67] | Su et al. | 2026 | Dense 5G Planning | Hierarchical MAPPO | Scalable network optimization/Proposes HMAPPO-RL to optimize base station placement and antenna beamwidth |
| [68] | Sun et al. | 2025 | UAV Networks/Resource Allocation | Hierarchical PPO (empirical simulation) | Hierarchical PPO for dynamic resource allocation in UAV networks, demonstrating improved spectrum efficiency and service quality. |
| [70] | He et al. | 2025 | O-RAN/Digital Twin | HAPPO | Heterogeneous-Agent PPO RL for xApps Coordination in Digital Twin Enabled O-RAN |
| [71] | Qazzaz et al. | 2026 | 6G O-RAN/Green Communications | PPO | A hierarchical rApp-xApp framework that jointly optimizes radio unit activation and multi-criteria user association weights to reduce power |
| [72] | Asemian et al. | 2026 | MEC-O-RAN/Anti-Jamming | PPO (Jamming Estimator) | DHRL framework using a PPO-based jamming estimator (ADPHRP-JE) to predict jammed slots and a Transformer-based task scheduler |
| [74] | Lee and Kim | 2024 | Hybrid RL Optimization | Hybrid PPO-DQN | Improved exploration and policy stability. Implements a dual-agent framework (PPO-DQN variant) for safety-critical navigation. |
| [78] | Luo et al. | 2024 | Video Streaming (ABR) | BC-PPO (Behavior Cloning + PPO) | Proposes BC-PPO ABR to address slow convergence in learning-based ABR algorithms, using PPO to handle severe network fluctuations in mmWave 5G environments. |
| [79] | Iqbal et al. | 2026 | RIS-Assisted Multi-User MISO Systems | PPO | Develops an on-policy PPO-based DRL algorithm to jointly optimize base station beamforming and RIS phase shifts, significantly reducing computational complexity compared to Fractional Programming (FP). |
| [80] | Wara et al. | 2025 | Full-Duplex RIS-Aided NOMA-ISAC | MAPPO | Uses Centralized Training with Decentralized Execution (CTDE) to maximize minimum beampattern gain by jointly controlling beamforming, RIS configuration, and power allocation. |
| [81] | Hu et al. | 2024 | Secure MmWave D2D Networks | Nested PPO/MAPPO | Employs a nested DRL structure using discrete PPO for RIS-user association and MAPPO for multi-agent |
| [87] | Xie et al. | 2025 | UAV-Assisted Semantic D2D Networks | GNN-enabled PPO | Integrates heterogeneous Graph Neural Networks with PPO to jointly optimize MU transmission power, channel allocation, and semantic symbol rates to maximize QoE under malicious jamming. |
| [88] | He et al. | 2024 | D2D Mobile-Edge Computing (MEC) | MAPPO | Dynamic partitioning scheme for idle/active devices using MAPPO to minimize long-term average task delay for delay-sensitive applications under strict deadline constraints. |
| [82] | Iqbal et al. | 2024 | RIS-Assisted MU-MISO Systems | PPO/DPPO | Jointly optimizes beamforming and phase shifts using PPO with surrogate clipping to maximize bit-per-joule energy efficiency in MU-MISO systems. |
| [83] | Wu et al. | 2025 | UAV-Assisted Vehicular MEC | MAPPO (joint caching/computation) | MAPPO for joint caching, computation, and resource management in UAV-vehicular MEC. |
| [86] | Zhang et al. | 2025 | Multi-UAV Cooperative Systems | AS-MAPPO (improved multi-agent PPO) | Enhanced MAPPO for cooperative dynamic target search in multi-UAV systems. |
| [89] | Hu et al. | 2026 | 6G-enabled IoV/VEC | MAPPO (Improved for CTDE decoupling) | Utilizes an improved MAPPO algorithm to decouple centralized training from distributed execution and a server-weighted scoring selection (SS) algorithm to optimize task offloading across cloud-edge-device collaborative layers while balancing Quality of Experience (QoE) and energy consumption. |
| [95] | Shen et al. | 2024 | Edge Computing/ISAC | PPO (with Robust Design) | Proposes a computationally robust PPO algorithm to optimize joint communication, perception, and task offloading in edge-assisted ISAC systems under computation uncertainty. |
| [96] | Sun et al. | 2025 | ISAC | MAPPO | MAPPO-driven resource allocation for ISAC in UAV/LEO scenarios. |
| [97] | Zhou et al. | 2026 | Industrial Wireless/Formation Control | Multi-Agent PPO | Solves the multi-agent networked formation control problem using MAPPO, leveraging global ISCC state information to reduce synchronization errors and latency |
| [100] | Ghomri et al. | 2024 | NOMA-UAV Networks/IoT | PPO | PPO-based DRL agent with a multi-action space to simultaneously optimize UAV 3D trajectory, transmit power, IoT node association, and power allocation factors to balance energy efficiency and far-near fairness |
| [101] | Aung et al. | 2024 | Aerial STAR-RIS-assisted MEC | PPO | Utilizes PPO for its sample efficiency and stability to minimize total energy consumption by jointly optimizing task offloading, aerial STAR-RIS trajectory, amplitude and phase shift coefficients, and power allocation |
| [102] | Wang et al. | 2024 | Multi-UAV Relay Communication | Mix-Greedy MAPPO | MAPPO for path planning and relay communication in air-ground UAV networks. |
| [103] | Chaudhary et al. | 2025 | STAR-RIS-assisted V2V Networks | PPO | Implements a PPO-based DRL algorithm to maximize system sum rates in STAR-RIS networks by jointly optimizing beamforming vectors, coefficient matrices, and symbol rates. |
| [104] | Zhou et al. | 2024 | IoRT/UAV-Satellite Integrated Networks | Compound-action PPO (CPPO) | CPPO to handle mixed continuous and discrete action spaces, optimizing UAV trajectories, sensor scheduling, and transmission decisions to balance AoI, energy, and costs |
| [106] | Nguyen & Kim | 2025 | 5G/6G Dynamic TDD Networks | PPO-TA (Actor-Critic PPO variant) | PPO-TA dynamically schedules TDD time slots to maximize the uplink/downlink sum rate using a KL-divergence-penalized objective function to ensure training stability and QoS compliance. |
| [107] | Adhikari et al. | 2025 | 6G/RIS-Assisted Hybrid Slicing | PPO | Implements a hybrid slicing technique (NOMA + puncturing) with PPO to intelligently schedule URLLC traffic on top of eMBB traffic |
| [108] | Raja et al. | 2026 | B5G O-RAN/Traffic Classification | PPO | Couples SVM-based classification with a PPO agent for slice-aware PRB allocation, including a silence-aware resource reclamation mechanism |
| [109] | Zhang et al. | 2024 | MEC/D2D Communication | PPO (Actor-Critic PPO) | Jointly optimizes task offloading and resource allocation in D2D-assisted MEC, outperforming DQN and A2C in balancing delay and energy consumption. |
| [110] | Hikmat & Sahib et al. | 2026 | Network Slicing, Service Function Chaining, and Orchestration | PPO-MDP | Dynamic resource allocation and network slicing in 5G, with empirical gains in throughput, energy efficiency, fairness, and QoS compared to traditional methods (GA, PSO, etc.). |
| [114] | Fu et al. | 2025 | RIS-assisted NOMA Satellite Networks | PPO-DQN | Merges PPO (for continuous phase-shift and power variables) and DQN (for discrete user pairing) to maximize the average sum rate in 6G satellite-ground communications |
| [115] | Zhang et al. | 2026 | LoRa-based Direct-to-Satellite IoT Networks | Self-Attention PPO (SAPPO) | Integrates a self-attention mechanism into the PPO algorithm to optimize uplink scheduling and fairness while avoiding packet collisions in LEO satellite IoT environments |
| [116] | Xu et al. | 2023 | Multibeam GEO Satellite Communications | PPO | Joint power and bandwidth allocation using PPO-based DRL; achieved USC performance comparable to optimized genetic algorithms with substantially lower computation time. |
| [117] | Li et al. | 2025 | Ultra-Dense LEO Satellite Networks/Packet Routing | MAPPO | MAPPO-based routing algorithm integrated with Graph Attention Networks (GATs) and an M/M/1/K queuing model to minimize communication delay and energy consumption across dynamic topologies. |
| [118] | Zhang et al. | 2026 | Satellite Communications/LEO Satellite Networks | Hybrid DQN-PPO/Dual-Agent PPO | Joint optimization of satellite handover and power allocation in LEO satellite systems |
| [119] | Sadiki et al. | 2023 | MIMO-based Multi-access Edge Computing (MEC) | PPO | Formulates the offloading problem in a massive MIMO-MEC system as an MDP and introduces a PPO-based algorithm to solve the limitations of discrete action spaces, specifically for continuous power allocation. |
| [120] | Li et al. | 2026 | Satellite-Terrestrial Integrated Networks (STINs) | MAPPO | Proposes a handover-oriented learning scheme for multi-user STINs using MAPPO under a centralized training |
| [121] | Zhao et al. | 2025 | Multi-LEO Satellite Networks | Hierarchical PPO (HPPO) | Digital Twin-empowered framework using Hierarchical PPO (HPPO) to jointly optimize beam hopping, bandwidth, and power allocation, effectively reducing satellite load imbalances. |
| [122] | Meng et al. | 2025 | LEO Satellite Communication/Hybrid Wide-Spot Beam Coverage | MAPPO | Cooperative MAPPO algorithm to jointly optimize power allocation and dynamic beam hopping to maximize throughput and minimize delay fairness among beam positions |
| [123] | Alharbi | 2025 | Smart City Systems | Deep Multi-Objective PPO | Review of deep multi-objective reinforcement learning and vision-based systems for smart cities. PPO as a key algorithm for balancing conflicting urban goals like congestion and energy management |
| [124] | Louati et al. | 2024 | Autonomous Vehicle Networks | Multi-Agent PPO | Cooperative autonomous vehicle coordination for sustainable smart city environments. Benchmark a novel Multi-Agent Actor-Critic (MA2C) algorithm against MAPPO for multi-AV lane-changing decisions, emphasizing passenger comfort and energy efficiency gains across varying traffic densities. |
| [127] | Zhang et al. | 2025 | Wireless Security | MAPPO | Proposes a multi-user anti-jamming algorithm using MAPPO for throughput optimization |
| [132] | Zhen et al. | 2025 | 5G Ultra-Dense Networks (UDN) | MAPPO + MSD | The base station control algorithm uses MAPPO, integrated with a novel Mode Switching Decision (MSD) algorithm that uses cosine similarity to minimize unnecessary sleep-mode transitions and energy waste. |
| [133] | Wang et al. | 2025 | Multi-Agent Navigation Systems | MAPPO | Proposes the Integrated Adaptive Communication Network (IACN) based on MAPPO, featuring dynamic topology adjustment via learnable graphs, content optimization for task relevance, and adaptive frequency modulation via Bayesian Networks. |
| [136] | Espinosa et al. | 2025 | Edge-Cloud ECC/V2X | PPO | Performs a comparative study between PPO and Particle Swarm Optimization (PSO) for energy-efficient task allocation in Kubernetes-orchestrated clusters, identifying PPO as a faster and more computationally lightweight option for resource-constrained devices. |
| [137] | Wu et al. | 2025 | Vehicular Edge Computing | PPO-Transformer (AHP-PPO) | Lightweight adaptive task offloading optimization in vehicular edge computing. PPO with a lightweight Transformer using dual-stream attention to optimize joint task offloading and proactive caching in dynamic IoV environments. |
| [138] | Li et al. | 2025 | Air-Ground Vehicular Edge Computing Networks | MAPPO/DRL-based PPO Framework | Joint task offloading and resource allocation in energy-harvesting vehicular edge networks. DC-MAPPO (Dual-Clip MAPPO) for task offloading in air-ground networks. |
| [139] | Wu et al. | 2026 | IoT Networks/Dynamic Spectrum Access (DSA) | Decentralized PPO | Fully decentralized framework that uses an interference-aware state-action representation and adaptive reward shaping to optimize throughput and reduce MAC-layer latency in dense IoT environments |
| [142] | Liu et al. | 2026 | mmWave Systems | PPO-SelfAttention | Proposes SAPPO to enhance beam tracking accuracy via self-attention mechanisms. |
| [143] | Liu et al. | 2025 | UAV-assisted ISAC | PPO | Implements PPO within a Federated Learning framework for ISCC resource optimization. |
| [144] | Hu et al. | 2025 | Wireless Powered IoT (WP-IoT)/AoI | HVF-based PPO (Hybrid Value Function PPO) | Integrates clipped and unclipped value functions to jointly optimize scheduling and power control for minimizing the average Age of Information (AoI) in hybrid-action spaces. |
| [145] | Chen et al. | 2026 | Semantic Communications | Hybrid PPO-DDPG | Integrates PPO for power management and DDPG for knowledge base adaptation in RSMA systems. |
References
- Ericsson. EMR June 2025 Highlights Growing Monetization Appeal of 5G Fixed Wireless Access. 2025. Available online: https://www.ericsson.com/en/press-releases/2025/6/emr-june-2025-highlights-growing-monetization-appeal-of-5g-fixed-wireless-access (accessed on 24 June 2026).
- Statista. 5G—Statistics & Facts. 2025. Available online: https://www.statista.com/topics/3447/5g/?srsltid=AfmBOoqwwLz1PwaNuxcOGXpe4chSvyttaEnxiMFzNriHlbev7lnecxh- (accessed on 24 June 2026).
- Hoang, D.T.; Huynh, N.V.; Nguyen, D.N.; Hossain, E.; Niyato, D. Deep Reinforcement Learning for Wireless Communications and Networking: Theory, Applications and Implementation, 1st ed.; Wiley: Hoboken, NJ, USA, 2023. [Google Scholar]
- Chen, A.C.H.; Jia, W.-K.; Hwang, F.-J.; Liu, G.; Song, F.; Pu, L. Machine Learning and Deep Learning Methods for Wireless Network Applications. EURASIP J. Wirel. Commun. Netw. 2022, 2022, 115. [Google Scholar] [CrossRef]
- Wang, Y.; Lei, J.; Shang, F.; Li, Y. A Comprehensive Survey of Multi-Agent Deep Reinforcement Learning for Wireless Spectrum Management. Neurocomputing 2025, 653, 131236. [Google Scholar] [CrossRef]
- Goel, A.; Masurkar, S.; Pathade, G.R. An Overview of Digital Transformation and Environmental Sustainability: Threats, Opportunities, and Solutions. Sustainability 2024, 16, 11079. [Google Scholar] [CrossRef]
- Safitra, M.F.; Lubis, M.; Kurniawan, M.T.; Alhari, M.I.; Nuraliza, H.; Azzahra, S.F.; Putri, D.P. Green Networking: Challenges, Opportunities, and Future Trends for Sustainable Development. In Proceedings of the 2023 11th International Conference on Computer and Communications Management, Nagoya Japan, 4–6 August 2023; ACM: New York, NY, USA, 2023; pp. 168–173. [Google Scholar]
- Kumar, R.; Gupta, S.K.; Wang, H.-C.; Kumari, C.S.; Korlam, S.S.V.P. From Efficiency to Sustainability: Exploring the Potential of 6G for a Greener Future. Sustainability 2023, 15, 16387. [Google Scholar] [CrossRef]
- Eskandari, M.; Zhu, H.; Shojaeifard, A.; Wang, J. Statistical CSI-Based Beamforming for RIS-Aided Multiuser MISO Systems via Deep Reinforcement Learning. IEEE Wirel. Commun. Lett. 2024, 13, 570–574. [Google Scholar] [CrossRef]
- Zhu, C.; Dastani, M.; Wang, S. A Survey of Multi-Agent Deep Reinforcement Learning with Communication. Auton. Agents Multi-Agent Syst. 2024, 38, 4. [Google Scholar] [CrossRef]
- Zheng, M.; Zhang, J.; Zhan, C.; Ren, X.; Lü, S. Proximal Policy Optimization with Reward-Based Prioritization. Expert Syst. Appl. 2025, 283, 127659. [Google Scholar] [CrossRef]
- Sha, S.; Liu, Y.; Huo, B. Dynamic Proximal Policy Optimization: Enhancing PPO with Adaptive Entropy and Smooth Clipping. Neurocomputing 2026, 674, 132861. [Google Scholar] [CrossRef]
- Cui, H. Evaluating the Performance Metrics of PPO, DQN, and DDPG in Continuous Control Tasks. ITM Web Conf. 2025, 78, 01009. [Google Scholar] [CrossRef]
- Sheikh, A.; Chong, E.K.P. Multi-Agent Reinforcement Learning Framework for Optimizing Smart Cities as System of Systems. Syst. Eng. 2026, 29, 3–19. [Google Scholar] [CrossRef]
- Cheng, P.; Chen, Y.; Ding, M.; Chen, Z.; Liu, S.; Chen, Y.-P.P. Deep Reinforcement Learning for Online Resource Allocation in IoT Networks: Technology, Development, and Future Challenges. IEEE Commun. Mag. 2023, 61, 111–117. [Google Scholar] [CrossRef]
- Jiao, L.; Shao, Y.; Sun, L.; Liu, F.; Yang, S.; Ma, W.; Li, L.; Liu, X.; Hou, B.; Zhang, X.; et al. Advanced Deep Learning Models for 6G: Overview, Opportunities, and Challenges. IEEE Access 2024, 12, 133245–133314. [Google Scholar] [CrossRef]
- Ericsson. Mobile Subscriptions Outlook. Available online: https://www.ericsson.com/en/reports-and-papers/mobility-report/dataforecasts/mobile-subscriptions-outlook (accessed on 24 June 2026).
- Markets and Markets 6G Market Size & Outlook, 2030–2036. 2025. Available online: https://www.marketsandmarkets.com/Market-Reports/6g-market-213693378.html (accessed on 24 June 2026).
- Haritwal, S.; Baul, S. 6G Market; NMSC: Boston, MA, USA, 2025. [Google Scholar]
- Latreche, S.; Bellahsene, H. A Comprehensive Survey on 6G: Enabling Technologies, Key Applications, and Future Challenges. Frankl. Open 2026, 15, 100559. [Google Scholar] [CrossRef]
- Fernando, X.; Lăzăroiu, G. Energy-Efficient Industrial Internet of Things in Green 6G Networks. Appl. Sci. 2024, 14, 8558. [Google Scholar] [CrossRef]
- Musaddiq, A.; Olsson, T.; Ahlgren, F. Reinforcement-Learning-Based Routing and Resource Management for Internet of Things Environments: Theoretical Perspective and Challenges. Sensors 2023, 23, 8263. [Google Scholar] [CrossRef] [PubMed]
- Puspitasari, A.A.; Lee, B.M. A Survey on Reinforcement Learning for Reconfigurable Intelligent Surfaces in Wireless Communications. Sensors 2023, 23, 2554. [Google Scholar] [CrossRef] [PubMed]
- Ismail, A.A.; Khalifa, N.E.; El-Khoribi, R.A. A Survey on Resource Scheduling Approaches in Multi-Access Edge Computing Environment: A Deep Reinforcement Learning Study. Clust. Comput. 2025, 28, 184. [Google Scholar] [CrossRef]
- Hady, M.A.; Hu, S.; Pratama, M.; Cao, Z.; Kowalczyk, R. Multi-Agent Reinforcement Learning for Resources Allocation Optimization: A Survey. Artif. Intell. Rev. 2025, 58, 354. [Google Scholar] [CrossRef]
- Amodu, O.A.; Althumali, H.; Mohd Hanapi, Z.; Jarray, C.; Raja Mahmood, R.A.; Adam, M.S.; Bukar, U.A.; Abdullah, N.F.; Luong, N.C. A Comprehensive Survey of Deep Reinforcement Learning in UAV-Assisted IoT Data Collection. Veh. Commun. 2025, 55, 100949. [Google Scholar] [CrossRef]
- Shen, Y.; Xie, L.; Li, M. Intelligent Path Selection Algorithm for Tactical Communication Networks Enhanced by Link State Awareness. Front. Commun. Netw. 2025, 6, 1635982. [Google Scholar] [CrossRef]
- Zhang, R.; Xiong, K.; Lu, Y.; Fan, P.; Ng, D.W.K.; Letaief, K.B. Energy Efficiency Maximization in RIS-Assisted SWIPT Networks With RSMA: A PPO-Based Approach. IEEE J. Sel. Areas Commun. 2023, 41, 1413–1430. [Google Scholar] [CrossRef]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [PubMed]
- Viana, J.; Farkhari, H.; Gil Jiménez, V.P. Securing 5G and Beyond-Enabled UAV Links: Resilience Through Multiagent Learning and Transformers Detection. IEEE Access 2025, 13, 153993–154007. [Google Scholar] [CrossRef]
- Zuo, P.; Miao, C.; Fu, C.; Wang, X.; Liu, X.; Liu, B. SMAPPO: A Security-Aware Multi-Agent Reinforcement Learning Framework for Secure Computation Offloading in SAGIN. J. King Saud Univ. Comput. Inf. Sci. 2025, 37, 336. [Google Scholar] [CrossRef]
- Ericsson. Growth of Mobile Network Data Traffic Persists. Available online: https://www.ericsson.com/en/reports-and-papers/mobility-report/dataforecasts/mobile-traffic-forecast (accessed on 24 June 2026).
- Bokobza, Y.; Dabora, R.; Cohen, K. Deep Reinforcement Learning for Simultaneous Sensing and Channel Access in Cognitive Networks. IEEE Trans. Wirel. Commun. 2023, 22, 4930–4946. [Google Scholar] [CrossRef]
- Bai, W.; Zheng, G.; Xia, W.; Mu, Y.; Xue, Y. Multi-User Opportunistic Spectrum Access for Cognitive Radio Networks Based on Multi-Head Self-Attention and Multi-Agent Deep Reinforcement Learning. Sensors 2025, 25, 2025. [Google Scholar] [CrossRef] [PubMed]
- Hong, S.; Jeong, Y.; Hwang, U.; Hong, S. QIPPO/CA: A Quantized Communication-Efficient MARL Framework for Fully Distributed Channel Access in Next-Generation Wireless Networks. IEEE Internet Things J. 2026, 13, 8615–8627. [Google Scholar] [CrossRef]
- Wang, Q.; Xu, W.; Chen, H.-H. A Heterogeneous-Agent Deep Reinforcement Learning Approach for Dynamic Spectrum Access in Cognitive Wireless Networks. IEEE Trans. Cogn. Commun. Netw. 2025, 12, 2221–2235. [Google Scholar] [CrossRef]
- Wang, T.; Cheng, K.; Du, X. Multi-Agent Independent PPO-Based Automatic ECN Tuning for High-Speed Data Center Networks. In Proceedings of the 2025 IEEE International Conference on Cluster Computing (CLUSTER), Edinburgh, UK, 2–5 September 2025; pp. 1–11. [Google Scholar]
- Lin, G.; Xiao, Y.; Yu, S.; Yu, K.; Liu, J. Constrained Probabilistic Routing with RouterRL: A General Packet-Level Network Simulation Framework. IEEE Trans. Netw. Sci. Eng. 2026, 13, 2604–2622. [Google Scholar] [CrossRef]
- Sefati, S.S.; Haq, A.U.; Nidhi; Craciunescu, R.; Halunga, S.; Mihovska, A.; Fratu, O. A Comprehensive Survey on Resource Management in 6G Network Based on Internet of Things. IEEE Access 2024, 12, 113741–113784. [Google Scholar] [CrossRef]
- Xu, Y.; Zhu, K.; Xu, H.; Ji, J. Deep Reinforcement Learning for Multi-Objective Resource Allocation in Multi-Platoon Cooperative Vehicular Networks. IEEE Trans. Wirel. Commun. 2023, 22, 6185–6198. [Google Scholar] [CrossRef]
- Alsahfi, T.; Badshah, A.; Alsini, R.; Shoie Alallah, F.; Bedewi, W.; Daud, A. Proximal Policy Optimization for Vehicular Big Data Offloading Across Edge, Regional, and Cloud Layers. J. Grid Comput. 2025, 23, 28. [Google Scholar] [CrossRef]
- Rehman, A.U.; Sualiheen, S.; Chang, K. Adaptive Resource Allocation in 6G A2G-TN Integrated System: A Cross-Layer Multi-Agent PPO Approach. Comput. Netw. 2025, 272, 111655. [Google Scholar] [CrossRef]
- Shang, C.; Yu, J.; Thai Hoang, D. Energy-Efficient and Intelligent ISAC in V2X Networks with Spiking Neural Networks-Driven DRL. IEEE Trans. Wirel. Commun. 2026, 25, 1182–1195. [Google Scholar] [CrossRef]
- Kahraman, İ.; Köse, A.; Koca, M.; Anarim, E. Age of Information in Internet of Things: A Survey. IEEE Internet Things J. 2024, 11, 9896–9914. [Google Scholar] [CrossRef]
- Wang, W.; Chen, Q.; Shen, Y.; Xiang, Z. Leakage Identification of Underground Structures Using Classification Deep Neural Networks and Transfer Learning. Sensors 2024, 24, 5569. [Google Scholar] [CrossRef] [PubMed]
- Ayyappan, V.; Bruno, M.A. Applying Machine Learning to Optimize Resource Allocation & Maritime Wireless Mobile Network. J. Wirel. Mob. Netw. Ubiquitous Comput. Dependable Appl. 2025, 16, 406–416. [Google Scholar] [CrossRef]
- Lin, X.; Liu, A.; Han, C.; Liang, X.; Sun, Y.; Ding, G.; Zhou, H. Intelligent Adaptive MIMO Transmission for Nonstationary Communication Environment: A Deep Reinforcement Learning Approach. IEEE Trans. Commun. 2025, 73, 5965–5979. [Google Scholar] [CrossRef]
- Khaskheli, M.B.; Zhao, Y.; Lai, Z. Sustainable Maritime Governance of Digital Technologies for Marine Economic Development and for Managing Challenges in Shipping Risk: Legal Policy and Marine Environmental Management. Sustainability 2025, 17, 9526. [Google Scholar] [CrossRef]
- Mao, Z.; Zhang, Z.; Lu, F.; Liu, X.; Xu, Z.; Pan, Y.; Kang, J.; You, Y. Dynamic Joint Resource Allocation in Maritime Wireless Communication Networks: A Meta-Reinforcement Learning Approach Based on Knowledge Embedding. Front. Inf. Technol. Electron. Eng. 2025, 26, 2672–2687. [Google Scholar] [CrossRef]
- Lin, T.; Du, J.; Zhang, H.; Nallanathan, A.; Wang, J. PPO-Based Energy-Efficient Power Control and Spectrum Allocation in In-Vehicle HetNets. In Proceedings of the GLOBECOM 2023—2023 IEEE Global Communications Conference, Kuala Lumpur, Malaysia, 4–8 December 2023; pp. 6334–6339. [Google Scholar]
- Ma, W.; Lin, B.; Pan, H.; Sun, G.; Shi, E.; An, J.; Yuen, C. SIM-Assisted Secure Mobile Communications via Enhanced Proximal Policy Optimization Algorithm. IEEE Trans. Wirel. Commun. 2026, 25, 11964–11979. [Google Scholar] [CrossRef]
- Fang, J.; Wang, X.; Liu, Y.; Tang, H.; Li, X. Multi-Agent Collaborative Inference Optimization for Large-Scale DNNs in IoT Edge Systems. IEEE Internet Things J. 2026, 13, 24938–24953. [Google Scholar] [CrossRef]
- Ferrag, M.A.; Friha, O.; Kantarci, B.; Tihanyi, N.; Cordeiro, L.; Debbah, M.; Hamouda, D.; Al-Hawawreh, M.; Choo, K.-K.R. Edge Learning for 6G-Enabled Internet of Things: A Comprehensive Survey of Vulnerabilities, Datasets, and Defenses. IEEE Commun. Surv. Tutor. 2023, 25, 2654–2713. [Google Scholar] [CrossRef]
- Mustafa, E.; Shuja, J.; Rehman, F.; Namoun, A.; Bilal, M.; Iqbal, A. Computation Offloading in Vehicular Communications Using PPO-Based Deep Reinforcement Learning. J. Supercomput. 2025, 81, 547. [Google Scholar] [CrossRef]
- Soni, L.; Taneja, A.; Alqahtani, N.; Alqahtani, J. Robust ISAC Based Framework for Location Estimation and Target Detection in 6G Networks. PLoS ONE 2026, 21, e0337050. [Google Scholar] [CrossRef] [PubMed]
- Jabeen, N.; Lei, H.; Muhammad, A.; Ali, A.; Khan, Z.U.; Pan, G. Localization in ISAC: A Review. IEEE Internet Things J. 2025, 12, 46526–46552. [Google Scholar] [CrossRef]
- Wu, K.; Wang, Z.; Chen, S.-L.; Zhang, J.A.; Guo, Y.J. ISAC: From Human to Environmental Sensing. IEEE J. Sel. Top. Electromagn. Antennas Propag. 2025, 1, 84–98. [Google Scholar] [CrossRef]
- Wang, K.; Sun, Y.; Liu, P.; Zhang, Y.; Shao, Z. Energy-Efficient Deep Reinforcement Learning RSMA in Multi-UAV-Assisted Wireless-Powered Communication Network. IEEE Trans. Netw. Sci. Eng. 2026, 13, 2420–2438. [Google Scholar] [CrossRef]
- Guan, W.; Zhang, H. Introduction. In Network Slicing for Future Wireless Communication; Wireless Networks; Springer Nature: Cham, Switzerland, 2024; pp. 1–12. [Google Scholar]
- Donatti, A.W.; Cristina Machado, M.; Alexander Lopez Martinez, M.; Rogério Antunes, S.S.; Carlos Figueiredo Souza, E.; Corrêa, S.L.; Ferreto, T.C.; Augusto Suruagy, J.; Martins, J.S.B.; Cristina Carvalho, T. Energy Efficiency in Network Slicing: Survey and Taxonomy. IEEE Access 2025, 13, 134570–134589. [Google Scholar] [CrossRef]
- Dubey, M.; Singh, A.K.; Mishra, R. AI Based Resource Management for 5G Network Slicing: History, Use Cases, and Research Directions. Concurr. Comput. Pract. Exp. 2025, 37, e8327. [Google Scholar] [CrossRef]
- Mahmood, A.; Abdallah, A.M.; Baharom, B.B.; Habilah, A.S.K. Ensuring Satellite Operational Integrity: A Power Budget Analysis for Next Generation Satellites. IEEE Access 2025, 13, 44901–44911. [Google Scholar] [CrossRef]
- Chen, R.; Long, W.-X.; Wang, B.; He, Y.; Sun, R.; Cheng, N.; Zheng, G.; Niyato, D. Multibeam High Throughput Satellite: Hardware Foundation, Resource Allocation, and Precoding. IEEE Commun. Surv. Tutor. 2026, 28, 5379–5415. [Google Scholar] [CrossRef]
- Zhao, D.; Wang, Y.; Song, B.; Zhou, Y.; Qin, P. Learning When and Where to Handover: A Hierarchical Reinforcement Learning Framework for Dense LEO Satellite Constellations. IEEE Trans. Wirel. Commun. 2026, 25, 12787–12801. [Google Scholar] [CrossRef]
- Marzuk, F.; Vejar, A.; Chołda, P. Deep Reinforcement Learning for Energy-Efficient 6G V2X Networks. Electronics 2025, 14, 1148. [Google Scholar] [CrossRef]
- Li, Y.; Chang, Y.; Fukawa, K.; Kodama, N. Reinforcement Learning-Based Cognitive Radio Transmission Scheduling in Vehicular Systems. In Proceedings of the 2023 IEEE 97th Vehicular Technology Conference (VTC2023-Spring), Florence, Italy, 20–23 June 2023; pp. 1–5. [Google Scholar]
- Su, W.; Liu, H.; Li, T.; Lv, X.; Rui, H.; Huang, W.; Wang, Z.; Li, Y. Jointly Optimizing Deployment and Antenna of Base Stations Using Hierarchical Reinforcement Learning. ACM Trans. Knowl. Discov. Data 2026, 20, 1–25. [Google Scholar] [CrossRef]
- Sun, K.; Yang, J.; Li, J.; Yang, B.; Ding, S. Proximal Policy Optimization-Based Hierarchical Decision-Making Mechanism for Resource Allocation Optimization in UAV Networks. Electronics 2025, 14, 747. [Google Scholar] [CrossRef]
- Li, C.; Tan, X.; Chen, C. Deep Reinforcement Learning–Driven Multi-Satellite Collaborative Observation Planning for Emergency and Disaster Monitoring. Int. J. Digit. Earth 2025, 18, 2554310. [Google Scholar] [CrossRef]
- He, Z.; Luo, Y.; Shojafar, M.; Mi, D. Heterogeneous-Agent PPO RL for xApps Coordination in Digital Twin Enabled O-RAN. In Proceedings of the 2025 IEEE/CIC International Conference on Communications in China (ICCC), Shanghai, China, 10–13 August 2025; pp. 1–6. [Google Scholar]
- Qazzaz, M.M.H.; Salama, A.; Hafeez, M.; Zaidi, S.A.R. OREO: Open RAN Energy Optimization via Deep Reinforcement Learning for 6G Networks. IEEE Open J. Commun. Soc. 2026, 7, 4165–4182. [Google Scholar] [CrossRef]
- Asemian, G.; Amini, M.; Kantarci, B. Anti-Jamming Task Scheduling in MEC-O-RAN With Hierarchical DRL and Transformer-Based Control. IEEE Internet Things J. 2026, 13, 7714–7729. [Google Scholar] [CrossRef]
- Di, Z.; Zhong, Z.; Pengfei, Q.; Hao, Q.; Bin, S. Resource Allocation in Multi-User Cellular Networks: A Transformer-Based Deep Reinforcement Learning Approach. China Commun. 2024, 21, 77–96. [Google Scholar] [CrossRef]
- Lee, T.H.; Kim, J. Hybrid PPO–DQN for Multi-Objective Adaptive Cruise Control in Eco-Driving: Reward Shaping Toward Safety and Sustainability (Student Abstract). In Proceedings of the AAAI Conference on Artificial Intelligence, Singapore, 20–27 January 2026. [Google Scholar]
- Bereketeab, L.; Zekeria, A.; Aloqaily, M.; Guizani, M.; Debbah, M. Energy Optimization in Sustainable Smart Environments with Machine Learning and Advanced Communications. IEEE Sens. J. 2024, 24, 5704–5712. [Google Scholar] [CrossRef]
- Song, J.; Gao, Y.; Wu, D.; Zhou, L. Adaptive Live Tactile Streaming with Scalable Coding for Immersive Communications. IEEE Trans. Mob. Comput. 2026, 25, 5133–5145. [Google Scholar] [CrossRef]
- Jihad, M.; Al Fahad, A.; Roy, P.; Razzaque, M.A.; Alelaiwi, A.; Hassan, M.R.; Hassan, M.M. Quality of Experience Aware Task Execution in Digital Twinning Vehicular Edge Computing: A Framework and A3C Algorithm. Future Gener. Comput. Syst. 2026, 176, 108144. [Google Scholar] [CrossRef]
- Luo, B.; Lu, X.; Lu, W.; Han, H.; Huang, G.; Zhang, Y. Neural Adaptive Video Streaming via Imitation Learning and Reinforcement Learning. In Proceedings of the 2024 IEEE 10th International Symposium on Microwave, Antenna, Propagation and EMC Technologies for Wireless Communications (MAPE), Guangzhou, China, 27–30 November 2024; pp. 1–4. [Google Scholar]
- Iqbal, A.; Al-Habashna, A.; Wainer, G.; Boudreau, G. Sum Rate Maximization in RIS-Assisted Multi-User MISO Systems: A Proximal Policy Optimization-Based Approach. Phys. Commun. 2026, 74, 102961. [Google Scholar] [CrossRef]
- Wara, N.; Paul, A.; Singh, K.; Kaushik, A.; Shin, W. Multi-Agent PPO-Based Resource Optimization for Full-Duplex RIS-Aided NOMA-ISAC Systems. IEEE Open J. Commun. Soc. 2025, 6, 9802–9820. [Google Scholar] [CrossRef]
- Hu, J.; Ju, Y.; Wang, H.; Liu, L.; Pei, Q.; Guo, Y.; Wu, C. Multi-RIS Intelligent Collaboration Empowered Secure MmWave D2D Communication. In Proceedings of the GLOBECOM 2024—2024 IEEE Global Communications Conference, Cape Town, South Africa, 8–12 December 2024; pp. 3243–3248. [Google Scholar]
- Iqbal, A.; Al-Habashna, A.; Wainer, G.; Boudreau, G.; Bouali, F. PPO-Based Energy Efficiency Maximization For RIS-Assisted Multi-User Miso Systems. In Proceedings of the 2024 IEEE 100th Vehicular Technology Conference (VTC2024-Fall), Washington, DC, USA, 7–10 October 2024; pp. 1–6. [Google Scholar]
- Wu, Y.; Huang, Y.; Wang, Z.; Xu, C. Joint Caching and Computation in UAV-Assisted Vehicle Networks via Multi-Agent Deep Reinforcement Learning. Drones 2025, 9, 456. [Google Scholar] [CrossRef]
- O’Connell, E.; O’Brien, W.; Bhattacharya, M.; Moore, D.; Penica, M. Digital Twins: Enabling Interoperability in Smart Manufacturing Networks. Telecom 2023, 4, 265–278. [Google Scholar] [CrossRef]
- Yang, B.; Wu, B.; You, Y.; Guo, C.; Qiao, L.; Lv, Z. Edge Intelligence Based Digital Twins for Internet of Autonomous Unmanned Vehicles. Softw. Pract. Exp. 2024, 54, 1833–1851. [Google Scholar] [CrossRef]
- Zhang, P.; Li, G. A Cooperative Dynamic Target Search Approach for Multi-UAV Systems Utilizing the MAPPO Algorithm. Discov. Artif. Intell. 2025, 5, 153. [Google Scholar] [CrossRef]
- Xie, W.; Yang, H.; Xiong, Z. Resource Allocation for UAV-Assisted Anti-Jamming Semantic D2D Networks: A Graph Reinforcement Learning Approach. Comput. Netw. 2025, 269, 111463. [Google Scholar] [CrossRef]
- He, H.; Yang, X.; Mi, X.; Shen, H.; Liao, X. Multi-Agent Deep Reinforcement Learning Based Dynamic Task Offloading in a Device-to-Device Mobile-Edge Computing Network to Minimize Average Task Delay with Deadline Constraints. Sensors 2024, 24, 5141. [Google Scholar] [CrossRef] [PubMed]
- Hu, F.; Fu, Q.; Zhang, S.; Huang, J. A Multi-Agent Deep Reinforcement Learning-Based Task Offloading Method for 6G-Enabled Internet of Vehicles with Cloud-Edge-Device Collaboration. Comput. Mater. Contin. 2026, 87, 1. [Google Scholar] [CrossRef]
- Hu, Z.; Liu, X.; Guo, M.; Liu, C. PPO-Based Joint Task Offloading and Resource Allocation for Vehicular Edge Computing Via V2I and V2V Communications. In Proceedings of the 2025 13th International Conference on Intelligent Computing and Wireless Optical Communications (ICWOC), Chengdu, China, 28–29 June 2025; pp. 316–321. [Google Scholar]
- Singh, R.; Kaushik, A.; Shin, W.; Renzo, M.D.; Sciancalepore, V.; Lee, D.; Sasaki, H.; Shojaeifard, A.; Dobre, O.A. Toward 6G Evolution: Three Enhancements, Three Innovations, and Three Major Challenges. IEEE Netw. 2025, 39, 139–147. [Google Scholar] [CrossRef]
- Mohammed, S.A.; Murad, S.S.; Albeyboni, H.J.; Soltani, M.D.; Ahmed, R.A.; Badeel, R.; Chen, P. Supporting Global Communications of 6G Networks Using AI, Digital Twin, Hybrid and Integrated Networks, and Cloud: Features, Challenges, and Recommendations. Telecom 2025, 6, 35. [Google Scholar] [CrossRef]
- Luo, X.; Lin, Q.; Zhang, R.; Chen, H.-H.; Wang, X.; Huang, M. ISAC—A Survey on Its Layered Architecture, Technologies, Standardizations, Prototypes, and Testbeds. IEEE Commun. Surv. Tutor. 2026, 28, 485–526. [Google Scholar] [CrossRef]
- Mata, L.; Sousa, M.; Vieira, P.; Queluz, M.P.; Rodrigues, A. Optimizing Energy and Spectral Efficiency in Mobile Networks: A Comprehensive Energy Sustainability Framework for Network Operators. IEEE Access 2025, 13, 22342–22364. [Google Scholar] [CrossRef]
- Shen, L.; Li, B.; Zhu, X. Robust Offloading for Edge Computing-Assisted Sensing and Communication Systems: A Deep Reinforcement Learning Approach. Sensors 2024, 24, 2489. [Google Scholar] [CrossRef] [PubMed]
- Sun, A.; Wu, P.; Jin, S.; Jiao, F. Deep Reinforcement Learning-Driven Multidomain Resource Allocation for Integrated Sensing, Communication, and Computing. In Proceedings of the Eighth International Conference on Artificial Intelligence and Pattern Recognition (AIPR 2025), Quanzhou, China, 19–21 December 2025; Tian, H., Ed.; SPIE: Bellingham, WA, USA, 2025; p. 239. [Google Scholar]
- Zhou, Y.; Feng, Z.; Wei, Z.; Ma, D.; Huang, D.; Meng, Z.; Fan, Y.; Xu, J.; Zhang, P. Integrated Sensing, Communication, and Control for Multi-Agent Networked Formation Control. Sci. China Inf. Sci. 2026, 69, 142301. [Google Scholar] [CrossRef]
- Kastwar, N.; Gupta, A. Drone Taxis and Drone Deliveries in Smart Logistics Security Surveillance: An Examination of the Drone Revolution from Legal and Ethical Dimensions. In Innovative Strategies in Aviation Management and Marketing; Sousa, B.B., Marques, M.I., Arantes, L., O’Neill, A., Eds.; IGI Global Scientific Publishing: Palmdale, PA, USA, 2025; pp. 149–180. [Google Scholar]
- Manda, V.K.; Christy, V.; Hlali, A. Current Trends, Opportunities, and Futures Research Directions in Geospatial Technologies for Smart Cities. In Advances in Geospatial Technologies; Darwish, D., Chemingui, H., Eds.; IGI Global: Hershey, PA, USA, 2024; pp. 239–270. [Google Scholar]
- Ghomri, B.I.-D.; Bendimerad, M.Y.; Bendimerad, F.T. DRL-Driven Optimization for Energy Efficiency and Fairness in NOMA-UAV Networks. IEEE Commun. Lett. 2024, 28, 1048–1052. [Google Scholar] [CrossRef]
- Aung, P.S.; Nguyen, L.X.; Tun, Y.K.; Han, Z.; Hong, C.S. Aerial STAR-RIS Empowered MEC: A DRL Approach for Energy Minimization. IEEE Wirel. Commun. Lett. 2024, 13, 1409–1413. [Google Scholar] [CrossRef]
- Wang, Y.; Cui, Y.; Yang, Y.; Li, Z.; Cui, X. Multi-UAV Path Planning for Air-Ground Relay Communication Based on Mix-Greedy MAPPO Algorithm. Drones 2024, 8, 706. [Google Scholar] [CrossRef]
- Chaudhary, S.; Budhiraja, I.; Chaudhary, R.; Garg, S.; Choi, B.J.; Alrashoud, M. Proximal Policy Optimization Based Sum Rate Maximization Scheme for STAR-RIS-Assisted Vehicular Networks Underlaying UAV. Alex. Eng. J. 2025, 118, 700–710. [Google Scholar] [CrossRef]
- Zhou, W.; Yi, M.; Zhang, Y.; Wang, X.; Liu, J. Satellite-Assisted UAV Data Collection for Information Freshness in IoRT Networks. In Proceedings of the 2024 IEEE Wireless Communications and Networking Conference (WCNC), Dubai, United Arab Emirates, 21–24 April 2024; pp. 1–6. [Google Scholar]
- Boufakhreddine, Z.; Nohra, A.; Haidar, G.A.; Achkar, R.; Owayjan, M. Exploring the Potential of AI in Network Slicing for 5G Networks: An Optimisation Framework. IET Commun. 2025, 19, e70116. [Google Scholar] [CrossRef]
- Nguyen, T.T.H.; Kim, T. Proximal Policy Optimization for Up/Downlink Time Slots Allocation in 5/6G Dynamic TDD Networks. KSII Trans. Internet Inf. Syst. 2025, 19, 259–278. [Google Scholar] [CrossRef]
- Adhikari, B.; Shaharyar Khwaja, A.; Jaseemuddin, M.; Anpalagan, A. DRL-Leveraged and RIS-Assisted Hybrid Network Slicing for eMBB and URLLC Co-Existence in 6G Systems. IEEE Open J. Commun. Soc. 2025, 6, 6156–6176. [Google Scholar] [CrossRef]
- Raja, G.; Sanjeev, A.; Ravishankar, K.; Arunachalam, K. TRIP-B5G: Traffic Classification and Resource Allocation Using Intelligent PPO in B5G O-RAN. In Proceedings of the 2026 IEEE 23rd Consumer Communications & Networking Conference (CCNC), Las Vegas, NV, USA, 9–12 January 2026; pp. 1–4. [Google Scholar]
- Zhang, C.; Wu, C.; Lin, M.; Lin, Y.; Liu, W. Proximal Policy Optimization for Efficient D2D-Assisted Computation Offloading and Resource Allocation in Multi-Access Edge Computing. Future Internet 2024, 16, 19. [Google Scholar] [CrossRef]
- Hikmat, F.A.; Sahib, M.A. PPO-Based Deep Reinforcement Learning Framework for Dynamic Resource Allocation and Network Slicing in 5G Mobile Networks. Int. J. Electron. Telecommun. 2026, 72, 1–9. [Google Scholar] [CrossRef]
- Cui, Z.; Qamar, F.; Kazmi, S.H.A.; Zainol Ariffin, K.A.; Safdar, G.A.; Ur Rehman, M.H. A Review of Multi-Agent Deep Reinforcement Learning for Resource Allocation in beyond 5G Network Slicing: Solutions, Challenges and Future Research Directions. PeerJ Comput. Sci. 2026, 12, e3728. [Google Scholar] [CrossRef]
- Yang, L.; Bi, Z.; Wang, Z.; Liang, X.; Zhang, J.; Wu, R. Resource Allocation for SFC Networks: A Deep Reinforcement Learning Approach. In Proceedings of the 2024 7th World Conference on Computing and Communication Technologies (WCCCT), Chengdu, China, 12–14 April 2024; pp. 210–215. [Google Scholar]
- Zhang, Y.; Joe, I. The Optimized Deployment of Service Function Chain Based on Deep Reinforcement Learning Algorithm. In Proceedings of the International Conference on Machine Learning, Pattern Recognition and Automation Engineering, Singapore, 7–9 August 2024; ACM: New York, NY, USA, 2024; pp. 24–28. [Google Scholar]
- Fu, S.; Wei, W.; Feng, X.; Yin, L. Average Sum Rate Optimization in RIS-Assisted NOMA Satellite Network: A Deep Reinforcement Learning Approach. IEEE Wirel. Commun. Lett. 2025, 14, 1772–1776. [Google Scholar] [CrossRef]
- Zhang, H.; Han, X.; Xing, C.; Chen, H.; Zhao, J. Analysis of Uplink Transmission Scheduling Strategies for LoRa-Based Direct-to-Satellite IoT Networks Using Deep Reinforcement Learning. IEEE Trans. Green Commun. Netw. 2026, 10, 1279–1292. [Google Scholar] [CrossRef]
- Xu, J.; Zhao, Z.; Wang, L.; Zhang, Y. A Novel Deep Reinforcement Learning Architecture for Dynamic Power and Bandwidth Allocation in Multibeam Satellites. Acta Astronaut. 2023, 204, 73–82. [Google Scholar] [CrossRef]
- Li, S.; Wu, Q.; Wang, R. Efficient Packet Routing in Ultra-Dense LEO Satellite Networks via Cooperative-MARL with Queuing Theory Model. In Proceedings of the 2025 IEEE Wireless Communications and Networking Conference (WCNC), Milan, Italy, 24–27 March 2025; pp. 1–6. [Google Scholar]
- Zhang, Q.; Fu, S.; Yang, Z. Jointly Optimizing Satellite Handover and Power Allocation in LEO Satellite Network: A Dual-Agent Framework. IEEE Trans. Veh. Technol. 2026, 1–6. [Google Scholar] [CrossRef]
- Sadiki, A.; Bentahar, J.; Dssouli, R.; En-Nouaary, A.; Otrok, H. Deep Reinforcement Learning for the Computation Offloading in MIMO-Based Edge Computing. Ad Hoc Netw. 2023, 141, 103080. [Google Scholar] [CrossRef]
- Li, Z.; Tian, J.; Zhang, H.; Shi, T.; Xu, B.; Zhou, T. A Multi-Agent Proximal Policy Optimization-Based Handover Scheme for Satellite-Terrestrial Integrated Networks. IEEE Wirel. Commun. Lett. 2026, 15, 2428–2432. [Google Scholar] [CrossRef]
- Zhao, R.; Cai, J.; Luo, J.; Ran, Y.; Gao, J.; Xu, Y. Joint Beam Hopping and Resource Allocation for Load Balancing and Interference Avoidance in Multi-LEO Satellite Networks. In Proceedings of the ICC 2025—IEEE International Conference on Communications, Montreal, QC, Canada, 8–12 June 2025; pp. 958–963. [Google Scholar]
- Meng, M.; Hu, B.; Chen, S.; Kang, S. Joint Beamforming and Dynamic Beam Hopping Based on MAPPO for LEO Satellite Communication System. IEEE Wirel. Commun. Lett. 2025, 14, 1461–1465. [Google Scholar] [CrossRef]
- Alharbi, S. A Review of Deep Multi-Objective Reinforcement Learning and Vision-Based Systems for Smart Cities. Informatica 2025, 49. [Google Scholar] [CrossRef]
- Louati, A.; Louati, H.; Kariri, E.; Neifar, W.; Hassan, M.K.; Khairi, M.H.H.; Farahat, M.A.; El-Hoseny, H.M. Sustainable Smart Cities through Multi-Agent Reinforcement Learning-Based Cooperative Autonomous Vehicles. Sustainability 2024, 16, 1779. [Google Scholar] [CrossRef]
- Wang, J.; Wang, R.; Zheng, Z.; Lin, R.; Wu, L.; Shu, F. Physical Layer Security Enhancement in AAV-Assisted Cooperative Jamming for Cognitive Radio Networks: A MAPPO-LSTM Deep Reinforcement Learning Approach. IEEE Trans. Veh. Technol. 2025, 74, 4713–4727. [Google Scholar] [CrossRef]
- Chen, M.; Chen, X.; Wang, R.; Ding, H. Reactive Jamming Resilient Power Allocation in Cognitive Radio Networks via Deep Reinforcement Learning. In Intelligent Networked Things; Zhang, L., Yu, W., Laili, Y., Qu, T., Eds.; Communications in Computer and Information Science; Springer Nature: Singapore, 2026; Volume 2624, pp. 327–335. [Google Scholar]
- Zhang, F.; Niu, Y.; Zhou, W. Intelligent Anti-Jamming Decision Algorithm for Wireless Communication Based on MAPPO. Electronics 2025, 14, 462. [Google Scholar] [CrossRef]
- Ma, H.; You, J.; Wu, H.; Xing, L.; Zhang, X. A Probabilistic Routing Algorithm Based on CNN and Q-Learning for Vehicular Edge Network. Trans. Emerg. Telecommun. Technol. 2025, 36, e70050. [Google Scholar] [CrossRef]
- Alvarado-Padilla, J.J.; Celaya-Padilla, J.M.; Martinez-Torteya, A.; Soto-Murillo, M.A.; Gamboa-Rosales, H.; Gamboa-Rosales, N.K. Optimizing Autonomous Vehicle Control Through Deep Q-Learning: A Simulation-Based Approach with CARLA. In Advanced Research in Technologies, Information, Innovation and Sustainability; Guarda, T., Portela, F., Augusto, M.F., Eds.; Communications in Computer and Information Science; Springer Nature: Cham, Switzerland, 2025; Volume 2348, pp. 141–154. [Google Scholar]
- Pimenow, S.; Pimenowa, O.; Prus, P. Challenges of Artificial Intelligence Development in the Context of Energy Consumption and Impact on Climate Change. Energies 2024, 17, 5965. [Google Scholar] [CrossRef]
- Samaniego, J.F. From Hyperconnectivity to AI: How Can We Tackle the Environmental Impact of the New Internet Era? Available online: https://www.uoc.edu/en/news/2023/080-sustainable-digitalization (accessed on 24 June 2026).
- Zhen, Y.; Tao, L.; Wu, D.; Tang, T.; Wang, R. Energy-Saving Control Strategy for Ultra-Dense Network Base Stations Based on Multi-Agent Reinforcement Learning. Digit. Commun. Netw. 2025, 11, 1007–1017. [Google Scholar] [CrossRef]
- Wang, J.; Li, Y.; Hong, Y.; Tang, Y. Integrated Adaptive Communication in Multi-Agent Systems: Dynamic Topology, Frequency, and Content Optimization for Efficient Collaboration. Neurocomputing 2025, 617, 129068. [Google Scholar] [CrossRef]
- Li, S.; Fang, B. AI Development Strategies in Countries Around the World. In Artificial Intelligence Security and Safety; Fang, B., Ed.; Springer Nature: Singapore, 2025; pp. 51–84. [Google Scholar]
- García-Pineda, V.; Valencia-Arias, A.; Patiño-Vanegas, J.C.; Flores Cueto, J.J.; Arango-Botero, D.; Rojas Coronel, A.M.; Rodríguez-Correa, P.A. Research Trends in the Use of Machine Learning Applied in Mobile Networks: A Bibliometric Approach and Research Agenda. Informatics 2023, 10, 73. [Google Scholar] [CrossRef]
- Espinosa, A.; Samos, X.; Ulied, D.; Marias, J.; Touma, R. Optimizing Energy Consumption of Edge-Cloud Environments: A Comparative Study Between PPO and PSO. Int. J. Comput. Intell. Syst. 2025, 19, 16. [Google Scholar] [CrossRef]
- Wu, Z.; Fang, H.; Tang, J.; Yang, X. Lightweight Adaptive PPO-AHP Enhanced Algorithm for Task Offloading in Vehicular Edge Computing. In Proceedings of the 2025 International Joint Conference on Neural Networks (IJCNN), Rome, Italy, 30 June–5 July 2025; pp. 1–9. [Google Scholar]
- Li, S.; Huang, Q.; Chen, H.; Jiang, R.; Dong, M.; Ota, K.; Quek, T.Q.S. DRL-Based Joint Task Offloading and Resource Allocation in Air-Ground Integrated Vehicular Edge Computing Network with Energy Harvesting. IEEE Trans. Veh. Technol. 2025, 75, 6658–6672. [Google Scholar] [CrossRef]
- Wu, C.-M.; Guan, S.-Z.; Yang, C.-C.; Lin, K.-T.; Kuang, M.-Y. Scalable and Interference-Aware Spectrum Access in IoT Networks via Decentralized Proximal Policy Optimization. Comput. Netw. 2026, 275, 111885. [Google Scholar] [CrossRef]
- Nzewi, O.I. Adaptive Governance for Resilient Local Service Delivery. J. Local Gov. Res. Innov. 2025, 6, a322. [Google Scholar] [CrossRef]
- Hu, Y.; Cong, R.; Matsumoto, T.; Li, Y. Environmental and Economic Impacts of V2X Applications in Electric Vehicles: A Long-Term Perspective for China. Energies 2025, 18, 3636. [Google Scholar] [CrossRef]
- Liu, X.; Tan, J.; Ren, X.; Dai, H. Self-Attention Proximal Policy Optimization for Beam Tracking in mmWave Communications. IEEE J. Sel. Areas Commun. 2026, 44, 4552–4569. [Google Scholar] [CrossRef]
- Liu, C.; Zhao, J.; Li, J.; Wang, D.; Yu, F.R. UAV Aided Integrated Sensing, Communication and Computing: Optimization via Federated Learning. IEEE Trans. Veh. Technol. 2025, 75, 6045–6058. [Google Scholar] [CrossRef]
- Hu, H.; Tang, H.; Zhang, R.; Jiang, F.; Ding, Z.; Niyato, D. Joint Scheduling and Power Control in AoI-Oriented WP-IoT Networks: An HVF-Based PPO Approach. IEEE Trans. Veh. Technol. 2025, 75, 6876–6881. [Google Scholar] [CrossRef]
- Chen, L.; Wu, W.; Tian, F. Efficient Resource Allocation for RSMA-Based Semantic Image Transmission with Shared Knowledge Base. IEEE Wirel. Commun. Lett. 2026, 15, 2154–2158. [Google Scholar] [CrossRef]


| Review | PPO Specific | Multi-Domain Coverage | Sustainability Analysis | PPO Variant Comparison | 6G/B5G Focus | Cross-Domain Taxonomy | Key Limitation |
|---|---|---|---|---|---|---|---|
| Cheng et al. (2023)—DRL for Wireless Resource Allocation [15] | ✗ | Partial | ✗ | ✗ | Partial | ✗ | Primarily focused on resource allocation and broad DRL methods. |
| Jiao et al. (2024)—AI-Enabled 6G Networking Review [16] | ✗ | ✓ | Partial | ✗ | ✓ | ✗ | PPO was discussed only briefly among multiple AI techniques |
| Musaddiq et al. (2023)—RL-Based Routing and Spectrum Management Survey [22] | ✗ | ✗ | ✗ | ✗ | Partial | ✗ | Limited to routing and spectrum-access applications |
| Puspitasari and Lee (2023)—DRL for RIS and Beamforming Optimization Review [23] | ✗ | ✗ | Partial | ✗ | ✓ | ✗ | Domain-specific focus on RIS and beamforming |
| Ismail et al. (2025)-—–MEC and Computation Offloading Using DRL Survey [24] | ✗ | ✗ | ✗ | ✗ | Partial | ✗ | Focused exclusively on MEC and offloading systems |
| Hady et al. (2025)—Multi-Agent Reinforcement Learning for Wireless Networks Review [25] | Partial | Partial | ✗ | Partial | ✓ | ✗ | Limited comparison of PPO variants across domains |
| Amodu et al. (2025)—RL-Based UAV Communication Systems Survey [26] | ✗ | ✗ | Partial | ✗ | Partial | ✗ | Restricted to UAV-assisted communication networks |
| PPO in Intelligent Communication Systems (This Review) | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | Provides integrated cross-domain synthesis of PPO architectures and deployment challenges |
| Feature | Benefit in Communications |
|---|---|
| Clipped Objective Function | Prevents large, destructive policy updates, ensuring stable learning even in volatile wireless environments [12,13]. |
| Continuous Action Spaces | Unlike DQN, PPO can directly optimize continuous variables, such as transmit power and phase shifts [11,12]. |
| Sample Efficiency | Allows for multiple epochs of updates on the same batch of data, reducing the need for massive real-time datasets [12,13]. |
| Ease of Tuning | Generally, requires less hyperparameter tuning than algorithms like DDPG or SAC, making it more practical for deployment [13]. |
| Energy Efficiency | PPO-based optimization improved energy efficiency by 15.8% over DDPG in RIS-assisted MU-MISO systems [9]. |
| Algorithm | Key Strength | Limitation | Typical Communication Use |
|---|---|---|---|
| SAC | High sample efficiency, entropy-driven exploration, and strong performance in continuous control | More hyperparameter tuning, higher implementation complexity | Power control, beamforming, dynamic spectrum management, UAV communications |
| TD3 | Mitigates overestimation bias, improved stability over DDPG | Sensitive to parameter selection, increased architectural complexity | Resource allocation, RIS optimization, continuous control problems |
| DDPG | Efficient continuous-action optimization, lower interaction cost | Training instability, exploration challenges, prone to local optima | Power allocation, computation offloading, and beamforming |
| PPO | Stable training, easy implementation, strong continuous-action support, robust convergence | Lower sample efficiency than off-policy methods | Resource allocation, beamforming, network slicing, MEC, V2X, satellite communications |
| TRPO | Strong theoretical guarantees and monotonic policy improvement | Computationally expensive, difficult implementation due to constrained optimization | Early wireless resource management and continuous-control optimization |
| A3C | Parallel learning, reduced training time, simple architecture | Less stable than PPO, lower sample efficiency | Routing, congestion control, adaptive network management |
| Communication Domain | Typical PPO Variant | Key Mechanism/Outcome |
|---|---|---|
| Resource Allocation | PPO, MAPPO, PPO-Transformer | Reduces unnecessary power transmission and enables greener network operation [35,50]. |
| Beamforming & RIS | MAPPO, IPPO | Hardware-efficient signal enhancement lowers overall energy requirements [9,28,51]. |
| MEC & Offloading | PPO, DPPO | Balances local energy use with transmission latency for improved trade-offs [52,53,54]. |
| Vehicular Edge Computing & Task Offloading | PPO, MAPPO | Joint optimization of computation offloading, resource allocation, latency reduction, and edge-resource utilization |
| Integrated Sensing (ISAC) | PPO, MAPPO, PPO-DQN | Reduces the need for dedicated sensing hardware in smart city infrastructure [55,56,57]. |
| UAV Communications | PPO, Dual-Clip MAPPO | Energy management extends flight time and balances secrecy with efficiency [51,58]. |
| Network Slicing | PPO, MAPPO | Enables adaptive multi-service resource orchestration [59,60,61]. |
| Satellite Communications | PPO, PPO-DQN Hybrid | Reduces the energy cost of providing connectivity to remote regions [62,63,64]. |
| V2X & Cognitive Radio | PPO, MAPPO | Reduces idle time, device battery drain, and overall transportation congestion [40,65,66]. |
| Comparison | Communication Domain | Metric/Outcome | Representative Result |
|---|---|---|---|
| PPO vs. DQN | Vehicular edge computing | Total delay reduction | PPO reduced total delay by 13.85% compared with DQN under dynamic vehicle speeds and task deadlines [54]. |
| PPO vs. DDQN | Vehicular edge computing | Total delay reduction | PPO reduced total delay by 11.24% compared with DDQN [54]. |
| PPO vs. DDPG | RIS-assisted MU-MISO beamforming | Energy efficiency | PPO improved energy efficiency by 15.8% relative to DDPG [9]. |
| PPO vs. Fractional Programming (FP) | RIS-assisted MU-MISO beamforming | Energy efficiency | PPO improved energy efficiency by 34.2% relative to FP [9]. |
| PPO vs. Heuristic Algorithms | Heterogeneous networks (HetNets) | Throughput and user adaptation | PPO outperformed heuristic resource-allocation methods under varying user loads and channel conditions. |
| PPO vs. Greedy Algorithms | Network slicing and MEC | Resource efficiency/QoS | PPO-based slicing achieved better resource utilization and QoS satisfaction than greedy approaches [111]. |
| PPO vs. Genetic Algorithm | GEO satellite power allocation | Runtime/unmet system demand | PPO achieved comparable unmet-demand performance while operating faster than an optimized genetic algorithm [116]. |
| PPO vs. ACER | O-RAN resource allocation | Energy–latency trade-off | PPO achieved a more favorable balance between energy consumption and user latency than ACER. |
| MAPPO vs. IPPO | 5G resource allocation/routing | Convergence speed, packet loss, and latency | MAPPO converged faster and achieved lower latency and packet loss than IPPO [30]. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Manda, V.K.; Madhu, B.; Tarnanidis, T. Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review. Future Internet 2026, 18, 340. https://doi.org/10.3390/fi18070340
Manda VK, Madhu B, Tarnanidis T. Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review. Future Internet. 2026; 18(7):340. https://doi.org/10.3390/fi18070340
Chicago/Turabian StyleManda, Vijaya Kittu, Bhukya Madhu, and Theodore Tarnanidis. 2026. "Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review" Future Internet 18, no. 7: 340. https://doi.org/10.3390/fi18070340
APA StyleManda, V. K., Madhu, B., & Tarnanidis, T. (2026). Proximal Policy Optimization in 5G, B5G, and 6G Communication Systems: A Systematic Review. Future Internet, 18(7), 340. https://doi.org/10.3390/fi18070340

