A Load-Balancing-Aware Learning Framework for Collaborative UAV-MEC Computation Offloading
Abstract
1. Introduction
2. Related Work
2.1. Latency Optimization for Computation Offloading
2.2. Energy-Efficient Optimization for Computation Offloading
2.3. Joint Optimization of Latency and Energy Consumption
2.4. Research Gaps and Proposed Positioning
- Optimization Paradigm Shift: While existing methods predominantly aim to “minimize total cluster energy consumption” or “individual node energy consumption,” we expand the optimization objective to “minimize the residual energy variance of the cluster.” This fundamental shift directly addresses the critical issue of cluster energy balancing that traditional methods overlook.
- Algorithmic Architecture Breakthrough: Breaking through the limitations of static linear scalarization in traditional MORL, we propose a hybrid Evolutionary-MORL framework. The outer evolutionary algorithm dynamically searches for Pareto-optimal weights, while the inner DRL agent trains policies accordingly, achieving an adaptive optimal trade-off between latency and energy consumption within a non-convex solution space.
- Physics-Informed Integration Mechanism: We design a physics-informed reward shaping mechanism that embeds the energy balance deviation () as a regularization term into the reward function. This mechanism strictly guides the RL agent to autonomously learn load-balancing strategies during the dynamic decision-making process.
2.5. Comparative Analysis
3. System Model and Problem Formulation
3.1. System Model
3.1.1. Offloading Delay Model
3.1.2. Cluster Job Duration Model
3.1.3. Cluster Energy Efficiency Model
3.2. Problem Formulation
4. Solution Approach
4.1. Method Framework
4.1.1. Environment and State Observation
- Action Space: For each time slot t, the offloading decision for each arriving task n involves selecting a target server from the available UAV-MEC cluster. Thus, the individual action for a single task can be defined as . Consequently, for a UAV-MEC node handling N tasks simultaneously within a time slot, the size of the joint action space is . As the action space grows exponentially with the number of tasks N, finding an exact optimal solution using deterministic algorithms (such as dynamic programming) in polynomial time is strictly limited to extremely small-scale scenarios. This severe exponential complexity inherently dictates and justifies our adoption of a Deep Reinforcement Learning (DRL) approach, which aims to efficiently discover high-quality approximate solutions with an acceptable computational overhead.
4.1.2. Reward Mechanism Design
4.1.3. Offloading Action Decision
4.1.4. Parameter Training and Update
4.2. Algorithm and Optimization Design
| Algorithm 1 MORL-LAPB Algorithm |
|
- (a)
- For all objective functions, the performance of A is no worse than that of B, i.e., (assuming the objectives are to be minimized);
- (b)
- There exists at least one objective j such that .
| Algorithm 2 Evolutionary Multi-Objective Optimization Algorithm |
|
5. Simulation Validation
5.1. Simulation Environment Setup
- Environment Settings: To emulate real-world task arrival patterns, the simulation considers a probabilistic task arrival model. At each time step, a task arrives with probability p, and its computational length is defined as L; if no task arrives, the corresponding workload is set to zero. The default environment parameters used in the experiments are listed in Table 3.
- Encoding Scheme: We employ real-parameter encoding for the weights, with the search space strictly bounded within .
- Crossover Operator: We utilize Simulated Binary Crossover (SBX) to generate offspring, with a crossover probability of 0.9 and a crossover distribution index of 10.
- Mutation Operator: Polynomial mutation is applied to introduce genetic diversity, with a mutation probability of 0.2 and a mutation distribution index of 20.
- Population Size: The population size is set to 60 individuals per generation.
- Selection Mechanism: A tournament selection method is used to select parent individuals for reproduction. Furthermore, an elitism strategy is incorporated to preserve the highest-performing individuals across generations, thereby preventing the loss of good solutions and ensuring stable algorithmic convergence.
- Performance Metrics: Following [21,27], the proposed method’s offloading performance is evaluated using the average completion time (Objective 2, denoted as ). In addition, cluster operational duration (Objective 1, denoted as ) and cluster energy efficiency ratio (Objective 3, denoted as ) are introduced as supplementary performance indicators for comprehensive evaluation.
5.2. Baseline Methods
- Random Single-Edge Server Offloading (RSO) [18]: The RSO method randomly selects an edge server to process each computation offloading task. This approach represents a classical offloading strategy and is widely adopted as a comparison baseline in MEC-related studies.
- Nearest Single-Edge Server Offloading (NSO) [12]: The NSO method selects the geographically nearest edge server for each computation task offloading. As another classical offloading scheme, it is frequently used for comparison in MEC performance evaluations.
- Deep Reinforcement Learning-Based Single-Edge Server Offloading (DRLSO) [27]: The DRLSO method utilizes a deep reinforcement learning framework to select the optimal edge server for each computation task. It achieves state-of-the-art performance among DRL-based offloading strategies in MEC environments.
5.3. Experimental Results
5.3.1. Convergence Analysis
5.3.2. Diversity
5.3.3. Impact of Different Task Processing Deadlines
5.3.4. Impact of Varying Numbers of Arriving Tasks
5.3.5. Impact of Dynamic Energy Consumption of Computing Units
5.3.6. Summary of Experimental Results
6. Discussion
6.1. Interpretation of Results and Comparison with Existing Works
- Task Deadlines (): As shown in Figure 8, even under tight deadline constraints (low ), our method maintains a lower failure rate compared to baselines. This indicates that the latency-aware reward component effectively guides the agent to prioritize urgent tasks.
- Task Load (): In high-traffic scenarios (Figure 9), while all methods experience degradation, MORL-LAPB exhibits the slowest rate of performance decline. This robustness is attributed to the load-balancing mechanism (e.g., queue length ), which prevents specific nodes from becoming bottlenecks during congestion.
- Dynamic Computing Energy Cost (): Figure 10 demonstrates that the framework adapts well to varying hardware energy specifications. By dynamically sensing the energy consumption rate , the algorithm shifts the optimization focus towards energy conservation when the cost of computing increases, thereby preserving the cluster’s operational lifespan.
6.2. Mechanism Analysis: Why MORL-LAPB Works
6.3. Limitations
- Scalability: The current approach employs a centralized training architecture. As the number of UAVs increases significantly (e.g., massive swarms), the state space dimension will grow exponentially, potentially leading to the “curse of dimensionality” and making convergence difficult.
- Communication Overhead: We assumed reliable communication for state information exchange. In practical, highly dynamic environments, exchanging state information (queue length, battery level) between UAVs and the central scheduler incurs signaling overhead and latency, which were simplified in this simulation.
- Mobility Model: The current study focuses primarily on computation offloading, with UAV hovering positions assumed to be relatively stable during task execution slots. The joint optimization of flight trajectory and task offloading was not fully explored in this specific framework.
6.4. Future Research Directions
- Multi-Agent Reinforcement Learning (MARL): To address scalability, the centralized framework can be extended to a decentralized MARL approach (e.g., MADDPG or MAPPO), allowing for UAVs to make local decisions while cooperating to achieve global energy balance.
- Joint Trajectory and Offloading Optimization: Integrating UAV trajectory control into the action space would allow for the system to physically approach user clusters to improve channel quality, thereby reducing transmission energy consumption alongside computational energy optimization.
- Real-World Prototype Validation: Moving from simulation to hardware-in-the-loop (HIL) testing or small-scale field experiments using varying UAV platforms (e.g., DJI Matrice series with onboard Jetson modules) to validate the algorithm’s robustness under real-world interference and battery characteristics.
6.5. Synergy of Evolutionary Optimization and Reinforcement Learning
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Hossain, A.R.; Liu, W.; Ansari, N.; Kiani, A.; Saboorian, T. AI-Native for 6G Core Network Configuration. IEEE Netw. Lett. 2023, 5, 255–259. [Google Scholar] [CrossRef]
- Han, Z.; Zhou, T.; Xu, T.; Hu, H. Joint User Association and Deployment Optimization for Delay-Minimized UAV-Aided MEC Networks. IEEE Wirel. Commun. Lett. 2023, 12, 1791–1795. [Google Scholar] [CrossRef]
- Zhang, Y.; Wang, W.; Ren, J.; Huang, J.; He, S.; Zhang, Y. Efficient Revenue-Based MEC Server Deployment and Management in Mobile Edge-Cloud Computing. IEEE/ACM Trans. Netw. 2023, 31, 1449–1462. [Google Scholar] [CrossRef]
- Chen, G.; Chen, Y.; Mai, Z.; Hao, C.; Yang, M.; Du, L. Incentive-Based Distributed Resource Allocation for Task Offloading and Collaborative Computing in MEC-Enabled Networks. IEEE Internet Things J. 2023, 10, 9077–9091. [Google Scholar] [CrossRef]
- Xu, J.; Ota, K.; Dong, M. Big data on the fly: UAV-mounted mobile edge computing for disaster management. IEEE Trans. Netw. Sci. Eng. 2020, 7, 2620–2630. [Google Scholar] [CrossRef]
- Kaleem, Z.; Yousaf, M.; Qamar, A.; Ahmad, A.; Duong, T.Q.; Choi, W.; Jamalipour, A. UAV-Empowered Disaster-Resilient Edge Architecture for Delay-Sensitive Communication. IEEE Netw. 2019, 33, 124–132. [Google Scholar] [CrossRef]
- Basharat, M.; Naeem, M.; Khattak, A.M.; Anpalagan, A. Digital-Twin-Assisted Task Offloading in UAV-MEC Networks with Energy Harvesting for IoT Devices. IEEE Internet Things J. 2024, 11, 37550–37561. [Google Scholar] [CrossRef]
- Zhang, Y.; Kuang, Z.; Feng, Y.; Hou, F. Task Offloading and Trajectory Optimization for Secure Communications in Dynamic User Multi-UAV MEC Systems. IEEE Trans. Mob. Comput. 2024, 23, 14427–14440. [Google Scholar] [CrossRef]
- Dai, P.; Hu, K.; Wu, X.; Xing, H.; Yu, Z. Asynchronous Deep Reinforcement Learning for Data-Driven Task Offloading in MEC-Empowered Vehicular Networks. In Proceedings of the IEEE INFOCOM 2021—IEEE Conference on Computer Communications, Vancouver, BC, Canada, 10–13 May 2021; pp. 1–10. [Google Scholar]
- Sun, L.; Liu, Z.; Ning, Z.; Wang, J.; Fu, X. Multi-Agent Q-Net Enhanced Coevolutionary Algorithm for Resource Allocation in Emergency Human-Machine Fusion UAV-MEC System. IEEE Trans. Autom. Sci. Eng. 2025, 22, 4473–4489. [Google Scholar] [CrossRef]
- Tang, M.; Wong, V.W.S. Deep Reinforcement Learning for Task Offloading in Mobile Edge Computing Systems. IEEE Trans. Mob. Comput. 2022, 21, 1985–1997. [Google Scholar] [CrossRef]
- Li, M.; Gao, J.; Zhao, L.; Shen, X. Deep Reinforcement Learning for Collaborative Edge Computing in Vehicular Networks. IEEE Trans. Cogn. Commun. Netw. 2020, 6, 1122–1135. [Google Scholar] [CrossRef]
- Gao, A.; Wang, Q.; Liang, W.; Ding, Z. Game Combined Multi-Agent Reinforcement Learning Approach for UAV Assisted Offloading. IEEE Trans. Veh. Technol. 2021, 70, 12888–12901. [Google Scholar] [CrossRef]
- Tang, F.; Hofner, H.; Kato, N.; Kaneko, K.; Yamashita, Y.; Hangai, M. A Deep Reinforcement Learning-Based Dynamic Traffic Offloading in Space-Air-Ground Integrated Networks (SAGIN). IEEE J. Sel. Areas Commun. 2022, 40, 276–289. [Google Scholar] [CrossRef]
- Basharat, M.; Naeem, M.; Anpalagan, A. Latency and Cost Minimization for Task Offloading with Energy Harvesting in UAV-MEC Network. In Proceedings of the 2023 IEEE Symposium on Wireless Technology & Applications (ISWTA), Kuala Lumpur, Malaysia, 15–16 August 2023; pp. 103–108. [Google Scholar]
- Wang, D.; Tian, J.; Zhang, H.; Wu, D. Task Offloading and Trajectory Scheduling for UAV-Enabled MEC Networks: An Optimal Transport Theory Perspective. IEEE Wirel. Commun. Lett. 2022, 11, 150–154. [Google Scholar] [CrossRef]
- Sun, G.; Wang, Y.; Sun, Z.; He, L.; Zheng, X. Joint Task Offloading and Trajectory Control for Multi-UAV-Assisted Mobile Edge Computing. In Proceedings of the ICC 2024—IEEE International Conference on Communications, Denver, CO, USA, 9–13 June 2024; pp. 2652–2657. [Google Scholar]
- Fan, W.; Zhao, L.; Liu, X.; Su, Y.; Li, S.; Wu, F.; Liu, Y.A. Collaborative Service Placement, Task Scheduling, and Resource Allocation for Task Offloading with Edge-Cloud Cooperation. IEEE Trans. Mob. Comput. 2024, 23, 238–256. [Google Scholar] [CrossRef]
- Wang, L.; Wang, K.; Pan, C.; Xu, W.; Aslam, N.; Nallanathan, A. Deep Reinforcement Learning Based Dynamic Trajectory Control for UAV-Assisted Mobile Edge Computing. IEEE Trans. Mob. Comput. 2022, 21, 3536–3550. [Google Scholar] [CrossRef]
- Xu, C.; Guo, J.; Li, Y.; Zou, H.; Jia, W.; Wang, T. Dynamic Parallel Multi-Server Selection and Allocation in Collaborative Edge Computing. IEEE Trans. Mob. Comput. 2024, 23, 10523–10537. [Google Scholar] [CrossRef]
- Liu, Y.; Yan, J.; Zhao, X. Deep Reinforcement Learning Based Latency Minimization for Mobile Edge Computing with Virtualization in Maritime UAV Communication Network. IEEE Trans. Veh. Technol. 2022, 71, 4225–4236. [Google Scholar] [CrossRef]
- Liu, Y.; Xiong, K.; Ni, Q.; Fan, P.; Letaief, K.B. UAV-Assisted Wireless Powered Cooperative Mobile Edge Computing: Joint Offloading, CPU Control, and Trajectory Optimization. IEEE Internet Things J. 2020, 7, 2777–2790. [Google Scholar] [CrossRef]
- Chen, H.; Qin, X.; Li, Y.; Ma, N. Energy-aware Path Planning for Obtaining Fresh Updates in UAV-IoT MEC systems. In Proceedings of the 2022 IEEE Wireless Communications and Networking Conference (WCNC), Austin, TX, USA, 10–13 April 2022; pp. 1791–1796. [Google Scholar]
- You, J.; Jia, Z.; Dong, C.; He, L.; Cao, Y.; Wu, Q. Computation Offloading for Uncertain Marine Tasks by Cooperation of UAVs and Vessels. In Proceedings of the ICC 2023—IEEE International Conference on Communications, Rome, Italy, 28 May–1 June 2023; pp. 666–671. [Google Scholar]
- Ning, Z.; Yang, Y.; Wang, X.; Guo, L.; Gao, X.; Guo, S.; Wang, G. Dynamic Computation Offloading and Server Deployment for UAV-Enabled Multi-Access Edge Computing. IEEE Trans. Mob. Comput. 2023, 22, 2628–2644. [Google Scholar] [CrossRef]
- Yue, S.; Ren, J.; Qiao, N.; Zhang, Y.; Jiang, H.; Zhang, Y.; Yang, Y. TODG: Distributed Task Offloading With Delay Guarantees for Edge Computing. IEEE Trans. Parallel Distrib. Syst. 2022, 33, 1650–1665. [Google Scholar] [CrossRef]
- Pervez, F.; Sultana, A.; Yang, C.; Zhao, L. Energy and latency efficient joint communication and computation optimization in a Multi-UAV-Assisted MEC network. IEEE Trans. Wirel. Commun. 2024, 23, 1728–1741. [Google Scholar] [CrossRef]
- Prakash, R.; Umrao, B.K.; Ranvijay. ReinforceAdapt: Multi-objective approach for environmental adaptation method. Connect. Sci. 2025, 37, 2523960. [Google Scholar] [CrossRef]










| Literature | Methodology | Objectives | Advantages | Limitations | ||
|---|---|---|---|---|---|---|
| Lat. | Ene. | Bal. | ||||
| Wang [15], Liu [16], Pervez [17] | Convex Opt., Game Theory | – | ✓ | – | Theoretically rigorous with optimality guarantees; suitable for static environments. | High computational complexity; difficult to adapt to highly dynamic environments. |
| Tang [13], Liu [14], Wang [18], Chen [19] | Q-learning, DQN, DDPG, AC, D3QN | ✓ | ✓ | – | Capable of handling high-dimensional continuous state spaces; enables real-time decisions. | Often single-node decision making lacking cluster coordination; struggles to balance conflicting objectives. |
| Gao [20] | MADDPG | ✓ | ✓ | – | Explicitly models multi-UAV interaction and coordination; suitable for distributed clusters. | High training complexity; focuses on individual energy, ignoring cluster energy distribution. |
| Sun [22] | BCD + SCA | ✓ | ✓ | – | Combines traditional optimization with heuristics to achieve effective objective trade-offs. | Offline optimization limits real-time capabilities; requires iterative solving with high overhead. |
| Ours (MORL-LAPB) | Hybrid Evolutionary-MORL | ✓ | ✓ | ✓ | Multi-objective balancing strategy with adaptive Pareto search; prevents “wooden barrel effect”. | Slightly higher training complexity, but hyperparameter tuning is automated via EA. |
| Symbol | Description |
|---|---|
| The set of time slots. | |
| The set of UAV-MECs. | |
| The set of tasks arriving at UAV-MEC during time slot . | |
| The length of each time slot . | |
| The deadline for processing task . | |
| The data size of task . | |
| The computation density of task . | |
| x | The decision variable for workload allocation. |
| v | The transmission rate between UAV-MECs. |
| f | The computation capacity of a UAV-MEC. |
| The processing delay of task that arrives at UAV-MEC and is offloaded to UAV-MEC during time slot . | |
| The processing make-span of task arriving at UAV-MEC during time slot . | |
| The power consumption of UAV-MEC b at time slot . | |
| denotes the battery energy state of a UAV-MEC at time t, while represents the battery capacity. | |
| refers to the dynamic power consumption of the UAV-MEC computing unit, while denotes the static power consumption. |
| Parameter | Value | Parameter | Value |
|---|---|---|---|
| B | 5 | T | 100 |
| 25 | Mbits | ||
| 0.1 s | GHz | ||
| cycles/bit | 1 s | ||
| Mbit/s | 0.3 | ||
| 6000 Wh | Wh | ||
| W | 3 W/GCycles |
| Category | Parameter | Value |
|---|---|---|
| Reinforcement Learning | Learning rate | 0.0005 |
| Reward decay () | 0.2 | |
| Exploration rate () | 0.99 | |
| increment | 0.0005 | |
| Batch size | 64 | |
| Target update interval | 200 | |
| Network layers | 4 | |
| Evolutionary Algorithm | Population size () | 60 |
| Max generations () | 10 | |
| Crossover prob. () | 0.9 | |
| Mutation prob. () | 0.2 |
| No. | Obj1 | Obj2 | Obj3 | Remark | No. | Obj1 | Obj2 | Obj3 | Remark |
|---|---|---|---|---|---|---|---|---|---|
| 1 | 55 | 0.16823 | 0.43721 | I | 6 | 71 | 0.20018 | 0.72175 | D |
| 2 | 58 | 0.16933 | 0.49521 | H | 7 | 77 | 0.20849 | 0.78549 | C |
| 3 | 63 | 0.17421 | 0.54796 | G | 8 | 85 | 0.22769 | 0.89698 | B |
| 4 | 67 | 0.18521 | 0.59875 | F | 9 | 89 | 0.25735 | 0.97312 | A |
| 5 | 69 | 0.19693 | 0.65820 | E |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, H.; Wang, Y.; Liu, H.; Li, J.; Wang, X.; Lei, Q.; Xiao, K.; Zhu, H. A Load-Balancing-Aware Learning Framework for Collaborative UAV-MEC Computation Offloading. Sensors 2026, 26, 1920. https://doi.org/10.3390/s26061920
Li H, Wang Y, Liu H, Li J, Wang X, Lei Q, Xiao K, Zhu H. A Load-Balancing-Aware Learning Framework for Collaborative UAV-MEC Computation Offloading. Sensors. 2026; 26(6):1920. https://doi.org/10.3390/s26061920
Chicago/Turabian StyleLi, Huafeng, Yuxuan Wang, Hengming Liu, Jiaxuan Li, Xu Wang, Qun Lei, Ke Xiao, and Hongliang Zhu. 2026. "A Load-Balancing-Aware Learning Framework for Collaborative UAV-MEC Computation Offloading" Sensors 26, no. 6: 1920. https://doi.org/10.3390/s26061920
APA StyleLi, H., Wang, Y., Liu, H., Li, J., Wang, X., Lei, Q., Xiao, K., & Zhu, H. (2026). A Load-Balancing-Aware Learning Framework for Collaborative UAV-MEC Computation Offloading. Sensors, 26(6), 1920. https://doi.org/10.3390/s26061920

