Graph-Attention Constrained DRL for Joint Task Offloading and Resource Allocation in UAV-Assisted Internet of Vehicles
Highlights
- A UAV-assisted IoV framework is proposed that jointly optimizes UAV hovering position selection, split task offloading between UAV and RSU, and multi-dimensional communication and computing resource allocation under dynamic vehicular environments.
- A graph-attention-based constrained TD3 (GAT-CTD3) algorithm is developed to solve the resulting CMDP, effectively capturing spatial correlations among vehicular tasks while explicitly enforcing long-term UAV energy constraints.
- The proposed GAT-CTD3 scheme can enhance the task completion rate, reduce end-to-end delay, and lower the energy consumption per task completed by the UAV, especially under dense traffic conditions.
- This provides a scalable and energy-efficient learning framework, offering practical guidance for deploying intelligent UAV-assisted IoV systems under strict latency and energy constraints.
Abstract
1. Introduction
- We propose a novel UAV-assisted IoV framework that jointly optimizes UAV positioning, split computing between UAV and RSU, and multi-dimensional communication and computing resource allocation.
- We develop a split computing task offloading model that leverages the UAV’s computing capability to alleviate backhaul congestion and RSU workload, thereby improving overall system performance.
- We design a graph-attention-based constrained TD3 algorithm to solve the formulated CMDP, enabling efficient and constraint-aware decision-making in dynamic vehicular environments.
- The simulation results demonstrate that the proposed scheme significantly outperforms the comparative scheme in terms of task completion rate, execution delay, and energy efficiency.
2. Related Work
2.1. UAV-Assisted Task Offloading and Resource Allocation for IoV
2.2. Task Offloading and Resource Allocation Scheme Based on Deep Reinforcement Learning
3. System Modeling
3.1. Motion Model
3.2. Communication Model
3.3. Computation Model
3.4. Queueing Dynamics and Delay Model
3.4.1. Vehicle-to-UAV Uploading
3.4.2. UAV Computing
3.4.3. UAV-to-RSU Offloading
3.4.4. RSU Computing
3.4.5. End-to-End Delay Calculation
3.5. Energy Consumption Model
3.5.1. Vehicle Transmission Energy
3.5.2. UAV Transmission, Computing, and Mobility Energy
3.6. Problem Formulation
4. Approach Design
4.1. CMDP Formulation
4.2. GAT Encoder
4.3. GAT-CTD3
| Algorithm 1 GAT-CTD3 for Joint UAV Positioning, Split Computing, and Resource Allocation |
|
4.4. Complexity and Stability
5. Experiments and Analysis
5.1. Base Settings
5.2. Comparative Algorithms
5.2.1. TD3-Based Joint Optimization (TD3-JO)
5.2.2. Deep Deterministic Policy Gradient (DDPG)
5.2.3. Heuristic Greedy Resource Allocation (HGRA)
5.2.4. GAT-TD3 with Heuristic Reward Penalty (GAT-TD3-RP)
5.3. Experimental Results and Analysis
5.4. Hyper-Parameter Ablation and Robustness Analysis
5.4.1. Ablation of UAV Energy Budget
5.4.2. Ablation of GAT KNN Neighbor Size
5.4.3. Ablation of Discount Factor
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Taslimasa, H.; Dadkhah, S.; Neto, E.C.P.; Xiong, P.; Ray, S.; Ghorbani, A.A. Security issues in Internet of Vehicles (IoV): A comprehensive survey. Internet Things 2023, 22, 100809. [Google Scholar] [CrossRef] [Scilit]
- Gao, X.; Zhang, X.; Lu, Y.; Huang, Y.; Yang, L.; Xiong, Y.; Liu, P. A Survey of Collaborative Perception in Intelligent Vehicles at Intersections. IEEE Trans. Intell. Veh. 2024, 1–20. [Google Scholar] [CrossRef] [Scilit]
- Dui, H.; Zhang, S.; Liu, M.; Dong, X.; Bai, G. IoT-enabled real-time traffic monitoring and control management for intelligent transportation systems. IEEE Internet Things J. 2024, 11, 15842–15854. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Shu, Q.; Lu, Y.; Zhang, Y.; Wang, Y. QCTF: A quantized communication and transferable fusion framework for multi-agent collaborative perception. IEEE Trans. Intell. Transp. Syst. 2025, 26, 15013–15027. [Google Scholar] [CrossRef] [Scilit]
- Dong, S.; Tang, J.; Abbas, K.; Hou, R.; Kamruzzaman, J.; Rutkowski, L.; Buyya, R. Task offloading strategies for mobile edge computing: A survey. Comput. Netw. 2024, 254, 110791. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Fang, X.; Min, G.; Chen, H.; Luo, C. Intelligent Offloading Balance for Vehicular Edge Computing and Networks. IEEE Trans. Intell. Transp. Syst. 2025, 26, 5792–5803. [Google Scholar] [CrossRef] [Scilit]
- Ren, R.; Zhao, J.; Zhang, Q. UAV-assisted collaborative sensing task offloading and resource allocation in IoV. IEEE Trans. Veh. Technol. 2025, 1–10. [Google Scholar] [CrossRef] [Scilit]
- Saeedi, I.D.I.; Al-Qurabat, A.K.M. A comprehensive review of computation offloading in UAV-assisted mobile edge computing for IoT applications. Phys. Commun. 2025, 72, 102810. [Google Scholar] [CrossRef] [Scilit]
- Ullah, I.; Singh, S.K.; Adhikari, D.; Khan, H.; Jiang, W.; Bai, X. Multi-Agent Reinforcement Learning for task allocation in the Internet of Vehicles: Exploring benefits and paving the future. Swarm Evol. Comput. 2025, 94, 101878. [Google Scholar] [CrossRef] [Scilit]
- Bakirci, M. Internet of Things-enabled unmanned aerial vehicles for real-time traffic mobility analysis in smart cities. Comput. Electr. Eng. 2025, 123, 110313. [Google Scholar] [CrossRef] [Scilit]
- Adil, M.; Song, H.; Jan, M.A.; Khan, M.K.; He, X.; Farouk, A.; Jin, Z. UAV-assisted IoT applications, QoS requirements and challenges with future research directions. ACM Comput. Surv. 2024, 56, 1–35. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Liang, H.; Hong, X.; Yao, Y.; Lin, B.; Zhao, D. DRL-Based resource allocation game with influence of review information for vehicular edge computing systems. IEEE Trans. Veh. Technol. 2024, 73, 9591–9603. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.; Cao, H.; Duan, J.; Qing, H.; Mohajer, A. Adaptive energy-efficient task offloading and resource management in UAV-assisted mobile edge networks using dynamic DRL. Int. J. Sens. Netw. 2025, 48, 212–226. [Google Scholar] [CrossRef] [Scilit]
- Velickovic, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio, P.; Bengio, Y. Graph attention networks. ICLR 2018, 6, 2. [Google Scholar]
- Mao, Y.; You, C.; Zhang, J.; Huang, K.; Letaief, K.B. A survey on mobile edge computing: The communication perspective. IEEE Commun. Surv. Tutor. 2017, 19, 2322–2358. [Google Scholar] [CrossRef] [Scilit]
- Wang, K.; Yin, H.; Quan, W.; Min, G. Enabling collaborative edge computing for software defined vehicular networks. IEEE Netw. 2018, 32, 112–117. [Google Scholar] [CrossRef] [Scilit]
- Dai, M.; Huang, N.; Wu, Y.; Gao, J.; Su, Z. Unmanned-aerial-vehicle-assisted wireless networks: Advancements, challenges, and solutions. IEEE Internet Things J. 2022, 10, 4117–4147. [Google Scholar] [CrossRef] [Scilit]
- Dai, X.; Xiao, Z.; Jiang, H.; Lui, J.C. UAV-assisted task offloading in vehicular edge computing networks. IEEE Trans. Mob. Comput. 2023, 23, 2520–2534. [Google Scholar] [CrossRef] [Scilit]
- Michailidis, E.T.; Miridakis, N.I.; Michalas, A.; Skondras, E.; Vergados, D.J.; Vergados, D.D. Energy optimization in massive MIMO UAV-aided MEC-enabled vehicular networks. IEEE Access 2021, 9, 117388–117403. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Ren, C.; Hu, Y.; Zhang, Y.; Lu, Y.; Li, Q.; You, T.; Rodrigues, J.J. Dual-centralized Q-network-based reinforcement learning for cooperative path planning of multiple UAVs. IEEE Trans. Intell. Transp. Syst. 2025, 26, 13232–13246. [Google Scholar] [CrossRef] [Scilit]
- Hui, M.; Chen, J.; Yang, L.; Lv, L.; Jiang, H.; Al-Dhahir, N. UAV-assisted mobile edge computing: Optimal design of UAV altitude and task offloading. IEEE Trans. Wirel. Commun. 2024, 23, 13633–13647. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Gao, L.; Ma, Z.; Su, J.; Li, F.; Yuan, Y.; Guan, X. Joint task offloading and resource allocation scheme with UAV assistance in vehicle edge computing networks. Comput. Netw. 2025, 273, 111746. [Google Scholar] [CrossRef] [Scilit]
- Zhao, L.; Zhao, Z.; Zhang, E.; Hawbani, A.; Al-Dubai, A.Y.; Tan, Z.; Hussain, A. A digital twin-assisted intelligent partial offloading approach for vehicular edge computing. IEEE J. Sel. Areas Commun. 2023, 41, 3386–3400. [Google Scholar] [CrossRef] [Scilit]
- Ahmed, M.; Fatima, N.; Raza, S.; Ali, H.; Qayum, A.; Khan, W.U.; Sheraz, M.; Chuah, T.C. Optimizing Resource Allocation and Task Offloading in Multi-UAV MEC Networks. IEEE Access 2025, 13, 68710–68725. [Google Scholar] [CrossRef] [Scilit]
- Kuang, Z.; Pan, Y.; Yang, F.; Zhang, Y. Joint task offloading scheduling and resource allocation in air–ground cooperation UAV-enabled mobile edge computing. IEEE Trans. Veh. Technol. 2023, 73, 5796–5807. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Gao, H.; Lv, T.; Lu, Y. Deep reinforcement learning based computation offloading and resource allocation for MEC. In Proceedings of the 2018 IEEE Wireless Communications and Networking Conference (WCNC), Barcelona, Spain, 15 April 2018; IEEE: New York, NY, USA, 2018; pp. 1–6. [Google Scholar]
- Tong, Z.; Deng, X.; Ye, F.; Basodi, S.; Xiao, X.; Pan, Y. Adaptive computation offloading and resource allocation strategy in a mobile edge computing environment. Inf. Sci. 2020, 537, 116–131. [Google Scholar] [CrossRef] [Scilit]
- Shang, C.; Sun, Y.; Luo, H.; Guizani, M. Computation offloading and resource allocation in NOMA–MEC: A deep reinforcement learning approach. IEEE Internet Things J. 2023, 10, 15464–15476. [Google Scholar] [CrossRef] [Scilit]
- Do, H.M.; Tran, T.P.; Yoo, M. Deep reinforcement learning-based task offloading and resource allocation for industrial IoT in MEC federation system. IEEE Access 2023, 11, 83150–83170. [Google Scholar]
- Yan, M.; Xiong, R.; Wang, Y.; Li, C. Edge computing task offloading optimization for a UAV-assisted internet of vehicles via deep reinforcement learning. IEEE Trans. Veh. Technol. 2023, 73, 5647–5658. [Google Scholar] [CrossRef] [Scilit]
- Bai, J.; Luo, J.; Chen, Y.; Tang, Y.; Jin, L.; Shi, Y.; Yang, B.; Ji, H. The DDPG-based joint optimization of task offloading and content caching in UAV assisted IoV. IEEE Internet Things J. 2025, 12, 40330–40346. [Google Scholar] [CrossRef] [Scilit]
- Liang, H.; Zhang, H.; Ale, L.; Hong, X.; Wang, L.; Jia, Q.; Zhao, D. Joint task partitioning and resource allocation in uav-enabled vehicular edge computing based on deep reinforcement learning. IEEE Internet Things J. 2025, 12, 15453–15466. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Wang, Y.; Zhang, Y.; Lu, Y.; Shu, Q.; Hu, Y. Extrinsic-and-intrinsic reward-based multi-agent reinforcement learning for multi-UAV cooperative target encirclement. IEEE Trans. Intell. Transp. Syst. 2025, 26, 17653–17665. [Google Scholar] [CrossRef] [Scilit]
- Uddin, A.; Sakr, A.H.; Zhang, N. Intelligent offloading in vehicular edge computing: A comprehensive review of deep reinforcement learning approaches and architectures. arXiv 2025, arXiv:2502.06963. [Google Scholar]
- Liu, J.; Ahmed, M.; Mirza, M.A.; Khan, W.U.; Xu, D.; Li, J.; Aziz, A.; Han, Z. RL/DRL meets vehicular task offloading using edge and vehicular cloudlet: A survey. IEEE Internet Things J. 2022, 9, 8315–8338. [Google Scholar] [CrossRef] [Scilit]
- Fujimoto, S.; Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. In Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden, 10 July 2018; pp. 1587–1596. [Google Scholar]
- Song, T.; Tan, X.; Ren, J.; Hu, W.; Wang, S.; Xu, S.; Wang, X.; Sun, G.; Yu, H. DRAM: A DRL-based resource allocation scheme for MAR in MEC. Digit. Commun. Netw. 2023, 9, 723–733. [Google Scholar] [CrossRef] [Scilit]










| Symbol | Description | Symbol | Description |
|---|---|---|---|
| Set of vehicles in the IoV system | k | Index of vehicle and its corresponding task | |
| t | Discrete time-slot index | Duration of one time slot | |
| Set of active vehicular tasks at time slot t | Position of vehicle k at time t | ||
| Position of the UAV at time t | Fixed position of the RSU | ||
| Set of candidate UAV hovering positions | Index of UAV hovering position at time t | ||
| Size of sensing data generated by vehicle k | Required CPU cycles per bit for task k | ||
| Maximum tolerable latency of task k | Total fused task data size of vehicle k | ||
| Total required CPU cycles of task k | Task split computing ratio between UAV and RSU | ||
| Bandwidth allocation ratio for vehicle–UAV uplink | Bandwidth allocation ratio for UAV–RSU offloading | ||
| Total available uplink bandwidth from vehicles to the UAV | Total available backhaul bandwidth from the UAV to the RSU | ||
| Transmit power of vehicle k at time slot t | Transmit power of the UAV at time slot t | ||
| CPU allocation ratio of UAV for task k | CPU allocation ratio of RSU for task k | ||
| Uplink transmission rate from vehicle k to UAV | Offloading transmission rate from UAV to RSU | ||
| Remaining uplink data of task k | Remaining UAV-side computing workload of task k | ||
| Remaining offloading data of task k | Remaining RSU-side computing workload of task k | ||
| Indicator of vehicle–UAV uploading stage | Indicator of UAV computing stage | ||
| Indicator of UAV–RSU offloading stage | Indicator of RSU computing stage | ||
| Instantaneous reward at time slot t | Instantaneous UAV energy cost at time slot t | ||
| Total energy consumption of the UAV | Maximum allowable UAV energy budget | ||
| Original system state at time t | GAT-encoded system state | ||
| Joint action at time t |
| Network Module | Core Parameter | Value |
|---|---|---|
| 1. GAT Encoder | ||
| GAT Encoder | Input feature dimension per task node | 7 |
| Number of GAT layers /attention heads | 2/4 | |
| Node embedding dimension | 64 | |
| Activation function | LeakyReLU | |
| KNN neighbor size /Dropout rate | 5/0.1 | |
| Final encoded state dimension | 67 | |
| 2. TD3 Actor–Critic Networks | ||
| Actor Network | Network structure | 3-layer fully connected MLP |
| Input dimension | 67 | |
| Hidden layer dimensions | 256, 128 | |
| Output activation | Softmax/Sigmoid | |
| Twin Critic Networks | Number of independent critics | 2 |
| Network structure per critic | 3-layer fully connected MLP | |
| Hidden layer dimensions | 256, 128 | |
| Parameter Category | Symbol | Value/Range | Unit |
|---|---|---|---|
| 1. System and Environment Parameters | |||
| UAV fixed flight altitude | H | 50 | m |
| Maximum horizontal speed of UAV | 25 | m/s | |
| Duration of one discrete time slot | 20 | ms | |
| Vehicle moving speed | - | 8–15 | m/s |
| Task data size generated by vehicle | 10–20 | Mbits | |
| UAV local fused sensing data size | 5–10 | Mbits | |
| CPU cycles required per bit of task data | 500–1000 | cycles/bit | |
| Maximum tolerable end-to-end latency of task | 0.5–1.0 | s | |
| UAV propulsion energy cost per unit distance | 8 | J/m | |
| Effective switched-capacitance coefficient of UAV processor | F | ||
| 2. Communication and Computing Parameters | |||
| Total uplink bandwidth (vehicle to UAV) | 10 | MHz | |
| Total backhaul bandwidth (UAV to RSU) | 20 | MHz | |
| Vehicle transmit power | 23 | dBm | |
| UAV transmit power | 30 | dBm | |
| System noise power spectral density | −174 | dBm/Hz | |
| Total CPU frequency of UAV | 10 | GHz | |
| Total CPU frequency of RSU MEC server | 50 | GHz | |
| 3. Algorithm and Training Hyperparameters | |||
| Training batch size | B | 128 | - |
| Actor–critic network learning rate | 0.001 | - | |
| Discount factor | 0.98 | - | |
| Experience replay buffer size | - | - | |
| Number of GAT layers | 2 | - | |
| GAT node embedding dimension | 64 | - | |
| Number of GAT attention heads | - | 4 | - |
| GAT KNN neighbor size | 5 | - | |
| TD3 delayed update frequency (Critic:Actor) | - | 2:1 | - |
| Target network soft update coefficient | 0.005 | - | |
| Lagrange multiplier learning rate | 0.01 | - | |
| Number of Vehicles K | Scheme | Task Completion Rate (Mean ± Std) | Average Execution Delay (Mean ± Std, Unit: s) | Energy Consumption per Task (Mean ± Std, Unit: J/Task) |
|---|---|---|---|---|
| 10 (Low Load) | GAT-CTD3 (Ours) | 0.98 ± 0.008 | 0.46 ± 0.021 | 115 ± 3.2 |
| TD3 | 0.95 ± 0.015 | 0.51 ± 0.028 | 130 ± 4.1 | |
| DDPG | 0.93 ± 0.017 | 0.58 ± 0.032 | 145 ± 4.5 | |
| HGRA | 0.88 ± 0.022 | 0.74 ± 0.037 | 161 ± 5.1 | |
| 20 (Medium Load) | GAT-CTD3 (Ours) | 0.94 ± 0.010 | 0.54 ± 0.023 | 122 ± 3.4 |
| TD3 | 0.85 ± 0.020 | 0.67 ± 0.032 | 160 ± 4.9 | |
| DDPG | 0.81 ± 0.024 | 0.79 ± 0.037 | 175 ± 5.3 | |
| HGRA | 0.72 ± 0.029 | 0.99 ± 0.042 | 204 ± 6.3 | |
| 30 (High Load) | GAT-CTD3 (Ours) | 0.90 ± 0.012 | 0.63 ± 0.024 | 130 ± 3.6 |
| TD3 | 0.78 ± 0.023 | 0.82 ± 0.035 | 191 ± 5.8 | |
| DDPG | 0.72 ± 0.028 | 0.98 ± 0.041 | 206 ± 6.2 | |
| HGRA | 0.55 ± 0.031 | 1.25 ± 0.045 | 252 ± 7.3 |
| Scheme | Task Completion Rate | Average Delay (s) | Energy Consumption per Task (J/Task) | Long-Term Constraint Satisfaction Rate (%) |
|---|---|---|---|---|
| GAT-TD3-RP | 0.84 ± 0.018 | 0.72 ± 0.031 | 162 ± 4.8 | 81.2 ± 3.5 |
| GAT-CTD3 (Ours) | 0.90 ± 0.012 | 0.63 ± 0.024 | 130 ± 3.6 | 99.6 ± 0.3 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, P.; Zheng, X.; Kostromitin, K.I.; Zhang, W.; Shi, H.; Tan, L. Graph-Attention Constrained DRL for Joint Task Offloading and Resource Allocation in UAV-Assisted Internet of Vehicles. Drones 2026, 10, 201. https://doi.org/10.3390/drones10030201
Zhang P, Zheng X, Kostromitin KI, Zhang W, Shi H, Tan L. Graph-Attention Constrained DRL for Joint Task Offloading and Resource Allocation in UAV-Assisted Internet of Vehicles. Drones. 2026; 10(3):201. https://doi.org/10.3390/drones10030201
Chicago/Turabian StyleZhang, Peiying, Xiangguo Zheng, Konstantin Igorevich Kostromitin, Wei Zhang, Huiling Shi, and Lizhuang Tan. 2026. "Graph-Attention Constrained DRL for Joint Task Offloading and Resource Allocation in UAV-Assisted Internet of Vehicles" Drones 10, no. 3: 201. https://doi.org/10.3390/drones10030201
APA StyleZhang, P., Zheng, X., Kostromitin, K. I., Zhang, W., Shi, H., & Tan, L. (2026). Graph-Attention Constrained DRL for Joint Task Offloading and Resource Allocation in UAV-Assisted Internet of Vehicles. Drones, 10(3), 201. https://doi.org/10.3390/drones10030201

