1. Introduction
In recent years, vehicular edge computing networks (VECNs) have been widely considered key enablers for latency-sensitive vehicular applications. They offload computation-intensive tasks from vehicles to nearby edge servers, especially in dense traffic or at the network edge [
1,
2]. However, the highly dynamic topology and limited roadside infrastructure coverage often degrade real-time responsiveness and service continuity [
3,
4]. To enhance coverage and agility, unmanned aerial vehicles (UAVs) have been introduced as aerial edge aggregators to provide flexible computing and communication support for vehicles [
5,
6]. Meanwhile, federated learning (FL) enables UAV-assisted VECNs to collaboratively train models without sharing raw data, alleviating privacy risks and reducing backhaul burden [
7]. It provides a promising solution for task offloading in highly dynamic and latency-sensitive VECN environments [
8].
Building on these advantages, a number of studies have proposed strategies to address practical challenges in implementing FL within UAV-assisted VECNs. In resource-constrained vehicular federated learning, the authors in [
9] developed a distributed incentive mechanism based on a multi-leader–follower Stackelberg game. This scheme balanced task demands while reducing communication energy consumption in multi-task FL scenarios. In [
10], the authors developed a hierarchical blockchain architecture for Internet of Vehicles (IoV) tasks, which employs a reputation-driven scheme for model selection and aggregation. This design enhances system robustness and resilience against malicious or unreliable participants. The authors in [
11] propose a UAV-assisted three-tier FL architecture to mitigate real-time computing and communication overhead in VECNs. An extended Kalman filter is designed to acquire real-time vehicle locations, thereby dynamically optimizing the allocation of communication and computational resources for enhanced efficiency.
In summary, FL enhances task real-time performance, data privacy, and system scalability in UAV-assisted VECNs [
12]. However, its application to high-mobility vehicle-to-everything (V2X) scenarios faces two major challenges. First, the tight coupling and limitation of vehicular resources necessitate a dynamic allocation strategy that jointly optimizes computational accuracy and offloading latency [
13]. Second, high vehicle mobility introduces uncertainty in channel state information (CSI), whereas most existing studies rely on the idealized assumptions of static or perfectly known CSI [
14].
Reinforcement learning (RL) has emerged as a novel and effective approach to tackle resource scheduling issues in VECNs [
15]. The authors in [
16] introduced a federated multi-agent scheme based on Deep Q-Network (DQN) for distributed sidelink resource allocation in VECNs, leveraging FL to maximize spectrum-energy efficiency under strict delay constraints. In [
17], a deep deterministic policy gradient (DDPG)-based algorithm was developed for multiuser computation offloading in VECNs, handling continuous decision spaces and reducing execution delay. In [
18], a Twin Delayed DDPG (TD3)-based scheme was presented for task offloading and resource allocation in vehicular fog computing, aiming to optimize power allocation and maximize network utility. The authors in [
19] proposed Multi-Task Multi-Agent Deep Deterministic Policy Gradient (MTMA-DDPG), a multi-agent reinforcement learning algorithm for multi-task environments. It enabled parallel training on distributed nodes with temporal decay-based parameter sharing, improving reward and convergence without centralized control.
While existing RL approaches are viable for resource allocation, their inherent sequential sampling and synchronous updating often result in slow training and low sample efficiency [
20]. This limitation fundamentally conflicts with the low-latency and rapid adaptation requirements of VECN applications. Consequently, asynchronous parallel RL has emerged as a key pathway to enhance training efficiency. Its core idea leverages multiple parallel actors for distributed data collection and asynchronous global model updates, which significantly reduces training time and improves real-time responsiveness in dynamic environments. The authors in [
18] proposed an Asynchronous Advantage Actor–Critic (A3C)-based energy-efficient scheme for Radio Access Networks (RANs) slicing, jointly allocating power and resource blocks while ensuring low computational complexity. Aiming at dense task processing and low-latency requirements in VECNs, the authors in [
21] developed a multi-threaded interactive A3C (MIA3C) algorithm, which lowers the computational burden per vehicle and enhances convergence.
Given the above, despite the notable progress of prior studies, the following limitations remain.
- (1)
They do not consider client selection to improve the efficiency and robustness of FL in UAV-assisted VECNs. Without an effective client selection mechanism, unreliable or low-quality participants may be chosen, resulting in slower convergence and lower model accuracy. Consequently, the overall learning efficiency of the system is significantly reduced.
- (2)
Research on asynchronous parallel resource scheduling strategies remains limited. Most existing studies rely on conventional RL algorithms such as DDPG, A3C [
22], and TD3, which typically have slow convergence and limited scalability in dynamic and continuous state-action spaces [
23]. As a result, when jointly optimizing multiple variables, their effectiveness will be constrained [
24,
25].
- (3)
They often overlook the dynamic nature of CSI, which is critical in practical wireless environments. In VECNs, CSI changes rapidly due to vehicle mobility, frequent topology variations, and multipath fading [
26]. These fluctuations reduce transmission reliability, cause frequent retransmissions, and lead to higher latency and energy consumption. In addition, rapidly changing CSI further degrades the convergence of traditional optimization-based scheduling methods.
Motivated by these, we design a robust scheme to achieve joint optimization of energy efficiency, communication latency, and learning accuracy under uncertain channel conditions. Specifically, the main contributions are summarized as follows.
- (1)
We develop a UAV-assisted VECN framework, which explicitly considers the dynamic characteristics of time-varying channels in FL resource scheduling. By modeling the trade-off among energy consumption, communication latency, and learning accuracy, a joint optimization problem reflecting the characteristics of the actual vehicle communication environment is formulated.
- (2)
To address the complexity of the joint optimization problem, we propose an asynchronous parallel deep deterministic policy gradient (APDDPG) algorithm along with a reputation-based client selection mechanism. The APDDPG algorithm dynamically allocates computational and communication resources, with the Actor–Critic network learning the optimal resource allocation strategy under rapidly varying CSI. Meanwhile, the reputation model assigns reliability scores to clients based on their historical behavior and model update quality.
- (3)
Field experiments were conducted to reflect real vehicular communication conditions, and the obtained data were adopted as simulation parameters for evaluating the developed APDDPG algorithm. Extensive simulations demonstrate that compared to state-of-the-art algorithms [
17,
18,
19], the developed APDDPG algorithm achieves 20% faster convergence, reduces energy consumption by 9%, and consistently attains an FL accuracy of 95.8%. Furthermore, the algorithm achieves the lowest standard deviation of performance jitter amplitude, verifying its robustness under uncertain channel conditions. These results highlight that the proposed strategy not only improves numerical performance metrics but also enhances the practical reliability and stability of FL in vehicular edge networks.
The rest of this paper is organized as follows. The system model and problem formulation are given in
Section 2.
Section 3 proposes the APDDPG algorithm. Simulation results and discussions are shown in
Section 4. Finally,
Section 5 draws the conclusions of this paper.
2. System Model and Problem Formulation
In this section, we present the system model, including the FL-enabled UAV-assisted vehicular network model, communication model, computation model, and reputation model. On this basis, we formulate a joint optimization problem, aiming to achieve efficient collaborative model training while reducing the overall system cost including energy consumption, latency, and FL loss.
2.1. FL-Enabled UAV-Assisted Vehicular Network Model
The FL-enabled UAV-assisted VECN with APDDPG is illustrated in
Figure 1. This network consists of
I vehicles and
U UAVs, where the set of vehicles is denoted as
, and the set of UAVs is denoted as
Vehicles act as edge nodes in the FL-enabled UAV-assisted VECN. The velocity of
i-th
vehicle is assumed to remain constant within the communication range of
u-th
UAV. Each vehicle is equipped with an edge computing module integrating local sensing, data collection, communication, and on-device training functions. In addition, each vehicle maintains an independent dataset
along with computational and storage resources, to support real-time data processing and collaborative model updating. By leveraging local data such as the global position system (GPS) location, driving speed, and the states of nearby vehicles, the
i-th vehicle performs local training tasks.
UAVs act as aerial edge aggregators that collect and process data from vehicle clusters within their coverage. In each FL training round, the
-th UAV hovers at a fixed altitude
over the traffic-dense region to provide stable coverage and reliable communication support [
27]. It first selects participating clients and coordinates resource allocation, and then aggregates the received model updates. We consider a solar-powered UAV to ensure sufficient energy supply and defer the issue of UAV replacement to future work [
28,
29]. The FL-enabled UAV-assisted VECN is then characterized by the following three models.
2.2. Communication Model
Under the three-dimensional Cartesian coordinate system, the planar coordinates of the
u-th UAV and the
i-th vehicle can be denoted as
and
, respectively, where
denote the horizontal and vertical coordinates of the
u-th UAV, while
represent the horizontal and vertical coordinates of the
i-th vehicle. To perform task scheduling and V2X communication, the
u-th UAV needs to establish communication links with the
i-th vehicle within its coverage area. First, the Euclidean distance
between the
u-th UAV and the
i-th vehicle is defined as
Considering the high mobility of vehicles, the wireless channel is subject to significant uncertainty. Specifically, high-speed vehicle motion causes frequent variations in relative positions and Doppler shifts, which in turn lead to rapid fluctuations in the propagation environment. These dynamics result in unstable channel conditions that are difficult to predict. To capture the stochastic nature of such fluctuations, we introduce a random small-scale fading component
We model the uplink channel gain
from the
i-th vehicle to the
u-th UAV as the product of a deterministic path loss component
and the small-scale fading term
The overall channel gain
can be expressed as
where
is modeled as a circularly symmetric complex Gaussian random variable with zero mean and unit variance, denoted by
The amplitude of the channel gain
is assumed to follow a Rayleigh distribution, which is widely used to characterize small-scale fading in non-line-of-sight (NLOS) scenarios. In VECNs, rapid topology changes and frequent signal blockage caused by buildings and other vehicles make Rayleigh fading a realistic assumption, especially in dense urban environments and V2V communication. The path loss
is given by
where
is the carrier frequency,
c is the speed of light, and
represents the additional average loss under NLOS conditions, defined as
Here
is a binary indicator variable representing the presence of obstacles between the
u-th UAV and the
i-th vehicle.
and
denote the average additional transmission loss under NLOS and LOS conditions, respectively. The noise power spectral density is
, where
is the Boltzmann constant and
is the temperature in Kelvin. The total noise power over bandwidth
is
. Therefore, the signal-to-noise ratio (SNR) of
i-th vehicle and the
u-th UAV is calculated as
where
is the transmission power from the
i-th vehicle to the
u-th UAV over the uplink. According to Shannon’s formula, the uplink transmission rate
for the
i-th vehicle in a given time slot can be calculated as
2.3. Computation Model
In general, vehicular tasks can be classified into two types: latency-sensitive tasks that ensure control and safety, and delay-tolerant tasks such as traffic analysis and infotainment. Since latency-sensitive tasks are typically handled by onboard processors or roadside edge units [
30], this study focuses on delay-tolerant tasks. For these tasks, the main objectives are to improve computational efficiency and optimize resource utilization [
31]. FL offers an effective framework to achieve these goals by enabling distributed computation across vehicles without centralized data aggregation [
32]. By exchanging model gradients instead of raw data, FL preserves data privacy while maintaining efficient joint learning [
33,
34].
In FL systems, the local training time and computational capacity of each vehicle significantly affect overall system efficiency [
35]. In the proposed model, computational tasks are generated and processed locally by each vehicle. As intelligent nodes with both computing and storage capabilities, vehicles construct local datasets based on their operational environment and application requirements. The training process is constrained by device computing power and influenced by workload, channel conditions, and available communication bandwidth.
Specifically, the
i-th vehicle produces a separate
, which is characterized by several key attributes, denoted as
where
signifies the task’s identifier,
represents the data size, and
the CPU cycles required to compute one bit of data. The total CPU cycles
required by the task can be calculated as
In addition, each vehicle client has limited storage and computation capacity. In the model, the
i-th vehicle can provide a maximum CPU frequency of
. After being scheduled by the UAV, the actual CPU frequency of vehicle
is
where
denotes the fraction of the vehicle’s own maximum CPU frequency allocated for local training in each FL round as determined by the UAV. Based on this allocated frequency, the time
required by the
i-th vehicle to complete local model training is
Since the downlink transmission latency from the edge center node to the vehicle is typically negligible, it is ignored in this model. The time consumed by the MEC server to complete one round of tasks consists of the transmission latency and computation latency. The transmission latency
can be expressed as
Furthermore, the power consumption
and energy consumption
of the
i-th vehicle for executing the task
can be expressed as
and
respectively, where
is the effective switching capacitance dependent on the chip architecture. For the
i-th vehicle, the energy consumption
during transmission is
where
denotes the transmission power of the
i-th vehicle for model uploading,
is the maximum transmission power of the
i-th vehicle, and
represents the upload power scheduling factor assigned by the
u-th UAV to the
i-th vehicle.
At time
t, the set of clients selected by the
u-th UAV through the reputation model
is
In this cycle, the total time cost
is determined by the maximum sum of transmission latency and computation latency among the selected clients. The total energy consumption
is the sum of resource consumption of all clients participating in computation, expressed as
and
2.4. Reputation Model
In FL-enabled UAV-assisted VECNs, client selection is guided by a reputation-based model that dynamically evaluates each client’s computing capability, data quality, and channel condition. To improve robustness under time-varying environments, the reputation model consists of direct and indirect components. The direct reputation is updated online based on the client’s recent participation outcomes, while the indirect reputation is refined using feedback from other UAVs together with historical records.
To update the direct reputation, the UAV evaluates each vehicle based on its recent task outcome. Since such observations are noisy in time-varying urban environments, we aggregate multiple past evaluations with historical decay factor
to improve robustness. Let
be the direct reputation value of the
i-th vehicle assessed by the
u-th UAV during the task cycle
, whose derivation is provided in
Appendix A. Therefore, the direct reputation score at time
is expressed as
To enhance the accuracy of reputation assessment, the indirect reputation model is introduced. In addition to evaluating clients directly, each UAV also considers reputation feedback from other UAVs. The similarity
between UAVs is calculated by comparing their historical reputation vectors. This similarity is then used to aggregate indirect reputation scores from all other UAVs. The time decay coefficient
is applied to discount past data. We can derive the indirect reputation value for a user at time
, denoted as
Finally, the reputation model of the
i-th vehicle evaluated by the
u-th UAV is denoted as
We adopt this reputation design for UAV-assisted VECNs to compensate for incomplete single-UAV observations via cross-UAV feedback. Accordingly, we leverage the resulting reputation score to design a client selection mechanism, leading to more reliable updates, faster convergence, and more accurate global aggregation.
For each selected client
, the local training performance is evaluated by a sample-wise loss function
. It measures the prediction error of the local model
relative to the expected input–output pairs
under the given reputation weight
. Accordingly, the global model loss of the
u-th UAV is expressed as
where
denotes the local dataset of the
i-th vehicle, and
denotes the global model parameters.
2.5. Problem Formulation
In FL-enabled UAV-assisted VECNs, the goal is to achieve efficient collaborative model training while minimizing the overall system cost (i.e., energy consumption, latency, and FL loss). Due to vehicle mobility, the communication topology between the UAV and vehicles evolves over time, which makes client participation highly dynamic. To jointly account for communication dynamics, computation cost, and learning performance, we formulate a joint client selection and resource allocation problem that minimizes the weighted sum of energy consumption, latency, and FL loss.
To capture the dynamic connectivity, we introduce a time-varying binary indicator
, which equals 1 if the
i-th vehicle lies within the UAV’s communication range
, and 0 otherwise. Accordingly, the optimization problem continuously updates the set of active clients over time, and the overall objective function is formulated as follows:
In this formulation, represents the temporal evolution of the vehicular topology, enabling adaptive optimization under dynamic VECN conditions. The weighting factors , , and balance the trade-off among energy efficiency, latency, and learning stability. Constraint (a) ensures the fairness-based selection of active clients, while (b)–(c) specify the allocation ranges of transmission power and computation resources. Constraint (d) captures link availability via the binary variable , and constraint (e) guarantees that each local training round meets the latency deadline . Constraint (f) imposes that the set of selected clients is selected by the reputation model , and constraint (g) limits the total computational resources within .
Due to the stochastic nature of wireless channels and fluctuations caused by multipath fading and interference, system performance is significantly impacted. Traditional resource allocation methods often rely on static or quasi-static modeling of the channel state, making them less effective when rapid fluctuations occur and resulting in reduced resource utilization. To address this challenge, this paper introduces the APDDPG algorithm. It is well-suited for continuous action spaces and enables continuous policy optimization under uncertain conditions. Furthermore, the asynchronous parallel learning mechanism of APDDPG can significantly improve training efficiency and enable dynamic adaptation to channel variations during resource allocation, thereby enhancing system stability and performance.
3. Asynchronous Parallel DDPG Algorithm for Resource Allocation
In this section, we first present the APDDPG algorithm framework based on the asynchronous parallel structure of A3C and the continuous deterministic strategy of DDPG. Next, we describe the algorithm implementation and training procedure. Finally, we analyze the convergence and stability of the APDDPG algorithm.
3.1. Algorithm Framework
In recent years, RL algorithms have been widely applied in edge computing. The A3C algorithm offers high training efficiency as an asynchronous parallel framework but employs a relatively simple Actor–Critic network. In contrast, DDPG excels in continuous action spaces but has weaker exploration capability. Therefore, this paper integrates the asynchronous parallel framework of A3C with the continuous deterministic control of DDPG, proposing an APDDPG algorithm to accelerate training convergence while enhancing continuous control ability.
The APDDPG algorithm parallelizes the training of local agents and integrates their results, improving model accuracy and training efficiency. As shown in
Figure 2, each agent includes three modules, Actor, Critic and Learner, while the global model contains only the Actor and Critic to store aggregated parameters. To support asynchronous updates, the Updater uploads local gradients to the global model, and the Puller periodically synchronizes the latest parameters back to the agents.
Each agent operates independently and in parallel within its own local environment. It makes autonomous decisions according to the observed environmental state. The Actor selects an appropriate action based on the current state, the Critic evaluates this action, and the Learner computes the gradient direction without directly updating the network parameters. This division of roles enables asynchronous optimization while maintaining consistency in information and parameters. In addition, the algorithm incorporates a global experience replay buffer. During interactions, each agent stores its state, action, reward, and next-state data in this buffer. This mechanism increases scenario diversity and mitigates instability caused by sample correlation. By accessing shared experiences, agents can improve policy stability and generalization. Overall, APDDPG combines deep deterministic policy optimization with parallel collaborative learning. The asynchronous update mechanism accelerates training, and the experience replays buffer supports knowledge reuse, enhancing adaptability in dynamic environments.
3.2. Algorithm Implementation
3.2.1. Markov Decision Process
To enable adaptive optimization of client selection and resource allocation under dynamic vehicular mobility, the proposed APDDPG framework models the interaction between the UAV and vehicular clients as a Markov Decision Process (MDP). The UAV serves as the central agent that maintains a global view of the network. At each time slot t, the environment state, the action of the agent, and the corresponding reward are defined as follows.
- (1)
State: At each time slot t, the state encapsulates both communication and computation conditions of all active vehicles, defined as
- (2)
Action: At each step, the agent executes an action that jointly controls the resource allocation and client participation strategy:
where
represents the computation resource scheduling factor, and
denotes the transmission power scheduling factor at time
t.
- (3)
State Transition: The environment evolves according to vehicular mobility and wireless dynamics. The next state depends on the current state and the action , as expressed by
where
denotes deterministic system evolution, and
represents random variations from channel fading and vehicle mobility.
- (4)
Reward Function: The reward function is designed to align with the optimization objective (21), jointly minimizing energy consumption and latency while improving learning performance:
3.2.2. Training Procedure
Multiple vehicles are selected as agents using the proposed client selection mechanism. Without loss of generality, the
i-th vehicle is treated as the
i-th agent. During each task cycle
t, the
i-th agent receives a state
and selects an action
The environment then computes a reward
and transitions to the next state
based on the agent’s decision. Each agent independently interacts with its local environment, while all agents share a global replay buffer for asynchronous experience aggregation. The detailed training steps can be found in
Appendix B. The overall implementation of the asynchronous parallel DDPG algorithm is summarized in Algorithm 1.
| Algorithm 1 Asynchronous Parallel Deep Deterministic Policy Gradient (DDPG) |
| Input: |
| | Client vehicles ; learning rate; dataset ; positions of vehicles and edge servers , ; vehicle’s own computation capacity and transmission power ; communication link blockage status ; hyper-parameters , , , ; soft-update rate ; local step number K; maximum sample delay ; the minimum replay buffer size for training |
| Output: |
| | Computation capacity scheduling parameters , and transmission power scheduling parameters |
| Initialization: |
| | For each client : |
| | Initialize local Critic network and Actor network |
| | Initialize global Critic network
and Actor network |
| | Initialize target networks: |
| | Create a shared experience replay buffer D storing |
| Steps: |
| | For each episode p = 1 to Max_Episodes: |
| | Receive initial environment state |
| | For each step t = 1 to Max_Length: |
| | | For each thread (client ) from 1 to n: |
| | | (1) Set exploration noise |
| | | (2) Observe state , where includes , and |
| | | historical computing power |
| | | (3) Select action , and decode it into |
| | | (4) Execute the action; environment returns: |
| | | Latency of each client |
| | | Energy consumption |
| | | Global model accuracy acc. |
| | | New local state |
| | | (5) Calculate channel uncertainty: |
| | | for all parallel clients. |
| | | (6) Calculate immediate reward for each i: |
| | | |
| | | (7) Store into D. |
| | | (8) If and : |
| | | | For each thread from 1 to n:
• Sample a mini-batch of size m from D, keeping samples with
(Bounded Staleness).
• Compute target value:
• Compute Critic loss:
• Compute Actor loss:
• Calculate local gradients:
• Apply gradient clipping if necessary.
• Perform local EMA averaging:
• Asynchronously push to global networks:
• Soft update (Polyak averaging) for stability:
|
3.3. Theoretical Convergence and Stability Analysis
To theoretically guarantee the stability of the proposed asynchronous framework, we analyze the convergence behavior of APDDPG based on stochastic approximation theory. Let
and
denote the parameters of the global Actor and Critic networks, respectively. At each global iteration
t, the asynchronous update rules can be expressed as
and
where
and
represent the exponentially averaged local gradients of the
i-th agent, and
is the learning rate satisfying
and
3.3.1. Convergence Analysis
Under the standard assumptions that the policy function
and the Q-function
are both Lipschitz continuous and have bounded gradients, the stochastic process forms a Robbins–Monro approximation with bounded noise. Taking the expectation of both sides yields
Following the convergence lemma of stochastic approximation, we obtain
which indicates that the APDDPG algorithm converges in expectation to a stationary point of the performance objective
3.3.2. Stability Mechanism
Unlike conventional DDPG, which is prone to oscillation under asynchronous or delayed updates, APDDPG ensures stable convergence through three complementary mechanisms.
First, the bounded-staleness sampling restricts the temporal delay of replayed experiences , preventing gradient bias caused by outdated samples.
Second, the local gradient exponential moving average (EMA) smooths stochastic updates and reduces gradient variance among parallel agents, thus aligning their optimization directions.
Finally, the Polyak soft update introduces a low-pass filtering effect on the target networks:
4. Performance Evaluation
In this section, we first present the collected experimental data and specify the key simulation parameters and comparison schemes. The UAV collects state information from all vehicles to form the global state , ensuring that the decision process satisfies the Markov property. Next, we analyze the generalization performance of the proposed APDDPG algorithm using two neural network architectures, AlexNet and ResNet-18, on two datasets, CIFAR-10 and KITTI. Based on this, we compare the algorithm with state-of-the-art methods and discuss the impact of network parameters on its performance.
4.1. Experimental Setup and Field Data Integration
We simulate the FL-enabled UAV-assisted VECN to validate the proposed reputation model and APDDPG algorithm. UAVs are fixed at an altitude of 100 m above the center of the road with 40 vehicles in the coverage area. In addition, we conducted field experiments to collect real-world channel and system data, which were preprocessed and used to train and validate the proposed APDDPG algorithm. The experiments were conducted in the Zhongguancun area of Beijing, which features a dense urban environment containing both open-air zones and urban canyons. This area provides diverse propagation conditions with frequent transitions between line-of-sight (LOS) and NLOS states due to building blockages and multipath effects.
As illustrated in
Figure 3, in the experiment, four vehicles (
,
,
, and
) were driven along a predefined route that included both high-rise and open-area segments. Each vehicle was equipped with a positioning device, a communication unit, and a laptop with storage capability, enabling real-time data collection, local task processing, and data recording for subsequent analysis and model validation. Along the route, signals from multiple edge service points (
,
,
, and
) operating across various frequency bands were collected, capturing channel power gains, frequency-dependent characteristics, and location-based variations.
During the experiment, real-world vehicular communication data were collected, and the raw data underwent comprehensive postprocessing, including Doppler effect compensation and smoothing filters, to generate realistic, time-synchronized channel condition parameters. These parameters were then used as environmental inputs for model training and performance evaluation. Key parameter settings are summarized in
Table 1.
Vehicles dynamically generated computational tasks ranging from 15 to 50 Mbits. Task urgency was classified based on environmental conditions and mapped to specific task types. Vehicle location, channel state information, battery levels, and other features were integrated as inputs to the model. Since the majority of tasks in VECNs are not safety-critical, they are generally considered latency-insensitive. Therefore, the simulations primarily focus on evaluating the computational efficiency and resource utilization for latency-insensitive tasks.
Subsequently, three benchmark methods are examined to assess performance:
Benchmark 1 (DDPG) [
17]: The DDPG algorithm is applied to handle continuous decision-making in UAV-assisted VECNs. It optimizes resource scheduling and transmission power control through deterministic policy gradients within an Actor-Critic framework.
Benchmark 2 (A3C) [
18]: The A3C algorithm employs multiple asynchronous agents to improve learning efficiency in large-scale UAV-assisted VECNs. Each agent interacts independently with the environment, enabling faster convergence and improved adaptability.
Benchmark 3 (TD3) [
19]: The TD3 algorithm addresses overestimation bias in Q-value estimation, enhancing decision stability in complex UAV-assisted VECNs. It effectively manages the trade-off between energy consumption and latency by using twin critics and delayed policy updates under dynamic network conditions.
4.2. Generalization Analysis
To further verify the generalization capability of the proposed algorithm, comparative experiments were conducted using two neural network architectures, AlexNet and ResNet-18, on two datasets, CIFAR-10 and KITTI, under identical FL configurations. The CIFAR-10 dataset contains 60,000 color images across 10 categories, with a balanced data distribution. In contrast, the KITTI dataset contains complex driving scenarios, resulting in highly heterogeneous data and greater learning challenges.
As illustrated in
Figure 4, the proposed algorithm demonstrates consistent convergence behavior across different datasets and architectures. On CIFAR-10, AlexNet converges rapidly, reaching over 90% accuracy within the first 70 communication rounds and stabilizing at 94.0
0.2%. In contrast, the deeper ResNet-18 converges more slowly but ultimately achieves a higher accuracy of 98.1
0.1%, owing to its greater representational capacity. On the KITTI dataset, both networks converge more gradually and reach lower final accuracies due to the increased data complexity. Specifically, AlexNet attains 79.2
0.3% after approximately 100 rounds, while ResNet-18 reaches 88.4
0.2% after around 180 rounds.
These results confirm that the proposed algorithm maintains stable convergence under different datasets and model architectures, and indicates strong generalization capability. Considering the practical constraints of VECNs, including limited computational and communication resources, AlexNet achieves an effective balance among accuracy, convergence speed, and computational efficiency. In addition, CIFAR-10 serves as a representative benchmark for evaluating learning performance and generalization. Taking both algorithmic generality and system constraints into account, AlexNet trained on the CIFAR-10 dataset is adopted as the default configuration for further performance verification and analysis.
4.3. Optimization of Weighting Factors
Figure 5 illustrates the three-dimensional surface of the optimization objective as a function of the weight parameters
and
, presenting the impact of different weight combinations on the overall optimization performance. The learning stability weight
is determined by the constraint
. The results indicate that the optimization function reaches its maximum value at
,
, and
, corresponding to the best weight configuration. Consequently, the optimized weights are used in the subsequent experiments to ensure fairness and stability in performance comparison.
4.4. Latency Optimization for Federated Learning Tasks
Figure 6 presents the variation in the average computation time during training for different algorithms. All four methods show high computational overhead at the initial stage and gradually converge as training proceeds. The proposed APDDPG algorithm consistently achieves the lowest computation latency and converges faster than the baseline methods, including DDPG, A3C, and TD3. After approximately 100 training rounds, the computation time of APDDPG stabilizes at around 80 ms, whereas DDPG and TD3 converge to about 110 ms and 125 ms, respectively. A3C remains at a higher level of approximately 200 ms throughout the process. The enlarged subfigure further indicates that APDDPG experiences smaller fluctuations during convergence, confirming its superior stability and computational efficiency compared with the other approaches.
4.5. Energy Consumption Optimization for Federated Learning Tasks
As shown in
Figure 7, the APDDPG algorithm achieves faster convergence in energy consumption optimization. In terms of convergence efficiency, the APDDPG algorithm clearly outperforms DDPG and TD3, requiring approximately 10% fewer training episodes to reach a stable state than TD3, and about 20% fewer than DDPG. Although its performance gain in resource consumption (approximately 5% to 9%) is slightly less pronounced than the optimization observed for computation time, APDDPG nonetheless achieves the highest resource utilization efficiency and the fastest convergence speed among all compared algorithms. These results fully demonstrate the superiority of APDDPG in achieving efficient and rational resource scheduling within FL-enabled UAV-assisted VECN.
4.6. Impact of Reputation-Based Client Selection on System Performance and Accuracy
Experimental results clearly demonstrate that the number of client vehicles selected by the reputation model significantly influences both the computation time and energy consumption within the VECN. As illustrated in
Figure 8, as the number of selected client vehicles increases, both the total computation time and total energy consumption of the system exhibit an upward trend. With the number of vehicles increasing from 4 to 8, all algorithms show a steady growth in computation time and energy consumption, reflecting the impact of increased system complexity on resource allocation efficiency. Among all algorithms, APDDPG consistently maintains superior performance, achieving lower computation time and energy consumption compared with other methods. Moreover, its performance degradation remains minimal as the number of vehicles increases, demonstrating its scalability and robustness. In contrast, DDPG and TD3 deliver moderate resource scheduling efficiency, incurring slightly higher computational and energy costs than APDDPG. The A3C algorithm demonstrates the poorest performance, showing a pronounced rise in both computation time and energy consumption in multi-vehicle scenarios. This result indicates that A3C faces challenges in coordinating multi-objective optimization under complex and dynamic environments.
In addition to resource optimization, the model performance under different client selection mechanisms is further analyzed to evaluate the learning robustness of the FL framework.
Figure 9 illustrates the accuracy trend of the FL model over training epochs under different client selection mechanisms. Overall, the accuracy of all four schemes gradually improves with training and stabilizes around 150 epochs. The model without malicious clients achieves the best performance in the ideal scenario. Its accuracy rises rapidly and eventually stabilizes above 95.8%. The reputation-based selection mechanism maintains a comparable level of performance even when malicious clients are present. This result indicates the robustness and stability of the reputation model. In contrast, random client selection scheme converges more slowly and reaches a lower accuracy, suggesting that unstructured selection is more vulnerable to the interference of unreliable clients. Selecting all clients in each round initially improves accuracy rapidly, but its final performance remains slightly lower than that of the reputation-based method. This outcome implies that full aggregation may reduce overall accuracy when malicious nodes exist in the network. The magnified view of the late training stage further shows that the reputation-based model maintains minimal fluctuations and high accuracy. Overall, the reputation-based client selection mechanism improves the resistance to adversarial attacks and enhances the global accuracy of the FL system.
4.7. Final Optimization Objective
As shown in
Figure 10, the evolution of the reward function during training reflects the differences in policy optimization and convergence performance among the algorithms. In the initial stage, all four algorithms show a clear upward trend, indicating that the agents gradually improve their policies through continuous interaction with the environment. The proposed APDDPG algorithm achieves the highest reward throughout the entire training process. It converges faster and more stably compared with the other algorithms. After approximately 100 training rounds, the reward of APDDPG stabilizes at around 5.5, suggesting that it can learn an optimal resource scheduling policy efficiently. In contrast, TD3 and DDPG converge more slowly, and their final reward values remain lower. The A3C algorithm performs the worst, showing the slowest convergence and the lowest steady-state reward of about 3.0, which indicates its limited efficiency in policy updates under continuous state space. Specifically, while DDPG exhibits a slightly faster initial learning rate, its performance eventually plateaus or fluctuates due to the accumulated overestimation bias. TD3, despite a slower start caused by its conservative Clipped Double Q-learning and delayed policy updates, achieves higher average rewards and better stability in the long run, confirming its effectiveness in handling the complex state space of UAV-assisted VECNs.
4.8. Algorithm Reliability and Robustness Evaluation
Figure 11 illustrates the fluctuations in computational latency and energy consumption of different algorithms under uncertain wireless channel conditions. It is observed that the standard deviations of APDDPG and TD3 are significantly lower than those of the other methods, indicating better stability and disturbance resistance in dynamic channel environments. APDDPG maintains consistently low fluctuation levels in both computational latency and energy consumption. Its standard deviation of energy consumption is only 3.798 J, the lowest among all schemes, demonstrating that the algorithm can sustain stable and efficient decision-making even under random channel variations. In contrast, DDPG and A3C show much larger fluctuations. DDPG reaches 14.730 ms in computational latency, revealing a higher sensitivity to channel uncertainty. The unscheduled scheme shows the highest level of instability, with standard deviations of 8.794 J in energy consumption and 25.131 ms in delay. This result suggests that the absence of intelligent scheduling leads to a substantial degradation of system stability.
Overall, the proposed APDDPG algorithm demonstrates the strongest robustness and scheduling stability under complex dynamic channel conditions, confirming its potential for real-world V2X applications.
5. Conclusions
This paper investigates resource scheduling in FL-enabled UAV-assisted VECNs under dynamic and uncertain wireless channel conditions. Traditional optimization scheduling strategies often suffer from slow convergence and the risk of getting trapped in local optima due to channel fluctuations. To overcome the limitations of conventional optimization methods, we integrate the APDDPG algorithm with a reputation-based client selection mechanism in an FL-enabled UAV-assisted VECN. The APDDPG algorithm jointly optimizes computation frequency and transmission power, enhancing adaptability to channel variations and improving energy efficiency. The reputation-based mechanism selects reliable clients, ensuring more accurate and stable federated training. Experimental results confirm that APDDPG has 20% faster convergence, 9% lower energy consumption, and 95.8% FL accuracy. In addition, by incorporating an asynchronous parallel mechanism, the APDDPG algorithm achieves the most stable performance under highly dynamic and uncertain wireless conditions. These results confirm that integrating a reputation-based client selection mechanism with the APDDPG algorithm enables adaptive resource scheduling and joint optimization of energy, latency, and learning accuracy in FL-enabled UAV-assisted VECNs.
Future research will focus on applying V2V relaying to enhance FL-enabled VECNs. Specifically, the proposed APDDPG algorithm can be extended to graph neural network-based DDPGs to model the dynamic vehicular topology and enhance cooperative decision. In addition, we intend to adopt real-world vehicle trajectory datasets to simulate more realistic traffic scenarios, thereby enhancing the reliability and applicability of the proposed scheme.