1. Introduction
The rapid proliferation of the Industrial Internet of Things (IIoT) has driven a paradigm shift in modern industrial automation, enabling the large-scale deployment of interconnected sensors in harsh, remote, and expansive environments such as smart mines, automated ports, and large-scale manufacturing plants [
1]. In these scenarios, continuous environmental monitoring, equipment diagnostics, and emergency anomaly reporting are essential for ensuring production safety and operational efficiency. However, industrial sensor nodes are often deployed in hazardous or hard-to-reach areas, which makes frequent battery replacement impractical or even infeasible. To overcome this energy bottleneck, the integration of unmanned aerial vehicles (UAVs) and wireless-powered communication networks (WPCNs) has emerged as a promising solution [
2]. Benefiting from their flexible three-dimensional (3D) mobility and adjustable operating altitude, UAVs can serve as mobile hybrid access points (HAPs) to establish favorable line-of-sight (LoS) links with ground industrial nodes. In this way, UAV-mounted HAPs can proactively broadcast radio-frequency energy to charge ground nodes and subsequently collect their sensory data.
Despite these advantages, the practical deployment of UAV-assisted WPCNs in IIoT environments still faces significant challenges due to differing service requirements and tightly coupled physical constraints [
3]. On the one hand, industrial sensing data are inherently heterogeneous. Critical alarm nodes usually require ultra-reliable transmission and strict quality-of-service (QoS) guarantees with specific throughput thresholds, whereas regular monitoring nodes mainly require best-effort data uploading. Traditional orthogonal multiple access (OMA) schemes, such as time-division multiple access (TDMA), suffer from low spectral efficiency when serving massive regular nodes. In contrast, pure non-orthogonal multiple access (NOMA) schemes may have difficulty providing deterministic QoS guarantees for critical emergency traffic. To address this trade-off, a hybrid approach can be utilized in which emergency nodes use exclusive channels (TDMA) for guaranteed QoS and regular nodes share the remaining spectrum via NOMA. On the other hand, the onboard energy capacity of UAVs is intrinsically limited. Most existing studies adopt simplified two-dimensional (2D) horizontal trajectory models with fixed flight altitudes, thereby neglecting the realistic aerodynamic propulsion energy consumption caused by 3D maneuvers. In complex industrial environments, the UAV needs to dynamically adjust its 3D hovering altitude to adapt to the varying LoS probability and channel conditions [
4], resulting in a highly coupled optimization problem involving spatial trajectory planning, altitude control, and wireless resource allocation.
Solving this multidimensional optimization problem is particularly challenging because of the strong coupling between macroscopic UAV trajectory planning and microscopic physical-layer resource allocation. Such continuous state inheritance introduces strong temporal dependency among consecutive visiting decisions and breaks the independence among local optimization subproblems. As a result, traditional mathematical optimization methods such as block coordinate descent (BCD) and successive convex approximation (SCA) may suffer from prohibitive computational complexity, and can hardly guarantee global optimality for the overall coupled problem. Conversely, directly applying an end-to-end deep reinforcement learning (DRL) approach to the entire problem may lead to the curse of dimensionality. Moreover, a purely data-driven neural network may fail to strictly satisfy rigid analytical constraints [
5] such as deterministic QoS thresholds and non-convex successive interference cancellation (SIC) feasibility conditions, especially during the exploration stage. Therefore, an effective optimization paradigm is required to decouple global UAV trajectory planning from local power and time resource scheduling while preserving the feasibility of varied service constraints.
Motivated by the above challenges (as shown in
Figure 1), this paper investigates the energy efficiency (EE) maximization problem in a UAV-assisted WPCN by jointly optimizing the UAV trajectory, hovering altitude, and NOMA/TDMA resource allocation.
To decouple the spatial–temporal decision variables and handle the high-dimensional system state, we propose a DRL-driven Hierarchical Energy-Efficient Collection scheme, referred to as DRL-HEEC. The proposed scheme adopts a two-tier offline-to-online decision mechanism to jointly address the 3D state inheritance of UAV movement and heterogeneous resource scheduling among nodes. Specifically, a hierarchical DRL framework is developed in which the upper layer employs a Double Deep Q-Network (DDQN) to determine the global UAV trajectory based on inherited physical states and the lower layer solves the intra-cell resource allocation problem by combining analytical optimization techniques. The main contributions of this paper are summarized as follows:
We propose the hierarchical DRL-HEEC scheme, which structurally decouples macroscopic UAV trajectory planning from microscopic wireless resource allocation. Within this architecture, an upper-layer DDQN agent dynamically determines the global UAV trajectory as a discrete spatial routing problem over hexagonal sub-cells according to the inherited UAV physical states, while a lower-layer optimization engine optimizes intra-cell time scheduling, access mode selection, and hovering altitude. This hierarchical design effectively reduces the dimensionality of the original coupled problem, improving both convergence performance and execution efficiency.
We integrate a rotary-wing aerodynamic propulsion energy model to characterize the realistic energy consumption of UAVs during 3D flight maneuvers. To capture the temporal dependency caused by altitude variation, we introduce a discretized altitude state inheritance mechanism in the Markov decision process (MDP). Specifically, the terminating altitude of the UAV in the previous grid cell is treated as part of the initial state for the subsequent grid cell, enabling the DDQN agent to learn EE-aware trajectory decisions under practical 3D mobility constraints.
To support heterogeneous QoS requirements, we formulate a lower-layer joint resource allocation problem for hybrid multiple access (HMA) transmission. Strictly orthogonal TDMA time slots are assigned to critical alarm nodes in order to guarantee their throughput requirements, while regular monitoring nodes are served through NOMA to improve spectral efficiency. By exploiting the Dinkelbach method together with BCD-SCA techniques, the proposed resource allocation engine jointly optimizes time allocation, SIC decoding order, and transmission scheduling, thereby improving the sum rate of regular nodes while satisfying the QoS constraints of critical nodes. Crucially, the optimized objective value derived from this lower-layer analytical engine is subsequently fed back as a deterministic reward signal to guide the upper-layer DRL trajectory exploration.
Extensive simulation results demonstrate that the proposed DRL-HEEC scheme outperforms other baselines. The results show that the proposed scheme effectively improves global energy efficiency, especially under limited UAV battery capacity and dense node deployment. In addition, DRL-HEEC maintains QoS satisfaction for critical alarm nodes, verifying its effectiveness and robustness in complex WPCNs.
The rest of this paper is organized as follows:
Section 2 reviews the related work;
Section 3 describes the system model;
Section 4 formulates the lower-level optimization problem and develops the corresponding solution method;
Section 5 presents the upper-level reinforcement learning algorithm;
Section 6 presents and discusses the simulation results; finally,
Section 7 concludes the paper with a summary and discussion of future work.
2. Related Work
UAV-assisted communication networks have been widely studied due to their flexible deployment and mobility-enabled performance gains. However, their optimization is challenging because UAV trajectory, channel conditions, energy consumption, and resource allocation are highly coupled [
6,
7,
8]. Existing works mainly address these issues through trajectory optimization and artificial intelligence-based methods, as reviewed in the following subsections.
2.1. Trajectory Optimization in UAV-Assisted Communication
In UAV-assisted communication systems, trajectory optimization is crucial for enhancing system-wide data collection performance [
9,
10]. To address this, the recent literature has extensively investigated various trajectory planning methodologies coupled with resource allocation strategies [
11,
12,
13]. For instance, Zhang et al. [
14] investigated the joint design of 3D UAV trajectory and time allocation for aerial data collection in hybrid-powered NOMA-IoT networks, seeking to maximize fair network throughput. Cao et al. [
15] proposed a novel reinforcement learning-based heuristic algorithm that utilizes an attention mechanism to extract dynamic network topologies and delay priorities, enabling online UAV trajectory optimization to jointly minimize energy consumption and average sensor delay. Focusing on total energy efficiency, Li et al. [
16] and Zhang et al. [
17] developed a joint resource allocation and 3D trajectory optimization scheme for UAV-enabled networks. By actively manipulating the UAV’s altitude and horizontal position to improve channel gain differentiation, their strategy was able to achieve superior NOMA decoding performance compared to rigid fixed-altitude flight patterns. Zuo et al. [
18] formulated a mixed-integer nonlinear programming (MINLP) problem to maximize the energy efficiency of fixed-wing UAVs by optimizing their 3D trajectories within multi-obstacle environments. Liu et al. [
19] minimized the average age of information (AoI) metric by jointly optimizing the sensor node transmission power, island clustering, and UAV flight trajectory. In [
20,
21], the authors explored multi-UAV enabled wireless networks in which trajectory and communication were jointly optimized to maximize the minimum throughput. Despite these significant advancements, a critical limitation persists in the current literature, as existing studies either simplify the problem by assuming a constant UAV flight altitude or fail to consider whether altitude affects communication links.
2.2. Applications of Artificial Intelligence in UAV Communications
The integration of artificial intelligence into UAV communications is profoundly reshaping the architecture and performance of future wireless networks, encompassing full-stack optimization from physical-layer signal processing [
22] to network-layer resource allocation [
23]. In [
24,
25], the authors utilized deep reinforcement learning algorithms to jointly optimize the 3D flight trajectories and transmitting powers of UAVs, significantly reducing energy consumption and mitigating co-channel interference while satisfying QoS constraints. Wang et al. [
26] investigated resource allocation in UAV-D2D networks by proposing a scalable heterogeneous multi-agent mean-field actor–critic framework. Karegar et al. [
27] proposed a deep Q-learning-based fuzzy UAV path optimization method for software-defined wireless sensor networks; their approach effectively minimized system energy consumption while significantly enhancing communication bit rates and packet delivery performance. To ensure threshold AoI for UAV-assisted mobile crowd sensing under limited energy supplies, Wang et al. [
28] proposed a decentralized multi-agent deep reinforcement learning framework, which they termed DRL-UCS(
). It incorporates a transformer-enhanced distributed architecture and an adaptive intrinsic reward mechanism to maximize data collection while minimizing AoI violation ratios. Xiao et al. [
29] proposed a proximal policy optimization (PPO)-based UAV-EH scheme with action space cropping to optimize meta-surface allocation, effectively enhancing energy harvesting and extending UAV endurance. Alam and Moh [
30] proposed a joint trajectory control, frequency allocation, and routing (JTFR) algorithm that maximizes link utility by considering link stability, queuing delay, and residual energy. Their approach integrates an adaptive distributed multi-agent deep deterministic policy gradient approach with swarm behavior to obtain the optimal solution. Artificial intelligence and especially deep reinforcement learning have shown great potential in UAV-assisted communication networks by enabling adaptive trajectory planning and resource allocation; however, practical deployment is still hindered by the curse of dimensionality, high training complexity, large sample requirements, limited onboard computation capability, and the sim-to-real performance gap.
Motivated by these limitations, this work investigates energy-efficient and QoS-aware UAV-assisted communication by considering differing QoS requirements of different nodes and altitude-dependent air-to-ground links. We develop a two-layer decoupled optimization framework that separates trajectory planning from communication resource optimization, aiming to improve energy efficiency while reducing problem complexity and preserving their coupling effects. Specifically, the maximized objective value from the lower layer serves as the exact reward signal guiding the upper layer’s spatial exploration. Unlike purely decoupled benchmarks that optimize these aspects in isolation, our framework reduces problem complexity while explicitly preserving the spatiotemporal trade-offs required to maximize overall EE.
3. System Model
3.1. Network Model
In this paper, we consider a mission scenario where a mobile UAV is dispatched to collect data across a WPCN. As illustrated in
Figure 2, the overall network consists of a stationary HAP, the mobile UAV, and
N sensor nodes distributed over the target region.
The HAP, located at the coordinates
, is primarily responsible for charging the UAV and receiving the data collected by the UAV, where
denotes the altitude of the HAP above ground level, which is also set as the initial flight altitude of the UAV. To model the practical scenario, we partition the target area into
M identical regular hexagonal grid cells, denoted by the set
. For any
, the center coordinates of the
m-th grid cell are represented as
. Let
denote the number of sensor nodes within the
m-th hexagonal grid cell. Specifically, the total number of sensor nodes
N satisfies
The
i-th sensor node in grid cell
m is denoted by
, with its coordinates represented as
, where
. These sensor nodes are equipped with energy harvesting (EH) modules. Operating under the harvest-then-transmit protocol, the nodes first harvest the radio frequency (RF) energy radiated by the UAV and use the accumulated energy to transmit their data back to the UAV. It is worth noting that the data collected by the nodes is categorized into regular data and emergency data; each node has a probability
of generating emergency data (e.g., anomaly monitoring data and critical data). Consequently, emergency data should be granted prioritized transmission so as to satisfy a predefined minimum throughput threshold.
In this network architecture, the UAV plays the dual role of both mobile energy transmitter (MET) and mobile data collector (MDC). Specifically, the UAV flies horizontally along a dynamic trajectory to the center of another grid cell in order to execute energy broadcasting and data collection. Once the UAV hovers at the center of a grid cell, it first executes the wireless energy transfer (WET) phase, broadcasting RF energy to all nodes within this grid cell in the downlink. During the wireless information transfer (WIT) phase in the uplink, it collects the data transmitted via these nodes by employing the HMA strategy (as detailed in
Section 3.5). To further enhance system performance, the hovering altitude of the UAV within each grid cell can be dynamically adjusted. Let
denote the hovering altitude of the UAV when it hovers at the center of grid cell
m. The adjustable altitude is bounded by
where
and
respectively represent the minimum and maximum safe flight altitudes of the UAV.
Furthermore, all operational energy of the UAV is drawn from its equipped battery with a capacity of
(in J). To safely sustain the multi-tour operation, the residual energy
must strictly satisfy the condition
where
is a predefined energy threshold guaranteeing a safe return flight. Once
drops to this threshold, the UAV will finish its current data collection and return to the HAP in the next step. After being fully recharged at the HAP, the UAV will resume its uncompleted data collection tasks.
3.2. Proposed Scheme Overview
To achieve data collection that balances the QoS for emergency data and the global energy efficiency in large-scale wireless sensor networks, this paper proposes a DRL-HEEC scheme. In this subsection, we first present the overall operational workflow of our proposed scheme.
As illustrated in
Figure 3, by adopting a two-tier optimization architecture, the proposed scheme decouples the highly complex joint trajectory and resource allocation problem into upper-level UAV trajectory planning and lower-level local resource allocation. We define each recurring operational cycle—where the UAV flies horizontally to a designated grid cell and subsequently adjusts to an optimal altitude to execute local tasks—as a “single step”. The execution workflow and mechanism of the scheme are detailed as follows:
- (1)
Upper-Level UAV Trajectory Planning
At the upper level, the UAV is responsible for horizontal traversing flights among multiple hexagonal target grid cells. To find the optimal visiting sequence that maximizes the global system energy efficiency, we model this routing process as an MDP and employ a DRL algorithm for autonomous solving. Prior to the execution of each single step, the agent evaluates the UAV’s residual energy to ensure that the next planned action satisfies the safe-return energy threshold (). If the energy falls below this warning line, the UAV must return to the HAP for recharging first before resuming the uncompleted tasks.
- (2)
Lower-Level Local Resource Allocation
When the UAV horizontally arrives above the center of a target grid cell, the lower-level local optimization mechanism is triggered for the current step. In this phase, the UAV executes a harvest-then-transmit workflow:
Wireless Energy Transfer: The UAV first broadcasts RF energy to all ground nodes in the downlink region.
Wireless Information Transfer: Following the WET phase, the sensor nodes adopt an HMA scheme for channel access to execute WIT. For emergency nodes carrying critical information (e.g., anomaly monitoring), the system allocates an exclusive orthogonal time slot (TDMA) to prioritize the QoS of their data transmission. For the remaining regular nodes, the system utilizes NOMA technology to concurrently receive data on the same time-frequency block, significantly enhancing the local throughput.
Local Joint Optimization: During the aforementioned interactions, the lower-level algorithm incorporating the probabilistic LoS channel model optimizes the UAV’s hovering altitude () and time allocation strategy in the current grid cell. This maximizes the local EE while satisfying the QoS requirements of emergency nodes.
Through the tight collaboration between the upper and lower layers, the DRL-HEEC scheme prioritizes emergency data transmission while optimizing underlying resource allocation. DRL-HEEC drives the intelligent evolution of the UAV’s global trajectory from the bottom up, ultimately realizing energy efficiency maximization throughout the entire lifecycle of the data collection mission.
3.3. Channel Model
Within the considered WPCN, the communication links between the UAV and the ground nodes are susceptible to blockages caused by environmental obstacles such as foliage or terrain variations. Consequently, these links comprise a mixture of LoS and non-line-of-sight (NLoS) propagation paths. To precisely capture this heterogeneous propagation environment, we adopt the probabilistic air-to-ground (A2G) channel model presented in [
31,
32], which explicitly integrates the occurrence probabilities of both LoS and NLoS links.
Based on the discrete hovering trajectory established above, the UAV communicates with the nodes exclusively while hovering at the grid cell centers. When the UAV hovers at the center of the
m-th grid cell with an altitude of
, its 3D coordinates are denoted by
. Consequently, the Euclidean distance
between the UAV and the
i-th sensor node
(
) located in grid cell
m is expressed as
The LoS probability is determined by the elevation angle (in degrees), which is defined as the angle formed between the horizontal plane and the line connecting the UAV to the ground node. The elevation angle
with respect to the node
can be calculated as
As in [
31,
33], the probability of the existence of an LoS link, denoted as
, is given by
where
a and
b are environment-dependent constants [
34] (e.g., for urban or suburban areas).
Thus, the average channel power gain
between the UAV and the node
is formulated as the expectation over both line-of-sight and non-line-of-sight links:
where
denotes the channel power gain of the LoS link, which follows the free-space path loss model given by
Here,
represents the channel power gain at a reference distance of 1 m and
denotes the path loss exponent. Furthermore, the channel power gain of the NLoS link, denoted by
, is given by
where
is the additional attenuation factor specific to the NLoS link.
3.4. Energy Model
In the considered UAV-assisted WPCN, the sustainability of the entire system is governed by the energy dynamics of both the aerial and ground terminals. We systematically analyze the energy model from two perspectives, the UAV and the ground nodes. For the UAV, its operational lifecycle is fundamentally bounded by the physical capacity of its onboard battery. Its total energy consumption consists of two main components: the mechanical propulsion energy required for spatial flying and hovering, and the communication energy consumed for downlink RF wireless charging. Conversely, for the ground nodes, their energy lifecycles operate strictly on a harvest-then-transmit protocol; they harvest energy from the UAV’s RF broadcasting and consume it for uplink data transmission. To provide a fine-grained analysis, we detail the specific energy consumption and harvesting models for each operational step as follows.
3.4.1. UAV Energy Consumption Model
Since the flight altitude of the UAV is dynamically adjustable, the propulsion energy is driven by horizontal flying, vertical ascent/descent, and hovering. Suppose that the UAV flies horizontally at a constant speed
V, requiring a horizontal flying power
. According to the rotary-wing energy theory in [
35],
is precisely modeled as
where
and
denote the blade profile power and induced power in hover, respectively,
is the tip speed of the rotor blade,
is the mean rotor induced velocity,
and
denote the fuselage drag ratio and air density, and
s and
A represent the rotor solidity and rotor disk area. When the flight speed
, this formula reduces to the hovering power, denoted as
, which is given by
Concurrently, the UAV ascends at a constant velocity
and descends at
. Considering the weight of the UAV and motor characteristics, the corresponding ascent power
and descent power
are defined in relation to the aforementioned hovering power
as
where
and
represent the power scaling coefficients for descending and ascending, respectively, which are determined by the UAV’s weight and motor efficiency.
3.4.2. Single-Step Energy Consumption
During the mission execution cycle, the UAV executes a sequence of operational steps. A single step
k involves moving from the previous way-point
to the target grid cell
, adjusting the altitude to the optimal hovering height, and performing local tasks. Here,
explicitly denotes the specific grid cell
m visited at step
k. Based on this discrete trajectory, the 2D horizontal flying distance
and the vertical altitude difference
at step
k are defined as follows:
The total energy consumption of the UAV for executing the single step
k, denoted as
, consists of three main components: the movement energy
, the hovering propulsion energy
, and the RF transmission energy
, that is,
Specifically,
accounts for both the horizontal flight at a constant speed
V and the vertical altitude adjustment, which is formulated as
Upon reaching the hovering position at grid cell
, the UAV initiates the WET phase by broadcasting RF energy with a constant transmission power
for a duration of
, then enters the WIT phase to collect data for a duration of
. Consequently, the hovering propulsion energy
and the RF charging energy
are respectively given by
where
denotes the hovering duration in grid cell
, satisfying
.
Note that if step k is the required return step to the HAP, then the UAV only performs physical movement without executing wireless power or information transfer tasks.
3.4.3. Ground Node Energy Harvesting Model
In the downlink WET phase, the UAV acts as a wireless energy transmitter broadcasting RF energy to the ground nodes. Specifically, when the UAV hovers at the center of the target grid cell m at step k (i.e., ), it transmits with a constant power for a dedicated energy broadcasting duration, denoted by .
Accordingly, the RF power received by the
i-th sensor node
(
) located in this target grid cell
m can be expressed as
where
is the channel power gain derived previously.
To quantify the wireless energy harvesting process at the ground nodes, the widely adopted linear EH model is employed in this section [
36]. Under this assumption, the harvested energy is strictly proportional to both the received RF power and the charging duration. Therefore, the total energy harvested by node
during this WET phase, denoted as
, is mathematically modeled as
where
denotes the wireless energy conversion efficiency of the ground sensor node.
3.5. Hybrid Multiple Access Data Transmission Model
During the WIT phase while the UAV is hovering, ground nodes suffer from a “double penalty” caused by the near–far effect [
37]: nodes located farther from the grid cell center harvest significantly less energy in the downlink WET phase, yet experience more severe channel attenuation during the uplink transmission. This creates a critical vulnerability: if such a distant edge node generates emergency data, its QoS is highly likely to be compromised; more specifically, if a pure NOMA protocol were employed, the extremely weak signal from this distant node would be deeply submerged in co-channel interference. Such a signal would be relegated to the bottom of the SIC decoding order, making it impossible to satisfy the strict minimum throughput threshold required for critical information.
To fundamentally address the aforementioned fairness bottleneck and QoS urgency issue, we propose a hybrid TDMA-NOMA protocol, termed the HMA strategy. Consistent with previous definitions, let m denote the target grid cell visited at step k (i.e., ) and let (equivalent to ) be the total WIT duration in this grid cell. Given that each ground node independently generates emergency data with a probability , multiple priority nodes may coexist within the same grid cell. Therefore, we let denote the set of high-priority emergency nodes in grid cell m, while the set of remaining regular nodes is denoted by .
As visually illustrated in
Figure 4, the total uplink WIT time
is partitioned into two consecutive sub-slots: a dedicated TDMA slot
for the emergency nodes, and a shared NOMA slot
for all other regular nodes, satisfying
Specifically, for the sake of simplicity and fairness among critical data, the dedicated TDMA phase is equally divided into individual orthogonal sub-slots for each emergency node . Therefore, the allocated transmission duration for each emergency node is , where denotes the total number of emergency nodes in grid cell m. Note that if no emergency node exists in the current grid cell (i.e., ), the TDMA phase is skipped entirely, yielding .
Operating under the harvest-then-transmit protocol, the ground nodes fully utilize their harvested energy to maximize the uplink throughput. Assuming
, the constant uplink transmission power for a given node
is calculated as follows:
3.5.1. Phase 1: Emergency Nodes TDMA Transmission
During the TDMA phase, each emergency node
sequentially performs exclusive orthogonal transmission within its equally allocated sub-slot. The signal-to-noise ratio (SNR)
at the UAV receiver and the corresponding data throughput
uploaded by emergency node
are given by
where
B is the system bandwidth and
represents the additive white Gaussian noise (AWGN) power.
3.5.2. Phase 2: Regular Nodes NOMA Transmission
During the subsequent shared phase
, all regular nodes
transmit their data simultaneously over the same frequency band using the NOMA protocol. To decouple the superimposed signals, the UAV receiver employs the SIC technique. To maximize the NOMA uplink sum-throughput, the UAV decodes the signals in descending order of their channel gains. Specifically, suppose thaat the regular nodes in
are sorted such that their channel gains satisfy
. According to the SIC principle, when decoding the signal of the
i-th regular node
, interference originates solely from the remaining nodes that have not yet been decoded (i.e., those with comparatively weaker channel conditions). Thus, the signal-to-interference-plus-noise ratio (SINR)
and the achievable data throughput
for regular node
are formulated as follows:
4. Single-Step Energy Efficiency Optimization
As outlined in the system overview, directly solving the deeply coupled global multi-tour optimization problem is computationally intractable. To address this, we adopt a two-tier decoupling architecture. This section focuses strictly on the local optimization executed at each hovering phase. Specifically, we first formulate the single-step energy efficiency optimization problem, then develop an efficient algorithm to obtain its optimal solution.
4.1. Problem Formulation
Suppose that at a given step
k, the UAV has been dispatched from the previous waypoint
to a target grid cell
m (i.e.,
) by the upper-level algorithm. The core objective of this lower tier is to jointly optimize the UAV’s continuous hovering altitude
and the time allocation strategy
within this grid cell. This local problem aims to maximize the single-step EE
, taking into account not only the local operations but also the physical flying energy expended to reach the grid cell. Based on our previous definitions, the single-step optimization problem (P1) is formulated as follows:
In this model, the objective function (
28a) evaluates the ratio of the total collected data (the sum of throughput from both emergency nodes and regular nodes) to the total single-step energy consumption
. Constraint C1 provides a QoS guarantee threshold
for all emergency nodes. Constraints C2 and C3 define the safe flight altitude range and the non-negativity of time variables, respectively. Constraint C4 represents the node’s transmit power constraint, while
denotes the maximum transmission power of the node. Note that the global UAV battery capacity limit is not included as a hard constraint in this local optimization; instead, it is inherently managed by the upper-level DRL, where the UAV will be forced to return to the HAP if its residual energy falls below a critical safety threshold (i.e.,
).
4.2. Dinkelbach-Based Fractional Transformation
For notational brevity in the subsequent derivations, we first define the total collected data throughput in grid cell
m as
. The objective of (P1) is a highly non-convex fractional function. Introducing an auxiliary parameter
, based on Dinkelbach’s theorem, the original problem is transformed into an equivalent subtractive form, denoted as (P2):
The optimal parameter
represents the global optimal energy efficiency if and only if
. To find
, an iterative procedure is employed. At the
n-th outer iteration with a given
, we solve (P2) to obtain the optimal variables
. Then, the parameter is updated as
until convergence.
4.3. BCD-SCA Decoupled Solution Algorithm
Despite eliminating the fractional structure, the altitude and time allocation remain deeply coupled in (P2). To guarantee strict convergence, we adopt the block coordinate descent (BCD) framework in the inner loop to decouple the problem into two alternating subproblems.
4.3.1. Optimizing with Fixed
With altitude fixed, the channel gains become constants. Let
denote the compound energy harvesting coefficient. The emergency node throughput
inherently exhibits a perspective function form, rendering it jointly concave with respect to time variables. However, the NOMA throughput
contains non-convex co-channel interference. We leverage the perspective function to reconstruct it into a strict difference of concave (DC) form:
At the
r-th SCA iteration point
, first-order Taylor expansion is applied to the strictly concave function
in order to obtain its global affine upper bound:
The precise partial derivatives are derived as follows:
where
. By replacing
with
, the NOMA throughput is lower-bounded by a concave function
. Thus, Subproblem 1 at the
r-th iteration is approximated as the following standard convex problem:
4.3.2. Optimizing with Fixed
With time fixed, the energy consumption
is convex with respect to altitude due to the convexity-preserving max operations in the UAV propulsion model. The throughput
is rigorously proven to be convex with respect to
within the physical altitude range (
). Utilizing SCA, the global concave lower bound for the throughput of any node
(including emergency nodes in C1) is derived via Taylor expansion at the local point
:
For any regular node
, we define the conditionally fixed power factor as
, where the composite numerator
perfectly absorbs the fixed time variables, transmit power, and fixed LoS probability
from the current iteration point. Under this decoupling strategy, the effective partial derivative of the regular node throughput is chained as
where
. By substituting the non-convex throughput expressions with their concave lower bounds, Subproblem 2 at the
r-th iteration is transformed into
This approximated subproblem is strictly convex and can be efficiently solved by standard convex optimization solvers such as the MATLAB R2021b-based CVX toolbox.
Through alternating iterations between (P2.1) and (P2.2), the inner BCD-SCA algorithm strictly converges to a locally optimal strategy for a given q. Consequently, the outer Dinkelbach loop updates q until convergence, successfully yielding the maximum single-step energy efficiency . This optimal analytical value seamlessly serves as the “immediate reward” for the UAV’s global trajectory action, providing mathematically exact feedback for the upper-level RL environment detailed in the following section.
5. Upper-Level Global Trajectory Design via DDQN
While the lower-level local optimization maximizes the single-step energy efficiency, the global trajectory heavily depends on the sequential visitation order. Since the UAV’s current flight decision strictly affects its remaining energy and available future paths, this sequential decision-making process is formulated as an MDP and solved using the DDQN algorithm.
5.1. MDP Modeling
The MDP is defined by a tuple , capturing the dynamic interaction between the UAV agent and the communication environment.
- (1)
State Space : At each step k, the state vector encompass the UAV’s physical status and the global mission progress, defined as . Here, is a tuple denoting the UAV’s spatial position at step k, where represents the location index ( corresponds to the HAP, while the remaining indices indicate the target service grid cells) and indicates the hovering altitude at step k. Furthermore, is the residual energy at step k and is a binary mask indicating the unserved status of target grid cells at step k (where means unvisited).
- (2)
Action Space : To simulate realistic flight kinematics while providing configurable trajectory flexibility, the UAV’s mobility is governed by a maximum allowed flight distance per step. Specifically, when hovering above a target grid cell, the UAV can transit to any grid cell within an
n-tier hexagonal neighborhood (i.e., up to
n hops away, where
) or directly return to the HAP (i.e.,
). However, once the UAV is at the HAP (e.g., after recharging), it can be dispatched to any grid cell to resume data collection. Thus, the dynamically available action space
is formulated as
where
denotes the set of adjacent neighboring grid cells up to
n geometric layers from grid cell
and
encompasses all location indices.
- (3)
Reward Function : The immediate reward
is the critical bridge coupling the upper and lower tiers. It dynamically evaluates the selected action
:
Specifically, visiting an unserved grid cell yields the scaled optimal local energy efficiency , where acts as a normalization factor to balance the reward magnitude. Conversely, traversing an already-visited grid cell or prematurely returning to the HAP yields a mild penalty . This minor penalty fundamentally discourages redundant flights that consume energy without throughput gains, yet is kept intentionally small because traversing a visited grid cell is sometimes geometrically inevitable for transiting to distant unserved areas. is a positive completion bonus awarded when all target areas are successfully visited, and is a massive penalty for crashing due to energy exhaustion.
5.2. DDQN Algorithm
Given the high-dimensional hybrid state space and the dynamically masked discrete action space, we employ DDQN, an advanced value-based reinforcement learning framework that effectively mitigates the overestimation bias inherent in standard DQN.
The framework maintains two neural networks: an online Q-network with weights to select actions, and a target Q-network with weights to evaluate the selected actions. To strictly prevent the UAV from selecting physically unreachable grid cells or executing actions that lead to inevitable energy depletion, an action mask is applied to the network’s output layer. This mask forces the Q-values of invalid actions to , ensuring they are never selected during the -greedy exploration and exploitation phases.
During training, whenever the UAV executes a valid action, the environment pauses to run the lower-level solver, computes
, and feeds it back. The transition tuple
is then stored in an experience replay buffer
. The online network is updated by sampling mini-batches from
and minimizing the mean squared error (MSE) loss:
where
is the DDQN target, defined as
with
denoting the discount factor. By decoupling action selection (via
) from target evaluation (via
), DDQN prevents the propagation of overly optimistic value estimates. Through iterative interactions, decoupled dual-tier updates, and periodic synchronization of
with
, the algorithm converges to a globally optimal energy-efficient trajectory and resource management policy.
Algorithm 1 outlines the proposed joint upper-lower tier optimization framework, where a DDQN agent manages global trajectory planning and action masking while an inner Dinkelbach-BCD-SCA method optimizes local altitude and resource allocation whenever an unserved target grid cell is visited. Specifically, starting from the HAP with a full battery
, the UAV iteratively selects valid actions
that account for physical energy consumption
and recharging mechanics at
, receiving immediate rewards
scaled by the optimal local energy efficiency
for new service visits, global completion bonuses
, crash penalties
, or mild penalties for grid cell revisiting
, while experience replay and periodic target network updates with frequency
ensure stable neural network training and convergence.
| Algorithm 1 Joint Upper-Lower Tier Optimization via DDQN |
- Require:
Set of all grid cells , Battery capacity , Replay buffer , Target update frequency , Learning rate , Discount factor . - Ensure:
Optimal global flight policy and local resource allocation . - 1:
Initialize Online Q-network and Target Q-network with identical weights . - 2:
for each training episode do - 3:
Reset environment: (HAP), , . - 4:
for each decision step do - 5:
Obtain state and valid action space . - 6:
Select action using -greedy policy based on masked . - 7:
Execute : update location and calculate energy consumption, yielding . - 8:
if and then - 9:
Recharge battery: . - 10:
end if - 11:
if and then - 12:
Inner Optimization: Execute Dinkelbach-BCD-SCA algorithm: - 13:
while Dinkelbach iteration n not converged do - 14:
while BCD-SCA iteration r not converged do - 15:
Update by solving (P2.1) via CVX. - 16:
Update by solving (P2.2) via CVX. - 17:
end while - 18:
Update . - 19:
end while - 20:
Set local EE , assign reward , update . - 21:
else if all target grid cells are served () then - 22:
Assign completion bonus and terminate episode. - 23:
else if then - 24:
Assign crash penalty and terminate episode. - 25:
else - 26:
Assign step penalty . - 27:
end if - 28:
Transition to , store experience in . - 29:
Sample a random mini-batch from . - 30:
Calculate DDQN target: . - 31:
Perform gradient descent on Loss: . - 32:
if step then - 33:
Update target network weights: . - 34:
end if - 35:
end for - 36:
end for
|
5.3. Computational Complexity Analysis
The overall per-step computational complexity of Algorithm 1 during the training phase is expressed as , where is the neural network computation cost and represents the mini-batch size for experience replay. The indicator denotes whether the UAV visits a target grid cell. When triggered, the lower-layer optimization requires Dinkelbach iterations and BCD-SCA iterations, with being the dimension of variables.
Furthermore, the proposed algorithm exhibits favorable computational efficiency [
38] and practical viability. Although this nested structure introduces significant overhead during offline training, our decoupled framework allows the lower-layer resource allocation to be pre-computed offline as an efficient lookup table. Consequently, the online execution avoids heavy real-time convex optimization, relying solely on lightweight network forward passes to ensure high efficiency for practical UAV deployments.
6. Simulation Results and Analysis
To evaluate the effectiveness and superiority of the proposed dual-tier DRL-HEEC framework, extensive numerical simulations are conducted using MATLAB. The local convex optimization subproblems are solved utilizing the CVX 2.2 toolbox. All simulations are executed on a desktop computer configured with a 12th Gen Intel(R) Core(TM) i7-12700KF CPU operating at 3.60 GHz, 32 GB of memory, and an NVIDIA GeForce RTX 3070 Ti GPU.
6.1. Simulation Setup
We consider a multi-cell target region partitioned into several regular hexagonal grid cells, as shown in
Figure 5. The side length of each hexagonal grid cell is set to
m. Within each grid cell,
N ground nodes are randomly distributed, with each node having a probability
of being an emergency node. The UAV is dispatched from the HAP (grid cell 1) to execute data collection tasks at other nodes, with its initial flight altitude set to
m. Following [
35] and [
19], the propulsion dynamics parameters of the UAV are specified as follows: the blade profile power in hover
, the induced power in hover
, the tip speed of rotor blade
, the mean rotor induced velocity in hover
, the fuselage drag ratio
, the air density
, the rotor solidity
, and the rotor disc area
. Furthermore, the horizontal flight speed
, the vertical ascending and descending speeds are
, and the corresponding power correction factors are
and
, respectively. The remaining parameters are summarized in
Table 1.
6.2. Baseline Schemes
To evaluate the effectiveness of the proposed DRL-HEEC scheme, we compare it with three categories of baseline schemes, including deep Q-network (DQN)-based schemes, PPO-based schemes, and a greedy algorithm (GA)-based scheme. The details of these baselines are described as follows.
DQN-based schemes [27]: The UAV trajectory planning is performed by a conventional DQN algorithm. In
DQN-Opt, the UAV trajectory is determined by the DQN algorithm while the remaining components, including the optimal hovering altitude and lower-level resource allocation, are optimized using the same optimization framework as in the proposed scheme. In
DQN-Fixed, UAV trajectory planning is performed by the DQN algorithm, while the UAV flies at a fixed altitude of 15 m during data collection. For the lower-level transmission strategy, all ground nodes are served using NOMA-based data transmission.
PPO-based schemes [29]: The UAV trajectory planning is performed by the PPO algorithm. In
PPO-Opt, the UAV visiting trajectory is determined by the PPO algorithm, while the optimal hovering altitude and lower-level resource allocation are solved using the same optimization framework as in the proposed scheme. In
PPO-Fixed, UAV trajectory planning is conducted by the PPO algorithm, while the UAV maintains a fixed flight altitude of 15 m. In the lower-level resource allocation, TDMA is adopted for data transmission among ground nodes.
GA-based scheme: In the GA scheme, both UAV trajectory planning and resource allocation are performed based on a greedy strategy. Specifically, at each decision step, the UAV selects the next visiting grid cell according to the locally best immediate reward, such as the shortest distance, largest amount of collected data, or highest instantaneous energy efficiency. The corresponding resource allocation is also determined greedily without considering long-term system performance.
6.3. Performance Comparison and Discussion
As shown in
Figure 6, the proposed DRL-HEEC scheme consistently achieves a significantly higher and more stable cumulative reward than DQN and PPO. While the upper-level action space (e.g., UAV trajectory selection) remains relatively stable, expanding the node scale to
significantly intensifies the complexity of underlying local resource allocation and data collection constraints. Under such dense and complex communication environments, traditional algorithms such as DQN and PPO often struggle to coordinate the tight coupling between trajectory control and lower-layer allocation, leading to performance degradation or suboptimal rewards. In contrast, DRL-HEEC effectively handles this heightened complexity through its robust reward design and coordination mechanism, maintaining superior reward performance across different scales.
As illustrated in
Figure 7, the proposed DRL-HEEC achieves the shortest total flight distance of 4212.85 m, outperforming DQN (4270.03 m), PPO (4339.64 m), and the randomized greedy baseline (4576.24 m). The greedy approach suffers from myopic decisions, leading to severe detours and local optima, while traditional DRL benchmarks such as DQN and PPO face limitations in value estimation or policy updates. On the other hand, the proposed DRL-HEEC framework effectively captures network dynamics through its optimized reward mechanism; specifically, the incorporation of dynamic action masking proactively prunes invalid spatial transitions, eliminating the redundant exploration steps that inflate trajectory lengths in unconstrained benchmarks. Furthermore, the double DQN architecture effectively mitigates upward overestimation bias, ensuring stable value convergence and preventing the agent from adopting suboptimal routing policies. Consequently, DRL-HEEC generates a smoother and more globally optimal path, which translates directly into significant energy savings and enhanced operational efficiency for the UAV system.
To evaluate the performance of the proposed scheme across varying node scales, we generated topologies with different numbers of nodes (
) and conducted ten randomized trials for each, yielding a total of
experimental results. By averaging these results, we determined the energy efficiency of various schemes at different node scales, as depicted in
Figure 8. It is evident that our DRL-HEEC consistently maintains the highest energy efficiency. Meanwhile, DQN and DDQN, which integrate the underlying optimization algorithm proposed in this paper, also achieve commendable energy efficiency, further validating the effectiveness and universality of the proposed optimization mechanism. As the number of nodes increases, higher network complexity causes energy efficiency fluctuations in most schemes. Nevertheless, DRL-HEEC maintains robust and superior performance. This excellent scalability fundamentally stems from our two-layer decoupled architecture, which separates UAV trajectory planning from local resource allocation. This design effectively prevents the exponential explosion of the state action space as the network expands. Consequently, DRL-HEEC makes adaptive decisions without being overwhelmed by dimensionality, whereas traditional tightly coupled baselines suffer severe performance degradation in large-scale scenarios.
In addition, we compare the ratio of emergency nodes satisfying the QoS requirements to the total number of emergency nodes under different emergency data generation probabilities, i.e.,
, as shown in
Figure 9. It can be clearly observed that the QoS requirements of emergency nodes are well satisfied under the proposed DRL-HEEC, DQN, and PPO schemes. In contrast, for the schemes employing a fixed flight altitude and a single TDMA- or NOMA-based access mechanism, an increasing number of emergency nodes fail to meet their QoS requirements as
increases. This observation further indicates that the QoS satisfaction of emergency nodes is less affected by the UAV trajectory, but is more closely related to the underlying resource allocation algorithm. Therefore, an effective resource allocation mechanism plays a critical role in guaranteeing the QoS performance of emergency services, especially when the probability of emergency data generation becomes higher.
Finally, we investigate the impact of the UAV battery capacity on energy efficiency, where the number of nodes in each grid cell is fixed at
. The corresponding results are presented in
Figure 10. It can be observed that the energy efficiency decreases significantly when the UAV battery capacity is relatively small. This is mainly because the UAV has to return to the HAP more frequently for recharging, which introduces additional flight energy consumption and reduces the effective service time. Nevertheless, the proposed DRL-HEEC scheme still achieves the highest energy efficiency among all compared methods. This performance gain can be attributed to the proposed bi-level mechanism, which jointly improves both UAV trajectory planning and resource allocation. As a result, even under limited battery capacity, the proposed scheme can utilize the available energy more efficiently and maintain superior energy-efficiency performance. Compared with the other reinforcement learning-based algorithms, our proposed DRL-HEEC improves energy efficiency by approximately
to
.
7. Conclusions and Future Work
This paper has studied an energy-efficient UAV-assisted data collection scheme for WPCN with both normal and emergency traffic. To maximize system energy efficiency while satisfying the QoS requirements of emergency nodes, we propose a DRL-based hierarchical energy-efficient control scheme, which we name DRL-HEEC. The proposed DRL-HEEC jointly optimizes UAV trajectory planning, charging decisions, access mode selection, and resource allocation through a bi-level decision-making structure. By coordinating upper-level UAV mobility control and lower-level TDMA/NOMA-based resource allocation, DRL-HEEC can effectively adapt to dynamic network states and improve energy utilization. Extensive simulation results verify the superiority of the proposed scheme over benchmark methods. Our proposed DRL-HEEC achieves the highest energy efficiency under different numbers of nodes, emergency data generation probabilities, and UAV battery capacities. Moreover, it can maintain a high QoS satisfaction ratio for emergency nodes even when the emergency traffic load increases. Compared with other reinforcement learning-based algorithms, the proposed scheme improves the energy efficiency by approximately to , demonstrating the effectiveness of the proposed hierarchical joint optimization framework.
In future work, we will extend the proposed framework to multi-UAV cooperative scenarios in which UAV coordination, collision avoidance, and distributed resource allocation need to be addressed. In addition, more realistic system models such as imperfect channel state information, dynamic obstacles, and mobile IoT nodes, will be considered. Finally, the design of lightweight reinforcement learning algorithms will be explored in order to reduce computational complexity and facilitate real-time implementation in practical UAV-assisted WPCNs.
Author Contributions
Conceptualization, S.G. and Q.C.; methodology, S.G., K.Q., and Q.C.; software, Q.W. and S.G.; validation, S.G., H.W., Z.Z., Q.W., Y.W., and Q.C.; formal analysis, S.G., Y.W., and Q.C.; writing—original draft preparation, S.G., K.Q., and H.W.; writing—review and editing, S.G. and Q.C.; supervision, S.G. and Z.Z.; project administration, S.G. and Q.C.; funding acquisition, S.G., K.Q., and Q.C. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported in part by Zhejiang Provincial Natural Science Foundation of China under Grant No. LQ23F020004, Wenzhou Municipal Science and Technology Bureau Project G2023002, Zhejiang Provincial Education Department Project FX2025126, the Natural Science Foundation of Henan Province under Grants 262300422562 and 262300420692, the Huaian Science and Technology Project under Grant No. HAB2024071, and the Natural Science Foundation of China under Grant 62502242.
Data Availability Statement
The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding authors.
Acknowledgments
During the preparation of this manuscript, the authors used Gemini 3.6-Flash for the purposes of language polishing and grammatical editing. The authors have reviewed and edited the output and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflicts of interest.
Abbreviations
The following abbreviations are used in this manuscript:
| IIoT | Industrial Internet of Things |
| UAV | Unmanned Aerial Vehicle |
| WPCN | Wireless Powered Communication Network |
| BCD | Block Coordinate Descent |
| SCA | Successive Convex Approximation |
| DRL | Deep Reinforcement Learning |
| EE | Energy Efficiency |
| EH | Energy Harvesting |
| RF | Radio Frequency |
| QoS | Quality of Service |
| AoI | Age of Information |
| PPO | Proximal Policy Optimization |
| DQN | Deep Q-Network |
| DDQN | Double Deep Q-Nnetwork |
| WET | Wireless Energy Transfer |
| WIT | Wireless Information Transmission |
| NOMA | Non-Orthogonal Multiple Access |
| SIC | Successive Interference Cancellation |
| TDMA | Time Division Multiple Access |
| LoS | Line-of-Sight |
| NLoS | Non-Line-of-Sight |
| A2G | Air-to-Ground |
References
- Afrin, S.; Rafa, S.J.; Kabir, M.; Farah, T.; Alam, M.S.B.; Lameesa, A.; Ahmed, S.F.; Gandomi, A.H. Industrial Internet of Things: Implementations, challenges, and potential solutions across various industries. Comput. Ind. 2025, 170, 104317. [Google Scholar] [CrossRef] [Scilit]
- Tang, Q.; Yang, Y.; Liu, L.; Yang, K. Minimal Throughput Maximization of UAV-enabled Wireless Powered Communication Network in Cuboid Building Perimeter Scenario. IEEE Trans. Netw. Serv. Manag. 2023, 20, 4558–4571. [Google Scholar] [CrossRef] [Scilit]
- Sharma, D.; Kumar, A.; Tyagi, N.; Chavan, S.S.; Gangadharan, S.M.P. Towards intelligent industrial systems: A comprehensive survey of sensor fusion techniques in IIoT. Meas. Sens. 2024, 32, 100944. [Google Scholar] [CrossRef] [Scilit]
- Zeng, Y.; Wu, Q.; Zhang, R. Accessing from the sky: A tutorial on UAV communications for 5G and beyond. Proc. IEEE 2019, 107, 2327–2375. [Google Scholar] [CrossRef] [Scilit]
- Donti, P.L.; Rolnick, D.; Kolter, J.Z. DC3: A Learning Method for Optimization with Hard Constraints. arXiv 2021, arXiv:2104.12225. [Google Scholar]
- Amodu, O.A.; Jarray, C.; Azlina Raja Mahmood, R.; Althumali, H.; Ali Bukar, U.; Nordin, R.; Abdullah, N.F.; Cong Luong, N. Deep Reinforcement Learning for AoI Minimization in UAV-Aided Data Collection for WSN and IoT Applications: A Survey. IEEE Access 2024, 12, 108000–108040. [Google Scholar] [CrossRef] [Scilit]
- Gu, X.; Zhang, G. A Survey on UAV-Assisted Wireless Communications: Recent Advances and Future Trends. Comput. Commun. 2023, 208, 44–78. [Google Scholar] [CrossRef] [Scilit]
- Luo, J.; Wang, Z.; Xia, M.; Wu, L.; Tian, Y.; Chen, Y. Path Planning for UAV Communication Networks: Related Technologies, Solutions, and Opportunities. ACM Comput. Surv. 2023, 55, 1–37. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Gao, Z.; Zhang, J.; Cao, X.; Zheng, D.; Gao, Y.; Ng, D.W.K.; Renzo, M.D. Trajectory Design for UAV-Based Internet of Things Data Collection: A Deep Reinforcement Learning Approach. IEEE Internet Things J. 2022, 9, 3899–3912. [Google Scholar] [CrossRef] [Scilit]
- Abou Arkoub, M.; Hamdi, R.; Qaraqe, M. Trajectory Optimization for UAV-based Communication Systems Powered by Energy Harvesting. In Proceedings of the 2024 IEEE 100th Vehicular Technology Conference (VTC2024-Fall), Washington, DC, USA, 7–10 October 2024; pp. 1–6. [Google Scholar]
- Baccarelli, E.; Scarpiniti, M.; Momenzadeh, A. Energy-minimizing 3D circular trajectory optimization of rotary-wing UAV under probabilistic path-loss in constrained hotspot environments. Veh. Commun. 2024, 46, 100730. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Qi, N.; Jiang, H.; Xiao, M.; Liu, H.; Jia, L.; Zhao, D. Trajectory Planning for UAV-Assisted Data Collection in IoT Network: A Double Deep Q Network Approach. Electronics 2024, 13, 1592. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Guo, Y.; Hao, J.; Du, Y. Energy Efficient UAV Trajectory Design for Hovering-Flying Data Collection. In Proceedings of the 2022 IEEE Globecom Workshops (GC Wkshps); IEEE: New York, NY, USA, 2022; pp. 886–890. [Google Scholar]
- Zhang, Z.; Xu, C.; Li, Z.; Zhao, X.; Wu, R. Deep Reinforcement Learning for Aerial Data Collection in Hybrid-Powered NOMA-IoT Networks. IEEE Internet Things J. 2023, 10, 1761–1774. [Google Scholar] [CrossRef] [Scilit]
- Cao, H.; Zhu, W.; Chen, Z.; Sun, Z.; Wu, D.O. Energy-Delay Tradeoff for Dynamic Trajectory Planning in Priority-Oriented UAV-Aided IoT Networks. IEEE Trans. Green Commun. Netw. 2023, 7, 158–170. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Zhang, H.; Long, K.; Jiang, C.; Guizani, M. Joint Resource Allocation and Trajectory Optimization With QoS in UAV-Based NOMA Wireless Networks. IEEE Trans. Wirel. Commun. 2021, 20, 6343–6355. [Google Scholar] [CrossRef] [Scilit]
- Zhang, L.; Celik, A.; Dang, S.; Shihada, B. Energy-Efficient Trajectory Optimization for UAV-Assisted IoT Networks. IEEE Trans. Mob. Comput. 2022, 21, 4323–4337. [Google Scholar] [CrossRef] [Scilit]
- Zuo, X.; Yang, Z.; Fan, Y.; Yao, R.; Xu, J.; Li, L. 3D Trajectory Optimization for Energy-Efficient UAV Communication With Obstacle Constraints. IEEE Trans. Veh. Technol. 2026, 75, 1073–1084. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Liu, H.; Zheng, K.; Liu, J.; Taleb, T.; Shiratori, N. AoI-Minimal Clustering, Transmission and Trajectory Co-Design for UAV-Assisted WPCNs. IEEE Trans. Veh. Tech. 2025, 74, 1035–1051. [Google Scholar] [CrossRef] [Scilit]
- Wu, Q.; Zeng, Y.; Zhang, R. Joint Trajectory and Communication Design for Multi-UAV Enabled Wireless Networks. IEEE Trans. Wirel. Commun. 2018, 17, 2109–2121. [Google Scholar] [CrossRef] [Scilit]
- Qi, H.; Wu, M.; Zhang, Z.; Zhao, M. Trajectory Design for Multi-UAV-Enabled Wireless Powered Communication Networks: A Multi-Agent DRL Approach. In Proceedings of the 2024 IEEE Wireless Communications and Networking Conference (WCNC), Dubai, United Arab Emirates, 21–24 April 2024; pp. 1–6. [Google Scholar]
- Amodu, O.A.; Busari, S.A.; Othman, M. Physical layer aspects of terahertz-enabled UAV communications: Challenges and opportunities. Veh. Commun. 2022, 38, 100540. [Google Scholar] [CrossRef] [Scilit]
- Pang, L.; Song, K.; Miao, P.; Huang, C.; Ji, B.; An, Z.; Chen, G. Dynamic Interference Management by Using the Enhanced Clustering and Deep Reinforcement Learning in VLC-Enabled UAV Communication. IEEE Trans. Consum. Electron. 2025, 71, 7584–7596. [Google Scholar] [CrossRef] [Scilit]
- Orikumhi, I.; Bae, J.; Park, H.; Kim, S. DRL-based Multi-UAV trajectory optimization for ultra-dense small cells. ICT Express 2023, 9, 1128–1132. [Google Scholar] [CrossRef] [Scilit]
- Cui, H.; Zhang, N.; Liu, P. Trajectory Optimization for 6G-UAV Based on Deep Reinforcement Learning. IEEE Trans. Veh. Technol. 2024, 73, 17935–17939. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Li, H.; Wang, X.; Xia, S.; Liu, T.; Wang, R. Resource Allocation in UAV-D2D Networks: A Scalable Heterogeneous Multi-Agent Deep Reinforcement Learning Approach. Electronics 2024, 13, 4401. [Google Scholar] [CrossRef] [Scilit]
- Karegar, P.A.; Al-Hamid, D.Z.; Chong, P.H.J. Deep Reinforcement Learning for UAV-Based SDWSN Data Collection. Future Internet 2024, 16, 398. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Liu, C.H.; Yang, H.; Wang, G.; Leung, K.K. Ensuring Threshold AoI for UAV-Assisted Mobile Crowdsensing by Multi-Agent Deep Reinforcement Learning With Transformer. IEEE ACM Trans. Netw. 2024, 32, 566–581. [Google Scholar] [CrossRef] [Scilit]
- Xiao, K.; Yu, Z.; Wang, J.; Gao, F. Proximal Policy Optimization Algorithm for Enhancing Energy Harvesting in UAV-Assisted Communications with RIS. In Proceedings of the 2024 IEEE Wireless Communications and Networking Conference (WCNC); IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar]
- Alam, M.M.; Moh, S. Joint Trajectory Control, Frequency Allocation, and Routing for UAV Swarm Networks: A Multi-Agent Deep Reinforcement Learning Approach. IEEE Trans. Mob. Comput. 2024, 23, 11989–12005. [Google Scholar] [CrossRef] [Scilit]
- Al-Hourani, A.; Kandeepan, S.; Jamalipour, A. Modeling air-to-ground path loss for low altitude platforms in urban environments. In Proceedings of the 2014 IEEE Global Communications Conference, Cape Town, South Africa, 8–12 December 2014; pp. 2898–2904. [Google Scholar]
- Jiang, Y.; Zhai, L.; Wu, T.; Li, B.; Zou, Y.; Yan, P. Energy-Efficiency Optimization for RIS-Assisted UAV-Enabled IoT Networks. IEEE Internet Things J. 2025, 12, 42599–42612. [Google Scholar] [CrossRef] [Scilit]
- Gong, S.; Qu, K.; Wang, H.; Wang, W.; Huang, H.; Qu, P.; Chen, Q. Spatio-Temporal Trajectory-Driven Dynamic TDMA Scheduling for UAV-Assisted Wireless-Powered Communication Networks. Electronics 2026, 15, 1861. [Google Scholar] [CrossRef] [Scilit]
- Shen, X.; Gu, L.; Yang, J.; Shen, S. Energy Efficiency Optimization for UAV-RIS-Assisted Wireless Powered Communication Networks. Drones 2025, 9, 344. [Google Scholar] [CrossRef] [Scilit]
- Zeng, Y.; Xu, J.; Zhang, R. Energy minimization for wireless communication with rotary-wing UAV. IEEE Trans. Wirel. Commun. 2019, 18, 2329–2345. [Google Scholar] [CrossRef] [Scilit]
- Rezaei, O.; Masjedi, M.; Naghsh, M.M.; Gazor, S.; Nayebi, M.M. Cooperative Throughput Maximization in a Multi-Cluster WPCN. IEEE Trans. Green Commun. Netw. 2024, 8, 1505–1520. [Google Scholar] [CrossRef] [Scilit]
- Tang, H.; Wu, Q.; Chen, W.; Wang, J.; Li, B. Mitigating the Doubly Near–Far Effect in UAV-Enabled WPCN. IEEE Trans. Veh. Technol. 2021, 70, 8349–8354. [Google Scholar] [CrossRef] [Scilit]
- Xie, X.; Dong, J.; Han, J.; Cheng, G. Does YOLO Really Need to See Every Training Image in Every Epoch? arXiv 2026, arXiv:2603.17684. [Google Scholar]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |