Skip to Content
ElectronicsElectronics
  • Article
  • Open Access

28 September 2026

17 Pages

PR-Nav: Potential-Shaped Reinforcement Learning with Physics-Modulated Manifolds for Agile Autonomous Navigation

,
,
,
,
and
1
China Yangtze Power Co., Ltd., Yichang 443002, China
2
State Key Laboratory of Robotics and Intelligent Systems, Shenyang Institute of Automation, Chinese Academy of Sciences, Shenyang 110016, China
*
Author to whom correspondence should be addressed.
This article belongs to the Section Artificial Intelligence

Abstract

High-speed autonomous navigation of micro aerial vehicles (MAVs) is essential for inspection tasks in restricted spaces (e.g., dense industrial facilities or narrow corridors), where reliable spatial perception and stable flight control are required. However, existing deep reinforcement learning-based navigation methods often suffer from delayed obstacle avoidance responses due to the low fitting efficiency of raw geometric distance representations for dynamic threats, trigger control oscillations that disrupt flight stability due to an over-reliance on external hard-clipping mechanisms, and frequently fall into local optima such as obstacle-edge hovering driven by heuristic penalties, thereby reducing overall navigation success rates and training efficiency. We propose PR-Nav, an agile navigation framework combining potential-field-augmented perceptual representation with a physics-modulated probabilistic action manifold. PR-Nav comprises three components: (1) a potential-field-augmented ray tensor representation that nonlinearly maps raw distance information into physical repulsive gradients to improve the policy network’s fitting efficiency for dynamic threat boundaries; (2) a physics-prior-modulated Beta action manifold that injects local potential-field gradients as residual biases into the probability density generation process, replacing external hard clipping with internal network constraints to reduce the action smoothness index (ASI); and (3) a potential-based reward shaping (PBRS) mechanism that replaces traditional heuristic penalties with global physical potential energy differences to improve the navigation success rate in complex environments and accelerate policy convergence. Experiments on the high-fidelity Isaac Sim simulator and a real-world physical flight platform show that PR-Nav achieves an overall navigation success rate of 96.5%, outperforming the strongest competing method by 10.8 percentage points. In quantitative evaluations of flight smoothness and training efficiency, its action smoothness index (ASI) is reduced to 4.25 m/s 3 and it achieves an approximate 54% reduction in convergence steps compared to the baseline, the best results among the compared methods. Ablation studies verify the contribution of each component to its corresponding metrics. These results demonstrate that combining physical potential field modeling with underlying probability distribution reconstruction is an effective route to robust and agile navigation in dynamic restricted spaces.

1. Introduction

High-speed autonomous navigation of micro aerial vehicles (MAVs) in cluttered, dynamic, and GPS-denied environments is a central challenge in embodied intelligence research. Recently, MAVs have been increasingly deployed in complex restricted spaces, such as the internal inspection of hydropower infrastructure, maritime vessels, and dense industrial facilities. Achieving agile, oscillation-free flight under such spatial constraints requires control policies that are responsive and aerodynamically stable.
Although deep reinforcement learning (DRL) provides an end-to-end navigation paradigm [1], current algorithms struggle to balance agile maneuvering with safety guarantees due to fragmented system designs. Mainstream methods typically treat perception, action constraints, and reward generation as isolated engineering tasks. For instance, relying on raw geometric ray-casting distances [2,3] fails to provide physical risk semantics, forcing the neural network to implicitly infer collision gradients. Furthermore, to ensure safe exploration, recent Safe RL frameworks often append external mathematical shields—such as Velocity Obstacles (VO) [4] or Control Barrier Functions (CBFs) [5]—for post hoc action clipping. However, a critical analysis reveals that this forced intervention disrupts the Markov chain continuity of the exploration manifold, inducing oscillations in low-level control commands and degrading flight stability.
Concurrently, the widely adopted distance-based heuristic penalties used to train these policies lack theoretical grounding. Such manually assembled multi-objective rewards readily trigger reward hacking [6,7], where policies learn to hover safely near obstacles rather than efficiently navigating towards the goal, trapping the agent in suboptimal local minima.
To address these deficiencies, we propose PR-Nav, an agile navigation framework that transcends fragmented engineering patches by performing a systematic reconstruction of the underlying Markov Decision Process (MDP). Rather than simply combining existing techniques, PR-Nav establishes a cohesive closed-loop system: it transforms raw geometric observations into a potential-field-augmented risk tensor to establish an explicit spatial gradient; it then injects this physical gradient directly into a continuous Beta action manifold as a residual bias, internalizing safety boundaries within the probability density generation to eliminate external hard-clipping; finally, it couples this continuous manifold with Potential-Based Reward Shaping (PBRS) to guarantee policy invariance. This integration aligns the MDP formulation with physical safety constraints. The primary contributions of this work are summarized as follows:
  • Theoretical MDP Reconstruction: We transcend heuristic multi-objective penalties by introducing a reward mechanism uniquely coupled with the continuous action manifold. By formulating PBRS, the framework guarantees policy invariance while leveraging global potential energy differences to preclude reward hacking.
  • Physics-Modulated Action Manifold: We reconstruct the probability flow of the policy network by injecting local potential-field gradients as probabilistic residual biases into a Beta distribution. This achieves internal safety constraint satisfaction, eliminating the need for post hoc external VO/CBF clipping and suppressing low-level action jitter.
  • Potential-Field-Augmented Perception: We validate a spatial tensor representation that nonlinearly converts raw range measurements into physical risk gradients. Extensive evaluations on a simulator and a physical UAV platform demonstrate navigation success rates and zero-shot sim-to-real transfer capabilities.

3. Methodology

Existing navigation frameworks are beset by structural deficiencies in state representation, action constraints, and reward design that collectively constrain agile continuous flight. In contrast to methods that rely on external patches or heuristic tuning, PR-Nav reconstructs the tensor and probability flow of reinforcement learning from first principles. The overall network architecture of PR-Nav is illustrated in Figure 1.
Figure 1. Overall architecture of the PR-Nav agile autonomous navigation framework. Left: high-dimensional LiDAR scans and low-dimensional kinematic state inputs. Center: the FE-Rep network extracts environmental risk features in parallel via a temporal GRU and an artificial potential field branch. Right: the Beta Manifold Modulation module injects physical gradients as residuals into the probability flow, substantially improving the continuity and safety of control commands at the lowest level.
The system first processes perceptual inputs through a Feature Extraction and Representation (FE-Rep) Network. Since a single static geometric ray-cast is incapable of inferring the velocity vectors or behavioral intent of moving obstacles in a dynamic environment, FE-Rep incorporates a Temporal Gated Recurrent Unit (GRU) to implicitly extract dynamic environmental features from the LiDAR sequence S laser ∈ R 37 × T —i.e., a 37-dimensional single-frame range vector downsampled at 5° resolution over a 180° forward field of view—with a sliding temporal window T = 3 . In comparison with large-parameter self-attention architectures such as Transformers, GRU is adopted for its substantially lower parameter count and microsecond-level inference latency, properties that satisfy the temporal-memory requirements of agile obstacle avoidance within the onboard computational budget of a Jetson Orin NX.
The extracted temporal LiDAR features are subsequently fused at the feature level with physical potential-field features. The Actor–Critic core network maps the fused representation to raw navigation actions, whereupon the Beta Manifold Modulation module imposes physical constraints on the output, facilitating smooth and safe control commands.

3.1. Potential-Field-Augmented Ray Tensor Representation

In the perceptual reconstruction designed for complex dynamic environments, the conventional practice of directly feeding raw geometric ray-casting distances as inputs is abandoned. Existing frameworks such as NavRL [3] typically stack the 3-D ray detection lengths d i j from onboard sensors into an observation matrix S stat ∈ R N h × N v . Such raw distance representations, however, lack physical gradient constraints on environmental risk, compelling CNNs to expend considerable sampling cost to implicitly infer the hazard level of obstacles.
To resolve this perceptual fitting delay, a potential-field-augmented tensor representation is proposed. A nonlinear mapping is defined that transforms the geometric length d i j of each ray into a scalar Φ i j reflecting the physical obstacle-avoidance risk:
Φ i j = η · 1 d i j − 1 d safe 1 d i j 2 if d i j ≤ d safe , 0 if d i j > d safe ,
where d safe is the preset safe influence radius and η is a gain coefficient. Specifically, this nonlinear mapping applies an artificial potential field formulation to each discrete sensing ray. As an obstacle approaches within the safe radius, the scalar risk value increases exponentially, effectively converting a purely geometric spatial distance into a physical repulsive intensity. Through this transformation, a novel risk prior tensor S stat is constructed whose elements directly encode the repulsive gradient of space.
This risk tensor is concatenated with the raw state prior to input into the policy network. By doing so, the network receives explicit guidance on risk distribution rather than merely geometric distance when processing sensor data, substantially relieving the decision layer’s perceptual burden and endowing the model with clearer spatial potential-energy awareness that accelerates convergence in cluttered dynamic scenarios.

3.2. Physics-Prior-Modulated Beta Action Manifold

To formally bridge the perceptual risk tensor S stat established in Section 3.1 with the low-level flight execution and to significantly mitigate control-command oscillation during high-speed flight, the action-output mechanism of the policy network is reconstructed. In baseline frameworks such as NavRL, the Actor network directly outputs the parameters α and β of a Beta distribution and samples from it to obtain a normalized action vector. Although this formulation accommodates continuous action spaces, the sampling process relies entirely on data-driven weight updates without responsiveness to the instantaneous physical environment, causing the UAV to produce extreme, discontinuous commands in complex boundary regions.
In lieu of appending an external safety clipper after the network output, a physics-prior-modulated action manifold mechanism is proposed. The local potential-field gradient vector ∇ Φ ∈ R 2 extracted by the front-end network is dynamically injected as a residual bias into the Beta distribution generation operator. Notably, this gradient vector is not obtained by analytically differentiating discrete radar rays; rather, it is a physical bias feature produced by the Actor core network upon mapping the risk prior tensor S stat into the robot’s 2-D kinematic coordinate system (X–Y plane). The corrected action-sampling distribution parameters are:
α ′ = α + ReLU w α · ∇ Φ ,
β ′ = β + ReLU w β · ( − ∇ Φ ) ,
where w α and w β are learnable mapping weights. In this manner, the guiding force of the physical potential field is directly encoded as the skewness of the probability density. When the UAV senses an increasing repulsive gradient, the mean of the Beta distribution shifts automatically toward the safe manifold at the probabilistic level.
The theoretical foundation for selecting the Beta distribution and the ReLU modulation operator stems from the enforcement of physical boundaries within the continuous action space. Unlike conventionally utilized Gaussian distributions that possess infinite support and necessitate post hoc hard clipping when sample values exceed actuator limits, the Beta distribution is natively bounded within a compact interval [ 0 , 1 ] . This mathematical property inherently aligns the network’s stochastic exploration flow with the maximum aerodynamic velocity and acceleration constraints of the rotorcraft. Furthermore, the shape parameters α ′ and β ′ of the Beta density function must satisfy the positive domain condition (i.e., α ′ , β ′ > 0 ) to remain mathematically well-defined. Since the potential field gradient ∇ Φ can yield negative directional values depending on the obstacle topology, the incorporation of the Rectified Linear Unit ( ReLU ) operator acts as a mathematical shield. It ensures that only positive repulsive components actively skew the action distribution toward the safe manifold, while precluding non-positive values from violating the shape parameter domain, thereby maintaining low-level decision continuity and numerical stability under intense dynamic perturbations.
The significance of this reconstruction lies in its transformation of the obstacle-avoidance constraint from external hard clipping to internal probabilistic guidance. This not only significantly enhances the mathematical continuity and smoothness of the output velocity command V ctrl , but also preserves the exploratory properties of reinforcement learning to the greatest extent. Experimental results confirm that this internalized action manifold constraint substantially suppresses system oscillations attributable to frequent VO triggering, substantially improving motion stability during high-speed UAV flight.

3.3. Potential-Based Policy-Invariant Reward Shaping

In an MDP, the design of the reward function directly determines the final performance of a reinforcement learning policy. In baseline frameworks such as NavRL, safety rewards—including a static safety reward r s s and a dynamic safety reward r d s —are typically formulated as a logarithmic sum of obstacle distances. This manually assembled heuristic multi-objective reward function r = λ 1 r vel + λ 2 r s s + λ 3 r d s + λ 4 r smooth + λ 5 r height not only demands laborious weight tuning but is also theoretically flawed: it readily triggers reward hacking, causing the UAV to hover or oscillate near obstacle edges in order to accumulate distance-based safety rewards, thereby deviating from the objective of rapid target-reaching and converging to suboptimal local minima.
To resolve this structural deficiency and accelerate network convergence, the heuristic penalty terms are replaced with rigorous Potential-Based Reward Shaping (PBRS) theory. The theory mathematically establishes that appending a global potential function Φ ( s ) in a specific difference form to the environment’s raw reward substantially improves learning efficiency without altering the Optimal Policy Invariance of the original task.
First, to align with the reward maximization objective of the MDP, the global physical potential function Φ ( s t ) is defined as a value potential (where higher values indicate safer and closer states):
Φ ( s t ) = − κ g ∥ P r − P g ∥ − κ o ∑ i ∈ O 1 ∥ P r − P o i ∥ 2 ,
where P r and P g are the positions of the robot and the goal, respectively; P o i is the coordinate of obstacle i; and κ g , κ o are the corresponding positive potential-energy coefficients. Note that the negative coefficient − κ g ensures that as the UAV approaches the goal ( ∥ P r − P g ∥ decreases), the potential value Φ ( s t ) increases (i.e., becomes less negative). This guarantees that progressing toward the target generates a positive shaping reward F ( s t , s t + 1 ) > 0 , aligning with the reward maximization objective.
Next, the potential difference between two consecutive states is defined as the internal shaping reward F ( s t , s t + 1 ) :
F ( s t , s t + 1 ) = γ Φ ( s t + 1 ) − Φ ( s t ) ,
where γ is the reward discount factor. The total reward R total used by PPO for advantage estimation and policy updates is then reconstructed as:
R total = R task + F ( s t , s t + 1 ) .
Under this rigorous mathematical framework, R task need only retain the purest task-objective reward (e.g., a sparse arrival reward or a basic velocity reward). Because F ( s t , s t + 1 ) takes a temporal-difference form, the potential contributions of intermediate states cancel via telescoping summation over a complete episode. This mechanism furnishes high-frequency gradient guidance toward the optimal path at each exploration step, directing the agent along the direction of increasing potential value. At the same time, it effectively prevents the policy network from stagnating or oscillating in pursuit of local safety rewards, thereby preserving the integrity of high-speed agile navigation. Furthermore, to ensure the theoretical validity of the Policy Invariance guarantee, boundary conditions are enforced. First, because the positions of dynamic obstacles ( P o i ) are implicitly encoded within the agent’s temporal LiDAR observation sequence, Φ ( s ) remains a valid potential function of the observable Markov state, even in non-stationary environments. Second, to properly handle episode truncation and terminal states, the potential of any terminal state (whether goal-reaching or collision) is defined as Φ ( s t e r m i n a l ) = 0 . This boundary handling prevents shaping-induced biases during value bootstrapping at the end of finite-horizon episodes, grounding the proposed framework in classical PBRS theory.

4. Experiments and Results

4.1. Experimental Setup and Hardware Configuration

To eliminate evaluation variance caused by random seeds, all reinforcement learning baselines and the proposed algorithm are independently trained and tested with five different random seeds; all reported metrics are presented in the form “mean ± std.”
Training and simulation evaluation are conducted entirely within the NVIDIA Isaac Sim high-fidelity physics engine, which faithfully simulates the aerodynamics and collision feedback of UAVs in dynamic cluttered environments. Hardware-wise, training is performed on a workstation equipped with a single NVIDIA GeForce RTX 4090 (24 GB), leveraging Isaac Sim’s tensorized parallel computation to instantiate 1024 independent environments simultaneously, substantially accelerating PPO sample throughput. Algorithm inference for physical flight testing is deployed on the onboard NVIDIA Jetson Orin NX edge computing platform. For real-world flight tests, a custom micro quadrotor (wheelbase: 250 mm) is utilized, equipped with a Pixhawk 6C flight controller (Holybro Tech Co., Ltd., Shenzhen, China) and a lightweight solid-state LiDAR (MID-360, Livox Technology Co., Ltd., Shenzhen, China) for real-time point cloud acquisition. The physical test environment consists of a 10 m × 10 m × 5 m indoor space with a Vicon motion capture system providing ground-truth state estimation. To ensure reproducibility, the training budget for all learning-based baselines is capped at 2.0 × 10 7 steps. The MDP termination criteria are defined as follows: an episode successfully terminates if the UAV’s distance to the goal is less than 0.5 m, and fails if a collision is detected by the Isaac Sim rigid-body physics engine or if the flight time exceeds the maximum episode truncation limit of 1000 steps. Baseline implementation details, environment definitions, and hyperparameter tuning procedures are documented in Appendix A.

4.2. Baselines and Quantitative Comparison

To evaluate PR-Nav’s performance comprehensively, it is compared quantitatively under identical Isaac Sim conditions against four recent state-of-the-art UAV navigation algorithms: ViGO [20]—a vision-assisted dynamic-obstacle-avoidance planner based on B-spline optimization; SAC-Nav [21]—a pure data-driven continuous-action planner based on maximum-entropy RL; PPO-CBF [5]—a safe RL algorithm integrating Control Barrier Functions for hard collision-avoidance clipping; and NavRL [3]—the direct baseline framework using heuristic rewards and a VO shield.
In the comparative evaluations, all algorithms were configured with the same observation spaces (a 114-dimensional concatenated vector of temporal LiDAR scans and kinematics) and continuous action constraints (maximum linear velocity v m a x = 2.0 m/s, maximum yaw rate ω m a x = π rad/s at a 20 Hz control frequency). These boundary values correspond to the aerodynamic limits of the 250 mm MAV. The baseline hyperparameters follow the settings reported in their respective original papers. For the NavRL baseline, a grid search over the multi-objective reward weights was performed for the dynamic environment. The dynamic safety penalty λ 3 was evaluated within the range [−20.0, −5.0], while the velocity reward λ 1 and static penalty λ 2 were evaluated in [1.0, 5.0] and [−10.0, −1.0], respectively. The resulting configuration used for NavRL in the tests is ( λ 1 = 2.5 , λ 2 = − 5.0 , λ 3 = − 15.0 ).
Five core metrics are extracted for evaluation: Navigation Success Rate (SR), Average Flight Speed (AFS), Action Smoothness Index (ASI), Convergence Steps (CS), and onboard Inference Latency (IL). The Action Smoothness Index, which quantifies the degree of abrupt command jumps during continuous execution, is defined as the root mean square of the first-order derivative of the linear acceleration (i.e., jerk) over the flight episode:
ASI = 1 T ∑ t = 1 T a t − a t − 1 Δ t 2 ,
where T is the total control steps of the episode, a t ∈ R 3 is the linear acceleration vector at step t derived by differencing the low-level flight control commands, and Δ t = 0.05 s is the policy decision period. A lower ASI value indicates smoother velocity-command transitions and stricter compliance with the aerodynamic constraints of the physical rotorcraft.
To demonstrate the performance improvements, the evaluation metrics were analyzed across 50 independent evaluation trials (distinct from the 5 training seeds). For continuous performance metrics, including Average Flight Speed (AFS), Action Smoothness Index (ASI), and Inference Latency (IL), a Welch’s t-test was employed to account for unequal variances between the baseline and the proposed method, reporting significance at p < 0.01 . For the Navigation Success Rate (SR), which represents a binary Bernoulli outcome (success or failure), statistical confidence was verified using 95% confidence intervals rather than a standard t-test. The statistical analysis across all evaluation scenarios (as reported in Table 1, Table 2 and Table 3) confirms that PR-Nav’s improvements over the baselines and ablated variants are statistically significant.
Table 1. Comprehensive performance comparison of navigation algorithms in a complex dynamic environment.
Table 2. Quantitative ablation results of the PR-Nav core architecture.
Table 3. Quantitative comparison in the dynamic crossing corridor scenario (50 independent flights per algorithm).
Quantitative results in a high-density static obstacle and randomly moving pedestrian environment are shown in Table 1.
As shown in Table 1 and Figure 2 and Figure 3, PR-Nav outperforms all baselines in both SR (96.5%) and AFS (1.82 m/s), demonstrating that the framework maintains high safety while attaining greater flight speed. With regard to macro path selection, ViGO tends to circumvent the central obstacle-dense region, resulting in longer overall flight distances, whereas RL-based frameworks generally opt to traverse the high-density central area. At the micro control level, SAC-Nav and NavRL (without physical constraints) exhibit local oscillations in narrow gaps, while PPO-CBF, though it avoids collisions via its safety shield, produces sharp-angle deflections and step-like trajectories at obstacle edges, degrading aerodynamic continuity. PR-Nav, guided by potential fields, consistently selects the shorter path through the map center, with its Beta manifold producing smooth transition arcs during continuous obstacle avoidance—an observation consistent with the ASI metric (4.25 m/s3) in Table 1. The PBRS potential shaping provides physics-informed gradient guidance that curtails aimless exploration in early training, and PR-Nav’s single-step inference latency of 8.5 ms satisfies the real-time control requirements of common onboard edge platforms.
Figure 2. Training convergence curves for PR-Nav and the RL baselines NavRL [3], PPO-CBF [5], and SAC-Nav [21]. PR-Nav reaches a stable convergence plateau at approximately 6.5 M steps, achieving a 54% reduction in convergence steps compared to the NavRL baseline (6.5 M vs. 14.2 M). This demonstrates that PBRS potential shaping provides effective physical-gradient guidance to reduce aimless exploration in early training.
Figure 3. 2-D flight trajectory visualization of the five algorithms: (a) ViGO; (b) SAC-Nav; (c) PPO-CBF; (d) NavRL; and (e) PR-Nav (Ours). Colored curves denote Task P1 (blue), Task P2 (orange), Task P3 (green), and Task P4 (red). Colored square and circular markers denote the start and goal positions, respectively. Black geometric shapes represent static obstacles with different geometries. The four representative topological tasks comprise diagonal crossing (P1/P4), wide-span traversal (P2), and lateral depth traversal (P3). PR-Nav trajectories exhibit smooth transition arcs during continuous obstacle avoidance, in contrast to the sharp deflections or local oscillations produced by baseline algorithms.
To further verify deployability on resource-constrained devices, the per-step inference computation time on the Jetson Orin NX is decomposed as follows: (1) FE-Rep risk feature extraction (including LiDAR tensor mapping and Temporal GRU): ≈2.8 ms; (2) Actor backbone network (MLP): ≈4.2 ms; (3) Physics-prior-modulated Beta manifold (residual bias injection): ≈1.5 ms. The proposed operator-level reconstruction incurs only microsecond-level additional overhead while delivering a substantial improvement in trajectory smoothness (ASI), striking an excellent balance between computational complexity and navigation safety—fully satisfying the real-time requirement (>100 Hz) of agile MAV low-level flight controllers.

4.3. Ablation Study

To systematically assess the individual contributions of each of the three low-level reconstruction modules—potential-field-augmented tensor representation, physics-modulated action manifold, and PBRS reward shaping—an ablation study is conducted under identical complex dynamic simulation environments with the following single-module removal variants:
  • w/o State: removes the potential-field-augmented tensor representation; degrades to raw geometric distance matrix input of the baseline.
  • w/o Action: removes the physics-bias-modulated Beta action manifold; degrades to the original post hoc VO-shield clipping mechanism.
  • w/o Reward: removes rigorous PBRS reward shaping; degrades to the heuristic log-distance reward of the original baseline.
As shown in Table 2, the three modules of PR-Nav are tightly coupled and mutually indispensable. PBRS reward shaping constitutes the core module for surmounting performance bottlenecks: its removal causes the model to revert to reward hacking behavior, with convergence time doubling to 13.5M steps and SR declining to 88.4%. Without the global gradient guidance of PBRS, the UAV struggles to escape local minima induced by heuristic penalties, leading to prolonged hovering or trajectory stagnation in highly cluttered regions.
Physics-modulated action manifold governs both smoothness and the performance ceiling: reversion to post hoc hard clipping causes ASI to surge to 11.75 m/s3, necessitating frequent deceleration. Specifically, the absence of this probabilistic internal modulation means the policy network generates unconstrained, aggressive commands that are subsequently sharply truncated by the external VO shield, resulting in the violent trajectory oscillations and instability observed in Figure 4b.
Figure 4. 3-D flight trajectory visualization of four model variants in a high-density pillar environment: (a) PR-Nav (w/o State); (b) PR-Nav (w/o Action); (c) PR-Nav (w/o Reward); and (d) PR-Nav (Ours). The white, red, yellow, and blue trajectories correspond to panels (a–d), respectively. The left and right white spherical markers denote the start and goal positions, respectively. The multicolored vertical pillars represent static obstacles; their colors are used only for visual distinction and do not encode additional variables. PR-Nav (Ours) produces the smoothest trajectory with appropriate clearance from obstacles, while variants without key modules exhibit oscillations or sharp-angle deflections.
Potential-field-augmented tensor representation effectively alleviates perceptual fitting pressure: the removal of this prior results in delayed reaction to dynamic threats, reducing SR to 91.2% and increasing interaction cost. By explicitly providing spatial repulsive gradients at the perception layer, this module frees the Actor network from implicitly inferring obstacle boundaries, thereby significantly enhancing the system’s dynamic responsiveness. Collectively, these three modules address the structural deficiencies of the baseline algorithm, yielding agile and safe UAV performance in complex environments.

4.4. Hyperparameter Sensitivity Analysis

To further assess the robustness of the framework with respect to key physical hyperparameters, and to rule out overfitting to specific parameter combinations, single-variable sensitivity analyses are conducted on the safe influence radius d safe and the goal attraction coefficient κ g . Experiments are conducted in the standard dynamic perturbation environment: all other network parameters are held fixed, d safe is swept over [ 0.4 , 2.0 ] m, and κ g is varied over [ 0.1 , 5.0 ] ; the average SR and convergence steps CS over 50 independent random tests are recorded. Results are presented in Figure 5.
Figure 5. Sensitivity analysis of core physical hyperparameters. (a) Effect of the safe influence radius d safe on navigation success rate and convergence steps. (b) Effect of the goal attraction coefficient κ g on performance. Both parameters exhibit wide stable plateaus, indicating that PR-Nav does not depend on extreme tuning. The red shaded regions indicate the variability in navigation success rate across the 50 independent random tests.
As shown in Figure 5a, when d safe is varied within the broad physical range [ 0.8 , 1.6 ] m, SR remains stably above 94%, indicating that the physics-modulated action manifold possesses strong adaptive buffering capacity. When d safe is reduced below 0.6 m, the network loses sufficient advance risk-representation guidance, causing a pronounced SR drop. Conversely, when d safe exceeds 1.8 m, the excessively wide repulsion range renders the policy overly conservative in narrow passages, prolonging exploration and substantially increasing CS.
Figure 5b illustrates the effect of κ g . A moderate attractive coefficient ( κ g ∈ [ 0.5 , 2.0 ] ) provides the optimal physical gradient for PBRS reward shaping, accelerating convergence to 6.5 M steps. A small κ g (e.g., 0.1) weakens goal-directedness and leads to convergence divergence, whereas a large κ g (e.g., 5.0) diminishes the relative weight of the safety repulsion, slightly reducing SR. Taken together, these results demonstrate that PR-Nav maintains high safety and rapid convergence across a reasonable range of physical parameters, without dependence on extreme tuning. In addition to physical parameters, the sensitivity of the framework to core algorithmic learning hyperparameters was evaluated, specifically focusing on the PPO learning rate and the discount factor γ . Empirical evaluations reveal that a learning rate of 3 × 10 − 4 provides the optimal balance between convergence speed and stability. Higher learning rates (> 5 × 10 − 4 ) induced severe policy gradient oscillations and degraded the ASI by causing abrupt velocity command shifts, whereas lower rates (< 1 × 10 − 4 ) significantly prolonged convergence beyond 10 M steps. Similarly, the discount factor γ = 0.99 proved crucial for long-horizon spatial potential planning. Reducing γ to 0.95 caused the policy network to act myopically—prioritizing immediate obstacle avoidance over global path optimality—which subsequently reduced the navigation success rate to 82.3% in dense clutter. These findings confirm that the chosen hyperparameter configuration in Appendix A is located within a robust and optimal performance plateau.

4.5. Complex Dynamic Scene Analysis

To concretely demonstrate the maneuverability advantage of PR-Nav following low-level MDP reconstruction, a “dynamic crossing corridor” scenario is constructed in Isaac Sim: the UAV must traverse a corridor filled with static pillars while simultaneously contending with multiple dynamic obstacles moving laterally at 1.0–1.5 m/s. Table 3 presents core quantitative data from 50 independent flights.
Whereas all baselines suffer severe performance degradation in this extreme scenario, PR-Nav demonstrates superior dynamic adaptability: sustaining high-speed flight at 1.78 m/s with near-zero emergency braking (0.1 events per episode) and a maximum deceleration of only 3.12 m/s 2 . This performance is directly attributable to PBRS reward shaping and the physics-modulated action manifold, through which the network learns compliant obstacle avoidance guided by physical gradients rather than repeatedly rebounding at safety boundaries.
Real-World Physical Flight Quantitative Evaluation. To fulfill the validation requirements in real physical environments, a series of 20 independent real-world flight trials were conducted in the 10 m × 10 m × 5 m Vicon-instrumented testing volume. The custom 250 mm MAV was tasked with navigating across randomly moving obstacles at a target speed of 2.0 m/s. Quantitatively, PR-Nav achieved a 95.0% navigation success rate with an average flight speed (AFS) of 1.65 ± 0.08 m / s and maintained a low action smoothness index (ASI) of 4.52 ± 0.35 m / s 3 . Compared with the high-fidelity simulation results in Table 1, the minor performance degradation (a 1.5% drop in success rate and a 9.3% reduction in AFS) is attributable to unmodeled real-world physical factors, such as aerodynamic ground effects, motor transient latency, and point cloud registration noise. Nevertheless, these physical flight statistics confirm that the operator-level MDP reconstruction of PR-Nav effectively preserves its safety and smoothness advantages during zero-shot Sim-to-Real transfer, without requiring post-deployment parameter tuning.

4.6. Dense Random Clutter Forest Benchmark

To validate PR-Nav’s zero-shot transfer capability in purely unknown, unstructured environments, a challenging Dense Random Clutter Forest benchmark is constructed in Isaac Sim, deliberately eschewing the regular geometric features of conventional test scenarios.
The environment generation rules are as follows: within a 20 m × 20 m effective flight space, 150–200 static obstacles are procedurally generated at random positions. Obstacle cross-sections are randomly assigned as cylinders (radius 0.2–0.5 m) or rectangular pillars (side 0.3–0.8 m), with heights uniformly distributed in [ 0 , 5 ] m. The resulting space exhibits non-convex, irregular, high-density occlusion characteristics, which negate simple geometry-inspired 2-D avoidance paths and impose extreme exploration pressure on the policy network’s underlying MDP.
In such non-convex dense environments, baselines relying on external VO or post hoc clipping readily become trapped in local minima owing to frequent safety boundary triggering, resulting in severe control jitter and stagnation. PR-Nav, by contrast, leverages its potential-field-augmented ray tensor to nonlinearly map raw range measurements into physical risk gradients; guided by PBRS reward shaping, the policy receives high-frequency physical-gradient guidance, while the physics-modulated Beta manifold mathematically enforces smooth, aerodynamically feasible velocity commands. As illustrated in Figure 6, these results validate PR-Nav’s agility and generalizability in compute-constrained settings under extreme clutter, stemming from a principled reconstruction of the spatial probability flow.
Figure 6. Dense Random Clutter Forest benchmark test environment and a representative PR-Nav flight trajectory. The non-convex irregular high-density layout imposes extreme exploration pressure on the policy network. PR-Nav successfully navigates through narrow gaps with smooth trajectories, demonstrating robust zero-shot generalization. The cyan curve denotes the representative flight trajectory, while the green and pink markers indicate the start and goal positions, respectively. The different obstacle colors are used only for visual differentiation and do not denote distinct obstacle categories.

4.7. Failure Case and Boundary Condition Analysis

Although PR-Nav achieves a navigation success rate of 96.5% in highly complex dynamic environments, the remaining 3.5% failure cases (collisions or timeouts) are systematically analyzed to characterize the performance limits of the algorithm under extreme boundary conditions.
Deep Dynamic Potential Well. When multiple randomly moving obstacles happen to form a closed or semi-closed “U-shape” encirclement directly ahead of the UAV, the PBRS generates a strong local repulsive gradient. If the target point lies beyond this encirclement, the local repulsion and global attraction may reach force equilibrium. Although PR-Nav’s RL exploration mechanism (Beta manifold perturbation) endows the UAV with a degree of escape capability, when the relative velocity of the encircling obstacles is sufficiently high, the UAV may nevertheless become trapped in a “dynamic potential well” and stall in place, ultimately triggering a timeout.
Beyond Kinematic Limits. In extreme test cases involving rapidly moving pedestrians, if a dynamic obstacle (e.g., an oncoming non-cooperative aircraft) establishes a head-on collision trajectory with the UAV at a relative distance below the physical braking limit, an unavoidable collision ensues. Log analysis reveals that in these failure instants, the PR-Nav Actor network has already issued the correct maximum evasive command; nonetheless, owing to the physical MAV’s maximum jerk constraint and motor response delay (∼tens of milliseconds), the platform is unable to complete the high-load maneuver before impact. Such failures originate from hardware actuator and aerodynamic physical limits, rather than from any deficiency in the algorithmic decision logic.
Physical Sensing Limitations. During real-world flight tests, a specific subset of failures originated from the inherent sensing conditions of the onboard LiDAR. When navigating through highly constrained physical environments, the limited vertical field of view (FOV) and occasional point cloud degradation (e.g., caused by highly reflective or transparent obstacles) resulted in transient losses of risk awareness. This intermittent perception delay occasionally caused the risk tensor S stat to underestimate local threats, triggering the UAV’s safety hovering mechanism and ultimately leading to task timeouts. Addressing such hardware-induced perceptual blind spots through multi-modal sensor fusion represents a critical area for future deployment.
Future work may consider incorporating velocity-vector forecasting from multi-frame historical obstacle sequences, thereby upgrading reactive potential field avoidance to a predictive “dynamic manifold reconstruction” paradigm that further alleviates the deep potential well constraint.

5. Conclusions

To address obstacle-avoidance jitter and slow convergence of UAVs in complex dynamic environments, this paper proposes PR-Nav, an agile navigation framework grounded in potential-function shaping and physics-modulated manifolds. The framework achieves a deep reconstruction of the low-level decision process: the potential-field-augmented tensor supplies explicit physical priors; the physics-modulated Beta manifold internalizes safety constraints to ensure continuous, smooth control commands; and PBRS reward shaping accelerates convergence while guaranteeing optimal policy invariance. Both simulation and onboard physical experiments demonstrate that PR-Nav comprehensively outperforms baseline algorithms in success rate, maneuverability, and smoothness, while exhibiting robust zero-shot generalization and effectively handling real-world tasks such as dam inspection. In future work, this general paradigm will be extended to 2-D manifold-constrained spaces [22], thereby providing fundamental algorithmic support for high-speed exploration and multi-agent collaboration on heterogeneous platforms such as wall-climbing robots.

Author Contributions

Conceptualization, T.W. and G.Q.; methodology, T.W. and Y.C.; software, G.Q.; validation, G.Q., H.Z. and C.Y.; formal analysis, T.W.; investigation, G.Q. and H.Z.; resources, T.W. and F.J.; data curation, F.J.; writing—original draft preparation, G.Q.; writing—review and editing, T.W. and C.Y.; visualization, G.Q.; supervision, T.W.; project administration, T.W.; funding acquisition, T.W. and C.Y. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China, grant number 52575042; and the Fundamental Research Project of SIA, grant number 2023JC2K04.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.The data are not publicly available due to confidentiality restrictions imposed by the collaborating company.

Acknowledgments

The authors thank the members of the State Key Laboratory of Robotics at Shenyang Institute of Automation (Chinese Academy of Sciences) for their support in the physical flight experiments.

Conflicts of Interest

Authors Ge Qin, Heng Zhong and Chang’e Yu were employed by the company China Yangtze Power Co., Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.

Abbreviations

The following abbreviations are used in this manuscript:
ASIAction Smoothness Index
AFSAverage Flight Speed
CBFControl Barrier Function
CNNConvolutional Neural Network
CSConvergence Steps
DRLDeep Reinforcement Learning
FE-RepFeature Extraction and Representation
GRUGated Recurrent Unit
ILInference Latency
MAVMicro Aerial Vehicle
MDPMarkov Decision Process
MLPMultilayer Perceptron
PBRSPotential-Based Reward Shaping
PPOProximal Policy Optimization
RLReinforcement Learning
RMSRoot Mean Square
SACSoft Actor-Critic
SRSuccess Rate
UAVUnmanned Aerial Vehicle
VOVelocity Obstacles

Appendix A. Network Architecture and Hyperparameters

To facilitate reproducibility, the detailed network architecture and PPO training hyperparameters are summarized in Table A1. The Actor and Critic networks share the same FE-Rep feature extraction backbone, while their output heads consist of independent multi-layer perceptrons (MLPs) with ELU activation functions.
Table A1. PPO Hyperparameters and Network Architecture Details.
To comprehensively address the computational overhead, the model parameter counts and runtime memory consumption of the learning-based baselines are compared in Table A2. PR-Nav introduces only a marginal increase in parameters (0.18 M) compared to the NavRL baseline (0.15 M) due to the lightweight Temporal GRU and residual bias layers. The inference memory footprint (24.5 MB) remains highly optimal for resource-constrained edge devices such as the Jetson Orin NX.
Table A2. Computational Complexity Comparison of Learning-based Baselines.

References

  1. Song, Y.; Shi, K.; Penicka, R.; Scaramuzza, D. Learning perception-aware agile flight in cluttered environments. In Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK, 29 May–2 June 2023; pp. 1989–1995. [Google Scholar]
  2. Huang, X.; Zhang, J.; Sun, P.; Tang, Y.; Liu, Y. A survey of path planning algorithms based on deep reinforcement learning. Robot 2026, 48, 196–216. [Google Scholar]
  3. Xu, Z.; Han, X.; Shen, H.; Jin, H.; Shimada, K. NavRL: Learning safe flight in dynamic environments. IEEE Robot. Autom. Lett. 2025, 10, 3668–3675. [Google Scholar] [CrossRef] [Scilit]
  4. He, T.; Zhang, C.; Xiao, W.; He, G.; Liu, C.; Shi, G. Agile but safe: Learning collision-free high-speed legged locomotion. In Proceedings of the Robotics: Science and Systems (RSS), Delft, The Netherlands, 15–19 July 2024. [Google Scholar]
  5. Liu, C.; Xiao, W.; Shi, G. PPO-CBF: Safe reinforcement learning via control barrier functions. IEEE Trans. Autom. Control 2024, 69, 456–468. [Google Scholar]
  6. Romero, A.; Song, Y.; Scaramuzza, D. Actor-Critic Model Predictive Control. In Proceedings of the IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; pp. 14777–14784. [Google Scholar]
  7. Xu, B.; Gao, F.; Yu, C.; Zhang, R.; Wu, Y.; Wang, Y. OmniDrones: An efficient and flexible platform for reinforcement learning in drone control. IEEE Robot. Autom. Lett. 2024, 9, 2838–2844. [Google Scholar] [CrossRef] [Scilit]
  8. Singla, A.; Padakandla, S.; Bhatnagar, S. Memory-based deep reinforcement learning for obstacle avoidance in UAV with limited environment knowledge. IEEE Trans. Intell. Transp. Syst. 2021, 22, 107–118. [Google Scholar] [CrossRef]
  9. Li, X.; Fang, J.; Du, K.; Mei, K.; Xue, J. UAV obstacle avoidance by human-in-the-loop reinforcement in arbitrary 3D environment. In Proceedings of the 2023 42nd Chinese Control Conference (CCC), Tianjin, China, 24–26 July 2023. [Google Scholar]
  10. Çetin, E.; Barrado, C.; Muñoz, G.; Macias, M.; Pastor, E. Drone navigation and avoidance of obstacles through deep reinforcement learning. In Proceedings of the IEEE/AIAA 38th Digital Avionics Systems Conference (DASC), San Diego, CA, USA, 8–12 September 2019; pp. 1–7. [Google Scholar]
  11. Sun, H.; Li, R.; Tian, M.; Shi, Y. Deep reinforcement learning for UAV navigation based on lightweight simulation environment. In Proceedings of the 2025 4th Conference on Fully Actuated System Theory and Applications (FASTA), Nanjing, China, 4–6 July 2025. [Google Scholar]
  12. Kochdumper, N.; Krasowski, H.; Wang, X.; Bak, S.; Althoff, M. Provably safe reinforcement learning via action projection using reachability analysis and polynomial zonotopes. IEEE Open J. Control Syst. 2023, 2, 79–92. [Google Scholar] [CrossRef] [Scilit]
  13. Emam, Y.; Notomista, G.; Glotfelter, P.; Kira, Z.; Egerstedt, M. Safe reinforcement learning using robust control barrier functions. IEEE Robot. Autom. Lett. 2025, 10, 2886–2893. [Google Scholar] [CrossRef] [Scilit]
  14. Kong, X.; Xia, Y.; Sun, Z.; Zhai, D.-H.; Deng, Y.; Zhang, S. Differential high order control barrier function-based safe reinforcement learning. IEEE Robot. Autom. Lett. 2025, 10, 7524–7531. [Google Scholar] [CrossRef] [Scilit]
  15. Guerrier, M.; Fouad, H.; Beltrame, G. Learning control barrier functions and their application in reinforcement learning: A survey. arXiv 2024, arXiv:2404.16879. [Google Scholar]
  16. Ng, A.Y.; Harada, D.; Russell, S. Policy invariance under reward transformations: Theory and application to reward shaping. In Proceedings of the 16th International Conference on Machine Learning (ICML), Bled, Slovenia, 1999; pp. 278–287. [Google Scholar]
  17. Berducci, L.; Aguilar, E.A.; Ničković, D.; Grosu, R. HPRS: Hierarchical potential-based reward shaping from task specifications. Front. Robot. AI 2025, 11, 1444188. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  18. Badnava, B.; Esmaeili, M.; Mozayani, N.; Zarkesh-Ha, P. A new potential-based reward shaping for reinforcement learning agent. In Proceedings of the 2023 IEEE 13th Annual Computing and Communication Workshop and Conference (CCWC), Las Vegas, NV, USA, 8–11 March 2023. [Google Scholar]
  19. Yang, C.; Xu, P.; Zhang, J. Learning individual potential-based rewards in multiagent reinforcement learning. IEEE Trans. Games 2025, 17, 334–345. [Google Scholar] [CrossRef] [Scilit]
  20. Xu, Z.; Xiu, Y.; Zhan, X.; Chen, B.; Shimada, K. ViGO: Vision-aided UAV navigation and dynamic obstacle avoidance using gradient-based B-spline trajectory optimization. IEEE Trans. Autom. Sci. Eng. 2024, 21, 1214–1225. [Google Scholar] [CrossRef] [Scilit]
  21. Legittimo, M.; Brilli, R.; Crocetti, F.; Costante, G. SAC-Nav: Maximum entropy reinforcement learning for agile continuous navigation. IEEE Robot. Autom. Lett. 2025, 10, 245–252. [Google Scholar]
  22. Hu, X.; Guo, R.; Huang, H.; Zhang, C. Design and analysis on an adaptive wall-climbing robot structure for variable curvature facade. Robot 2024, 46, 576–590. [Google Scholar]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Article Metrics

Citations

Article Access Statistics

Multiple requests from the same IP address are counted as one view.