4.1. Simulation Platform and Experimental Setup
4.1.1. Simulation Environment
To evaluate the proposed hierarchical PPO route-planning method in dynamic fire and smoke environments, this study constructs a simulation platform that includes the dynamic fire and smoke environment model, local grid observations, fixed-wing flight dynamics, and multiple high-level route-planning algorithms. The high-level route-planning policy and conventional planning algorithms are implemented in Python 3.8 and PyTorch 2.4.1, and flight dynamics are computed using the C172x six-degree-of-freedom model in JSBSim [
39]. All experiments were conducted on a workstation equipped with an Intel Core i7-14700 CPU and an NVIDIA GeForce RTX 3090 GPU. The online decision times reported in the comparative experiments were measured on this hardware platform.
The task area is 5 km × 5 km, with a terrain-grid resolution of 50 m. The low-level control period is 0.2 s, and the maximum task duration per episode is 300 s. In route-planning experiments, fire and smoke threat states update every 1 s. At each high-level decision time, the system reconstructs the local grid observation from the current UAV position and the latest fire and smoke threat state.
The experiments include valley, ridge, and multi-peak dynamic fire and smoke scenarios. The valley scenario evaluates local route selection in constrained passages. The ridge scenario evaluates detour capability under large-scale elevation changes and directional fire and smoke dispersion. The multi-peak scenario combines complex terrain relief with fragmented threat regions and evaluates overall planning capability in a multi-constraint environment. During training, each episode randomly selects one of the three terrain types and reinitializes the terrain, fire and smoke state, and UAV state. The terrain-resolved evaluation was conducted using 30 independent episodes for each model in each terrain scenario, giving 90 test episodes per model. The same random seeds (42–71) were used across models within each scenario, and the terrain, fire and smoke state, and UAV state were reinitialized for each episode.
4.1.2. Evaluation Metrics
To evaluate the planning algorithms comprehensively, this study reports task success rate, mean terrain clearance, fire and smoke threat-avoidance success rate, average decision time, average task duration, average route length, average executed-trajectory curvature, and maximum roll angle. Task success rate measures the ability to reach the target region safely within the time limit. Mean terrain clearance is the time average of the vertical distance between the UAV and terrain along each executed trajectory. The minimum safety margin is defined as the trajectory minimum of the smaller of the terrain clearance and the fire/smoke-threat clearance. The fire and smoke threat-avoidance success rate is the proportion of flights that do not enter burning-threat or high-concentration smoke-threat regions. Average decision time is the time required for one high-level policy inference or one replanning operation, excluding the offline training cost of reinforcement learning. Average task duration and average route length reflect execution efficiency, whereas average executed-trajectory curvature and maximum roll angle evaluate trajectory smoothness and the executability of the fixed-wing UAV.
When an algorithm fails to complete a task, the reported task duration, route length, executed-trajectory curvature, and maximum roll angle are computed from the actual executed trajectory from the start point to the termination time of that episode. Episodes terminate when the target is reached, a collision occurs, or the task time limit is reached.
For the terrain-resolved ablation-model evaluation, each model was evaluated in 90 episodes, comprising 30 independently seeded episodes in each of the valley, ridge, and multi-peak scenarios. The same terrain and random-seed combinations were used for all four models. Continuous metrics are reported as the mean ± standard deviation for each terrain and over all 90 episodes. For the planning-algorithm comparison, each algorithm was evaluated in 100 independently seeded episodes, with the three scenarios assigned cyclically, and continuous metrics are reported as arithmetic means over the 100 episodes. Because task success and fire and smoke threat avoidance are binary outcomes, their sampling uncertainty was quantified using Wilson 95% confidence intervals. Success differences between successive ablation models were assessed using two-sided exact paired McNemar tests because their episodes were matched by terrain and random seed. Planning-algorithm success proportions were compared using two-sided Fisher’s exact tests. Holm correction was applied separately to the ablation-model and planning-algorithm comparison families, and statistical significance was defined as an adjusted p value below 0.05.
4.1.3. Experimental Parameter Settings
To ensure reproducible simulation settings and to clarify the time scales and constraints linking the fire and smoke environment, the low-level flight controller, and the high-level route planner, key parameters are listed in
Table 4,
Table 5 and
Table 6. All comparison algorithms operate under the same task area, dynamic fire and smoke environment, local grid observation, low-level flight controller, and flight-dynamics conditions. A* and RRT* differ only in the high-level route-planning mechanism.
Table 4 lists the spatial discretization, update period, and safety thresholds for the fire and smoke environment. Factors such as wind, slope, fuel type, and fuel moisture are computed online from terrain and environmental inputs for each scenario according to the model in
Section 2 and are therefore not listed as globally fixed constants.
Table 5 and
Table 6 report the training parameters for the pretrained low-level flight controller and the complete H-PPO high-level route-planning model, respectively.
The environmental settings in
Table 4 are simulation-design parameters rather than universally calibrated wildfire constants. The 5 km × 5 km task area and 50 m surface-grid resolution were selected to retain kilometer-scale terrain and route-selection characteristics while maintaining a tractable grid size for repeated dynamic updates. The 1 s fire and smoke update period synchronizes the cellular-automaton and smoke-dispersion calculations with the simulation clock. The initial 100 propagation steps were used to generate a developed, nontrivial threat distribution before each route-planning episode. The smoke threshold
is a relative simulation-concentration threshold rather than a health-exposure standard, and the 100 m safety-clearance threshold corresponds to two surface-grid cells. Wind, slope, fuel type, and fuel moisture are scenario-dependent inputs calculated online according to the model in
Section 2 rather than globally fixed parameters.
The principal PPO and GAE parameters, including
,
, and the PPO clipping coefficient of 0.2 follow the standard PPO and generalized advantage-estimation formulations [
35,
38]. The remaining optimization and network settings in
Table 5 and
Table 6, including the learning rates, rollout-buffer capacities, minibatch sizes, update epochs, network width, number of parallel environments, and training budgets, and LSTM configuration, were selected through preliminary training trials. A single-layer LSTM with a hidden-state dimension of 128 was used to provide recurrent representation capacity while keeping the online network compact. The final values were retained when they provided stable learning behavior without excessive computational cost. Separate settings were used for the low-level controller and the high-level recurrent route-planning policy because they operate at different decision frequencies and solve different control and planning tasks. These settings are experimental configurations for the present simulation framework and are not claimed to be universally optimal.
4.2. Analysis of Dynamic Fire and Smoke Threat Evolution
To examine the temporal and spatial evolution generated by the fire-spread and smoke-dispersion models,
Figure 3 presents the simulated states over representative ridge terrain at 3, 10, 27, and 29 min. The colored surface represents the DEM elevation. The cumulative burned region, including both currently active and burned-out cells, is displayed as a dark-gray footprint on the terrain surface. Active-burning cells are overlaid as red vertical prisms, while high-concentration smoke cells satisfying
are overlaid as translucent light-gray vertical prisms. Both types of threat prisms extend from the local terrain surface to 50 m above ground, in accordance with the threat-height settings used in this simulation. The cumulative burned footprint has no vertical threat extent because it records the surface fire history rather than a current airborne threat.
At 3 min of simulation, the active burning area is 121.5 ha and the cumulative burned area is 121.5 ha. The fire is still in its early development stage, with active burning concentrated near the initial ignition point. The overall form is compact and locally clustered. High-concentration smoke threats mainly cover the space above the fire source and the adjacent downwind area, with a relatively limited extent. This result indicates that fire and smoke threats are still expanding locally.
At 10 min of simulation, the active burning area increases to 571.0 ha and the cumulative burned area reaches 834.8 ha. Compared with the state at 3 min, the fire coverage expands substantially and evolves from an initially local cluster into a continuous threat belt extending rapidly along the dominant spread direction. The active fire front advances visibly outward, while high-concentration smoke threats spread over a much larger area and already cover more space than the active burning region. The horizontal footprints of both regions therefore expand rapidly. As they spread across the ridge terrain, the corresponding 50 m high threat prisms occupy a wider range of absolute elevations because their lower boundaries follow the local DEM, while their prescribed above-ground vertical extent remains unchanged. This rapid horizontal expansion indicates that the fire has entered a rapid growth stage.
At 27 min of simulation, the active burning area decreases to 132.8 ha, whereas the cumulative burned area increases to 1488.2 ha. The horizontal projected area of the high-concentration smoke-threat region is 1225.00 ha, calculated by multiplying the number of cells satisfying by the area of each 50 m × 50 m cell, i.e., 0.25 ha. The smoke-threat area is therefore 9.23 times the active burning area and 0.82 times the cumulative burned area. For comparison, the projected smoke-threat areas at 3 and 10 min are 470.50 and 775.25 ha, respectively, and their ratios to the active burning area are 3.87 and 1.36. From 10 to 27 min, the active burning area decreases by 76.8%, while the smoke-threat area increases by 58.0%. This quantitative comparison shows that the remaining active fire fronts are distributed over a broader perimeter and continue to produce a large downwind smoke-threat footprint even though the total active burning area has declined.
At 29 min of simulation, the active burning area further decreases to 122.2 ha, the cumulative burned area increases to 1532.5 ha, and the projected high-concentration smoke-threat area reaches 1254.50 ha. The smoke-threat area is 10.26 times the active burning area and 0.82 times the cumulative burned area. Compared with the state at 27 min, the active burning area decreases by a further 7.9%, whereas the smoke-threat area increases by 2.4%. Active burning remains concentrated along the outer fire front, while the broad spatial distribution of the remaining sources maintains an extensive downwind smoke-threat region. These results quantitatively demonstrate that a reduction in active burning area does not immediately produce a corresponding reduction in smoke-threat coverage.
Overall, the fire and smoke evolution exhibits a typical staged pattern of early local expansion, mid-stage rapid spread, and late-stage active burning at the perimeter with burnout in the interior. Active burning and cumulative burned regions separate clearly in the middle and late stages, while smoke threats show greater spatial extent and persistence. Therefore, route planning in dynamic fire and smoke environments cannot assess risk only from visible flames and must also consider high-risk airspace created by smoke dispersion.
4.3. Analysis of Low-Level Flight Controller Training
To verify that the low-level flight controller provides stable and repeatable dynamic execution for high-level route planning, the PPO low-level control policy is first trained in an independent command-tracking environment. During training, the policy takes tracking errors in reference altitude, speed, and heading, together with attitude, angular-motion, and aerodynamic states, as inputs and outputs continuous throttle, aileron, elevator, and rudder commands.
Figure 4 shows the training–return curve of the PPO low-level flight controller over 5.0 × 10
6 training steps.
The figure shows that returns are low and somewhat variable early in training, indicating that the low-level control policy has not yet learned stable tracking of multiple reference-command channels. After approximately 0.6 × 106 training steps, the return rises clearly. Near 1.0 × 106 training steps, the policy already obtains a high and sustained return. The return then improves gradually and enters a high-return range after approximately 2.5 × 106 training steps.
From 3.0 × 106 to 5.0 × 106 training steps, the return fluctuates mainly around a high level, with no persistent degradation or training collapse. These small fluctuations are mainly associated with stochastic action sampling and mini-batch policy updates during PPO training, differences among the states sampled by the five parallel environments, and the nonlinear flight dynamics simulated by JSBSim. The total of 5.0 × 106 training steps was specified in advance as a fixed training budget. This budget allowed the controller to continue training for approximately 2.5 × 106 steps after it first entered the high-return regime, thereby providing a sufficiently long interval for evaluating convergence stability. This indicates that, in terms of training return, the PPO-based low-level flight controller reached a relatively stable converged state and can serve as a fixed flight-dynamics tracking module for high-level route planning.
4.4. Experimental Analysis of the High-Level Route-Planning Model
To examine the effects of hierarchical decision making, temporal memory, and composite reward design on route selection in dynamic fire and smoke environments, four high-level route-planning configurations are compared after independent pretraining of the low-level flight controller. Model A is a single-level PPO baseline without a hierarchical architecture and represents direct end-to-end route decisions. Model B introduces high-level route-subgoal planning and fixed low-level flight-control tracking in the same task environment. Model C adds LSTM temporal memory to Model B. Models A–C were trained using only the base reward in Equation (39). Model D further adds a safety potential function and route-subgoal progress reward, forming the complete H-PPO high-level route-planning model. Models B, C, and D reuse the same trained low-level flight controller, so performance differences mainly reveal the contribution of the high-level route-planning mechanism. The configurations are listed in
Table 7.
Figure 5 shows the training–return curves of the four high-level route-planning models. The return is the cumulative reward of the high-level policy over a complete task episode and reflects learning of route-subgoal selection and long-horizon task progress in dynamic fire and smoke environments. Model A maintains a low return for a long period, indicating that direct learning of long-horizon route decisions is difficult without a hierarchical route-planning structure. After the hierarchical structure is introduced, Model B gradually improves, indicating that decoupling high-level route-subgoal generation from low-level flight tracking reduces the learning difficulty of high-level route planning. Model C improves further, indicating that the LSTM helps the high-level policy model fire and smoke threat evolution and flight-state changes from sequential observations. The complete H-PPO Model D converges faster and reaches the highest final return, indicating that the composite reward provides clearer safety constraints and goal guidance. Overall, high-level training returns follow D > C > B > A, consistent with task success rates in subsequent independent tests.
For the complete H-PPO Model D, the ten consecutive 20-episode blocks covering Episodes 1801–2000 have a mean return of 290.77 with a standard deviation of 1.79. The corresponding mean training success and collision rates are 99.9% and 0.2%, respectively, and no persistent upward or downward trend is observed during this interval. These results indicate that Model D reached a relatively stable high-return plateau by the end of the fixed 2000-episode training budget. In contrast, Model A remained in a low-return regime and was therefore not regarded as having learned a feasible route-planning policy.
Table 8 and
Table 9 report the performance of the four high-level route-planning models separately for the valley, ridge, and multi-peak scenarios, together with the overall results over all 90 test episodes. Model A, the single-level PPO baseline without a hierarchical architecture, fails to learn a feasible long-horizon route policy and has a task success rate of 0% in all three terrain scenarios. Its overall fire and smoke threat-avoidance success rate is only 11.1%. Although its overall task duration and route length are smaller than those of the hierarchical models, this outcome mainly reflects early failure or termination rather than higher planning efficiency. Model A also exhibits substantially larger executed-trajectory curvature and roll angles in every terrain scenario, indicating poor fixed-wing flight feasibility.
After high-level route-subgoal planning is introduced, Model B achieves an overall task success rate of 77.8% and an overall fire and smoke threat-avoidance success rate of 93.3%. Its task success rates are 70.0%, 73.3%, and 90.0% in the valley, ridge, and multi-peak scenarios, respectively. Compared with Model A, the overall executed-trajectory curvature and maximum roll angle of Model B decrease by 68.5% and 69.0%, respectively. Model B also maintains large mean terrain clearances in all three scenarios. However, it has the longest task duration and route length among Models B–D in each terrain scenario, indicating that the route-subgoal policy still favors conservative detours.
After adding the LSTM to Model B, Model C increases the overall task success rate to 94.4% and the overall fire and smoke threat-avoidance success rate to 98.9%. It reaches a 100% task success rate in both the valley and ridge scenarios, while its success rate is 83.3% in the multi-peak scenario. The latter is slightly lower than the 90.0% obtained by Model B, showing that the advantage of temporal memory is not uniform for every individual terrain. Nevertheless, compared with Model B, Model C reduces the overall task duration, route length, executed-trajectory curvature, and maximum roll angle by 18.6%, 15.8%, 27.8%, and 21.1%, respectively, and the same efficiency and smoothness advantages are observed in each terrain scenario. The overall mean terrain clearance decreases from 784.98 m to 693.15 m, indicating that the efficiency gain is accompanied by some reduction in terrain-clearance redundancy.
After adding the safety potential function and route-subgoal progress reward to Model C, the complete H-PPO Model D reaches a 100% task success rate and a 100% fire and smoke threat-avoidance success rate in all three terrain scenarios. Relative to Model C, its mean terrain clearance increases from 576.13 m to 644.52 m in the valley scenario, from 676.10 m to 750.57 m in the ridge scenario, and from 827.21 m to 911.27 m in the multi-peak scenario. The overall mean terrain clearance increases by 10.9%, from 693.15 m to 768.79 m. Model D also restores the task success rate from 83.3% to 100% in the multi-peak scenario. Its overall task duration, route length, curvature, and maximum roll angle increase to varying degrees relative to Model C but remain substantially better than those of Model B. These results show that the complete H-PPO model provides the most consistent reliability and terrain-clearance performance across the three terrain types, while retaining acceptable flight efficiency and feasibility.
For the task-success rates in
Table 8, the Wilson 95% confidence intervals are [0.0%, 4.1%], [68.2%, 85.1%], [87.6%, 97.6%], and [95.9%, 100.0%] for Models A, B, C, and D, respectively. The corresponding intervals for fire and smoke threat avoidance are [6.1%, 19.3%], [86.2%, 96.9%], [94.0%, 99.8%], and [95.9%, 100.0%]. After Holm correction, the increases in task success from Model A to Model B and from Model B to Model C were statistically significant, with adjusted
p values of 5.08 × 10
−21 and 0.00815, respectively. The increase from Model C to Model D, from 94.4% to 100.0%, did not reach statistical significance at the 0.05 level (adjusted
p = 0.0625). These findings provide statistical support for the contributions of the hierarchical architecture and LSTM temporal representation. Although the C-to-D increase in task success was not statistically significant, Model D achieved 100% task success in all three terrain scenarios and increased the overall mean terrain clearance by 10.9%, indicating that the composite reward primarily improved terrain-clearance redundancy and cross-terrain consistency in the present evaluation.
4.5. Comparison and Discussion of Planning Algorithms
This study uses A* and RRT* as conventional planning baselines and compares them with the proposed H-PPO method [
11,
12]. To ensure fairness, all three algorithms use the same 41 × 41 four-channel local grid observation, the same 10 s high-level decision period, the same pretrained low-level flight controller, and the same JSBSim dynamics model [
39]. The conventional planners cannot access future fire and smoke threat states outside the observation range and can only replan periodically from the current local observation. The comparative reliability, terrain-clearance, and online decision results are reported in
Table 10, while the corresponding execution-efficiency and flight-feasibility results are reported in
Table 11.
A* uses a 26-neighbor 3D grid search with an isotropic grid spacing of 100 m in the online local planning space. Over the altitude range [100, 6000] m, this setting produces 60 nominal altitude levels before collision and validity filtering. RRT* samples continuous 3D positions and uses a maximum of 5000 samples, an extension step of 180 m, a goal-bias probability of 0.2, and a rewiring radius of 360 m. The 180 m extension step limits the growth of the RRT* tree and does not represent a vertical-layer spacing. The local route waypoints generated by A* and RRT* are tracked by the same low-level flight controller.
The baseline-specific settings were selected through preliminary planning trials, with route feasibility and single-replanning time considered jointly. For A*, the 26-neighbor connectivity includes all face-, edge-, and corner-adjacent moves of the shared isotropic 100 m grid, enabling diagonal motion in three dimensions and reducing axis-aligned directional bias. For RRT*, the 180 m extension step is of the same spatial order as the 100 m local-grid resolution. Each candidate edge is checked for collision at 50 m intervals, and the 360 m rewiring radius is twice the extension step. The goal-bias probability of 0.2 directs 20% of the samples toward the local goal. Among the remaining samples, 65% are generated around the start-to-goal corridor and 35% are sampled over the bounded sampling region; direct and a limited set of guided-detour connections are also checked before tree expansion to improve local replanning efficiency. The maximum of 5000 samples limits the online search effort relative to the 10 s replanning period. These settings were applied unchanged to all scenarios and random seeds.
Table 10 shows that H-PPO achieves a 100% task success rate over 100 tests, exceeding RRT* and A* by 21 and 62 percentage points, respectively. Its mean terrain clearance is 771.92 m, which is 44.2% and 86.2% higher than those of RRT* and A*, respectively. H-PPO and RRT* both achieve a 100% fire and smoke threat-avoidance success rate, but RRT* has a lower overall task success rate because local random sampling causes planning failures or timeouts in some complex-terrain and fragmented-threat conditions. The fire and smoke threat-avoidance success rate of A* is 63%, indicating that it is more likely to lose safety margin as fire and smoke threats continue to evolve.
For the task-success rates in
Table 10, the Wilson 95% confidence intervals are [96.3%, 100.0%], [70.0%, 85.8%], and [29.1%, 47.8%] for H-PPO, RRT*, and A*, respectively. The corresponding intervals for fire and smoke threat avoidance are [96.3%, 100.0%], [96.3%, 100.0%], and [53.2%, 71.8%]. Pairwise comparisons of task success remain statistically significant after Holm correction for H-PPO versus RRT*, H-PPO versus A*, and RRT* versus A*, with adjusted
p values of 2.95 × 10
−7, 9.39 × 10
−25, and 1.09 × 10
−8, respectively. Thus, the differences in task reliability reported in
Table 10 are unlikely to result solely from random variation among the 100 test episodes.
In terms of online computational performance, H-PPO has an average decision time of 1.290 ms, which is 99.77% and 99.91% lower than those of RRT* and A*, respectively. This metric includes only post-training policy inference time or the time for a single replanning operation by a conventional algorithm. It excludes the offline training cost of reinforcement learning. For each high-level decision, H-PPO performs forward inference through the three-layer convolutional encoder, feature-fusion layer, single-layer LSTM, and Actor output layer. Its computational complexity can be expressed as , where the first term represents the convolutional operations and and denote the LSTM input and hidden dimensions, respectively. Because the local-grid input size and network dimensions are fixed in this study, the online inference cost of H-PPO is effectively constant for each decision. By comparison, A* has a complexity of , which reduces to for the sparse local grid, while the expected complexity of RRT* is approximately when efficient nearest-neighbor search is used, where is the number of samples. Therefore, the computation times of A* and RRT* are more sensitive to the number and distribution of searchable nodes and samples. The results indicate that H-PPO has low and nearly fixed online computational cost during deployment, whereas the computation time of search-based planners is more sensitive to local grid structure, obstacle distribution, and the sampling process.
Table 11 shows that H-PPO does not achieve the shortest task duration, shortest route, or lowest executed-trajectory curvature. Its average task duration and route length are 211.196 s and 7.564 km, respectively, both larger than those of RRT* and A*. This result indicates that the complete H-PPO policy prefers longer detours to obtain a higher task success rate and greater mean terrain clearance. Its maximum roll angle is 15.464°, below the 30° flight-envelope limit used in this study, indicating that the generated trajectories satisfy the basic feasibility requirements of a fixed-wing UAV. However, H-PPO has higher average executed-trajectory curvature than RRT* and A*. It is therefore not valid to claim that its trajectory smoothness is superior to that of conventional planners.
Conventional planners can produce shorter routes in some successful episodes because their objectives place greater weight on the geometrically shortest route under the current local observation. In a dynamic fire and smoke environment, however, a shorter route does not necessarily preserve sufficient future safety margin. A* tends to select locally shortest corridors with low cost, and its original route can quickly approach a threat boundary as flames and smoke continue to spread. RRT* has stronger spatial exploration through random sampling, but its results remain affected by sampling randomness, limited computation time, and the local observation range. By contrast, H-PPO learns a more conservative subgoal-selection policy through LSTM and safety-prioritized rewards, thereby improving task reliability under dynamic fire and smoke threats.
Figure 6 presents representative dynamically executed trajectories in valley, ridge, and multi-peak fire and smoke scenarios. Each scenario includes a 3D executed trajectory and a top-down projection. The blue, orange, and green curves represent the executed trajectories of H-PPO, A*, and RRT*, respectively, while the green circles and purple stars mark the start and goal positions. Red regions denote active fire, gray regions denote high-concentration smoke threats, and the background map shows terrain elevation. The trajectories are representative single-run results for the corresponding scenarios and are used mainly to illustrate route-selection differences among the planning algorithms under complex terrain and dynamic fire and smoke threats.
The representative trajectories show clear safety-prioritized behavior by H-PPO in all three scenarios. Its executed trajectories generally climb proactively early in the mission and make lateral detours near active fire and high-concentration smoke threats, thereby maintaining large threat clearances. By contrast, RRT* trajectories are generally closer to the start-to-goal direction and are relatively shorter, but in some scenarios they pass closer to fire and smoke threats. A* is more likely to proceed along a geometric route with low current local cost and is therefore more likely to lose safety margin or terminate early when threat regions continue to expand or terrain constraints are strong.
In the multi-peak scenario, terrain relief and fragmented fire and smoke threats jointly shape route selection. H-PPO avoids major threat regions by flying at a higher altitude and making wider lateral detours. Although this behavior increases route length and task duration, it is consistent with its safety-prioritized reward design and with the higher task success rate and mean terrain clearance in
Table 10. Overall,
Figure 6 further shows that the advantage of H-PPO is not the shortest route or the lowest executed-trajectory curvature but more reliable safe avoidance and task completion under dynamic fire and smoke threats.
To complement the mean terrain-clearance statistics in
Table 10 and further compare the dynamic safety redundancy of the three planning methods,
Figure 7 presents the safety-margin histories corresponding to representative successful runs in the valley, ridge, and multi-peak scenarios. At each simulation step, the safety margin is defined as the smaller of the terrain clearance and the fire/smoke-threat clearance.
In the valley scenario, all three methods have approximately the same initial minimum safety margin of 225.5 m because they start from the same state. After departure, H-PPO establishes and maintains a larger safety margin, with a trajectory-averaged margin of 591.4 m, compared with 399.0 m for RRT* and 395.7 m for A*. In the ridge scenario, the minimum margins of H-PPO, RRT*, and A* are 290.1, 80.1, and 37.1 m, respectively, with RRT* and A* falling below the prescribed 100 m threshold. In the multi-peak scenario, the corresponding minimum margins are 325.8, 280.9, and 38.4 m, respectively, and only A* falls below the threshold. These representative safety-margin histories show that H-PPO maintains greater dynamic safety redundancy, whereas
Table 10 separately reports aggregate task outcomes and mean terrain clearance over 100 episodes.
Overall, the main advantages of H-PPO are task success rate, mean terrain clearance, and online decision time. Its cost is longer task duration, longer route length, and larger executed-trajectory curvature. This result indicates that route-planning performance in dynamic fire and smoke environments should not be judged only by route length or trajectory smoothness. It should be evaluated jointly with task reliability, dynamic safety margin, mean terrain clearance, and online computational cost.
4.6. Discussion, Limitations, and Practical Implications
4.6.1. Sensitivity Analysis of Environmental Inputs
To evaluate the robustness of the trained route-planning policy to environmental-input variations, a one-factor-at-a-time sensitivity analysis was conducted using the frozen complete H-PPO Model D without additional training. Wind speed, base fire-spread rate, and smoke-release rate were varied independently to 0.7, 1.0, and 1.3 times their baseline values. The corresponding wind speeds were 3.5, 5.0, and 6.5 m/s; the base fire-spread rates were 0.35, 0.50, and 0.65 m/s; and the smoke-release rates were 7, 10, and 13 kg/s. The base fire-spread rate was used as the controllable fire-intensity input because it directly determines the expansion rate of burning cells in the cellular-automaton model. Terrain sensitivity was evaluated separately using the valley, ridge, and multi-peak scenarios.
For each parameter level, 20 paired random seeds were evaluated in each of the three terrain scenarios, giving 60 episodes per environmental-input level. The same seed was retained across the low, baseline, and high levels of each factor to reduce variation unrelated to the tested input. Together with the shared baseline condition, the sensitivity analysis comprised 420 deployment episodes. Unlike the mean terrain clearance reported in
Table 8 and
Table 10, this analysis recorded task success rate, trajectory-level minimum safety margin, and task duration.
Figure 8 reports the absolute results for the three terrain types and the paired changes relative to the baseline condition for wind speed, fire-spread rate, and smoke-release rate.
All seven evaluated conditions achieved 60 successful missions out of 60 episodes, corresponding to a task-success rate of 100% with a Wilson 95% confidence interval of [94.0%, 100.0%]. The environmental perturbations nevertheless produced measurable changes in the simulated threat fields. When the base fire-spread rate decreased from 0.50 to 0.35 m/s, the mean active-fire area decreased from 63.50 to 46.71 ha, whereas increasing it to 0.65 m/s increased the active-fire area to 85.35 ha. The corresponding high-concentration smoke-threat areas were 386.21, 419.18, and 458.48 ha. Changing the smoke-release rate from 7 to 13 kg/s increased the mean high-concentration smoke-threat area from 405.66 to 428.34 ha. Thus, the tested inputs altered the fire and smoke environment rather than merely changing parameter labels.
Despite these environmental changes, the mean minimum safety margin remained between 287.77 and 289.18 m across all wind, fire-spread, and smoke-release levels, compared with 288.47 m under the baseline condition. The corresponding mean task durations ranged from 152.86 to 155.37 s, compared with the baseline value of 154.24 s. Relative to the baseline, the largest change in minimum safety margin was less than 0.25%, and the largest change in task duration was less than 0.90%.
Terrain produced a more visible change in safety margin. The mean minimum safety margins in the valley, ridge, and multi-peak scenarios were 225.46 ± 0.00, 290.07 ± 0.00, and 349.87 ± 20.50 m, respectively, whereas their mean task durations remained similar at 154.59 ± 1.32, 154.08 ± 1.72, and 154.04 ± 1.05 s. These results indicate that terrain geometry affects the available clearance, but the frozen H-PPO policy retained stable mission completion and task duration across the three tested terrain types.
Overall, the sensitivity results show that the trained H-PPO policy remains stable under the tested ±30% variations in wind speed, base fire-spread rate, and smoke-release rate, as well as across the three terrain scenarios. This conclusion is limited to the parameter ranges and one-factor-at-a-time design considered here and should not be interpreted as robustness to arbitrary wildfire or atmospheric conditions.
4.6.2. Limitations and Practical Implications
Several limitations should be considered when interpreting the results. First, the proposed framework was evaluated only in simulation using a specific fixed-wing flight-dynamics model and three terrain scenarios; its generalization to other aircraft platforms and real wildfire environments has not been established. Second, the cellular-automaton fire-spread model and Gaussian plume smoke model provide computationally efficient threat estimates but simplify the spatial and temporal variability of wind, fuel conditions, atmospheric stability, and fire–atmosphere interactions. The current experiments also use simulation-generated observations and do not explicitly consider sensing errors, observation delays, communication interruptions, or uncertainty in threat-state estimation.
For practical deployment, onboard or ground-based observations, meteorological measurements, terrain data, and fuel information could be used to update the fire and smoke threat maps. The planner could then generate route subgoals for execution by the onboard flight-control system, while forest managers supervise the mission through an interactive interface displaying threat regions, recommended routes, safety margins, and task status. Such deployment would require reliable data fusion, uncertainty-aware environmental updating, platform-specific integration, and validation through controlled flight tests and field experiments under realistic fire conditions.
Advanced deep learning methods may complement the proposed framework as upstream perception modules. FireDM generates fire-scene images and segmentation masks, FireSeg performs fire segmentation using pretrained latent-diffusion features, and FireSegNASUNet and FireSegUNet focus on computationally efficient fire segmentation [
40,
41,
42,
43]. Their outputs could provide image-derived fire-region observations for updating the local threat representation used by the planner. Data-driven prediction models could similarly provide short-term estimates of threat evolution. These perception and prediction components were not implemented in the present study, and their integration into a complete perception–prediction–planning–control chain remains future work.