3.2. Multi-Scenario Testing Under Tubulent Inflow Conditions
Before presenting the turbulent-inflow comparison results, it should be clarified that the RL-Fuzzy-PID parameters used in each scenario were obtained through an independent offline tuning process under the corresponding inflow condition. Specifically, the NTM results were obtained using the parameter set tuned under NTM, whereas the ETM results were obtained using the parameter set tuned under ETM. Accordingly, the following comparisons are intended to evaluate the effectiveness of the proposed RL-assisted offline tuning framework under different representative IEC-compliant turbulent wind conditions, rather than to demonstrate strict cross-condition generalization of a single parameter set.
The step-response characteristics of the different yaw control schemes are illustrated in
Figure 8. As can be observed, the proposed RL-based fuzzy PID controller exhibits faster convergence, reduced overshoot, and smoother yaw adjustment than the other controllers. These results indicate superior transient performance and actuation efficiency under deterministic conditions, thereby justifying the selection of the RL-Fuzzy-PID scheme as the reference controller for the subsequent turbulent-inflow evaluations.
Under IEC-compliant turbulent wind conditions, two representative inflow scenarios were considered in this study. Scenario 1 corresponds to an NTM case with a mean wind speed of 15 m/s and turbulence intensity class IEC-B. Scenario 2 corresponds to an ETM case with a mean wind speed of 18 m/s and turbulence intensity class IEC-A. For both scenarios, the turbulent wind fields were generated at a sampling frequency of 50 Hz with a horizontal grid size of 228 m × 228 m at each time step.
Figure 9 and
Figure 10 present the responses of the different controllers in terms of output power, yaw position, cumulative yaw energy, and yaw rate under NTM conditions.
Figure 11 and
Figure 12 present the corresponding comparative results under ETM conditions.
It should be noted that the PSO-Fuzzy-PID and GWO-Fuzzy-PID curves partially overlap in
Figure 9 and
Figure 11. This overlap is not caused by plotting error, but reflects the highly similar dynamic responses of the two optimization-based controllers under the tested turbulent wind conditions; this observation is further supported by the quantitative comparisons summarized in
Table 3 and
Table 4.
The control performance is quantitatively evaluated using the root mean square (RMS), standard deviation (STD), and mean absolute error (MAE) indices.
In addition to RMS, STD, and MAE, yaw-actuation energy consumption is quantified by integrating the absolute instantaneous yaw power over the simulation horizon:
where
denotes the instantaneous yaw power and
is the evaluation duration. In the numerical implementation, this quantity is approximated in discrete form as
.
For force-, rotation-, and transmission-related variables, the three-dimensional discrete components are first combined into an equivalent resultant magnitude using the Euclidean norm. The corresponding RMS or STD values are then evaluated from the resultant signals over the entire simulation horizon.
Table 3 and
Table 4 summarize the quantitative performance of different yaw controllers under both NTM and ETM conditions, respectively.
Under NTM conditions, compared with the conventional PID, Fuzzy-PID, PSO-optimized Fuzzy-PID, GWO-optimized Fuzzy-PID, and MPC controllers, the proposed RL-based Fuzzy-PID reduces yaw-rate RMS by 33.4%, 39.8%, 36.0%, 36.2%, and 21.0%, respectively, while decreasing total yaw-energy consumption by 38.0%, 41.1%, 41.2%, 41.6%, and 19.3%. In addition, the yaw-position MAE is reduced by 8.0% relative to PID, indicating improved yaw alignment together with reduced actuation activity.
Under ETM conditions, compared with PID, Fuzzy-PID, PSO-optimized Fuzzy-PID, GWO-optimized Fuzzy-PID, and MPC, the proposed method reduces yaw-rate RMS by 25.6%, 3.9%, 4.4%, 3.1%, and 10.5%, respectively, while lowering total yaw-energy consumption by 32.6%, 6.6%, 5.4%, 3.7%, and 5.9%. In addition, a 28.9% reduction in yaw-position MAE is achieved relative to MPC, demonstrating improved yaw alignment accuracy and the effective suppression of unnecessary actuation under severe wind-direction fluctuations.
All performance indicators follow a “smaller-is-better” criterion. After normalization, a radar chart is used to provide an overall comparison of the different yaw control strategies, as shown in
Figure 13. The results indicate that the RL-based controller attains relatively smaller normalized values for most indices, suggesting a more favorable trade-off among yaw accuracy, actuation effort, and power-related performance.
Figure 14 further illustrates the quantitative comparison of the different yaw control strategies under NTM and ETM conditions.
To more directly quantify the ability of the proposed RL-Fuzzy-PID controller to suppress unnecessary yaw activity, additional event-level yaw-motion statistics were introduced. To exclude the startup transient, all statistics were evaluated over the interval from 50 s to the end of each simulation. First, the yaw-rate signal
was smoothed using a moving-average window of 0.5 s. Based on the smoothed signal, a yaw-motion state variable
was defined as follows:
where
is the smoothed yaw rate and
is the yaw-rate threshold. A yaw-motion event was then defined as a sign-consistent nonzero segment of
, after merging short zero gaps and removing segments with insufficient duration or amplitude.
Based on the detected yaw-motion events, the following indicators were calculated:
The yaw action count is defined as the total number of detected yaw-motion events:
The direction-reversal count is defined as the number of sign changes between two consecutive yaw-motion events:
where
denotes the direction sign of the
-th yaw-motion event.
The cumulative yaw travel is defined as the total absolute yaw displacement over all detected yaw-motion events:
where
and
are the start and end times of the
-th yaw-motion event.
For the
-th yaw-motion event, the event amplitude is defined as:
where
is the yaw position.
The mean action amplitude is then calculated as follows:
The resulting statistics under NTM and ETM conditions are summarized in
Table 5 and
Table 6.
The event-level yaw-motion statistics in
Table 5 and
Table 6 provide direct quantitative evidence of the yaw-actuation behavior of the compared controllers. Under NTM conditions, RL-Fuzzy-PID exhibits the most balanced yaw-motion pattern, characterized by the lowest direction-reversal count and the lowest cumulative yaw travel among all controllers. Compared with PID, Fuzzy-PID, PSO-Fuzzy-PID, GWO-Fuzzy-PID, and MPC, its cumulative yaw travel is reduced by 33.1%, 44.5%, 36.1%, 36.4%, and 22.3%, respectively. Although Fuzzy-PID yields fewer yaw actions, its much larger cumulative yaw travel and mean action amplitude indicate a heavier actuation burden. Under ETM conditions, RL-Fuzzy-PID still maintains a relatively low yaw-action count, the lowest direction-reversal count, and a comparatively low cumulative yaw travel, with reductions of 23.5%, 1.8%, 0.2%, and 8.7% relative to Fuzzy-PID, PSO-Fuzzy-PID, GWO-Fuzzy-PID, and MPC, respectively. These results indicate that RL-Fuzzy-PID suppresses unnecessary reversal-dominated yaw activity under NTM and preserves a more organized correction pattern under ETM.
3.3. Robustness and Sensitivity Evaluation
3.3.1. Sensitivity to Communication Delay and Packet Loss
To further evaluate the applicability of the proposed distributed soft-real-time co-simulation framework under impaired communication conditions, a sensitivity analysis was conducted under the ETM condition by introducing fixed delay and packet loss into the wind-direction measurement channel. To exclude startup transients, all indices were evaluated over a common post-startup window from 50 s to the end of the simulation.
The communication conditions are denoted in the form , where represents the fixed delay in milliseconds and represents the packet-loss rate in percent. Accordingly, D0L0 denotes the no-impairment baseline, D0L10 denotes the case with 10% packet loss only, D100L0 denotes the case with 100 ms delay only, and D300L10 denotes the combined severe case with 300 ms delay and 10% packet loss. These cases were selected to represent the baseline condition, single-factor degradation, and a representative severe communication-impairment scenario.
Representative results are summarized in
Table 7. The output-power coefficient of variation, denoted as Output Power (CV), is calculated as follows:
where
and
denote the standard deviation and mean value of output power, respectively, within the common evaluation window.
To further quantify the sensitivity of each controller to communication degradation, the relative changes in the key indices with respect to each controller’s own baseline case (D0L0) are shown in
Figure 15. In this figure, the horizontal axis represents the communication-impairment cases, whereas the vertical axis represents the relative percentage change in each metric with respect to the corresponding no-impairment baseline. Accordingly, a positive value indicates an increase relative to the baseline, whereas a negative value indicates a decrease. This figure reflects the sensitivity of each controller to communication degradation rather than a direct comparison of their absolute performance levels.
As shown in
Figure 15, the relative variations in the CV of output power, Cumulative yaw travel, and Cumulative yaw energy remained small across all tested cases, indicating that the overall closed-loop performance was only mildly affected by the introduced delay and packet loss. By contrast, Yaw action count exhibited the largest variation, suggesting that communication degradation affected yaw-actuation behavior more noticeably than power-related performance. Moreover, the RL-Fuzzy-PID controller generally showed smaller relative deviations from its own baseline in several representative cases, indicating better robustness to communication impairment in relative-performance terms.
Overall, these results show that the proposed framework can serve as a soft-real-time, network-aware co-simulation platform for controller evaluation under communication delay and packet loss. Within the tested impairment range, the closed-loop system remained stable, the average generated power was nearly unaffected, and the main impact of communication degradation was observed in yaw-actuation behavior rather than mean power-capture capability.
3.3.2. Sensitivity Analysis Under Corrupted Input Data
Additional robustness tests were conducted under the ETM condition by injecting corrupted signals into the controller-side wind-direction input. In addition to the baseline case (0), additive noise (N), impulsive outliers (O), signal distortion (D), and the combined case (C) were considered. To eliminate startup effects, all metrics were evaluated over the interval from 50 s to the end of the simulation. To provide a more intuitive comparison of the robustness trends under different input perturbations, the key performance metrics of the two controllers are further summarized in
Figure 16. The corresponding numerical results are listed in
Table 8 for detailed quantitative comparison.
Table 8 summarizes the corresponding results. Both controllers remained stable under all tested perturbation conditions, and the yaw-tracking error varied only slightly. However, their degradation patterns differed markedly in actuation-related metrics. For the conventional Fuzzy-PID controller, the noise case increased the power-fluctuation coefficient, yaw-action count, cumulative yaw displacement, and cumulative yaw energy by 11.1%, 79.3%, 70.3%, and 144.3%, respectively, relative to its own baseline. Under the combined case, the corresponding increases reached 12.4%, 85.4%, 71.7%, and 146.7%
In contrast, the proposed RL-Fuzzy-PID controller showed much smaller degradation under the same perturbation conditions. Relative to its own baseline, the noise case changed these four metrics by −1.7%, +30.1%, +3.0%, and +3.9%, respectively, whereas the combined case yielded −2.6%, +19.2%, +3.9%, and +5.0%. Under the outlier and distortion cases, both controllers showed only limited degradation. Overall, the robustness benefit of the proposed RL-based tuning is mainly reflected in mitigating excessive yaw activity and energy deterioration under persistent input corruption.
3.3.3. Sensitivity Analysis of Learning Rate, Reward Architecture, and Initialization
To assess the algorithm-level robustness of the proposed RL-assisted offline tuner, additional sensitivity analyses were performed in the reduced-order training environment derived from the yaw-system state-space model. In this fast-screening setting, a pulse-type yaw-reference input was adopted to isolate the influence of key tuning-related factors, whereas the higher-fidelity turbulent-inflow validation of the final controller was retained in the main simulation section. In this setting, the reduced-order screening objective was constructed from normalized tracking-related and actuation-related terms, so that the optimization process could be compared under different hyperparameter choices in a unified manner.
First, learning-rate sensitivity was examined using
, 0.05, 0.10, 0.20, and 0.50, each tested under two random seeds. As shown in
Figure 17a, the optimization trajectories remain highly similar across most of the tested practical learning-rate range, and the final best-objective values also remain close to one another. Only the largest tested value,
, shows a slightly more noticeable deviation for one seed. These results indicate that the proposed offline RL tuner exhibits limited sensitivity to the learning-rate choice in terms of the final tuning outcome, whereas the learning rate mainly affects the convergence path and the degree of path-dependent variation.
Second, reward-architecture sensitivity was evaluated using three reward formulations: a normalized fitness-improvement reward, a raw fitness-improvement reward, and a regularized reward with an additional parameter-update penalty.
Figure 17b shows that the normalized and raw rewards produce nearly overlapping optimization trajectories in the present screening environment, suggesting that the search behavior is mainly governed by the direction of fitness improvement rather than the exact reward scaling. By contrast, the regularized reward leads to a slightly different convergence path and a different final parameter combination, indicating that the reward architecture mainly affects the detailed trade-off among performance terms rather than the overall tuning conclusion.
Third, initialization sensitivity was examined using two distinct starting points, namely a lower-biased initial vector
and a higher-biased initial vector
, under a fixed learning-rate setting. As shown in
Figure 17c, initialization exerts a more visible influence on the optimization trajectory than the learning-rate and reward variations. The lower-biased initialization tends to converge toward a solution with lower tracking-related indices, whereas the higher-biased initialization tends to yield lower actuation-related terms at the cost of slightly worse tracking performance. Thus, initialization mainly influences which local performance trade-off is selected, while both tested initializations still converge to stable tuned solutions.
Overall, these results show that the proposed RL-assisted tuning framework is reasonably stable in the reduced-order offline tuning environment. The main conclusion does not depend on a single isolated choice of learning rate, reward formulation, or initialization, although moderate differences may still appear in convergence behavior and detailed parameter evolution.
3.4. Online Execution-Time Profiling on the Controller-Side Processor
To evaluate online computational efficiency on the target controller-side processor, an additional controller-only replay benchmark was conducted. Two prerecorded input signals, namely nacelle yaw position and hub-height wind direction, were extracted from a representative co-simulation run and replayed to each controller under the same software and hardware environment. In this benchmark, only controller-side computation was timed, whereas OpenFAST (version 4.0.0) plant integration and inter-process communication overhead were excluded. For the proposed RL-assisted method, the reported execution time refers only to the online execution of the final tuned Fuzzy-PID controller, while the offline Q-learning stage is excluded from the per-step timing comparison. Each benchmark was repeated five times, and the average, maximum, and 95th-percentile per-step execution times were recorded.
Table 9 summarizes the controller-side execution times obtained from the replay benchmark. As expected, the classical PID controller exhibits the lowest computational burden, with an average execution time of 0.071 ms per step. The conventional Fuzzy-PID controller requires a higher computational cost, reaching an average of 4.066 ms per step. The proposed RL-Fuzzy-PID controller shows a similar online computational level, with an average execution time of 3.782 ms per step and a maximum value of 3.946 ms.
These results show that the proposed method is not faster than classical PID in raw online computational speed. However, its online execution cost remains close to that of conventional Fuzzy-PID, since RL is used only for offline parameter tuning and does not participate in per-step online learning. Under the present replay benchmark, the effective input update interval was 0.02 s, and the measured execution times of both Fuzzy-PID and RL-Fuzzy-PID remained below this interval, indicating that the proposed controller satisfies the timing requirement of the current co-simulation configuration.
In addition, the online computational burden of the benchmark MPC was profiled under the final controller settings. With a controller update period of 20 ms, the measured average, 95th-percentile, and maximum optimization times were 0.545 ms, 0.838 ms, and 3.343 ms, respectively. Since the worst-case optimization time remained well below the update period, the benchmark MPC could complete online optimization within the timing requirement of the present distributed soft-real-time co-simulation framework.