1. Introduction
With the advancement of carbon peak and carbon neutrality goals, the installed capacity of renewable energy sources continues to rise, significantly increasing the uncertainty of the power system [
1,
2]. To ensure the stable operation of the power grid, thermal power units are gradually shifting from base-load operation to deep peak-shaving operation, frequently undertaking rapid load change tasks. As the core power equipment of thermal power units, steam turbines operating under deep peak-shaving conditions are subjected to long-term variable-load states. The thermodynamic parameters fluctuate dramatically, and critical components face severe thermal stress concentration and fatigue damage, directly impacting the safety of the unit [
3,
4]. Therefore, conducting research on high-precision modeling of steam turbines for deep peak-shaving conditions holds significant engineering application value for enhancing the flexibility of thermal power units.
Steam turbines, as typical complex energy equipment, require high-fidelity modeling as the foundation for performance analysis, condition monitoring, and optimization control. Current research primarily focuses on reverse geometric modeling, high-precision mesh generation, digital twin construction, and hybrid modeling methods. In reverse modeling and geometric reconstruction, reverse engineering technology based on 3D-scanned point clouds has become a crucial means of obtaining precise geometric models of complex components. For complex curved surfaces, researchers have achieved rapid establishment of high-precision geometric models through point cloud fusion, surface patch recognition, and reconstruction [
5,
6,
7,
8]. For instance, Liu et al. [
5] established a five-dimensional framework based on digital twins for the testing process of large complex equipment components, where constructing a high-fidelity geometric model was a prerequisite for achieving virtual–real mapping. Zeng et al. [
6] proposed a geometric deviation modeling method for design and manufacturing based on the skin model concept, providing theoretical support for the geometric fidelity of digital twins. In mesh generation and simulation calculation, high-quality meshes are key to achieving high-precision flow and temperature field simulations [
9,
10]. For complex structures, researchers have significantly enhanced the boundary layer resolution capability of simulations through refined mesh parameter settings [
9]. However, the computational burden of full-order finite element models is substantial, making it difficult for them to meet real-time monitoring requirements [
10]. Consequently, reduced-order model techniques have gained widespread attention [
11,
12]. Cai et al. [
11] proposed a hybrid reduced-order model based on singular value decomposition and temporal convolutional networks, effectively capturing the thermal stress evolution law in critical rotor areas. Huang et al. [
12] combined adaptive mesh refinement with a hyper-reduced-order model to achieve fast prediction of thermo-mechanical coupling behavior in valves while maintaining accuracy.
In the construction of digital twins, numerous studies have been conducted by scholars worldwide [
13,
14,
15,
16]. Cheng et al. [
13] proposed a digital twin-enhanced industrial internet framework, illustrating the application modes of digital twins throughout the product lifecycle. Yu et al. [
14] proposed a hybrid modeling method based on operational data for the steam turbine control stage, achieving online performance monitoring under variable operating conditions. Zhang et al. [
15] proposed a digital twin method for gas turbines based on deep multi-model fusion, significantly improving performance prediction accuracy by constructing bidirectional data flow between mechanism models and data-driven models. Chen et al. [
16] established a digital twin model for a steam turbine system, evaluating system energy efficiency by combining mechanism-driven and data-driven methods, and achieved main steam pressure optimization. In model optimization and parameter identification, intelligent optimization methods like genetic algorithms are widely used for data-driven parameter identification of critical parameters that are difficult to measure directly in physical models [
17,
18]. Kalmykov et al. [
17] improved the standard fuel consumption calculation method for hydrogen co-production scenarios in thermal power plants and analyzed different operating modes using a digital twin model. Khoshgoftar et al. [
18] employed a gradient boosting algorithm to build a surrogate model for a steam thermal power plant and combined it with multi-objective optimization algorithms to achieve operational parameter optimization.
Despite significant progress in steam turbine modeling and digital twin applications, the following deficiencies remain in achieving high-precision, real-time modeling for deep peak-shaving conditions. Firstly, existing research often treats reverse geometric modeling and simulation mesh generation as separate steps, lacking a systematic method to directly map point cloud data and surface reconstruction results onto high-quality simulation meshes, leading to a loss of geometric accuracy during the simulation phase [
19,
20]. Secondly, full-order models can simulate flow and stress fields with high accuracy but are computationally expensive, failing to meet real-time monitoring needs [
21,
22]. Thirdly, existing hybrid models often focus on using data to compensate for the output errors of mechanism models, without fully exploiting the potential of data to correct internal parameters of the mechanism model itself [
23,
24]. Lastly, although optimization methods like genetic algorithms can effectively identify unmeasurable parameters, the iterative calculation process is typically performed offline, making online real-time updates difficult and limiting the adaptability of the digital twin [
25,
26].
In contrast to existing digital twin frameworks, this study presents three key innovations. First, we establish a seamless integration from reverse geometric modeling to CFD mesh generation, directly converting 3D-scanned point clouds into high-fidelity simulation meshes without geometric fidelity loss. Second, we propose a genetic algorithm-based online calibration method for unmeasurable physical parameters (flow coefficient, friction coefficient, and heat transfer coefficient) using real-time DCS data, which significantly improves dynamic simulation accuracy under deep peak-shaving conditions. Third, we implement a reduced-order model based on proper orthogonal decomposition that achieves real-time (sub-second) reconstruction of key physical fields while maintaining accuracy comparable to full-order CFD. This combination of reverse modeling, GA-driven parameter updating, and ROM-based real-time monitoring has not been previously reported for steam turbine digital twins under deep peak-shaving operations.
Based on the above analysis, leveraging the completed reverse modeling and mesh generation results of the full steam turbine model from the project, this paper proposes a digital twin modeling method for steam turbines tailored for deep peak-shaving conditions. The main research approach is as follows. Firstly, using 3D-scanned point cloud data, a full-model geometric twin, including 13 stages of high-pressure rotor blades, 9 stages of intermediate-pressure stator blades, and the cylinder, is established through fusion and surface reconstruction. Secondly, for critical parameters in the physical model that are difficult to measure directly, an optimization problem is formulated, aiming to minimize the error between the simulation output and monitoring data. Thirdly, based on methods like proper orthogonal decomposition, full-order simulation results are reduced to establish fast reconstruction models for key physical fields. Lastly, typical deep peak-shaving conditions are selected to compare the prediction accuracy of the model before and after optimization, verifying the effectiveness of the proposed method. The scope of this study is threefold: (i) to develop a seamless reverse modeling pipeline that converts 3D-scanned point clouds of a 600 MW subcritical steam turbine into high-fidelity CFD-ready meshes without geometric fidelity loss; (ii) to calibrate unmeasurable physical parameters (flow coefficient, friction coefficient, and heat transfer coefficient) using a genetic algorithm driven by real-time DCS data, thereby improving dynamic simulation accuracy under deep peak-shaving conditions (30% and 50% loads); and (iii) to construct a proper orthogonal decomposition (POD)-based reduced-order model that enables sub-second reconstruction of key temperature fields for online monitoring. The proposed digital twin framework is validated using field data from a 600 MW unit operating under typical deep peak-shaving scenarios.
2. Establishment of Geometric Model
The construction of the steam turbine digital twin is critical for monitoring the operating status and performing the optimization simulation. The structure of the steam turbine is complex, we propose a digital twin construction method integrating reverse modeling and refined mesh generation. The geometric model is established through four steps, including scanning, point cloud preprocessing, surface reconstruction, and model assembly. Firstly, to scan the key components, including 13 stages of high-pressure (HP) rotor blades and 9 stages of intermediate-pressure (IP) stator blades, a handheld laser 3D scanner with a single-shot accuracy of 0.05 mm is adopted.
Figure 1 shows the scanning of rotor blades and stator blades.
Affected by ambient light, vibration, and residual steam, some point clouds contain noise and outliers. Therefore, point cloud preprocessing is necessary; it comprises denoising, smoothing, and registration. Firstly, the statistical filtering algorithm is adopted to remove outliers, setting the number of neighboring points to 50 and the standard deviation multiplier to 2.0. Secondly, the moving least square algorithm is adopted to smooth the point cloud surface, eliminating scanning noise. For point clouds of different blades, the iterative closest point algorithm is adopted for multi-view registration, fusing multiple partial point clouds into a complete point cloud. After the point cloud preprocessing, the number of points is reduced from approximately 150,000 to 80,000, while preserving geometric features.
Figure 2 shows the rotor blade model before and after the point cloud preprocessing.
The surface reconstruction is the core of the reverse modeling. Firstly, based on the normal information and curvature variation of the point cloud, a region-growing algorithm is adopted to identify surface patches, segmenting the point cloud into several sub-regions with consistent geometric features. For each surface patch, a smooth-surface model is reconstructed with the non-uniform rational B-splines. Subsequently, the smooth-surface model is discretized into triangular meshes, generating an STL-format geometric model suitable for simulations.
Figure 3 shows the surface reconstruction of the rotor blade and stator blade.
The reconstructed models of the rotor, rotor blades, stator blades, control stage, and cylinder are spatially positioned and combined according to the actual assembly relationships. To ensure that the relative positions between blade stages and flow paths are consistent with the design drawings, the turbine centerline and bearing housing-locating surfaces are selected as assembly benchmarks.
Figure 4 shows the model assembly of the steam turbine, comprising approximately 3.2 million triangular facets, providing a high-fidelity geometric input for subsequent mesh generation.
The 3D scanning system has a specified point accuracy of ±0.05 mm. To assess the impact of scanning uncertainty on the simulation results, we performed a sensitivity analysis by perturbing the point cloud coordinates within ±0.05 mm and regenerating the mesh. The resulting variation in predicted HP exhaust temperature was less than ±0.3 °C, indicating that the scanning accuracy is sufficient for the target engineering application.
4. Unmeasurable Parameter Optimization
The construction of the high-precision physical model is a prerequisite for the digital twin being able to accurately reflect the operating status of the steam turbine. However, in the CFD model, there are several key parameters which are difficult to measure directly through theoretical calculation and experiment, including the flow coefficient, friction coefficient, and heat transfer coefficient, which affect the simulation accuracy significantly. If replaced by empirical constants, the model will produce large deviations under variable load conditions. Therefore, this paper proposes a method for identifying and optimizing unmeasurable parameters based on the genetic algorithm. By fusing monitoring data with the simulation model, the data-driven calibration of unmeasurable parameters is achieved, thereby enhancing the dynamic simulation accuracy and generalization capability of the digital twin.
Denote the vector of unmeasurable parameters as
, where
m is the number of unmeasurable parameters. The optimization problem is formulated with the objective of minimizing the root mean square error (RMSE) between the simulation output and monitoring data, namely,
where
N is the number of sampling points,
is the simulation output at the
i-th sampling point, and
is the corresponding monitoring data. The optimization variable
should satisfy physical constraints, including flow coefficient
, friction coefficient
, and heat transfer coefficient
.
As a global optimization algorithm simulating the natural evolution process, the genetic algorithm (GA) is suitable for solving nonlinear problems without requiring gradient information [
31,
32,
33,
34]. The genetic algorithm is composed of three parts, including encoding, the fitness function, and genetic operations.
Figure 6 shows the genetic algorithm optimization process. Firstly, real-number encoding is adopted, where each individual corresponds to a set of parameter vectors
. The population size is set to
. The initial population is randomly generated within the parameter constraint ranges. Secondly, the fitness function is defined as the reciprocal of the RMSE between the simulation output and monitoring data, ensuring that a larger fitness value indicates a better individual, namely,
where
is a small constant to prevent division by zero.
While other optimization methods exist, the genetic algorithm (GA) is selected here for three main reasons. First, the objective function is evaluated via CFD simulations, which are non-differentiable and may contain numerical noise; gradient-based methods (e.g., ADAM with surrogate neural networks) would require smooth approximations and are sensitive to initial guesses. Second, Bayesian optimization, though sample-efficient, typically performs poorly in moderate-dimensional (here ) problems with non-stationary response surfaces, and its surrogate (e.g., Gaussian process) scales cubically with the number of samples, becoming comparable to GA’s cost. Third, surrogate-assisted optimization was not employed because the CFD model (≈10,000 evaluations) proved computationally feasible given our parallel computing resources (64 cores, ≈3 s per evaluation). GA’s robustness to noise and its global search capability under physical constraints make it well suited for this unmeasurable parameter calibration problem.
In addition to the encoding and fitness function, the genetic operations are significant, including selection, crossover, and mutation. For the selection operation, tournament selection is adopted, namely,
individuals are randomly selected each time, and the one with the highest fitness is chosen to enter the next generation. This is repeated until the population size returns to
P. For the crossover operation, simulated binary crossover is adopted, with crossover probability
and distribution index
. For two parent individuals
and
, the offspring are generated by
where
is a random number determined by the distribution index. The third genetic operation is mutation, where polynomial mutation is adopted, with mutation probability
and distribution index
. After mutation, the offspring individual is checked against the parameter constraints. If it exceeds the bounds, re-mutation is performed. When the maximum number of generations
is met, the algorithm terminates.
In this study, the GA was executed directly with the CFD model, requiring approximately 10,000 function evaluations (50 individuals × 200 generations). Each CFD simulation (steady-state RANS) took about 3 s on a 64-core workstation, totaling around 8.3 CPU-hours, which is acceptable for an offline calibration task. However, for industrial deployment where CFD may be more expensive, a surrogate-assisted GA can be readily implemented. For example, a Gaussian process or radial basis function surrogate can be built from initial CFD samples, and the GA optimizes the surrogate, with periodic refinement using high-fidelity CFD evaluations. This reduces the number of CFD runs by 70–90% while retaining accuracy. Such an extension is straightforward and can be integrated into the proposed framework without altering the core methodology.
To validate the effectiveness of the unmeasurable parameter optimization, the optimized parameters
are substituted back into the CFD model. To cover the main operating range of the steam turbine during the deep peak-shaving, 30% load and 50% load are selected as typical conditions.
Table 5 lists the main operating parameters under typical deep peak-shaving conditions, which are extracted from the distributed control system (DCS) historical database. The temperature response is recalculated for 30% load and 50% load conditions and compared with the simulation results before optimization (using empirical constants).
Figure 7 shows the convergence curve of the genetic algorithm. As the number of generations increases, the fitness of both the average population and best individual increases monotonically. When the generation exceeds 50, the fitness of both the average population and best individual approaches steady state; the genetic algorithm reaches the convergence region.
The root mean square error (RMSE) is adopted to measure the performance of the genetic algorithm optimization, namely,
where
N is the number of sampling points,
is the simulation output at the
i-th sampling point, and
is the corresponding monitoring data.
Table 6 compares the temperature prediction errors before and after the genetic algorithm optimization. Compared with the RMSE of the temperature before the optimization, the RMSE of the temperature after the optimization is significantly lower, representing an accuracy improvement of over 72%. Furthermore,
Figure 8 compares the temperature prediction curves before and after optimization under the 50% load condition. Before optimization, the deviation between the simulation output and monitoring data is obvious, while the temperature prediction curve after optimization shows excellent agreement with the monitoring data. In summary, the genetic algorithm-based method can effectively calibrate the key parameters in the CFD model, significantly improving the simulation accuracy and generalization capability.
To evaluate the robustness of the GA optimization framework, a sensitivity analysis was conducted on three key parameters: population size
P, crossover probability
, and mutation probability
. The baseline values were set as
,
, and
. Each parameter was varied while keeping the others fixed. The resulting RMSE of temperature prediction under the 50% load condition is summarized in
Table 7. The results show that the RMSE varies within a narrow range (≤0.22 °C) across the tested values, indicating that the GA optimization is robust and not highly sensitive to parameter choices. The selected baseline values provide near-optimal performance.
5. Validation Under Deep Peak-Shaving Conditions
To achieve real-time monitoring, this paper employs the proper orthogonal decomposition (POD) method to establish a reduced-order model (ROM), enabling second-level reconstruction of the blade temperature fields. The specific steps include offline training and online mapping. In the offline training, the steady-state conditions at 30%, 50%, 75%, and 100% load are selected, totaling 32 snapshots. A snapshot matrix
is constructed, where
N is the number of grid nodes (approximately 200,000) and
M is the number of snapshots (32). Singular value decomposition (SVD) is performed on
, retaining the first
modes (capturing 99.2% of the energy) to obtain the reduced basis
. In the online mapping, the real-time boundary parameters collected from DCS (main steam pressure, temperature, pressure after control stage) are used as input. Radial basis function (RBF) interpolation yields the projection coefficient
for the current operating condition in the reduced space. The reconstructed physical field is calculated by
It is important to note that the POD method, as applied here, uses only steady-state snapshots (30%, 50%, 75%, 100% load); this captures the dominant spatial modes but does not explicitly model transient dynamics. Under deep peak-shaving conditions, rapid load ramps (e.g., 5% load/min) can introduce transient effects not fully represented by a steady-state POD basis. An alternative approach is Dynamic Mode Decomposition (DMD), which extracts spatio-temporal coherent structures from time-series data and can better capture evolving dynamics. A preliminary comparison using DMD on the same dataset yielded similar accuracy for quasi-steady operation but slightly higher errors (≈3.5% vs. 2.5%) during transient ramps due to noise sensitivity. Future work will combine POD and DMD in a hybrid ROM: using POD for spatial compression and DMD for temporal prediction, thereby improving fidelity under rapid load changes.
Figure 9 compares the temperature field reconstructed by the reduced-order model with the full-order CFD calculation result for the first-stage HP rotor blade. The maximum relative error between the full-order CFD results and reduced-order model reconstruction results is less than 2.5%; though the accuracy of the reduced-order model is a little worse than the full-order model, the computational efficiency is significantly accelerated: the single reconstruction with the reduced-order model is less than 1 s, meeting the real-time monitoring requirement.
To verify the accuracy and reliability of the monitoring method described above, experimental validation is conducted under two typical deep peak-shaving conditions, namely, 30% load and 50% load. The validation data is from the DCS historical database of a 600 MW subcritical steam turbine unit, with a sampling period of 1 s and a validation period length of 2 h for each condition.
Figure 10 shows a comparison between the monitored and measured values of HP exhaust temperature under the 50% load condition. The monitored values are reconstructed based on the reduced-order model, while the measured values come from the HP exhaust temperature sensor in the DCS. An excellent agreement between the reconstructed and measured temperatures throughout the entire period is shown, with a maximum absolute deviation of 3.2 °C, an average absolute deviation of 1.8 °C, and a relative error of less than 0.5%. Furthermore,
Table 8 summarizes the monitoring errors for the main temperature measurement points under the 30% and 50% loads. The average absolute deviation for all measurement points is less than 3 °C, and the relative error is below 0.6%, indicating that the reduced-order model possesses good accuracy for temperature field reconstruction.
The DCS temperature sensors have an accuracy of ±1 °C. To evaluate how sensor noise affects the ROM reconstruction, we added Gaussian noise ( = 0.5 °C, 1.0 °C) to the input boundary conditions (main steam temperature) and recomputed the reconstructed fields. The output temperature deviation was within ±0.6 °C, which is comparable to the ROM’s intrinsic error. Therefore, the overall uncertainty in monitored temperature is dominated by the ROM approximation error rather than sensor noise. A variance-based sensitivity analysis further revealed that the heat transfer coefficient contributed approximately 68% of the output variance in temperature prediction, highlighting the importance of its GA calibration.
To further contextualize the performance of our approach, a brief quantitative comparison with two representative digital twin modeling studies is provided. Yu et al. [
14] developed a hybrid model for a steam turbine control stage and reported a root mean square error of approximately 5.8 °C for temperature prediction under variable loads. Zhang et al. [
15] proposed a deep multi-model fusion digital twin for gas turbines, achieving a relative error of about 1.2%. In our study, the proposed method achieves an RMSE of 3.12 °C and 2.34 °C under 30% and 50% deep peak-shaving loads, respectively, with relative errors below 0.6%. Notably, our approach combines reverse modeling for geometric fidelity, GA-driven calibration of unmeasurable parameters, and real-time ROM reconstruction, providing a comprehensive framework that directly addresses the challenges of deep peak-shaving conditions, whereas previous studies focused on either geometric modeling or data-driven correction but not on the full integration.