A Multi-Source Cross-Domain Data Fusion Framework for Ordinal Health-State Assessment: A Reproducible Surrogate Benchmark Motivated by Hydrogen-Cooled Turbogenerators
Abstract
1. Introduction
2. Datasets
- (1)
- SKAB (mechanical cooling loop). SKAB v0.9 [27], collected from a laboratory water-pump testbed, covers the mechanical cooling loop, since it shares the pump-motor hydraulic dynamics that govern the stator cooling water circuit. Eight channels sampled at 1 Hz (accelerometer RMS, motor current/voltage, pump pressure, coolant/motor temperatures, and flow rate) are divided into 60 s windows with stride 30 s and channel-wise z-score standardization, yielding 1511 windows of shape . The per-window ordinal labels are derived from the anomaly ratio via the thresholding described below.
- (2)
- UCI-WWT (water chemistry). UCI-WWT [28] covers the water-chemistry domain with 527 daily physico-chemical records from a wastewater treatment facility, of which twelve features (conductivity, pH, biochemical oxygen demand (BOD), chemical oxygen demand (COD), suspended solids, sediment, zinc, and flow) are retained. Conductivity and pH overlap qualitatively with the indicators specified in DL/T 801-2010 [26], although the absolute conductivity of municipal wastewater is about two orders of magnitude higher than that of a deionized stator-coolant loop; the per-feature z-score standardization therefore retains the relative variation that drives the ordinal label, as is standard for heterogeneous-scale fusion, while the remaining features have no direct counterpart in a deionized loop (discussed in the validity analysis). Quantities such as BOD, COD, and suspended solids have no physical counterpart in an ultra-pure, deionized stator water loop; UCI-WWT therefore serves as a mathematical surrogate that supplies a heterogeneous tabular stream with an ordinal degradation signal, rather than a replica of the water-chemistry subsystem. Missing values are median-imputed, yielding 527 tabular samples of shape . The per-sample ordinal label is derived from the BOD removal efficiency . Specifically, the four ordinal levels are obtained by binning η at the fixed thresholds 0.91/0.87/0.80 listed in Table 1, which yields a 205/219/84/19 class distribution for this domain (Figure 1b). Since the label shares information with the two BOD channels that are also model inputs, UCI-WWT contributes mainly as one heterogeneous tabular stream within the fused task, where the max-aggregation rule bounds the influence of any single domain on the final grade.
- (3)
- CARE (electrical-thermal SCADA). CARE Wind Farm A [29] covers the electrical-thermal side with 22 SCADA recordings from an onshore Portuguese Wind Farm. The 46 channels available as 10 min averages are processed into 24-point sliding windows (stride 6, 4 h segments) with z-score standardization, yielding 8358 windows of shape . As the dataset is released with anonymized channel identifiers, CARE is used as a high-dimensional rotating-machinery SCADA stream that supplies a realistic multivariate electrical-thermal covariance structure and timestamped anomaly ground truth. This is the information the fusion model exploits, rather than a channel-level replica of a turbogenerator. Its event-level Normal/Anomaly labels provide the binary ground truth from which the ordinal levels are derived through the temporal overlap between each window and a labeled event; as with the other proxy domains, this ordinal information enters the fused task under the max-aggregation rule that bounds any single domain’s contribution. Because the native annotation is binary, the CARE domain is strongly concentrated on the Normal and Serious grades (5577/19/21/2741 windows); the two intermediate grades receive only 0.2% and 0.3% of the windows, so this domain contributes to them mainly indirectly, via the max-aggregation rule in the fused label.
- (4)
3. Methodology
3.1. Data Labeling and Splits
3.2. Multi-Stream Encoder–Fusion Network
- (1)
- Encoders. The five encoder candidates comprise three time-series variants (for SKAB, CARE, and Synth H2) and two tabular variants (for UCI-WWT). For the three time-series domains, the following three architectures are considered. The first, CNN-LSTM, applies two 1D-convolutional layers (kernel size 3, 32 channels, ReLU) followed by a one-layer bidirectional LSTM and attention pooling, projecting the output to . The second, Multi-Scale, extends this design with three parallel convolutions of kernel sizes , a Squeeze-and-Excitation channel-recalibration block [30], a two-layer BiLSTM, and a dual-pooling aggregation defined in Equation (4):
- (2)
- Fusion heads. Given the four domain embeddings , four fusion strategies of increasing expressiveness are evaluated. The first head, Concat, concatenates the four vectors and passes them through a two-layer MLP to produce . The second head, the Gated multi-modal unit [32], projects each branch through a tanh nonlinearity and combines them with learned softmax-normalized weights as in Equation (5):
3.3. Dual-Head Output and Hybrid Ordinal Loss
3.4. Training, Ensemble Calibration, and Evaluation Procedure
4. Results and Discussion
4.1. Comparison with State-of-the-Art Methods
4.2. Ablation Study on the Loss Function
4.3. Ablation Study on the Encoder–Fusion Configuration
4.4. Calibrated Ensemble and t-SNE Visualization
4.5. Scope and Limitations
5. Conclusions
- (1)
- A reproducible four-domain data collection has been constructed from SKAB (mechanical cooling loop), UCI-WWT (water chemistry), CARE Wind Farm A (electrical-thermal SCADA), and a physics-informed hydrogen-side stream derived from Henry’s law and a CSTR mass balance. It is released to provide public data for standards-aligned ordinal diagnosis. A multi-stream fusion network with a dual-head classifier, a hybrid CORN + EMD ordinal loss, and a calibrated ensemble based on per-model temperature scaling is introduced as the modeling counterpart to this data collection.
- (2)
- On the 600-sample test set, under a unified three-seed protocol against full-scale baselines, the calibrated ensemble achieves F1-macro = 0.5349, Accuracy = 0.6717, κ = 0.4713, and QWK = 0.5948, outperforming all thirteen baselines, with a +22.4 pp F1-macro gain over the strongest full-scale single-stream model (InceptionTime on CARE), +30.4 pp on κ and +36.2 pp on QWK. The margin over the strongest cross-entropy fusion baseline is +8.9 pp (bootstrap 95% CI [+5.1, +12.6] pp). The largest gains appear on the rank-aware indicators, which confirms the methodological value of pairing the CORN + EMD ordinal supervision with calibrated ensembling.
- (3)
- The ordinal-aware supervision shifts the error structure so that 60% of the errors fall on neighboring severity levels, producing the gap between = 0.4713 and QWK = 0.5948. The remaining 13% of all 600 predictions (40% of the errors) still differ from the ground truth by two or more levels, including safety-relevant underestimations of Serious samples. Reducing these distant underestimations is therefore the priority for the planned real-plant validation, since a severity-grading tool intended for a safety-critical asset can tolerate near-miss confusions far more readily than gross underestimations of the most degraded state.
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Klempner, G.; Kerszenbaum, I. Operation and Maintenance of Large Turbo-Generators; Wiley-IEEE Press: Hoboken, NJ, USA, 2004. [Google Scholar]
- Stone, G.C.; Culbert, I.; Boulter, E.A.; Dhirani, H. Electrical Insulation for Rotating Machines, 2nd ed.; Wiley-IEEE Press: Hoboken, NJ, USA, 2014. [Google Scholar]
- Khan, M.A.; Asad, B.; Kudelina, K.; Vaimann, T.; Kallaste, A. The Bearing Faults Detection Methods for Electrical Machines—The State of the Art. Energies 2023, 16, 296. [Google Scholar] [CrossRef]
- GB/T 43188-2023; Guide for Condition Evaluation of Turbogenerators. National Administration for Market Regulation and National Standardization Administration: Beijing, China, 2023.
- GB/T 7064-2017; Specific Requirements for Cylindrical Rotor Synchronous Machines. Standards Press of China: Beijing, China, 2017.
- Tshiloz, K.; Djurović, S. State-of-the-Art Techniques for Fault Diagnosis in Electrical Machines: Advancements and Future Directions. Energies 2023, 16, 6345. [Google Scholar] [CrossRef]
- Radiuk, P.; Rusyn, B.; Melnychenko, O.; Perzynski, T.; Sachenko, A.; Svystun, S. Criticality Assessment of Wind Turbine Defects via Multispectral UAV Fusion and Fuzzy Logic. Energies 2025, 18, 4523. [Google Scholar] [CrossRef]
- Filina, O.A.; Martyushev, N.V.; Malozyomov, B.V.; Tynchenko, V.S.; Kukartsev, V.A.; Bashmur, K.A. Increasing the Efficiency of Diagnostics in the Brush-Commutator Assembly of a Direct Current Electric Motor. Energies 2024, 17, 17. [Google Scholar] [CrossRef]
- Li, J.; Li, Q.; Zhu, J. Health Condition Assessment of Wind Turbine Generators Based on Supervisory Control and Data Acquisition Data. IET Renew. Power Gener. 2019, 13, 1343–1350. [Google Scholar] [CrossRef]
- Wu, H.; Hu, T.; Liu, Y.; Zhou, H.; Wang, J.; Long, M. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar] [CrossRef]
- Nie, Y.; Nguyen, N.H.; Sinthong, P.; Kalagnanam, J. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In Proceedings of the International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar] [CrossRef]
- Fawaz, H.I.; Lucas, B.; Forestier, G.; Pelletier, C.; Schmidt, D.F.; Weber, J. InceptionTime: Finding AlexNet for Time Series Classification. Data Min. Knowl. Discov. 2020, 34, 1936–1962. [Google Scholar] [CrossRef]
- Zhang, L.; Zhang, H.; Cai, G. The Multiclass Fault Diagnosis of Wind Turbine Bearing Based on Multisource Signal Fusion and Deep Learning Generative Model. IEEE Trans. Instrum. Meas. 2022, 71, 1–12. [Google Scholar] [CrossRef]
- Ding, J.; Zhao, Z. Diagnosis of Stator Inter-Turn Short Circuit Faults in Synchronous Machines Based on SFRA and MTST. Energies 2025, 18, 2142. [Google Scholar] [CrossRef]
- Xiao, Y.; Shao, H.; Liu, B. Evaluating Calibration of Deep Fault Diagnostic Models under Distribution Shift. Comput. Ind. 2025, 171, 104334. [Google Scholar] [CrossRef]
- Yang, Y.; Zhang, S.; Su, K.; Fang, R. Early Warning of Stator Winding Overheating Fault of Water-Cooled Turbogenerator Based on SAE-LSTM and Sliding Window Method. Energy Rep. 2023, 9, 199–207. [Google Scholar] [CrossRef]
- Zhang, Y.; Ji, J.C.; Ren, Z.; Ni, Q.; Gu, F.; Feng, K. Digital Twin-Driven Partial Domain Adaptation Network for Intelligent Fault Diagnosis of Rolling Bearing. Reliab. Eng. Syst. Saf. 2023, 234, 109186. [Google Scholar] [CrossRef]
- Yan, S.; Zhong, X.; Shao, H.; Ming, Y.; Liu, C.; Liu, B. Digital Twin-Assisted Imbalanced Fault Diagnosis Framework Using Subdomain Adaptive Mechanism and Margin-Aware Regularization. Reliab. Eng. Syst. Saf. 2023, 239, 109522. [Google Scholar] [CrossRef]
- Yang, S.; Yang, P.; Yu, H.; Bai, J.; Feng, W.; Su, Y. A 2DCNN-RF Model for Offshore Wind Turbine High-Speed Bearing-Fault Diagnosis under Noisy Environment. Energies 2022, 15, 3340. [Google Scholar] [CrossRef]
- Memari, M.; Shekaramiz, M.; Masoum, M.A.S.; Seibi, A.C. Data Fusion and Ensemble Learning for Advanced Anomaly Detection Using Multi-Spectral RGB and Thermal Imaging of Small Wind Turbine Blades. Energies 2024, 17, 673. [Google Scholar] [CrossRef]
- Jin, Z.; Xu, Q.; Jiang, C.; Wang, X.; Chen, H. Ordinal Few-Shot Learning with Applications to Fault Diagnosis of Offshore Wind Turbines. Renew. Energy 2023, 206, 1158–1169. [Google Scholar] [CrossRef]
- Shi, X.; Cao, W.; Raschka, S. Deep Neural Networks for Rank-Consistent Ordinal Regression Based on Conditional Probabilities. Pattern Anal. Appl. 2023, 26, 941–955. [Google Scholar] [CrossRef]
- Sander, R. Compilation of Henry’s Law Constants (Version 4.0) for Water as Solvent. Atmos. Chem. Phys. 2015, 15, 4399–4981. [Google Scholar] [CrossRef]
- Atkins, P.; de Paula, J.; Keeler, J. Atkins’ Physical Chemistry, 11th ed.; Oxford University Press: Oxford, UK, 2018. [Google Scholar]
- Fogler, H.S. Elements of Chemical Reaction Engineering, 5th ed.; Pearson: Upper Saddle River, NJ, USA, 2016. [Google Scholar]
- DL/T 801-2010; Water Quality and System Technical Requirements for Inner Cooling Water of Large Generators. China Electric Power Press: Beijing, China, 2010.
- Katser, I.; Kozitsin, V. Skoltech Anomaly Benchmark (SKAB); Kaggle: Moscow, Russia, 2020. [Google Scholar] [CrossRef]
- Poch, M. Water Treatment Plant; UCI Machine Learning Repository: Irvine, CA, USA, 1993. [Google Scholar] [CrossRef]
- Guck, C.; Roelofs, C.M.A.; Faulstich, S. CARE to Compare: A Real-World Benchmark Dataset for Early Fault Detection in Wind Turbine Data. Data 2024, 9, 138. [Google Scholar] [CrossRef]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 7132–7141. [Google Scholar] [CrossRef]
- Gorishniy, Y.; Rubachev, I.; Khrulkov, V.; Babenko, A. Revisiting Deep Learning Models for Tabular Data. In Advances in Neural Information Processing Systems; NIPS Foundation: San Diego, CA, USA, 2021. [Google Scholar] [CrossRef]
- Arevalo, J.; Solorio, T.; Montes-y-Gómez, M.; González, F.A. Gated Multimodal Units for Information Fusion. In ICLR Workshop; OpenReview.net: Toulon, France, 2017. [Google Scholar] [CrossRef]
- Hou, L.; Yu, C.-P.; Samaras, D. Squared Earth Mover’s Distance-Based Loss for Training Deep Neural Networks. arXiv 2016, arXiv:1611.05916. [Google Scholar] [CrossRef]
- Cui, Y.; Jia, M.; Lin, T.-Y.; Song, Y.; Belongie, S. Class-Balanced Loss Based on Effective Number of Samples. In 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 9268–9277. [Google Scholar] [CrossRef]
- Ren, J.; Yu, C.; Sheng, S.; Ma, X.; Zhao, H.; Yi, S. Balanced Meta-Softmax for Long-Tailed Visual Recognition. In Proceedings of the Advances in Neural Information Processing Systems, virtual, 6–12 December 2020. [Google Scholar] [CrossRef]
- Cao, K.; Wei, C.; Gaidon, A.; Aréchiga, N.; Ma, T. Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss. In Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar] [CrossRef]
- Zhang, H.; Cissé, M.; Dauphin, Y.N.; Lopez-Paz, D. Mixup: Beyond Empirical Risk Minimization. In Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, 30 April–3 May 2018. [Google Scholar] [CrossRef]
- Wolpert, D.H. Stacked Generalization. Neural Netw. 1992, 5, 241–259. [Google Scholar] [CrossRef]








| Level | Label | SKAB (r) | CARE (r) | UCI-WWT (η) | Synth (ṁ, μg/h) |
|---|---|---|---|---|---|
| Normal | 0 | r = 0 | r = 0 | η ≥ 0.91 | ṁ < 50 |
| Attention | 1 | 0 < r ≤ 0.30 | 0 < r ≤ 0.30 | 0.87 ≤ η < 0.91 | 50 ≤ ṁ < 500 |
| Abnormal | 2 | 0.30 < r ≤ 0.70 | 0.30 < r ≤ 0.70 | 0.80 ≤ η < 0.87 | 500 ≤ ṁ < 5000 |
| Serious | 3 | r > 0.70 | r > 0.70 | η < 0.80 | ṁ ≥ 5000 |
| Split | Class 0 | Class 1 | Class 2 | Class 3 | |
|---|---|---|---|---|---|
| Train | 3000 | 178 (5.9%) | 625 (20.8%) | 490 (16.3%) | 1707 (56.9%) |
| Validation | 600 | 50 (8.3%) | 111 (18.5%) | 94 (15.7%) | 345 (57.5%) |
| Test | 600 | 38 (6.3%) | 126 (21.0%) | 92 (15.3%) | 344 (57.3%) |
| Method | Params | F1-Macro | Acc | κ | QWK |
|---|---|---|---|---|---|
| InceptionTime [12] (CARE) | 0.51 M | 0.3112 ± 0.0149 | 0.4472 | 0.1674 | 0.2331 |
| TimesNet [10] (CARE) | 4.69 M | 0.3056 ± 0.0059 | 0.4378 | 0.1620 | 0.2299 |
| InceptionTime (SKAB) | 0.50 M | 0.2979 ± 0.0049 | 0.4550 | 0.1156 | 0.1613 |
| TimesNet (SKAB) | 4.69 M | 0.2829 ± 0.0124 | 0.3917 | 0.0864 | 0.1254 |
| PatchTST [11] (CARE) | 4.17 M | 0.2711 ± 0.0165 | 0.3739 | 0.0479 | 0.0782 |
| InceptionTime (Synth) | 0.50 M | 0.2411 ± 0.0076 | 0.3944 | 0.0113 | 0.0398 |
| XGBoost (4-domain) | – | 0.3847 ± 0.0108 | 0.6394 | 0.3366 | n/a |
| Random forest (4-domain) | – | 0.3826 ± 0.0147 | 0.5794 | 0.3134 | n/a |
| Logistic regression (4-domain) | – | 0.3691 ± 0.0000 | 0.4617 | 0.2144 | n/a |
| Transformer (CARE, CORN + EMD) | 0.10 M | 0.2905 ± 0.0068 | 0.5189 | 0.1289 | 0.1753 |
| MLP (UCI, CORN + EMD) | 0.03 M | 0.2061 ± 0.0063 | 0.5639 | 0.0182 | 0.0381 |
| CNN-LSTM + Cross-Attn + CE (4-domain) | 0.29 M | 0.4463 ± 0.0151 | 0.5317 | 0.3155 | n/a |
| CNN-LSTM + Cross-Attn + CORN (4-domain) | 0.29 M | 0.4602 ± 0.0233 | 0.6111 | 0.3561 | 0.4558 |
| Ours single (Multi-Scale + FTT + Gated, CORN + EMD) | 1.62 M | 0.4770 ± 0.0183 | 0.5950 ± 0.0341 | 0.3794 ± 0.0300 | 0.4929 ± 0.0173 |
| Ours ensemble (combined-32 + T) | 32 × (0.22–1.62 M) | 0.5349 | 0.6717 | 0.4713 | 0.5948 |
| Δ vs. InceptionTime (CARE) | +22.4 pp | +22.5 pp | +30.4 pp | +36.2 pp | |
| Δ vs. strongest CE fusion | +8.9 pp |
| Sub-Ensemble | T-Only (Val) | T-Only (Test) | T + Offsets (Val) | T + Offsets (Test) |
|---|---|---|---|---|
| ours-only (26 models) | 0.4657 | 0.5149 | 0.4677 | 0.5197 |
| legacy-only (6 models) | 0.4603 | 0.5235 | 0.4729 | 0.5360 |
| combined (32 models) | 0.4624 | 0.5349 | 0.4699 | 0.5217 |
| Perturbation | Ensemble F1 | ΔF1 (pp) | Ensemble QWK | ΔQWK (pp) | Single-Model ΔF1 (pp) |
|---|---|---|---|---|---|
| Nominal | 0.5349 | 0.00 | 0.5948 | 0.00 | 0.00 |
| Henry’s constant ×0.8/0.9/1.1/1.2 | 0.5349 | 0.00 | 0.5948 | 0.00 | 0.00 |
| Solution enthalpy ×0.8/×1.2 | 0.5349 | 0.00 | 0.5948 | 0.00 | 0.00 |
| Loop time constant ×0.8/×1.2 | 0.5349 | 0.00 | 0.5948 | 0.00 | 0.00 |
| Inlet temperature −2 °C/+2 °C | 0.5069/0.5065 | −2.80/−2.84 | 0.5704/0.5629 | −2.44/−3.19 | +0.18/−0.71 |
| H2 pressure ×0.9/×1.1 | 0.5251/0.5044 | −0.98/−3.06 | 0.5425/0.5661 | −5.22/−2.87 | −0.15/−2.35 |
| Sensor noise ×1.5 | 0.5303 | −0.46 | 0.5841 | −1.06 | −0.25 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zheng, C.; Huang, X.; Zhang, G. A Multi-Source Cross-Domain Data Fusion Framework for Ordinal Health-State Assessment: A Reproducible Surrogate Benchmark Motivated by Hydrogen-Cooled Turbogenerators. Appl. Sci. 2026, 16, 7764. https://doi.org/10.3390/app16157764
Zheng C, Huang X, Zhang G. A Multi-Source Cross-Domain Data Fusion Framework for Ordinal Health-State Assessment: A Reproducible Surrogate Benchmark Motivated by Hydrogen-Cooled Turbogenerators. Applied Sciences. 2026; 16(15):7764. https://doi.org/10.3390/app16157764
Chicago/Turabian StyleZheng, Changjun, Xuancheng Huang, and Guodong Zhang. 2026. "A Multi-Source Cross-Domain Data Fusion Framework for Ordinal Health-State Assessment: A Reproducible Surrogate Benchmark Motivated by Hydrogen-Cooled Turbogenerators" Applied Sciences 16, no. 15: 7764. https://doi.org/10.3390/app16157764
APA StyleZheng, C., Huang, X., & Zhang, G. (2026). A Multi-Source Cross-Domain Data Fusion Framework for Ordinal Health-State Assessment: A Reproducible Surrogate Benchmark Motivated by Hydrogen-Cooled Turbogenerators. Applied Sciences, 16(15), 7764. https://doi.org/10.3390/app16157764

