Federated Fault Diagnosis for Heterogeneous Satellite Constellations Using Adaptive Dual Knowledge Distillation
Abstract
1. Introduction
- (1)
- Satellite heterogeneity and data heterogeneity. Constellations typically comprise multiple satellite types with different sensor configurations, system architectures, and mission objectives, resulting in complex failure modes and diverse data formats. Traditional federated learning frameworks, designed for scenarios with homogeneous data and models, struggle to achieve effective knowledge transfer and exhibit poor model generalization capabilities.
- (2)
- Data transmission constraints under dynamic inter-satellite communications. Operating in complex and harsh space environments, satellite constellations rely on inter-satellite links characterized by limited bandwidth, high latency, and intermittent connectivity. Traditional federated learning requires frequent global model synchronization and data exchange, consuming substantial inter-satellite communication resources and hindering real-time model updates to fault diagnosis.
- (3)
- Dynamic constellation configuration changes and scarcity of fault samples. Due to real-time variations in constellation operations and environmental conditions, constellation configurations dynamically evolve. The global server cannot maintain real-time connections with all clients, necessitating local client models capable of diagnosing all faults across identical satellite models. However, individual satellites experience few faults during operation, leading to insufficient training samples. This results in local models overfitting and poor generalization, hindering accurate diagnosis of unknown faults.
- (1)
- To address model and data heterogeneity, a multimodal heterogeneous federated architecture and a feature alignment mechanism are designed. Multimodal local client models are established based on different satellite models. Local feature extractors capture personalized client features, overcoming the rigidity of unified modeling in traditional federated learning. A local feature alignment mechanism is proposed to aggregate multiple clients’ local features at the global server, enabling knowledge transfer across satellite models.
- (2)
- To address excessive inter-satellite communication pressure, a lightweight bidirectional knowledge distillation strategy is proposed. Clients perform local knowledge distillation via sparse feature aggregation, transmitting local feature vectors to the server. The server applies soft-target distillation to clients, using the server model’s diagnosed class probability distribution as a supervision signal. This approach transmits only lightweight log probabilities instead of raw data to clients, significantly reducing communication volume per round.
- (3)
- To address dynamic constellation topology and scarce fault samples, we propose a two-stage dynamic knowledge distillation strategy. The collaborative bidirectional distillation stage optimizes the server’s global diagnostic capability for known faults, and the subsequent task-specific stage enhances local models’ generalization to unseen faults using a frozen server teacher. Trained models can autonomously diagnose all faults across identical satellites without real-time communication.
2. Related Work
2.1. Constellation Fault Diagnosis
2.2. Federated Learning
2.3. Research Gaps
- (1)
- Lack of a heterogeneity-aware FL framework that can simultaneously handle both satellite platform heterogeneity and data heterogeneity.
- (2)
- Insufficient attention to communication efficiency in FL-based fault diagnosis methods, making them impractical for resource-constrained satellite systems.
- (3)
- Limited research on enhancing local model generalization to unseen faults, which is essential for autonomous constellation operations.
3. Adaptive Federated Dual Knowledge Distillation Framework
3.1. Constellation Heterogeneous Federated Learning Model Architecture
3.2. Algorithm Workflow
- Collaborative Bidirectional Distillation Phase. The optimization objectives are to enhance the server’s global diagnostic effectiveness and improve clients’ learning of known fault feature knowledge. Each client trains models based on locally seen data and extracts personalized features. Their output distributions and personalized features are uploaded to the server for knowledge alignment, enabling server-side aggregation updates. The updated server then downloads soft labels to clients, minimizing the discrepancy between client predictions and server soft labels using KL divergence. This feature knowledge transfer between clients and servers is termed bidirectional knowledge distillation.
- Task-specific Distillation Phase. When the server achieves the required global diagnostic rate, the task-specific distillation phase commences. The optimization goal is to enhance the client’s generalization capability for unseen faults. The server parameters are frozen as a fixed teacher to guide the client’s unknown data training. Simultaneously, the client trains based on locally known data to maintain its diagnostic capabilities for known faults. The client adjusts its training focus through a dynamic weighting mechanism, concentrating on enhancing diagnostic capabilities for unknown fault categories.
3.3. Collaborative Bidirectional Distillation Phase
3.3.1. Client Local Training
3.3.2. Server Knowledge Distillation
3.3.3. Client Reverse Knowledge Distillation
3.4. Task-Specific Distillation Phase
| Algorithm 1. Complete Two-Stage Training Procedure of AFDL |
| Input: Private datasets for k = 1, 2…N; Initial global server model ; Initial local client models ; Phase transition thresholds , ; Initial loss weights for Phase 2: , ; Weight update step size , weight bounds ; Maximum communication rounds T1 (Phase1), T2 (Phase1); Number of selected clients per round K; Output: Trained global server model ; Trained local client models ; I Phase 1: Collaborative Bidirectional Distillation For round t = 1, 2, …, T1 do: 1. Client Local Training Process For each client k = 1, 2, …, N in parallel do For each batch do ① Extract local features and predictions and ② Compute classification Loss ③ Update client models ④ Upload to server: Local features and predictions 2. Server Knowledge Distillation Process For server do ① Aggregate uploaded knowledge: Global feature center ; Average soft labels ② Compute total server loss : Server classification loss ; Distillation loss ; Feature alignment loss ③ Update server model 3. Client Reverse Distillation For each client k = 1, 2, …, N do ① Download from server: Updated global soft labels ; Server feature Vector ② Compute total client loss: Client classification loss ; Distillation loss ; Feature alignment loss ③ Update client models 4. Check phase transition condition If server accuracy ≥ and all clients’ seen accuracy ≥ : Break and proceed to Phase 2. II: Phase 2: Task-specific Distillation 1. Freeze server model parameters . 2. Server distributes teacher guidance to each client k: server latent features , server soft predictions . 3. Initialize: . For round t = 1, 2, …, T2 do: For each client k = 1, 2, …, N do: ① Compute seen loss via cross-entropy classification loss. ② Compute unseen loss . ③ Evaluate current unseen accuracy on local validation set. ③ Calculate accuracy improvement: ② Update loss weights dynamically. ② Compute total training loss ② Update local model parameters via gradient descent. |
4. Experiment
4.1. Proprietary In-Orbit Satellite Control System Dataset Experiment
4.1.1. Experimental Data
4.1.2. Parameter Settings
4.1.3. Experimental Results
4.1.4. Ablation Experiments
4.1.5. Comparison of Baseline Methods
4.2. Public ESA Anomaly Dataset Experiment
4.2.1. Experimental Data
4.2.2. Parameter Settings
4.2.3. Experiment Result
4.2.4. Ablation Experiments
4.2.5. Comparison of Baseline Methods
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Kulu, E. Satellite constellations—2024 survey, trends and economic sustainability. In Proceedings of the International Astronautical Congress, IAC, Milan, Italy, 14–18 October 2024; pp. 14–18. [Google Scholar]
- Hedayati, M.S.; Barzegar, A.; Rahimi, A. Fault diagnosis and prognosis of satellites and unmanned aerial vehicles: A review. Appl. Sci. 2024, 14, 9487. [Google Scholar] [CrossRef]
- Teng, F.; Zhu, Y.; Zhang, E.; Hu, X.; Sun, Q.; Feng, L. A knowledge service framework for fault diagnosis of low-earth orbit satellite constellation. In 2023 IEEE International Conference on Web Services (ICWS); IEEE: New York, NY, USA, 2023; pp. 669–676. [Google Scholar]
- Liu, C.; Chen, L.; Ding, J.; Shangguan, D. Modeling of satellite constellation in modelica and a PHM system framework driven by model data hybrid. Electronics 2022, 11, 2155. [Google Scholar] [CrossRef]
- Han, Y.; Fan, W.; Han, F.; Cheng, H. Low-Orbit Giant Constellation Fault Propagation and Fault Source Detection Algorithm. In 2024 43rd Chinese Control Conference (CCC); IEEE: New York, NY, USA, 2024; pp. 5002–5007. [Google Scholar]
- Zhu, Z.; Lei, Y.; Qi, G.; Chai, Y.; Mazur, N.; An, Y.; Huang, X. A review of the application of deep learning in intelligent fault diagnosis of rotating machinery. Measurement 2023, 206, 112346. [Google Scholar] [CrossRef]
- Wang, C.; Sun, Y.; Wang, X. Image deep learning in fault diagnosis of mechanical equipment. J. Intell. Manuf. 2024, 35, 2475–2515. [Google Scholar]
- Kumar, D.; Addula, S.R.; Lind, M.; Brown, S.; Odion, S. AI-Driven Hybrid Deep Learning and Swarm Intelligence for Predictive Maintenance of Smart Manufacturing Robots in Industry 4.0. Electronics 2026, 15, 715. [Google Scholar] [CrossRef]
- Ding, Y.; Ma, L.; Ma, J.; Suo, M.; Tao, L.; Cheng, Y.; Lu, C. Intelligent fault diagnosis for rotating machinery using deep Q-network based health state classification: A deep reinforcement learning approach. Adv. Eng. Inform. 2019, 42, 100977. [Google Scholar] [CrossRef]
- Li, Z.; Jiang, H.; Wang, X. A novel reinforcement learning agent for rotating machinery fault diagnosis with data augmentation. Reliab. Eng. Syst. Saf. 2025, 253, 110570. [Google Scholar] [CrossRef]
- Zhang, C.; Wang, Y.; You, X. Fault diagnosis in rotating machinery with discretized signal representation leveraging large language models. Appl. Soft Comput. 2025, 189, 114487. [Google Scholar] [CrossRef]
- Xu, C.; Wang, Z.; Jin, Y.; Nong, W. An adaptive industrial large language model for mechanical fault diagnosis under variable operating conditions. Adv. Eng. Inform. 2026, 74, 104821. [Google Scholar] [CrossRef]
- Wen, J.; Zhang, Z.; Lan, Y.; Cui, Z.; Cai, J.; Zhang, W. A survey on federated learning: Challenges and applications. Int. J. Mach. Learn. Cybern. 2023, 14, 513–535. [Google Scholar] [CrossRef] [PubMed]
- Boobalan, P.; Ramu, S.P.; Pham, Q.V.; Dev, K.; Pandya, S.; Maddikunta, P.K.R.; Gadekallu, T.R.; Huynh-The, T. Fusion of federated learning and industrial Internet of Things: A survey. Comput. Netw. 2022, 212, 109048. [Google Scholar] [CrossRef]
- Younis, R.; Fisichella, M. FLY-SMOTE: Re-balancing the non-IID IoT edge devices data in federated learning system. IEEE Access 2022, 10, 65092–65102. [Google Scholar] [CrossRef]
- Zhang, W.; Li, X.; Ma, H.; Luo, Z.; Li, X. Federated learning for machinery fault diagnosis with dynamic validation and self-supervision. Knowl.-Based Syst. 2021, 213, 106679. [Google Scholar] [CrossRef]
- Kang, S.; Sun, Y.; Li, X.; Wang, Y.; Wang, Q.; Liang, X. Unsupervised fault diagnosis method for rolling bearings based on federated universal domain adaptation. Eng. Appl. Artif. Intell. 2025, 162, 112518. [Google Scholar] [CrossRef]
- Yan, P.; Hu, Y.; Wen, W.; Hsu, L.-T. Multiple Faults Isolation for Multi-Constellation GNSS Positioning through Incremental Expansion of Consistent Measurements. IEEE Sens. J. 2025, 25, 6967–6981. [Google Scholar] [CrossRef]
- Meng, Q.; Liu, J.; Zeng, Q.; Feng, S.; Xu, R. Impact of one satellite outage on ARAIM depleted constellation configurations. Chin. J. Aeronaut. 2019, 32, 967–977. [Google Scholar] [CrossRef]
- Huang, G.; Xu, C.; Zhao, J.; Song, D. Bayesian fault-tolerant protection level for multi-constellation navigation from integrity perspective. Aerosp. Sci. Technol. 2022, 130, 107954. [Google Scholar] [CrossRef]
- Zhang, K.; Song, X.; Zhang, C.; Yu, S. Challenges and future directions of secure federated learning: A survey. Front. Comput. Sci. 2022, 16, 165817. [Google Scholar]
- Dai, M.; Xu, A.; Huang, Q.; Zhang, Z.; Lin, X. Vertical federated DNN training. Phys. Commun. 2021, 49, 101465. [Google Scholar] [CrossRef]
- Guan, J.; Cai, J.; Bai, H.; You, I. Deep transfer learning-based network traffic classification for scarce dataset in 5G IoT systems. Int. J. Mach. Learn. Cybern. 2021, 12, 3351–3365. [Google Scholar] [CrossRef]
- Jiang, S.; Wang, B.; Zhang, X.; Jiang, Y.; Liu, S.; Zhao, Z.; Li, R.; Chen, X. ObsBattery: Position-Aware Federated Learning with Dueling DQN Clustering and Training Adaptation for Satellite Battery Prediction. Electronics 2025, 14, 4697. [Google Scholar] [CrossRef]
- Zhang, P.; Wang, C.; Jiang, C.; Han, Z. Deep reinforcement learning assisted federated learning algorithm for data management of IIoT. IEEE Trans. Ind. Inform. 2021, 17, 8475–8484. [Google Scholar] [CrossRef]
- Lu, S.; Gao, Z.; Xu, Q.; Jiang, C.; Zhang, A.; Wang, X. Class-imbalance privacy-preserving federated learning for decentralized fault diagnosis with biometric authentication. IEEE Trans. Ind. Inform. 2022, 18, 9101–9111. [Google Scholar] [CrossRef]
- Lin, T.; Kong, L.; Stich, S.U.; Jaggi, M. Ensemble distillation for robust model fusion in federated learning. Adv. Neural Inf. Process. Syst. 2020, 33, 2351–2363. [Google Scholar]
- Wu, C.; Wu, F.; Lyu, L.; Huang, Y.; Xie, X. Communication-efficient federated learning via knowledge distillation. Nat. Commun. 2022, 13, 2032. [Google Scholar] [CrossRef] [PubMed]
- Wu, S.; Chen, J.; Nie, X.; Wang, Y.; Zhou, X.; Lu, L.; Peng, W.; Nie, Y.; Menhaj, W. Global prototype distillation for heterogeneous federated learning. Sci. Rep. 2024, 14, 12057. [Google Scholar] [CrossRef] [PubMed]
- Kotowski, K.; Haskamp, C.; Andrzejewski, J.; Ruszczak, B.; Nalepa, J.; Lakey, D.; Collins, P.; Kolmas, A.; Bartesaghi, M.; Martinez-Heras, J.; et al. European space agency benchmark for anomaly detection in satellite telemetry. arXiv 2024, arXiv:2406.17826. [Google Scholar]











| Method | Heterogeneity Support | Distillation Mechanism | Unseen Fault Generalization |
|---|---|---|---|
| FedDF | Only model architecture heterogeneity | One-way output-level distillation; requires proxy data | None |
| FedKD | Only model heterogeneity | Local mentor–mentee mutual distillation | None |
| FedGPD | Only data distribution heterogeneity (non-IID) | One-way feature-level prototype distillation | None |
| AFDL | Platform + data dual heterogeneity | Bidirectional feature & output distillation, lightweight | Two-stage adaptive dynamic weighting |
| Small Client | Large Client | Server | |
|---|---|---|---|
| Input Dimension | 26 | 33 | 33 |
| Feature Extractor | Linear (26, 256) Linear (256, 128) | Linear (33, 512) Linear (512, 128) | Linear (33, 512) Linear (512, 128) |
| Feature Alignment | Linear (128, 128) | Linear (128, 128) | Linear (128, 128) |
| Classifier | Linear (128, 20) | Linear (128, 20) | Linear (128, 20) |
| Dropout | 0.3 | 0.3 | 0.3 |
| Feature Dimension | 128 | 128 | 128 |
| Output Category | 20 | 20 | 20 |
| Total parameter count | 92 K | 145 K | 145 K |
| Parameters | Phase 1 | Phase 2 | Instruction |
|---|---|---|---|
| Learning Round | 60 | 60 | Early Stopping with Conditions |
| Client selection per round | 10/32 | 10/32 | Random Uniform Sampling |
| Local training epoch | 2 | 2 | Client Local Update |
| Batch size | 64 | 64 | Data Loading Batch |
| Client learning rate | 0.01 | 0.01 | Adam Optimizer |
| Server learning rate | 0.001 | 0.001 | Adam Optimizer |
| Feature Alignment Weight | 0.3 | 0.3 | MSE Loss Coefficient |
| Distillation Loss Weight | 0.5 | Dynamic Adjustment | KL Divergence Loss Coefficient |
| Temperature Parameters τ | 2.0 | 2.0 | Knowledge Distillation Temperature |
| Weight Adjustment | Fixed | Dynamic Adjustment | Improvement on Unseen Classes |
| Indicator | M1 | M2 | M3 | AFDL |
|---|---|---|---|---|
| Server Accuracy Rate | 67.8% ± 3.1% | 90.1% ± 3.6% | 91.9% ± 2.0% | 92.3% ± 1.4% |
| Mean Client Seen Accuracy | 99.3% ± 0.2% | 99.5% ± 0.3% | 98.9% ± 0.7% | 99.5% ± 0.3% |
| Mean Client Unseen Accuracy | 52.6% ± 7.1% | 88.4% ± 3.5% | 54.7% ± 5.1% | 84.8% ± 2.6% |
| Training Time (s) | 61.3 ± 8.5 | 86.4 ± 12.3 | 50.4 ± 7.9 | 69.7 ± 9.8 |
| Communication Volume (MB) | 365.2 ± 38.4 | 681.3 ± 68.7 | 285.3 ± 31.2 | 412.6 ± 42.3 |
| Indicator | FedDF | FedKD | FedGPD | AFDL |
|---|---|---|---|---|
| Server Accuracy Rate | 81.1% ± 3.2% | 69.2% ± 4.6% | 78.4% ± 3.1% | 92.3% ± 1.4% |
| Mean Client Seen Accuracy | 99.3% ± 0.4% | 99.1% ± 0.5% | 98.7% ± 0.9% | 99.5% ± 0.3% |
| Mean Client Unseen Accuracy | 53.2% ± 5.7% | 55.5% ± 6.9% | 63.6% ± 4.3% | 84.8% ± 2.6% |
| Training Time (s) | 57.9 ± 2.8 | 49.3 ± 7.4 | 52.4 ± 7.6 | 69.7 ± 9.8 |
| Communication Volume (MB) | 286.5 ± 30.3 | 263.4 ± 28.7 | 315.2 ± 33.9 | 412.6 ± 42.3 |
| Fault Mode | Faulty Channel Combination | Number | Total | |
|---|---|---|---|---|
| Normal | 753,472 | 753,472 | ||
| Fault Type 1 | Fault Mode 1 | Channel 22,30,31,38,39 | 20,560 | 69,391 |
| Fault Mode 2 | Channel 15,22,23,31,39 | 18,561 | ||
| Fault Mode 3 | Channel 15,16,22,23,24,31,32,33,39,40 | 8139 | ||
| Fault Mode 4 | Channel 22,31,39 | 3909 | ||
| Fault Mode 5 | Others | 18,222 | ||
| Fault Type 2 | Fault Mode 6 | Channel 16,24,32,33,40 | 21,595 | 55,538 |
| Fault Mode 7 | Channel 16,24,32,33,40,47,48,49,50,51,52,57,58,59,60,66 | 6953 | ||
| Fault Mode 8 | Channel 12,13,17,18,19,20,25,26,27,28,34,35,36,37,41,42,43,44,45,46,47,48,49,50,51,52 | 5638 | ||
| Fault Mode 9 | Channel 16,24,32,33,40,47,48,49,50,51,52,66 | 4287 | ||
| Fault Mode 10 | Others | 22,703 |
| Small Client | Large Client | Server | |
|---|---|---|---|
| Input Dimension | 50 | 76 | 76 |
| Feature Extractor | Linear (50, 256) Linear (256, 128) | Linear (76, 512) Linear (512, 128) | Linear (76, 512) Linear (512, 128) |
| Feature Adapter | Linear (128, 128) | Linear (128, 128) | Linear (128, 128) |
| Classifier | Linear (128, 11) | Linear (128, 11) | Linear (128, 11) |
| Dropout | 0.3 | 0.3 | 0.3 |
| Feature Dimension | 128 | 128 | 128 |
| Output Category | 11 | 11 | 11 |
| Total parameter count | 107 K | 183 K | 183 K |
| Parameters | Phase 1 | Phase 2 | Instruction |
|---|---|---|---|
| Learning Round | 50 | 50 | Early Stopping with Conditions |
| Client selection per round | 12/32 | 12/32 | Random Uniform Sampling |
| Local training epoch | 2 | 2 | Client Local Update |
| Batch size | 32 | 32 | Data Loading Batch |
| Client learning rate | 0.001 | 0.001 | Adam Optimizer |
| Server learning rate | 0.0005 | 0.0005 | Adam Optimizer |
| Feature Alignment Weight | 0.2 | 0.3 | MSE Loss Coefficient |
| Distillation Loss Weight | 0.6 | Dynamic Adjustment | KL Divergence Loss Coefficient |
| Temperature Parameters τ | 2.0 | 2.0 | Knowledge Distillation Temperature |
| Weight Adjustment | Fixed | Dynamic Adjustment | Improvement on Unseen Classes |
| Indicator | M1 | M2 | M3 | AFDL |
|---|---|---|---|---|
| Server Accuracy Rate | 63.4% ± 3.8% | 84.2% ± 2.6% | 84.7% ± 2.3% | 85.6% ± 1.9% |
| Mean Client Seen Accuracy | 99.1% ± 0.5% | 99.3% ± 0.4% | 98.6% ± 0.8% | 99.4% ± 0.3% |
| Mean Client Unseen Accuracy | 46.8% ± 6.4% | 74.1% ± 4.2% | 51.6% ± 5.7% | 72.9% ± 3.3% |
| Training Time (s) | 7812 ± 806 | 11,264 ± 1187 | 6358 ± 742 | 8967 ± 874 |
| Communication Volume (MB) | 148.7 ± 19.3 | 286.41 ± 34.86 | 119.58 ± 16.24 | 172.36 ± 22.42 |
| Indicator | FedDF | FedKD | FedGPD | AFDL |
|---|---|---|---|---|
| Server Accuracy Rate | 75.7% ± 3.3% | 64.8% ± 4.8% | 71.6% ± 2.7% | 85.6% ± 1.9% |
| Mean Client Seen Accuracy | 99.2% ± 0.4% | 99.1% ± 0.6% | 98.7% ± 0.9% | 99.4% ± 0.3% |
| Mean Client Unseen Accuracy | 45.6% ± 6.1% | 46.8% ± 6.4% | 54.2% ± 5.3% | 72.9% ± 3.3% |
| Training Time (s) | 7423 ± 815 | 6810 ± 769 | 6972 ± 798 | 8967 ± 874 |
| Communication Volume (MB) | 126.84 ± 17.53 | 118.62 ± 15.94 | 132.27 ± 18.68 | 172.36 ± 22.42 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Xing, X.; Wang, S.; Liu, W.; Liu, C.; Liang, H.; Zhang, Y. Federated Fault Diagnosis for Heterogeneous Satellite Constellations Using Adaptive Dual Knowledge Distillation. Electronics 2026, 15, 3056. https://doi.org/10.3390/electronics15143056
Xing X, Wang S, Liu W, Liu C, Liang H, Zhang Y. Federated Fault Diagnosis for Heterogeneous Satellite Constellations Using Adaptive Dual Knowledge Distillation. Electronics. 2026; 15(14):3056. https://doi.org/10.3390/electronics15143056
Chicago/Turabian StyleXing, Xiaoyu, Shuyi Wang, Wenjing Liu, Chengrui Liu, Hanyu Liang, and Yan Zhang. 2026. "Federated Fault Diagnosis for Heterogeneous Satellite Constellations Using Adaptive Dual Knowledge Distillation" Electronics 15, no. 14: 3056. https://doi.org/10.3390/electronics15143056
APA StyleXing, X., Wang, S., Liu, W., Liu, C., Liang, H., & Zhang, Y. (2026). Federated Fault Diagnosis for Heterogeneous Satellite Constellations Using Adaptive Dual Knowledge Distillation. Electronics, 15(14), 3056. https://doi.org/10.3390/electronics15143056

