1. Introduction
During the in-orbit operation of spacecraft, it is impossible to avoid the hypervelocity impact (HVI) caused by micrometeoroids and space debris (MMOD) [
1,
2,
3]. Such impacts typically act on the surface and internal units of structures at extremely high incident speeds and transient kinetic energy, causing a wide variety of damages (such as craters, cracks, perforations, etc.) with a large scale range. Moreover, the development patterns and maintenance strategies of different damage types vary significantly [
4,
5,
6,
7]. Therefore, precise detection of HVI damage is fundamental, and the development of classification technologies capable of accurately distinguishing different types of damage has become a key technical requirement for supporting the long-term stable service of spacecraft [
8,
9]. In this context, detection methods based on vibration response analysis and infrared thermography offer promising technical paths to address these challenges, due to their advantages of non-contact, remote, and full-field coverage [
10,
11,
12,
13]. Katam et al. [
10] proposed a damage detection method based on STFT time–frequency feature extraction and autoencoder dimensionality reduction, achieving high-accuracy damage classification under limited data conditions through time–frequency analysis of vibration signals. Zhang et al. [
11] introduced a vibration-based damage detection approach utilizing phase-based motion estimation and convolutional neural networks, which enables precise localization of bolt looseness damage with single-sample training via pixel-level vibration signal extraction from videos. Gao et al. [
12] proposed a hypervelocity impact (HVI) damage detection method based on infrared data extraction and multi-objective optimization algorithms. By extracting representative infrared features and reconstructing defect regions, the method effectively improved detection accuracy and efficiency. Yin et al. [
13] introduced the Dynamic Multi-Objective Feature Extraction Optimization (DM-FEO) method, which enhances the precision of HVI damage detection through temperature point extraction in infrared thermography and multi-directional prediction algorithms, successfully distinguishing the thermal characteristics of different damage types. Vibration detection is sensitive to changes in the structural frequency response function, making it effective for identifying local stiffness degradation and delamination damage, as well as macro-scale defects [
14]. Recent research has further demonstrated that coupling finite element models with metaheuristic optimization algorithms can significantly enhance the precision of damage localization in composite structures [
15]. Wang et al. [
16] comprehensively reviewed the physical principles, excitation modalities, and applications of multimode infrared thermal-wave imaging in non-destructive testing. In particular, the integration of active infrared imaging with deep learning techniques has been identified as a highly accurate solution for defect detection in complex industrial components [
17]. However, it is important to note that these methods primarily serve the detection and preliminary identification of “damage or no damage.” When the task objective shifts to “high-precision classification” of damage types, the inherent limitations of these unimodal methods become evident.
Accurate damage classification of spacecraft, such as identifying crack, perforation, and delamination, is crucial for supporting precision on-orbit maintenance. This task faces three major challenges: first, different damage types exhibit high feature similarity under a single modality, making them difficult to distinguish accurately; second, vibration signals are susceptible to noise interference while infrared features are influenced by imaging distance, leading to poor model stability—a problem also prominent in complex industrial settings, as Tian et al. [
18] demonstrated that variable working conditions hinder stable feature extraction; third, the scarcity of annotated on-orbit samples coupled with class imbalance limits the training effectiveness of data-driven models, prompting Yin et al. [
19] to employ generative adversarial networks to enhance scarce sample representation. Thus, relying on a single modality fails to achieve both comprehensive coverage and discriminative robustness, particularly under sample-limited conditions required for high-accuracy classification.
To overcome the limitations of unimodal methods, recent research has shifted toward multimodal data fusion for detection [
20,
21]. These studies demonstrate that integrating complementary multimodal information is an effective strategy for improving performance in complex tasks. However, existing fusion methods primarily focus on feature-level or data-level fusion, requiring precise alignment of heterogeneous data, and they do not explicitly address the critical issue of uncertainty in classification tasks. As a result, their effectiveness is limited when handling challenges such as distinguishing highly similar damage types in fine-grained classification. Data-driven automatic feature learning and classification provide a new technological path for HVI damage detection [
22,
23,
24]. Models, particularly convolutional neural networks (CNNs) and residual networks (ResNet), can learn discriminative features across scales and forms in an end-to-end manner, alleviating the degradation of deep models and enhancing training stability [
25]. Ye et al. [
26] developed a multi-sensor residual fusion network utilizing double-ring residual and global interactive modules to achieve robust diagnosis under noisy and small-sample conditions; Le-Xuan et al. [
27] constructed a hybrid 1DCNN-LSTM-ResNet architecture that effectively captures long-term temporal dependencies and enhances damage detection accuracy. Yan et al. [
28] proposed a QCNN-based method for bearing fault diagnosis by fusing audio and vibration signals for high-precision detection in noisy environments.
Furthermore, the combination of multimodality and deep learning shows even better prospects. Xu et al. [
29] proposed a multimodal neural network fusion method that integrates large-kernel networks with an attention mechanism, achieving highly robust damage identification in noisy environments. Meshram et al. [
30] designed an integrated model utilizing multimodal transformers and hybrid deep learning to optimize pothole detection and repair in diverse environments. Cao et al. [
31] proposed a multimodal feature fusion model combining ResNet and GRU to distinguish pseudo defects in ultra-thick stainless-steel welds using Phased Array Ultrasonic Testing. Peng et al. [
32] developed a multimodal hybrid neural network, significantly improving damage classification accuracy and data efficiency under complex operational conditions. Research by Zaman et al. [
33] and Chen et al. [
34] also demonstrated the advantages of multimodal fusion in fault diagnosis. However, the performance of these methods is highly dependent on the scale and quality of training data. In HVI damage scenarios, real labeled data is scarce, especially for infrared images, where image quality is highly dependent on the shooting distance. Distance uncertainty leads to feature distribution shifts, which has become a major bottleneck in applying deep learning methods.
This study aims to establish a robust multimodal fusion framework for high-precision HVI damage classification. Key objectives include: (1) overcoming unimodal physical limitations, such as vibration insensitivity and thermal blurring; (2) mitigating feature degradation caused by noise and variable imaging distances; and (3) resolving inter-modal conflicts via uncertainty quantification for reliable diagnosis. By integrating physics-informed signal enhancement with evidence-theoretic fusion, this work provides an accurate solution for on-orbit structural health monitoring. As depicted in
Figure 1, the proposed framework targets the challenge of high-precision HVI damage classification. Key contributions include:
A vibration-infrared multimodal fusion framework that systematically integrates vibration time–frequency analysis with infrared distance-aware enhancement, providing a robust solution for feature extraction in variable environments.
A dual-stream deep residual network architecture developed to extract and fuse modality-specific features. It combines spectral subtraction and STFT for vibration signals and utilizes hierarchical enhancement to address infrared feature shifts caused by imaging distance variations.
A decision-level fusion method based on D-S evidence theory. By quantifying epistemic uncertainty, this approach suppresses inter-modal conflicts and maximizes complementary advantages, ensuring reliable classification for spacecraft maintenance.
Figure 1.
Overall flowchart of the vibration-infrared multimodal fusion classification method.
Figure 1.
Overall flowchart of the vibration-infrared multimodal fusion classification method.
2. Problem Statement
Spacecraft inevitably encounter hypervelocity impacts (HVIs) from micrometeoroids and orbital debris (MMOD). To ensure survivability, mission requirements extend beyond detection to the precise classification of damage types (e.g., cracks, delamination, or perforation) for targeted maintenance. However, transitioning to high-precision classification faces fundamental challenges of intrinsic ambiguity and feature aliasing. As illustrated in
Figure 2, these challenges propagate from the physical perception layer to the decision layer.
First, signal degradation in harsh environments limits sensing reliability. Vibration responses, critical for identifying micro-damage, are often masked by broadband mechanical noise, resulting in a low Signal-to-Noise Ratio (SNR) that renders time-domain analysis ineffective. Similarly, infrared feature representation is highly sensitive to imaging distance. Variations in detection distance cause severe “domain shift” in thermal feature distributions, hindering model generalization under variable operating conditions.
Second, data heterogeneity and sample scarcity create a “semantic gap” in feature learning. Synergizing 1D temporal vibration signals (global stiffness) with 2D thermal images (local flow blockage) requires aligning incompatible geometric manifolds. Furthermore, real-world data exhibits an extreme long-tail distribution: scarce high-risk samples (e.g., perforations) are overwhelmed by abundant minor damage samples. This imbalance biases decision boundaries toward the majority class, leading to critical missed detections.
Finally, sensor conflicts compromise deterministic decision-making. In complex scenarios, a “Paradox State” arises when sensors yield contradictory diagnoses (e.g., internal delamination triggering vibration anomalies while remaining thermally invisible). In such cases, traditional hard-voting mechanisms fail. Therefore, establishing a mathematical framework to quantify “epistemic uncertainty” is imperative to manage conflicts and preserve high-confidence evidence.
To systematically address these challenges—signal degradation (Physical Layer), feature alignment gaps (Feature Layer), and evidential conflicts (Decision Layer)—this paper proposes a multimodal fusion framework (
Figure 2, right panel) that achieves robust classification through multi-level processing mechanisms. To satisfy on-orbit real-time constraints, the framework prioritizes computational efficiency through a lightweight design. By adopting decision-level fusion over high-dimensional feature concatenation, computational overhead is minimized. The decoupled dual-stream topology supports hardware parallelism, ensuring system latency is determined by the longest single branch rather than cumulative processing time. Additionally, input dimension optimization reduces floating-point operations (FLOPs), guaranteeing rapid response capabilities for sudden impacts.
5. Discussion
5.1. Physical Interpretation of Misclassification Mechanisms
The misclassification patterns observed in the confusion matrices stem from the inherent physical limitations of single sensing modalities and the strong physical co-occurrence of HVI damage. In the vibration modality, the misidentification of micro-cracks as normal states arises from the global nature of modal responses; micro-scale defects often fail to induce statistically significant global stiffness changes or frequency shifts, creating a detection blind spot. Similarly, the infrared modality is constrained by thermal diffusion effects, where ablation and delamination generate indistinguishable surface temperature gradients due to similar thermal resistance boundaries under transient heating.
These physical ambiguities confirm that relying on a single sensor inevitably leads to decision uncertainty. By leveraging Dempster–Shafer (D-S) evidence theory, the proposed framework effectively resolves these conflicts. Unlike traditional weighted averaging, which risks propagating errors in high-conflict scenarios, the D-S rule utilizes the “uncertainty” mass to discount conflicting evidence, thereby achieving superior reliability in distinguishing physically coupled damage types.
5.2. Framework Scalability and Material Universality
The validity of this study is underpinned by high-fidelity physical experiments rather than numerical simulations. From a mechanistic perspective, the proposed detection framework exhibits theoretical scalability across different material systems. In vibration analysis, damage-induced local mass loss and stiffness degradation inevitably lead to eigenfrequency shifts, a phenomenon governing both metallic and composite structures. Similarly, infrared detection relies on thermal wave scattering at defect interfaces, a fundamental heat transfer behavior applicable to various materials. Therefore, while signal amplitudes may vary, the underlying feature extraction and D-S fusion logic remain transferable. Future applications on other spacecraft materials would primarily require parameter fine-tuning rather than architectural reconstruction.
5.3. Limitations and Future Work
Despite the demonstrated efficacy of the proposed framework in identifying damage categories, this study focuses primarily on discrete classification and has not yet addressed the quantitative assessment of damage severity (e.g., hole size or crack depth). The current model provides a qualitative diagnosis suitable for triggering alarms but lacks the capability for precise lifespan prediction. Future work will extend this architecture from classification to regression tasks, aiming to establish a continuous mapping between multimodal features and quantitative damage metrics, thereby providing a more comprehensive solution for spacecraft structural health monitoring.
6. Conclusions
This study proposes a physics-informed multimodal fusion framework for reliable spacecraft HVI damage monitoring. By integrating vibration spectral subtraction and distance-aware infrared enhancement within a dual-stream ResNet architecture, the method overcomes unimodal physical limitations such as stiffness insensitivity and thermal ambiguity. A core innovation is the Dempster–Shafer (D-S) evidence-theoretic mechanism, which resolves inter-modal conflicts at the decision level.
Experiments on multi-class HVI specimens demonstrate that the fusion strategy achieves a mean accuracy of 99.01%, significantly outperforming unimodal baselines (92.96% and 97.11%). The D-S mechanism effectively corrects conflicting evidence, achieving perfect precision in four categories. Framework stability is confirmed by a low standard deviation of 0.74% in 5-fold random seed tests, while “Leave-One-Distance-Out” cross-validation proves robust generalization across imaging distances. Furthermore, ablation studies verify that the shallow-layer Channel Attention Mechanism is critical for capturing micro-scale damage features.
In summary, this work establishes a high-precision baseline for on-orbit diagnosis. Future research will extend the framework from discrete classification to continuous regression for quantitative damage assessment and further optimize the model for satellite edge computing deployment.