1. Introduction
As the bedrock of ensuring operational reliability and continuous airworthiness, the structural integrity of the aircraft skin remains a focal point in civil aviation maintenance and inspection. Under complex and volatile operational conditions, fuselage surfaces are inevitably subjected to the synergistic effects of cyclic aerodynamic stresses, environmental corrosion, and foreign object damage (FOD). These factors subsequently induce various types of degradation, including crack propagation, localized indentations, and coating delamination [
1,
2]. Research indicates that if such incipient micro-scale damages are not accurately identified during early-stage preventive maintenance, they can escalate into catastrophic structural failures [
3]. Nevertheless, contemporary maintenance frameworks remain heavily reliant on conventional non-destructive testing (NDT) techniques, primarily manual visual inspection (MVI) and acoustic-based tap testing. While these methods offer operational flexibility in engineering practice, their diagnostic outcomes are profoundly influenced by the subjective expertise of inspectors, leading to significant inter-operator variability. Furthermore, when tasked with large-scale inspections, these manual approaches encounter severe efficiency bottlenecks and expose personnel to the inherent hazards of high-altitude operations [
4,
5].
Modern aviation industry’s pursuit of inspection efficiency and objectivity has catalyzed the widespread application of various NDT technologies. However, the performance of existing techniques in practical hangar environments remains constrained by inherent limitations. For instance, ultrasonic testing (UT) exhibits high sensitivity to surface roughness and the quality of coupling agents; even minute interfacial air gaps can cause severe acoustic energy attenuation, thereby compromising the accuracy of defect characterization [
6]. Although digital shearography can capture sub-micron surface displacement gradients, its imaging quality is highly susceptible to interference from rigid-body motion and environmental instabilities when inspecting large-scale skin structures [
7,
8,
9]. More critically, the interpretation of results from these techniques still largely depends on expert heuristics, making a truly closed-loop, fully automated system difficult to achieve [
10,
11]. Consequently, developing a real-time, precise, and robust automated surface inspection system—such as deep learning-driven visual detection—has emerged as a pivotal research focus in the field of intelligent aviation maintenance [
12,
13,
14].
Over the past five years, the evolution of artificial intelligence has spearheaded a paradigm shift in computer vision (CV) within the realm of industrial quality inspection, propelling automated NDT into a new stage of intelligence [
15]. By integrating high-resolution imaging terminals with advanced algorithms, modern inspection systems have achieved a leap from coarse-grained monitoring to precision-driven analysis. In the landscape of modern intelligent manufacturing, deep learning-based visual inspection solutions—leveraging their distinct advantages of non-contact operation, high precision, and real-time online processing—are progressively superseding traditional hand-crafted feature extraction methods. Consequently, these technologies have emerged as the core decision-making engine for aircraft Maintenance, Repair, and Overhaul (MRO) systems [
16,
17,
18,
19], as illustrated in
Figure 1.
The technical evolution of aircraft skin inspection systems is essentially a paradigm shift in feature representation. To surmount imaging challenges in high-altitude and occluded regions, a hardware matrix centered around unmanned aerial vehicles (UAVs) and intelligent guidance platforms has emerged as the standardized acquisition infrastructure [
20,
21]. Nevertheless, the core efficacy of the inspection system ultimately hinges upon the backend algorithm’s capacity to deconstruct and interpret massive volumes of visual data. Consequently, the research nexus has shifted from manual, rule-based feature engineering to data-driven, hierarchical representation learning. Conventional detection paradigms are often constrained by the limited adaptability of predefined filters; when confronted with low-contrast defects such as subtle scratches or indentations, rigid parameterization frequently results in high miss rates. In stark contrast, deep learning architectures leverage multi-scale feature fusion strategies to spontaneously learn semantic-rich feature descriptors directly from pixel sequences [
22]. This end-to-end mechanism facilitates the adaptive decoupling of non-linear noise interference under complex operational conditions. Consequently, in real-world hangar scenarios characterized by variable illumination and severe background clutters, it demonstrates generalization resilience and recognition accuracy that are significantly superior to those of traditional heuristic algorithms.
Despite substantial advancements in visual inspection, achieving fully automated and precise identification within the domain of ASD still encounters significant bottlenecks. For instance, the majority of surface impairments, such as microscopic cracks, occupy an extremely low pixel ratio within a macroscopic field of view; their attenuated feature responses are frequently eclipsed during successive downsampling operations. Moreover, exacerbated by the high specular reflectivity of aluminum alloy skins, volatile ambient illumination, and complex fuselage curvatures, defects tend to amalgamate with background textures—such as rivet boundaries and coating particulates—resulting in exceptionally low contrast [
23].
Critically, ASD samples exhibit pronounced intra-class heterogeneity (diverse morphologies within a single defect category) and inter-class similarity (e.g., superficial scratches resembling background textures), posing a rigorous challenge to feature decoupling and the demarcation of precise classification boundaries. Furthermore, the extreme scarcity of severe defect samples in real-world datasets creates a distribution imbalance that predisposes models to overfit majority classes, thereby compromising the sensitivity toward critical yet rare anomalies [
24].
To address the aforementioned challenges, this study proposes DEE-Net, a high-precision defect detection framework based on an improved YOLO11. This framework adheres to the core design principles of information fidelity and selective enhancement, aiming to solve the problem of accurately identifying minute damages in complex industrial settings. The primary contributions of this paper are summarized as follows:
We incorporate a feature reassembly strategy based on the existing SPD-Conv mechanism to mitigate information loss during downsampling for tiny defects. By replacing traditional strided convolutions with space-to-depth transformations, this approach helps preserve fine-grained spatial information and maintains sub-pixel structural details. As a result, it provides a more informative feature representation for subsequent stages of the network.
We designed a dual-domain coordinated ASE module for feature enhancement. The module combines multi-scale perception paths with an edge enhancement component to better capture fine-grained defect features. By incorporating the existing DSM [
25], the module jointly considers spatial localization and frequency-domain information. This design helps the model distinguish defect-related features from redundant background textures under complex conditions, such as metallic specular reflections.
We incorporate a quality-aware regression optimization strategy based on the existing Wise-CIoU loss function. By integrating Wise-CIoU into the proposed framework, the model retains the comprehensive penalty terms of CIoU (including overlap area, center distance, and aspect ratio), ensuring stable localization performance. In addition, the non-monotonic focusing mechanism of Wise-CIoU is utilized to adaptively adjust the gradient contributions of samples with different qualities. This design helps mitigate the influence of low-quality samples and reduces training instability in complex industrial scenarios. As a result, it contributes to improved bounding box regression performance, particularly for subtle defects.
The remainder of this paper is organized as follows:
Section 2 provides a comprehensive review of the related work in the field of ASD detection, highlighting the evolution of methodologies.
Section 3 elaborates on the architectural design and technical mechanisms of the proposed DEE-Net, detailing the synergy between its core modules.
Section 4 presents the experimental setup, results, and a rigorous performance analysis, including comparative studies with state-of-the-art models. Finally,
Section 5 summarizes the research findings and outlines potential avenues for future work.
2. Related Works
In the early stages, aircraft skin defect detection predominantly hinged upon MVI and fundamental NDT methodologies. While MVI is inherently straightforward, it is highly susceptible to inspector fatigue, subjective expertise, and the rigorous constraints of high-altitude working environments, often leading to elevated missed detection rates [
26]. To bolster objectivity and reliability, techniques such as eddy current testing, UT, and digital shearography have been extensively deployed for identifying metallic surface impairments. However, these methods exhibit significant practical constraints: UT demonstrates poor adaptability to complex curvilinear geometries and relies heavily on the quality of the acoustic coupling agent; meanwhile, shearography remains extremely sensitive to ambient vibrations [
27]. Most critically, these conventional technologies struggle to facilitate large-scale, real-time autonomous operations.
With the rise of convolutional neural networks (CNNs), deep learning-based object detection algorithms have achieved milestone advancements in the field of industrial quality inspection. Currently, mainstream architectures are categorized into two primary paradigms: two-stage algorithms, epitomized by Faster R-CNN, and one-stage algorithms, represented by YOLO and SSD. Although two-stage algorithms offer advantages in localization precision, their inference speeds often fall short of engineering standards in aircraft inspection scenarios with stringent real-time requirements (e.g., UAV-based hangar inspections). In contrast, the YOLO series achieves an effective balance between detection velocity and accuracy by reformulating the detection task into a unified regression problem.
However, general-purpose detection models often suffer from domain shift when directly migrated to ASD tasks. The presence of intense spectral reflections on the skin surface, irregular rivet textures, and microscopic crack features poses a significant challenge. Specifically, during successive downsampling operations, universal models tend to lose critical, subtle signals, leading to both missed detections and false alarms [
23], as illustrated in
Figure 2.
To address the challenges of attenuated features and complex backgrounds in ASD tasks, recent academic efforts have explored three pivotal dimensions: feature fidelity, edge reinforcement, and loss function optimization [
28,
29]. First, conventional networks predominantly employ strided convolutions for downsampling, which inevitably leads to the excessive compression of spatial information. To mitigate the loss of tiny object features, the recently proposed SPD-Conv mechanism utilizes space-to-depth transformations for lossless feature reassembly, establishing a new theoretical paradigm for preserving fine-grained spatial topological information [
30]. Furthermore, since aircraft skin defects—such as scratches and cracks—typically manifest as linear edge signals, researchers have attempted to integrate attention mechanisms to bolster edge representation [
21,
31]. However, achieving precise discriminative decoupling of valid defect boundaries under intense background noise (e.g., brushed metal textures) remains a formidable challenge in current multi-scale perception research [
32]. Finally, traditional IoU loss functions, such as CIoU, are largely based on static geometric constraints. To handle the uneven distribution of sample quality in industrial datasets, the newly developed Wise-IoU (WIoU) offers a dynamic weighting scheme to mitigate the interference from low-quality samples [
33].
In summary, although existing studies have improved ASD detection, several challenges remain, particularly in preserving fine-grained features, suppressing background interference, and improving localization stability. Some approaches focus mainly on feature enhancement, while others emphasize geometric constraints or sample-quality modeling. In this study, DEE-Net is developed as a YOLO11-based framework that integrates feature preservation, selective feature enhancement, and regression optimization. By combining and adapting existing techniques with the proposed ASE module, the framework aims to improve defect representation under complex industrial backgrounds. To facilitate a clear understanding of the overall design, a high-level schematic of DEE-Net is presented in
Figure 3.
4. Experimental Results Analysis
4.1. Experimental Setup
4.1.1. Dataset
This study utilizes the publicly available aircraft skin defect dataset (ASDD). The dataset contains a total of 3007 high-resolution images sourced from the website (
https://universe.roboflow.com/project-5lf3h/dataset2-69an0, accessed on 18 December 2025), covering five typical types of surface damage in aircraft inspection: Crack, Dent, Missing-head, Paint-off, and Scratch. To ensure that the experimental results are comparable with existing public benchmarks, this experiment does not adopt the conventional 8:1:1 split but strictly follows the original distribution protocol of the dataset, as shown in
Table 1.
To address the issues of imbalanced sample distribution and complex geometric variations, this study employs a dual-strategy framework combining offline and online data augmentation. Specifically, offline augmentation is first applied to the original training data to enhance sample diversity, followed by online augmentation techniques during training. It is important to note that all data augmentation operations are applied exclusively to the training data. The validation and testing sets, as well as the cross-validation partitioning, are strictly based on the original (non-augmented) images to prevent data leakage and ensure a fair evaluation.
To ensure a fair and reliable evaluation, special attention was given to the data partitioning process and potential similarity between training and testing samples. All data splits were performed at the level of original images, and no augmented samples were used for validation or testing. In particular, augmented variants derived from the same original image were strictly confined to the training set, ensuring that no visually similar samples appeared across different subsets. This design effectively prevents potential data leakage. For robustness evaluation, a 5-fold cross-validation protocol was adopted on the entire set of original (non-augmented) images. The dataset was randomly shuffled and partitioned into five subsets of approximately equal size. In each fold, four subsets were used for training, and one subset was used for evaluation. To preserve the data distribution, a stratified sampling strategy was applied.
To facilitate reproducibility, all experiments were conducted using fixed training settings, including a consistent number of training epochs, learning rate, and batch size across all folds. The best-performing model weights were selected based on validation performance. In addition, a fixed random seed was used during data partitioning to ensure consistent results across repeated runs.
4.1.2. Implementation Details
The experiments are implemented based on the PyTorch 2.5.1 deep learning framework using Python 3.7. All models are trained and evaluated on a workstation running Windows 11. The detailed hardware and software configurations are summarized in
Table 2.
To ensure optimal convergence and facilitate a fair comparison, all models are trained from scratch without loading any pre-trained weights. This strategy is adopted to rigorously validate the feature extraction capability of the proposed method specifically for the aircraft skin defect detection task. The input images are resized to
pixels. The stochastic gradient descent (SGD) optimizer is employed for weight updates. The detailed training hyperparameters are summarized in
Table 3.
4.1.3. Evaluation Metrics
To comprehensively evaluate the performance of DEE-Net in aircraft skin defect detection, we selected evaluation metrics across three dimensions: detection accuracy, localization precision, and model efficiency. Regarding detection accuracy, we employ Precision (P), Recall (R), and their harmonic mean, the F1-Score, which are defined as
Localization precision is primarily assessed using the mean average precision at an IoU threshold of 0.5 (mAP@0.5) and mAP@0.5:0.95, the latter of which averages mAP values across IoU thresholds from 0.5 to 0.95 with a step size of 0.05. This metric imposes stricter requirements on defect boundary localization and provides a more realistic reflection of the model’s spatial positioning capabilities in complex industrial environments. Furthermore, model efficiency is measured by the number of the computational cost represented by GFLOPs, where the latter signifies the floating-point operations during a single forward pass and is critical for real-time detection on UAV-mounted edge terminals.
4.2. Cross-Validation Protocol and Results
To further evaluate the robustness of the proposed method and reduce the influence of a single fixed data partition, we conducted a 5-fold cross-validation experiment on all original non-augmented images. The dataset was randomly divided into five subsets of approximately equal size. In each fold, four subsets were used for training, and the remaining subset was used for evaluation. Data augmentation was applied only to the training subset in each fold, while the evaluation subset remained unchanged. This setting prevents augmented variants of the same original image from appearing in both training and evaluation subsets, thereby reducing the risk of data leakage.
The final cross-validation results are reported as the mean and standard deviation across the five folds. The overall results are shown in
Table 4, while the category-level cross-validation results for Scratch are presented in
Table 5.
4.3. Ablation Study
Ablation experiments are conducted to systematically verify the performance contribution of each core improvement module within DEE-Net. This study utilizes YOLO11 as the baseline model and progressively incorporates the SPD-Conv, ASE, and Wise-CIoU loss function to evaluate their respective impacts. The comprehensive comparative results of the ablation study for each improvement module are presented in
Table 6.
Quantitative Analysis of Ablation Study
The baseline model (YOLO11) exhibits a relatively high recall rate of 84.91% in the skin detection task, demonstrating its architectural potential for candidate object discovery. However, its precision is limited to 71.15%. A fine-grained analysis of the prediction results reveals that the baseline model suffers from significant perceptual deficiencies when processing high-frequency micro-features, such as cracks. Moreover, it is susceptible to artifact interference generated by the highly reflective background of the aircraft skin, which leads to a higher rate of false positives (FPs) and significant localization deviations.
By incorporating the SPD-Conv structure to replace traditional strided convolutions, the model’s precision significantly improved from 71.15% to 78.33%, while mAP@0.5 reached 81.37%. This mechanism utilizes a space-to-depth conversion strategy to reorganize sub-pixel structural features into the channel dimension, effectively alleviating the information loss inherent in conventional downsampling operations. Experimental evidence demonstrates that SPD-Conv bolsters the model’s perceptual sensitivity to defect edge signals, thereby suppressing the generation of invalid predictions at the source of feature extraction.
Following the integration of the ASE module, the mAP@0.5 steadily climbed to 83.25%. The core performance gain is manifested in the enhanced capture of geometrically complex categories, such as “Dent,” and defects with blurred edges. This improvement is primarily attributed to the high-frequency signal compensation provided by the internal EB component of the ASE module, as well as the filtering efficacy of the dual-domain selection mechanism (DSM). Functioning as a selective information bottleneck, DSM adaptively decouples target features from background noise, allowing the model to accurately lock onto discriminative semantics highly relevant to the detection task even within high-interference environments.
With the final introduction of the Wise-CIoU loss function, DEE-Net achieves its optimal comprehensive performance. The mAP@0.5:0.95 reaches a peak of 51.80%, representing an increase of 4.31 percentage points compared to the version incorporating only the ASE module. Notably, for the highly challenging “Crack” category, the mAP@0.5 eventually reaches 69.78%, significantly outperforming the 60.87% achieved by the baseline model. This validates the superiority of the non-monotonic focusing mechanism when handling uneven sample quality distributions, successfully guiding the model to converge toward sub-pixel level optimal bounding boxes.
4.4. Comparison of Detection Performance
To comprehensively evaluate the performance of the proposed DEE-Net algorithm in complex defect detection tasks, we conducted a quantitative comparison with current mainstream object detection algorithms under identical experimental settings. To ensure the fairness and objectivity of the comparison, all models were trained and optimized using a unified hyperparameter configuration within the same hardware environment. The comparative models include the classic two-stage detector Faster R-CNN (ResNet50) and a series of state-of-the-art one-stage detectors: YOLOv8s, YOLOv10s, YOLO11s, and YOLO12s.
Table 7 summarizes the comprehensive performance of each model across key metrics, including Precision, Recall, mAP@0.5, mAP@0.5:0.95, GFLOPs, and FPS. Furthermore,
Table 8,
Table 9,
Table 10,
Table 11,
Table 12 and
Table 13 provide a detailed breakdown of the specific detection results for each individual defect category.
From the perspective of overall detection performance, DEE-Net achieves competitive performance across most evaluation metrics. Specifically, DEE-Net obtains an mAP@0.5 of 87.47% and an mAP@0.5:0.95 of 51.80%, corresponding to improvements of 7.15% and 2.43%, respectively, compared to the YOLO11s baseline. These results suggest that the proposed modifications contribute to improved feature extraction and multi-scale representation. While maintaining a relatively high recall (83.82%), DEE-Net also improves precision to 82.54%, compared to 71.15% for YOLO11s. This indicates that the model is better able to reduce false positives under complex background conditions. Such characteristics are beneficial for industrial inspection scenarios where both detection accuracy and reliability are important.
Compared with the two-stage detector Faster R-CNN, DEE-Net achieves higher detection performance while maintaining significantly lower computational cost. Although Faster R-CNN employs a ResNet50 backbone, it achieves an mAP@0.5 of 65.47% with a computational cost of 948.18 GFLOPs, which is substantially higher than that of DEE-Net. These results indicate that the proposed method provides a more favorable trade-off between accuracy and efficiency. This significant disparity in computational cost stems from fundamental architectural differences: the GFLOPs of Faster R-CNN are substantially higher than those of the YOLO series, primarily because the large number of proposals generated by the region proposal network (RPN) in its two-stage architecture must be processed individually by the RoI Head. Notably, the reported GFLOPs correspond to the full inference cost including all proposal processing steps. In contrast, as a one-stage detector, the computational cost of the YOLO series is primarily concentrated in the convolutional operations of the backbone and neck and is typically reported as MACs. By comparison, DEE-Net reduces computational cost by approximately 98.8% (to only 11.5 GFLOPs), while improving inference speed by 68% (from 50.7 FPS to 85.23 FPS) and achieving a 22% increase in accuracy, fully demonstrating the dual superiority of the proposed method in both efficiency and accuracy.
A category-level analysis shows that detection performance varies among different defect types. Although the “Scratch” category achieves very high performance under the original fixed split, this result should be interpreted cautiously due to the limited number of test samples and the relatively distinguishable visual patterns of some scratch instances. The 5-fold cross-validation results reported in
Section 4.2 provide a more conservative estimate, suggesting that the near-100% performance under the original split may be influenced by the specific data partition. Therefore, the conclusions of this study are mainly based on the aggregated performance across all folds and categories. The performance on underrepresented categories will be further investigated in future work using larger and more balanced datasets.
In terms of computational efficiency, the introduction of additional feature enhancement modules increases the computational cost of DEE-Net to 11.5 GFLOPs, compared to approximately 6 GFLOPs for the YOLO11 baseline. This results in a reduction in inference speed from 98.18 FPS to 85.23 FPS. Nevertheless, the processing speed remains above the commonly required threshold for real-time industrial inspection. This trade-off between computational cost and detection performance reflects the improved representation capability of the model, particularly for detecting fine-grained defects. Overall, the results suggest that DEE-Net provides a balanced compromise between accuracy and efficiency for practical inspection scenarios.
Overall, the proposed method demonstrates consistent performance improvements over baseline models while maintaining real-time capability, indicating its potential applicability in aircraft surface inspection tasks.
4.5. Robustness Evaluation Under Environmental Perturbations
To further evaluate the generalization ability of DEE-Net under degraded imaging conditions, we conducted robustness experiments using an expanded challenge test set. In practical UAV-assisted aircraft inspection, image quality may be affected by factors such as insufficient illumination and camera motion. Therefore, the challenge test set (N = 147) was constructed to simulate two common degradation factors: low-light conditions and motion blur.
The overall performance comparison is summarized in
Table 14, while the detailed detection results for each specific defect category are presented in
Table 15 and
Table 16.
The results show that both the baseline YOLO11s and DEE-Net experience performance degradation under environmental perturbations. Specifically, the mAP@0.5 of YOLO11s decreases from 0.8032 to 0.5895, corresponding to an absolute drop of 21.37 percentage points. In comparison, DEE-Net decreases from 0.8747 to 0.7736, with an absolute drop of 10.11 percentage points. This smaller degradation suggests that DEE-Net maintains better detection performance under low-light and motion-blur conditions.
It should be noted that the challenge test set is generated through simulated environmental perturbations and therefore does not fully replace validation on independent real-world datasets. Nevertheless, these experiments provide additional evidence regarding the environmental robustness of the proposed method and complement the 5-fold cross-validation results reported in
Section 4.2.
Furthermore, we conducted supplementary evaluations of Faster R-CNN on the challenge test set, as shown in
Table 17. Compared with both Faster R-CNN and YOLO11s, DEE-Net achieves higher mAP@0.5 under the same degraded imaging conditions while maintaining real-time inference capability. These results suggest that the proposed method provides a more favorable balance between robustness and efficiency in simulated low-quality inspection scenarios.
4.6. Visualization Analysis
4.6.1. Visual Comparison of Detection Results
To provide a visual validation of the effectiveness of DEE-Net, this study conducted inference tests on aircraft skin surface images captured under complex backgrounds. DEE-Net demonstrates a superior continuous detection capability compared to the baseline, particularly for shallow surface scratches. It effectively avoids the fragmentation phenomenon where a single long scratch is erroneously identified as multiple disjointed targets.
As shown in the visualization comparison in
Figure 7, DEE-Net exhibits a significant technological generational gap in terms of robustness when handling extremely small targets. In the detection results of the baseline model, there is a prominent risk of missing detections for subtle scratches and early-stage micro-cracks. This is primarily due to the loss of pixel-level details during the traditional convolutional downsampling process, which leads to the “information annihilation” of critical defect features in deeper network layers.
In stark contrast, leveraging the introduced SPD-Conv lossless downsampling structure, DEE-Net can effectively preserve the spatial topological information of the original image. Even when faced with damage at extremely small scales that is difficult to discern with the naked eye, the model still achieves precise feature triggering and capture.
4.6.2. Heatmap-Based Interpretability Analysis
To further validate the feature extraction advantages of DEE-Net from the perspective of visual interpretability, this study employs heatmap analysis to compare the feature attention regions of the proposed model with those of the baseline algorithm, as illustrated in
Figure 8. The results indicate that the baseline model (YOLO11s) exhibits a diffuse feature response distribution when processing subtle skin defects. Due to its limited capability in capturing sub-pixel features, the baseline is highly susceptible to interference from metallic textures and complex lighting conditions. This leads to significant recognition biases and lower confidence levels, particularly for defects such as “Dent”.
In contrast, the activated regions in the DEE-Net heatmaps are highly consistent with the ground-truth physical contours of the defects. The model demonstrates a superior ability to penetrate background noise and precisely lock onto the core representations of “Paint-off” and “Dent.” This remarkable visual focusing capability provides strong empirical evidence for the high-frequency edge signal compensation provided by the ASE module, as well as the filtering efficacy of the DSM mechanism in suppressing irrelevant information across both spatial and frequency domains. Consequently, the model is able to accurately decouple discriminative features from complex industrial backgrounds, fundamentally underpinning its exceptional performance in detecting extremely fine defects such as “Crack” and “Scratch”.
5. Conclusions
Addressing the urgent requirements for structural integrity monitoring and flight safety assurance of aircraft in complex operating environments, this paper proposes DEE-Net, a high-efficiency algorithm for aircraft skin defect detection. The core contributions of this methodology are manifested through several integrated dimensions: first, the introduction of the SPD-Conv lossless feature reorganization mechanism utilizes space-to-depth conversion to reduce feature scales while preserving critical sub-pixel information, effectively addressing the information annihilation of micro-defects during the downsampling process at the source. Building upon this, the designed ASE module leverages high-frequency differential logic to compensate for the attenuation of subtle signals—such as cracks and scratches—inherent in traditional convolutions. Furthermore, the integration of the DSM enables the model to simultaneously focus on discriminative features in both the spatial and frequency domains, significantly suppressing interference from metallic reflections and complex background noise. To further refine the training process, the Wise-CIoU loss function, incorporating a non-monotonic focusing mechanism, was introduced to dynamically adjust gradient weights for samples of varying quality, thereby achieving a deep coupling of geometric constraints and sample quality awareness while markedly improving bounding box regression accuracy and convergence stability.
Comparative experiments show that DEE-Net achieves an mAP@0.5 of 87.47% on the ASDD dataset, representing an improvement over the YOLO11s baseline. The results suggest that the proposed method provides enhanced capability for detecting fine-grained defects. Heatmap analysis based on Grad-CAM further indicates that the model tends to focus on defect-related regions, supporting its effectiveness in capturing relevant features. In terms of efficiency, DEE-Net achieves an inference speed of 85.23 FPS while maintaining competitive detection performance. This suggests that the proposed method has the potential to meet real-time requirements in UAV-assisted inspection scenarios, providing a reasonable balance between accuracy and computational cost.
Despite the excellent performance of DEE-Net at this stage, its robustness under complex and variable climatic conditions—such as rain, fog, and intense glare—remains to be further validated. Consequently, subsequent research will focus on several critical avenues to enhance the practical utility of the model. Specifically, we aim to investigate cross-scene generalization capabilities by employing transfer learning and domain adaptation techniques to ensure detection consistency across different aircraft models and varying illumination environments. Furthermore, research into knowledge distillation and extreme lightweighting will be conducted to explore advanced model compression techniques, with the goal of achieving lower-latency deployment on power-constrained embedded inspection devices without compromising precision. In addition, the research focus will shift from qualitative recognition toward the deep quantitative assessment of defects, such as the precise measurement of defect length and area, to provide more valuable structural damage evaluation data for aircraft maintenance decision-making.
In conclusion, this study presents a deep learning-based approach for aircraft surface defect detection. The experimental results indicate that the proposed DEE-Net achieves improved detection performance while maintaining real-time inference capability. Within the scope of the current dataset and experimental settings, the proposed method shows potential for reducing manual inspection workload and supporting automated inspection processes in aircraft maintenance.