1. Introduction
In the context of the rapid development of intelligent transportation and industrial automation, mountain rail transportation systems have become essential infrastructure for mineral extraction, forestry production, and plateau logistics. However, prolonged operation under rugged terrain, harsh environments, and complex working conditions makes track bolts susceptible to defects such as loosening, bolt loss, and nut loss. If left undetected, these structural defects can compromise track stability and may lead to equipment damage, cargo loss, or even casualties. Therefore, accurate and efficient bolt defect detection is critical for ensuring the safety and reliability of mountain rail transportation systems.
At present, bolt inspection mainly relies on manual inspection, supplemented by sensing technologies such as ultrasonic testing, magnetic particle inspection, and infrared imaging. Although these methods are effective in specific applications, they generally suffer from low inspection efficiency, limited coverage, strong dependence on human expertise, and poor adaptability to complex environments, making them unsuitable for the real-time monitoring required in mountain rail systems. Furthermore, because these railways are typically deployed in remote mountainous regions with limited communication infrastructure, conventional cloud-based inspection frameworks are difficult to deploy efficiently.
With the advancement of computer vision and deep learning, object detection techniques based on convolutional neural networks have gradually replaced traditional methods that rely on handcrafted features, becoming a primary approach for addressing detection tasks in complex environments. From the R-CNN series to Fast R-CNN and Faster R-CNN, and subsequently to single-stage detectors such as YOLO and SSD [
1], both detection accuracy and inference speed have been continuously improved, thereby promoting the widespread application of vision-based detection technologies. In recent years, the YOLO series has demonstrated significant advantages in industrial defect inspection and traffic monitoring due to its end-to-end architecture, fast inference speed, and strong real-time performance. However, directly applying these models to mountainous rail scenarios remains challenging. On the one hand, bolts, as small-scale targets, are easily confused with background noise in complex environments. On the other hand, the lack of high-quality annotated datasets for bolt defects constrains model generalization and limits practical deployment performance. Therefore, constructing a representative bolt defect dataset tailored to real hilly and mountainous rail scenarios has become a fundamental prerequisite for advancing research and practical applications in this field [
2,
3].
Bolts are the most common fastening components in rail structures, and their integrity is directly related to track stability and operational safety. In mountain railway environments, tracks are often installed on steep slopes, sharp curves, and geologically complex terrain, where long-term cyclic vibration, rain and snow erosion, and substantial day–night temperature variations accelerate bolt loosening, corrosion, and nut loss. Because these defects usually develop locally and progressively, they are difficult to identify visually at an early stage. Conventional manual inspection is not only inefficient but also strongly influenced by inspectors’ subjective experience, making it difficult to meet the practical safety-monitoring requirements of mountain railways. Consequently, automated bolt defect detection has become an important research direction in rail engineering and intelligent sensing.
In early studies, rail and fastener defect detection primarily relied on various non-destructive testing techniques, including ultrasonic inspection, eddy current testing, and fiber Bragg grating sensing methods. Relevant studies have shown that ultrasonic inspection can evaluate bolt loosening by analyzing echo signal characteristics [
4]; eddy current testing exhibits high sensitivity in detecting surface cracks in metallic fasteners; and fiber Bragg grating-based structural health monitoring methods can capture strain variations at critical track locations. These methods demonstrate certain advantages in theoretical analysis and experimental validation; however, their detection performance is highly dependent on sensor deployment accuracy, environmental stability, and system maintenance conditions. In the actual operating environment of mountain railways, complex lighting conditions, strong noise interference, and variable climatic factors often lead to degraded detection performance. Simultaneously, the high costs associated with system construction and maintenance impose certain limitations on engineering implementation.
With the advancement of computer vision and pattern recognition technologies, image-based methods for detecting rail defects have gradually gained attention. Early visual inspection methods primarily relied on manually designed features, such as edge features, texture features, and morphological features. For instance, texture information extracted from rail components using a gray-level co-occurrence matrix can facilitate the identification of certain surface defects [
5]. Combining Canny edge detection with the Hough transform facilitates the localization and analysis of rail fastener structures [
6]. However, such artificial feature-based visual methods remain sensitive to lighting variations and background interference, exhibiting limited robustness under complex environmental conditions.
In recent years, breakthroughs in deep learning methods for object detection have opened new avenues for research in rail fastener defect detection. Deep learning models, exemplified by convolutional neural networks, can automatically learn discriminative features of targets through multiple layers of nonlinear mapping. They have achieved significant progress in tasks such as rail crack detection, fastener missing detection, and foreign object intrusion detection. From the perspective of detection frameworks, Faster R-CNN proposed by Girshick [
7] and Mask R-CNN proposed by He et al. [
8] demonstrate outstanding accuracy. Some researchers have applied these models to rail fastener detection for precise localization and instance segmentation [
9]. However, due to their complex model structures and relatively slow inference speeds, their application in real-time inspection scenarios faces certain limitations.
In contrast, the YOLO series of algorithms proposed by Redmon et al. [
10] is known for its end-to-end detection and efficient inference, making it more suitable for the real-time requirements of railway inspection. In research on bolt and related component detection, several studies have provided valuable insights. For example, a YOLOv3-based approach was used for railway fastener detection, where the introduction of a multiscale feature pyramid structure enhanced the ability to identify small targets [
11]. Subsequently, the integration of an attention mechanism into the YOLOv4 framework effectively strengthened the model’s focus on fine-grained bolt features [
12]. In domestic research, improvements were made to the loss function of YOLO-series models to address the issue of imbalanced sample distribution, thereby reducing the missed detection rate. Collectively, these studies offer useful references for the detection of rail bolt defects.
With the continuous evolution of the YOLO series, network architectures have been progressively optimized in terms of feature extraction efficiency and detection accuracy. As a new generation single-stage object detector, YOLOv11 further improves feature fusion, network module design, and inference efficiency, providing a more effective solution for detecting small and densely distributed objects in complex environments. However, its application to mountain rail scenarios remains challenging. First, track bolts are typically small and densely distributed, making them highly susceptible to interference from sleepers, rails, and surrounding structures. Second, the drastic variations in illumination, together with frequent shadows, reflections, and motion blur, place higher demands on model robustness and generalization. In addition, most existing studies primarily emphasize detection accuracy, whereas comprehensive evaluations of model lightweightness, real-time inference performance, and suitability for edge deployment remain limited.
In summary, although considerable progress has been made in track bolt and related defect detection, systematic studies targeting the complex operating conditions of mountain railways remain limited. In particular, further research is needed on dedicated dataset construction, task-specific model optimization, and validation in practical engineering applications. To address these challenges, this study investigates a track bolt defect detection method based on the YOLOv11 framework for hilly and mountainous railways. A dedicated dataset containing three representative defect categories, including missing bolts, loose bolts, and missing nuts, is constructed. Furthermore, network architecture optimization and structured pruning are incorporated to improve both detection performance in complex environments and model lightweightness, facilitating deployment on resource-constrained edge devices.
Based on the research background described above, this study proposes a lightweight detection model, termed LHFSE-YOLOv11, for track bolt defect detection in rail transporters operating in hilly and mountainous areas. The model integrates high-frequency detail enhancement, adaptive frequency-domain modulation, and structured pruning. Using YOLOv11m as the baseline, HFSE-YOLOv11, referring to High-Frequency and Spectral Enhanced YOLOv11, is first developed through coordinated improvements to the backbone and feature fusion network. Channel-level structured pruning is subsequently applied to further compress the model, yielding the final lightweight model LHFSE-YOLOv11, referring to Lightweight High-Frequency and Spectral Enhanced YOLOv11. The proposed method is designed to improve the recognition of small-scale bolt defects, local structural anomalies, and interference from complex backgrounds, while achieving a favorable balance among detection accuracy, overall processing speed, and model lightweightness. It thereby provides technical support for intelligent inspection and edge deployment of rail transporters in hilly and mountainous areas. The main contributions of this study are as follows:
(1) A high-resolution bolt defect image dataset was constructed for rail transportation scenarios in hilly and mountainous areas. The dataset contains three representative defect categories, namely missing bolts, loose bolts, and missing nuts, and includes samples collected under different weather conditions, illumination levels, imaging distances, viewing angles, and complex backgrounds. It provides a data foundation for model training, performance evaluation, and validation in engineering applications.
(2) To address the limited ability of the original YOLOv11m backbone to extract high-frequency details such as the edges, textures, and local structural anomalies of small-scale defects, the standard bottleneck structure in the C3K2 module was replaced with a High-Frequency Enhancement Residual Block, HFERB, to construct the HFERBC3K2 module. Through the coordinated operation of a local feature extraction branch and a high-frequency enhancement branch, this module improves the perception and preservation of fine-grained features associated with bolt defects.
(3) To reduce interference from complex vegetation, sleepers, gravel, shadows, and illumination variations, a Spectral Enhanced Feed-Forward Module, SEFF, was introduced into the C2PSA module of YOLOv11m to replace the original feed-forward network in the PSABlock, thereby constructing the SEFFNC2PSA module. This module integrates multiscale spatial modeling, adaptive frequency-domain modulation, and gated fusion to strengthen the representation of defect-related features while suppressing complex background interference. It is intended to improve detection performance under occlusion, blur, and complex texture conditions.
(4) The HFERBC3K2 and SEFFNC2PSA modules were jointly integrated into YOLOv11m to construct HFSE-YOLOv11, after which channel-level structured pruning was performed. By comparing detection accuracy, parameter count, computational cost, and model size under different pruning ratios, the final LHFSE-YOLOv11 model was obtained. While maintaining high detection accuracy, the model effectively reduces parameter scale and computational overhead, improves real-time inference performance, and provides a model foundation for subsequent deployment on edge devices.
3. Experiments
3.1. Dataset
To address the limited availability of bolt defect data for rail transporters operating in hilly and mountainous areas, image acquisition was conducted at a mountain rail transportation demonstration site in Songyang, Zhejiang Province, and a dedicated track bolt defect dataset was constructed. The acquisition route covered representative sections, including slopes, curves, and forest boundary areas. The collected images encompassed diverse weather conditions, illumination levels, viewing angles, and complex backgrounds. In total, 1448 original images with a resolution of 1920 × 1080 pixels were obtained.
The dataset consists of 1199 defect images and 249 normal bolt images. The defect images contain three representative defect categories: missing bolts, loose bolts, and missing nuts. The normal bolt images include structurally intact bolts under complex background conditions involving vegetation, soil, gravel, shadows, and reflections. These images were included to simulate practical continuous inspection conditions, in which normal bolt images account for a relatively high proportion of the acquired data.
The three defect categories were annotated with bounding boxes using CVAT [
23], and the annotations were cross-checked and reviewed by two annotators. Normal bolts were not treated as an independent detection category. Therefore, the corresponding YOLO label files for normal bolt images were left empty, and these images were used as complex background samples during model training and testing. Representative defect and normal bolt images are shown in
Figure 9.
As shown in
Figure 10, the investigated line is neither part of the national mainline railway network nor a dedicated material transport line for mining operations. Instead, it is a lightweight specialized rail transportation system designed to support agricultural production in hilly and mountainous areas, primarily transporting agricultural products, fertilizers, and farming equipment. The system has a maximum payload capacity exceeding 800 kg and can operate on gradients of up to 42°. The line consists of lightweight steel rails, supporting structures, and bolted connections, and includes straight sections, curved sections, and steep gradient sections.
Because the line operates outdoors over extended periods, the track and its fastening components are continuously exposed to transporter-induced vibration, rainfall erosion, and environmental contamination, making defects such as loose bolts, missing bolts, and missing nuts more likely to occur. In addition, dense vegetation surrounds the line, while soil, gravel, shadows, reflections, and occlusions are present in certain areas, further increasing the difficulty of defect feature extraction and localization.
3.2. Dataset Partitioning and Data Augmentation
To prevent data leakage caused by distributing an original image and its augmented derivatives across different subsets, the dataset was first partitioned at the level of the original images. The 1448 original images were divided into training, validation, and test sets at an approximate ratio of 7:2:1, containing 1011, 291, and 146 images, respectively. Each subset included both defect images and normal bolt images.
After dataset partitioning, offline data augmentation was applied only to training images with defect annotations. The augmentation methods included horizontal flipping, random cropping and scaling, small-angle rotation, brightness and contrast adjustment, and Gaussian noise perturbation. The random scaling factor ranged from 0.8 to 1.2, the rotation angle ranged from minus 15° to 15°, and the Gaussian noise had a mean of 0 and a standard deviation of 10. During augmentation, identical geometric transformations were applied to the images and their target bounding boxes, and any out-of-bounds or invalid annotations were removed.
Normal bolt images in the training set were assigned empty YOLO label files and were not subjected to offline augmentation. After reading the label file, the augmentation program directly skipped the corresponding image when the label content was empty and retained only the original image in the training set. Approximately four to five augmented samples were generated only for training images containing at least one annotated defect target. The validation and test sets were retained in their original form without any augmentation, as summarized in
Table 1.
This procedure increased the diversity of the defect samples while avoiding repeated expansion of normal background images, thereby preserving the authenticity and a reasonable proportion of negative samples. All augmented samples were generated exclusively from the training set, ensuring that no original image or its derived samples were distributed across different data subsets. Representative augmentation results are shown in
Figure 11.
3.3. Experimental Setup
The experiments were conducted on a platform running Ubuntu 20.04, with PyTorch 1.11.0 used as the deep learning framework and Python 3.8 as the programming language. The computing hardware included a 14-core Intel(R) Xeon(R) Gold 6330 CPU operating at 2.00 GHz and an NVIDIA GeForce RTX 3090 GPU with 24 GB of memory.
To improve reproducibility, the model version, training configuration, and experimental environment were explicitly recorded. All experiments were implemented using the Ultralytics YOLOv11 framework, with YOLOv11m selected as the baseline model and further modified by incorporating the HFERBC3K2 and SEFFNC2PSA modules. During training, the input image size was set to 640 × 640, the batch size to 16, and the number of training epochs to 200. SGD was used as the optimizer, with an initial learning rate of 0.01, a momentum of 0.9, a weight decay coefficient of 0.0005, and three warm-up epochs.
To ensure fair comparisons among different models under identical random conditions and to improve the reproducibility of each individual experiment, the random seed was fixed at 0 for all experiments. The same dataset split files, training parameters, and data augmentation settings were also used throughout. It should be noted that the fixed random seed was used only to control the experimental conditions. The statistical variability of the models has not yet been evaluated through independent repeated training with multiple random seeds. Therefore, all detection metrics reported in the tables correspond to single runs conducted under the fixed random setting. The detailed hyperparameter configuration is provided in
Table 2.
To ensure a fair comparison of inference efficiency across different models, the frame rate of each model was evaluated under identical software and hardware conditions. All tests were conducted using an NVIDIA GeForce RTX 3090 GPU and the PyTorch framework. The input image size was uniformly set to 640 × 640, and the batch size was set to 1. Inference was performed with FP32 precision, without test-time augmentation, TensorRT, or any other additional acceleration method. The overall processing frame rate of each model was measured on the same independent test set using identical confidence thresholds and nonmaximum suppression parameters.
The Ultralytics validation program was used to calculate the average processing time per image, including image preprocessing time, model inference time, and detection result postprocessing time. The overall processing frame rate was calculated as follows:
In the above equation, Tpre, Tinfer and Tpost denote the average preprocessing time, model inference time, and postprocessing time per image, respectively, all measured in milliseconds. The reported frame rate excludes the time required for image acquisition, disk input and output, result storage, and visualization. All FPS values in the subsequent comparative experiments were independently retested on the local experimental platform using the unified evaluation protocol described above.
3.4. Evaluation Metrics
Considering the characteristics of the object detection task, precision (P), recall (R), average precision (AP), and mean average precision (mAP) were selected as the primary metrics for evaluating model detection performance [
24,
25]. In addition, the number of parameters (Parameters), floating-point operations (FLOPs), and frames per second (FPS) were used to evaluate model complexity and inference efficiency. The metrics are defined as follows:
In the above equations, TP, or True Positive, denotes the number of positive instances correctly detected by the model; FP, or False Positive, denotes the number of negative instances incorrectly detected as positive; and FN, or False Negative, denotes the number of actual positive instances that are not correctly detected. Precision represents the proportion of correct positive predictions among all instances predicted as positive, whereas recall represents the proportion of correctly detected instances among all actual positive instances.
At different confidence thresholds, the corresponding precision and recall values can be obtained and used to plot the precision and recall curve, referred to as the PR curve. Average precision,
AP, represents the area under the PR curve over the recall interval [0, 1] and is mathematically defined as:
where
R denotes recall, and
Pinterp(
R) denotes the interpolated precision function, defined as follows:
This interpolation ensures that precision decreases monotonically as recall increases, thereby providing a more stable estimate of the area under the PR curve. In practice, discrete precision and recall points are obtained at different confidence thresholds, and AP is approximated using numerical integration. A higher AP indicates better overall detection performance across different recall levels.
For a detection task containing n object classes, mean average precision,
mAP, is defined as the arithmetic mean of the
AP values for all classes:
In the above equation, APi denotes the average precision for the ith object class, and n denotes the total number of object classes. In this study, mAP@0.5 and mAP@0.5:0.95 are primarily used to evaluate detection performance. Specifically, mAP@0.5 represents the mean AP across all classes at an Intersection over Union threshold of 0.5. mAP@0.5:0.95 represents the mean mAP calculated over multiple IoU thresholds ranging from 0.5 to 0.95 in increments of 0.05. Compared with mAP@0.5, mAP@0.5:0.95 imposes stricter requirements on object localization accuracy.
In addition, the parameter count is used to measure model size, FLOPs are used to quantify the computational cost of a single forward inference, and FPS is used to evaluate real-time inference speed. Lower parameter counts and FLOPs indicate greater suitability for deployment on resource-constrained devices, whereas a higher FPS indicates stronger real-time detection capability. Together, these metrics provide a comprehensive evaluation of model performance in terms of detection accuracy, model complexity, and inference efficiency.
3.5. Thermal Map Visualization for Bolt Defect Detection
To further investigate the feature attention of the proposed model during bolt defect detection, Grad-CAM++ [
26] was employed to visualize and compare the baseline YOLOv11m and the proposed LHFSE-YOLOv11. The visualization results are presented in
Figure 12. In the heatmaps, red and yellow regions indicate areas that contribute more strongly to the model predictions, whereas blue regions represent weakly activated areas. By comparing the activation distributions of the two models, the differences in their attention to defect regions and complex backgrounds can be intuitively analyzed.
As shown in
Figure 12, YOLOv11m is able to respond to bolt defects; however, its highly activated regions are relatively dispersed, with part of the activation extending to rail components, metallic surfaces, and surrounding vegetation. In the first sample, the baseline model responds to a relatively large area surrounding the defect, whereas LHFSE-YOLOv11 exhibits more concentrated activation around the missing circular component and its boundary, providing more complete coverage of the defect structure. In the second sample, which contains multiple defects and a complex background, the activation of YOLOv11m is relatively scattered and more susceptible to interference from vegetation and rail structures. In contrast, LHFSE-YOLOv11 produces more continuous activation over the primary defect regions while reducing responses to the surrounding background. In the third sample, which contains small-scale defects, the activation of the baseline model spreads along the metallic components, whereas the proposed model focuses more precisely on the defect location, with stronger attention to local edges and structural anomalies.
Overall, LHFSE-YOLOv11 exhibits greater spatial consistency between its highly activated regions and the actual defect locations across all three samples. The contrast between the responses of the target regions and the background is also more pronounced than that of the baseline model. These observations suggest that the joint integration of the HFERBC3K2 and SEFFNC2PSA modules enhances the representation of defect-related features, including edges, texture variations, and local structural anomalies, while reducing interference from vegetation and rail components. Although Grad-CAM++ provides qualitative rather than quantitative evidence, the visualization results offer intuitive support for the improved feature extraction capability and background suppression ability of the proposed model, thereby enhancing the interpretability of the detection results.
3.6. Ablation Experiments
All ablation models were trained using the same training set and evaluated consistently on the validation set. The validation results were used solely to assess the effectiveness of each module and determine the optimal network architecture, while the test set was not involved in model selection during the ablation study. To evaluate the effects of the HFERBC3K2 and SEFFNC2PSA modules on bolt defect detection in hilly and mountainous rail environments, the original YOLOv11m was adopted as the baseline model. The model incorporating only HFERBC3K2 was denoted as HFERB-YOLOv11, the model incorporating only SEFFNC2PSA was denoted as SEFF-YOLOv11, and the model integrating both modules was designated HFSE-YOLOv11. The individual contributions and combined effects of the two modules were analyzed by comparing detection accuracy, parameter count, computational cost, and overall processing frame rate, as reported in
Table 3.
As shown in
Table 3, the baseline YOLOv11m achieved a precision of 88.1%, a recall of 91.2%, an mAP@0.5 of 90.1%, and an mAP@0.5:0.95 of 77.1%. It contained 20.0 M parameters, required 67.7 GFLOPs, and reached an overall processing frame rate of 64.4 FPS. Although the baseline model already demonstrated a certain capability for bolt defect detection, further improvement was still needed in feature extraction and target localization under complex background conditions.
After introducing the HFERBC3K2 module, HFERB-YOLOv11 achieved a precision of 90.3%, a recall of 92.3%, an mAP@0.5 of 90.7%, and an mAP@0.5:0.95 of 78.5%, corresponding to improvements of 2.2, 1.1, 0.6, and 1.4 percentage points, respectively, over the baseline model. Meanwhile, the parameter count and computational cost were reduced to 18.9 M and 62.7 GFLOPs, respectively, and the overall processing frame rate increased to 65.2 FPS. These results indicate that HFERBC3K2 can reduce model complexity while improving the extraction of high-frequency details, including bolt edges, textures, and local structural anomalies.
After introducing the SEFFNC2PSA module, SEFF-YOLOv11 achieved a precision of 91.5%, a recall of 92.3%, an mAP@0.5 of 90.8%, and an mAP@0.5:0.95 of 79.4%, representing improvements of 3.4, 1.1, 0.7, and 2.3 percentage points, respectively, over the baseline model. Its parameter count and computational cost were 20.1 M and 67.8 GFLOPs, respectively, which were comparable to those of the baseline model, while the overall processing frame rate was 63.3 FPS. This suggests that the complete SEFFNC2PSA module improves overall detection performance through the combined effects of multiscale spatial modeling, adaptive frequency-domain modulation, and gated feature fusion, while reducing responses to complex backgrounds.
When HFERBC3K2 and SEFFNC2PSA were introduced simultaneously, HFSE-YOLOv11 achieved a precision of 91.7%, a recall of 93.4%, an mAP@0.5 of 91.3%, and an mAP@0.5:0.95 of 80.2%, yielding the best results across all four detection metrics in the ablation study. Compared with the baseline model, these metrics increased by 3.6, 2.2, 1.2, and 3.1 percentage points, respectively. The parameter count and computational cost were reduced to 19.0 M and 62.8 GFLOPs, respectively, although the overall processing frame rate decreased to 62.5 FPS. These results demonstrate that the two modules provide complementary benefits in high-frequency detail extraction, complex background suppression, and target localization, thereby improving detection accuracy while reducing both parameter count and computational cost.
It should be noted that the ablation study was conducted at the module level and was intended primarily to verify the effectiveness of the complete HFERBC3K2 and SEFFNC2PSA modules. Because SEFFNC2PSA jointly incorporates multiscale spatial convolution, FFT and IFFT operations, learnable frequency-domain modulation, and gated fusion, the current experiments cannot strictly isolate the independent contribution of each component. Therefore, the observed performance gains are attributed only to the combined effect of the complete SEFFNC2PSA module rather than to any individual frequency-domain operation.
3.7. Pruning Experiments and Lightweight Analysis of HFSE-YOLOv11
To further reduce the computational complexity and storage overhead of HFSE-YOLOv11 and improve its deployment efficiency on vehicle-mounted inspection terminals and edge devices, channel-level structured pruning experiments were conducted using the YOLO Pruning RKNN framework. Magnitude Pruner was adopted as the pruning method, with the L2 norm of convolutional channel weights used as the importance criterion. Redundant channels with relatively small weight magnitudes and limited contributions to the model output were preferentially removed. Meanwhile, the associated structures were adjusted according to the channel dependencies among network layers to preserve the integrity of the pruned network and ensure valid forward propagation. Models with different pruning ratios were independently fine-tuned on the training set, and the final pruning scheme was determined according to detection accuracy and model complexity on the validation set. The test set was not used for pruning ratio selection.
HFSE-YOLOv11, identified as the optimal model in the ablation study, was used as the pruning baseline. It contained 19.0 M parameters and required 62.8 GFLOPs. To investigate the effects of pruning intensity on model compression and detection performance, pruning ratios of 0.1, 0.4, and 0.5 were evaluated. All three experiments started independently from the same original optimal HFSE-YOLOv11 weights. The pruned model from one experiment was not used to initialize the next experiment; therefore, the procedure did not involve cumulative or progressive pruning.
After structured pruning, each model was independently fine-tuned using the same training configuration to allow the remaining parameters to adapt to the modified network structure and recover as much of the accuracy loss caused by channel removal as possible. No weight inheritance or accumulation of pruning results occurred across the three experiments, enabling a more objective assessment of the independent effect of each pruning ratio on model performance.
Because the detection head, attention modules, and several customized modules involve complex channel dependencies, directly pruning these components may cause feature dimension mismatches or structural damage. Therefore, modules without stable pruning support retained their original structures. In practice, pruning was applied primarily to standard convolutional layers and other prunable modules that could be correctly identified and processed by the framework. The specified pruning ratio denotes the target compression ratio of the prunable channels and does not correspond directly to the actual reduction in the total parameter count or computational cost of the complete model.
To ensure comparability among the pruning experiments, all three models were fine-tuned using the same strategy. The input image size was set to 640 × 640, the batch size to 16, and the maximum number of fine-tuning epochs to 120. Automatic mixed-precision training and a cosine annealing learning rate schedule were enabled, together with an early stopping patience of 15 epochs. The lightweight performance and detection capability of the pruned models were comprehensively evaluated using parameter count, GFLOPs, model size, mAP@0.5, and mAP@0.5:0.95. The structured pruning configuration is summarized in
Table 4.
As shown in
Table 5, the parameter counts, computational cost, and model size decreased continuously as the pruning ratio increased, demonstrating that structured pruning can effectively remove redundant channels from HFSE-YOLOv11. Before pruning, HFSE-YOLOv11 contained 19.0 M parameters, required 62.8 GFLOPs, and occupied 36.8 MB, with an mAP@0.5 of 91.3% and an mAP@0.5:0.95 of 80.2%.
At a pruning ratio of 0.1, the parameter counts, computational cost, and model size were reduced to 17.9 M, 60.2 GFLOPs, and 36.1 MB, corresponding to reductions of 5.8%, 4.1%, and 1.9%, respectively, relative to the unpruned model. The resulting mAP@0.5 and mAP@0.5:0.95 were 91.1% and 80.1%, representing decreases of only 0.2 and 0.1 percentage points. These results indicate that a low pruning ratio can achieve preliminary model compression while largely preserving detection accuracy.
When the pruning ratio was increased to 0.4, the parameter counts, computational cost, and model size decreased to 16.2 M, 54.3 GFLOPs, and 33.3 MB, corresponding to reductions of 14.7%, 13.5%, and 9.5%, respectively. The mAP@0.5 and mAP@0.5:0.95 were 91.1% and 79.8%, decreasing by 0.2 and 0.4 percentage points relative to the unpruned model. These results show that moderate pruning can substantially reduce model complexity while maintaining relatively stable detection performance.
At a pruning ratio of 0.5, the resulting LHFSE-YOLOv11 further reduced the parameter count, computational cost, and model size to 15.9 M, 53.8 GFLOPs, and 32.5 MB, corresponding to reductions of 16.3%, 14.3%, and 11.7%, respectively, compared with HFSE-YOLOv11. Its mAP@0.5 and mAP@0.5:0.95 were 91.0% and 79.3%, representing decreases of 0.3 and 0.9 percentage points. Although detection accuracy declined slightly, the model retained strong overall detection performance and achieved the greatest compression among the three pruning schemes.
Overall, the model pruned at a ratio of 0.4 achieved a favorable balance between detection accuracy and model complexity, whereas the model pruned at a ratio of 0.5 had fewer parameters, lower computational cost, and a smaller model file. Considering the deployment requirements of vehicle-mounted inspection terminals and resource-constrained edge devices, the model with a pruning ratio of 0.5 was ultimately selected as the lightweight deployment model and designated LHFSE-YOLOv11.
3.8. Comparative Experiments
To comprehensively evaluate the performance of LHFSE-YOLOv11 for bolt defect detection in rail transporters operating in hilly and mountainous areas, five representative detection models, namely YOLOv8m, SSD, YOLOv10m, YOLOv8s, and Faster R-CNN, were selected for comparison. These models represent both single-stage and two-stage object detection frameworks. All models were trained and evaluated using the same dataset partition under identical hardware and software environments. To ensure a fair comparison, the input image size, data augmentation strategy, and training configuration were kept as consistent as possible across all models. The evaluation metrics included precision (P), recall (R), mAP@0.5, mAP@0.5:0.95, parameter count, computational cost, and overall processing frame rate. After the network architecture and pruning strategy had been finalized, the configurations of all models were fixed, and their final performance was evaluated on the independent test set. The comparison results are presented in
Table 6.
As shown in
Table 6, LHFSE-YOLOv11 achieved a precision of 91.7%, a recall of 93.4%, an mAP@0.5 of 91.0%, and an mAP@0.5:0.95 of 79.3%, obtaining the best results across all four detection metrics among the compared models. These results indicate that the proposed model can identify more true defect targets, reduce the probability of missed detections, and maintain strong localization performance across different intersection over union thresholds.
YOLOv8m achieved a recall of 83.2% and an mAP@0.5 of 84.4%, indicating a certain risk of missed detections under complex background conditions. SSD required 287 GFLOPs but achieved an mAP@0.5:0.95 of only 62.1%, resulting in relatively low detection accuracy and inference efficiency. YOLOv8s had the fewest parameters and the lowest computational cost, reaching an overall processing frame rate of 93.6 FPS, but its precision and both mAP metrics were substantially lower than those of the proposed model. Although Faster R-CNN achieved relatively high precision, it contained 41.3 M parameters, required 178 GFLOPs, and reached an overall processing frame rate of only 32.5 FPS.
Compared with YOLOv10m, which had a similar parameter count, LHFSE-YOLOv11 improved recall, mAP@0.5, and mAP@0.5:0.95 by 1.7, 1.2, and 1.3 percentage points, respectively. It also reduced the parameter count and computational cost by 0.4 M and 5.1 GFLOPs, respectively, while achieving a higher overall processing frame rate. Overall, LHFSE-YOLOv11 maintained relatively low parameter and computational requirements while achieving high detection accuracy and an overall processing frame rate of 81.5 FPS, providing a favorable balance among detection performance, model complexity, and real-time capability.
To further analyze the detection performance of different models in practical hilly and mountainous rail environments, defect images containing complex vegetation backgrounds, partial occlusion, and illumination variations were selected for visual comparison, as shown in
Figure 13. YOLOv8m was susceptible to interference from vegetation and rail structures, resulting in false detections. SSD showed relatively weak localization capability for defect regions. YOLOv10m produced relatively low detection confidence for small targets and defects with indistinct features. Although YOLOv8s and Faster R-CNN were able to localize the defect regions, their predicted bounding boxes still showed deviations from the actual defect boundaries.
LHFSE-YOLOv11 accurately identifies defect targets under complex background conditions, produces bounding boxes that closely align with the actual target regions, and maintains high detection confidence. The HFERBC3K2 module improves the extraction of high-frequency details, including bolt edges, textures, and local structural anomalies. Through the combined effects of multiscale spatial modeling, adaptive frequency-domain modulation, and gated fusion, the SEFFNC2PSA module enhances the representation of defect-related features in complex backgrounds. Structured pruning further removes redundant channels, thereby reducing computational and storage overhead while largely preserving detection capability.
In summary, both the quantitative evaluation and visualization results demonstrate that LHFSE-YOLOv11 provides strong defect recognition capability, accurate target localization, and favorable real-time performance in complex hilly and mountainous rail environments. These findings validate the effectiveness of the integrated design combining HFERBC3K2, SEFFNC2PSA, and structured pruning.
3.9. Discussion
The ablation experiments in this study primarily evaluated the effectiveness of HFERBC3K2 and SEFFNC2PSA at the complete module level. The individual contributions of the multiscale spatial convolution, adaptive frequency-domain modulation, and gated fusion within SEFF were not further isolated. Therefore, the current findings mainly reflect the combined benefit of the complete SEFFNC2PSA module, and the observed performance improvement cannot be attributed entirely to a single frequency-domain operation. Future work will construct a purely spatial control module with matched parameter count and computational cost and will separately investigate the effects of FFT and IFFT operations, learnable frequency-domain weights, and gated fusion on detection performance.
All models were trained using a fixed random seed to ensure reproducibility and comparability under consistent experimental conditions. However, independent repeated training with multiple random seeds has not yet been conducted. Consequently, the current results primarily represent the relative performance of the models under a fixed experimental setting and do not fully quantify the variations caused by random initialization, training order, and stochastic data augmentation. Future studies will perform repeated experiments using multiple random seeds and report the mean, standard deviation, and confidence interval to further assess the statistical stability of model performance.
Although complete field deployment is beyond the scope of this study, the proposed lightweight LHFSE-YOLOv11 model was designed for edge deployment in rail transportation systems operating in hilly and mountainous areas. In practical applications, an industrial camera could be installed at the front of an inspection vehicle to continuously acquire images of track fastening components. The captured images could then be processed directly by an embedded computing platform without relying on network communication. Owing to its reduced parameter count and computational complexity, the proposed model has the potential to satisfy real-time detection requirements on resource-constrained edge devices. Future work will integrate the model into a practical inspection platform and evaluate its performance under different vehicle speeds, image acquisition frame rates, and inference latency conditions.
4. Conclusions
To address the challenges of small bolt defect targets, severe interference from complex backgrounds, and limited deployment capability on edge devices in rail transporters operating in hilly and mountainous areas, this study proposed a lightweight detection model named LHFSE-YOLOv11 and constructed an image dataset containing three representative defect categories: missing bolts, loose bolts, and missing nuts. Based on YOLOv11m, the standard bottleneck structure in C3K2 was replaced with HFERB to construct the HFERBC3K2 module. Meanwhile, SEFF was introduced into C2PSA to construct the SEFFNC2PSA module, with the aim of enhancing the extraction of high-frequency defect details and suppressing interference from complex backgrounds.
The ablation results showed that, after jointly introducing HFERBC3K2 and SEFFNC2PSA, HFSE-YOLOv11 achieved a precision of 91.7%, a recall of 93.4%, an mAP@0.5 of 91.3%, and an mAP@0.5:0.95 of 80.2%. Compared with the original YOLOv11m, these four metrics increased by 3.6, 2.2, 1.2, and 3.1 percentage points, respectively. These results indicate that the two complete modules provide complementary benefits in high-frequency detail extraction, background interference suppression, and target localization.
Channel-level structured pruning was then applied to HFSE-YOLOv11, and the model with a pruning ratio of 0.5 was ultimately selected and designated LHFSE-YOLOv11. Its parameter counts, computational cost, and model size were reduced to 15.9 M, 53.8 GFLOPs, and 32.5 MB, corresponding to reductions of 16.3%, 14.3%, and 11.7%, respectively, relative to the model before pruning. Its mAP@0.5 and mAP@0.5:0.95 were 91.0% and 79.3%, representing decreases of only 0.3 and 0.9 percentage points. These results demonstrate that structured pruning can effectively reduce model complexity while largely preserving detection performance.
In the comparative experiments, LHFSE-YOLOv11 achieved a precision of 91.7%, a recall of 93.4%, an mAP@0.5 of 91.0%, and an mAP@0.5:0.95 of 79.3%, obtaining the best results across all four detection metrics among the compared models. Its overall processing frame rate reached 81.5 FPS. The heatmaps and detection results showed that the model could more accurately attend to and localize bolt defects under conditions involving vegetation occlusion, overlapping rail structures, and complex backgrounds. Overall, LHFSE-YOLOv11 achieves a favorable balance among detection accuracy, model lightweightness, and real-time performance, indicating its potential for deployment on vehicle-mounted inspection terminals and resource-constrained edge devices.
It should be noted that this study has not yet quantitatively isolated the individual contributions of frequency-domain modulation, spatial convolution, and gated fusion within SEFFNC2PSA. In addition, the current results are based on a single training run with a fixed random seed, and the statistical variability of the model has not yet been quantified through independent repeated experiments using multiple random seeds. Future work will therefore construct component-level control experiments with matched parameter counts and computational costs and conduct repeated training with multiple random seeds to further investigate the contributions of the internal components and the statistical stability of the experimental results. The dataset will also be expanded to include more seasons, weather conditions, track sections, and defect types. Moreover, model quantization, TensorRT acceleration, and adaptation to heterogeneous hardware will be investigated to further evaluate the operational performance of the proposed model on practical edge detection platforms.