FF-DEIM: DEIM with Image Dehazing and Self-Supervised Pretraining for Catenary Support Component Detection
Abstract
1. Introduction
- (1)
- A dual-channel fusion network (DCFNet) image dehazing method is proposed, which significantly improves the quality of input images by removing foreground interferences (such as fog, raindrops, and dynamic blur).
- (2)
- A pretraining method based on contrastive learning for mask image modeling (MIMCL) is designed, which enhances the model’s ability to focus on key regions of the catenary network components, optimizing the feature extraction ability of the backbone and improving overall detection performance and model convergence speed.
- (3)
- A feature-focusing pyramid network (FFPN) module is proposed, which uses the focus feature module to fuse cross-layer contextual features, enhancing the ability to capture local details and improving the detection performance of multi-scale components.
- (4)
- A UAV-based catenary network image dataset is constructed, which includes images with fog, rain, and low light conditions. Various experiments demonstrate the effectiveness of the proposed method.
2. The Proposed Method
2.1. Guided Image Dehazing with Dual-Color Feature Fusion
2.2. Masked Image Modeling with Contrastive Learning
2.3. DETR with Feature-Focused Pyramids and Matching-Aware Design
2.4. Learning Strategy
2.4.1. Loss of DCF-NET
2.4.2. Loss of MIMCL
2.4.3. Loss of DEIM
3. Experiments and Result Analysis
3.1. Experimental Settings
3.2. Introduction of Dataset
3.3. Evaluation Indicator
3.4. Detection Performance Metrics
3.5. Experimental Analysis
- (1)
- The Analysis of Image Dehazing Methods: To evaluate the effectiveness of the proposed DCFNET in image dehazing tasks, comparative experiments were conducted using images resized to 640 × 512 for both training and testing. Two image quality metrics—PSNR and SSIM—were adopted for comprehensive evaluation, where higher values of PSNR and SSIM indicate better image quality. As presented in Table 1, DCFNET achieves the best performance across all quality metrics (PSNR: 22.56, SSIM: 0.826), significantly outperforming other methods. Furthermore, DCFNET maintains a lightweight computational footprint (13.32M parameters and 53.40G FLOPs), demonstrating an optimal balance between visual quality and efficiency.
- (2)
- The Analysis of Ablation Study Analysis: To further validate the effectiveness of the proposed components, we perform ablation studies on MIMCL, FFPN, and Dense O2O. Notably, all ablation experiments are conducted on dehazed images to ensure consistency with the full detection pipeline. The corresponding results are summarized in Table 2. The FPS is measured on input images with a resolution of 640 × 640. An analysis of the results shown in the ablation table indicates the following findings: (1) The integration of MIMCL, FFPN, and Dense O2O each contributes to noticeable improvements in detection accuracy. (2) Specifically, MIMCL alone improves mAP50–95 from 0.652 to 0.703, while FFPN and Dense O2O individually raise it to 0.671 and 0.683, respectively. When combining MIMCL with FFPN or Dense O2O, the mAP50–95 further increases to 0.734 and 0.727. Upon enabling all modules, the best performance is achieved with an mAP50–95 of 0.755, an mAP50 of 0.992, and a mAR of 0.991. (3) In terms of inference speed, the model maintains real-time performance, achieving 20.9 FPS even when all modules are enabled. The additional computational cost introduced by each component is minimal. These results demonstrate that the proposed modules are complementary and effective, striking a favorable balance between detection accuracy and speed.
- (3)
- The Analysis of Self-Supervised Learning: It is also important to note that all subsequent experiments are conducted on dehazed images to ensure consistency with real-world UAV inspection scenarios. The evaluation results of the network utilizing different pre-trained weights from various self-supervised learning methods are presented in Table 3. Clearly, our proposed method, MIMCL, consistently outperforms comparative self-supervised methods across multiple critical metrics. Specifically, MIMCL achieves the highest mAP of 75.53%, AP50 of 99.27%, and of 80.23%. In comparison, other self-supervised methods such as MAE, IGPT, SimMIM, BEIT, and DINO achieve relatively lower performance. For example, MAE achieves a mAP of 67.91%, SimMIM reaches 72.18%, and BEIT attains 74.31%. Particularly notable is MIMCL’s significant improvement in mAP and AP75, surpassing baseline methods by approximately 1.22–7.62%. These results demonstrate that our proposed MIMCL framework effectively leverages self-supervised strategies, leading to enhanced feature representation capabilities and improved detection performance.
- (4)
- The Analysis of Multi-Scale Feature Fusion Module: By analyzing the results in Table 4, we observe the following findings: (1) Our proposed metised hod, FFPN, achieves the best overall detection performance, with an mAP of 67.43%, an AP50 of 95.54%, an AP75 of 75.26%, and an APs of 65.88%, demonstrating its strong ability to handle both small and large object detection tasks. (2) The CFP method performs well, reaching the highest performance on AP75 and APl, with values of 73.96% and 83.95%, respectively, benefiting from its cross-feature pyramid design that helps preserve local details in high-level features. (3) CR-FPN and Bi-FPN show comparable results, but FFPN slightly outperforms them, especially in mAP and AP75, highlighting the effectiveness of FFPN in aggregating multi-level features and improving robustness in complex environments. At the core of FFPN is the focus feature module, which introduces a more focused and detailed approach to multi-scale feature handling. By refining features at multiple scales and capturing spatial patterns at different granularities, the focus feature module enhances the model’s ability to handle complex, multi-scale objects and improve overall detection performance. This focus on detailed feature refinement makes FFPN particularly effective in scenarios where object size and scale vary significantly, outperforming other FPN variations.
- (5)
- The Analysis of Classic Detection Methods: Following the FPN experiments, we further evaluated the performance of our method on object detection tasks. Specifically, we conducted comparative experiments with several classical object detection approaches, as shown in Table 5. Both training and testing images were resized to 640 × 640. The comparison experiments were structured as follows: (1) For one-stage detectors, YOLOv11-l, YOLOv12-l, and ATSS served as baseline methods. (2) For two-stage detectors, Cascade R-CNN and Dynamic R-CNN were selected as comparison methods. (3) Among state-of-the-art DETR-based approaches, Align-DETR, DDQ (DETR), DINO, HDINO, and RT-DETR were also included in the evaluations. As demonstrated in Table 5, the proposed method achieves an excellent balance between detection accuracy and inference speed, highlighting its advantages as follows: (1) Our proposed method achieves optimal performance across multiple detection metrics, obtaining mAP50–95 = 75.53%, AP50 = 99.27%, AP75 = 80.23%, APs = 65.71%, and APl = 85.34%, while maintaining a competitive inference speed (20.9 FPS). (2) Notably, in small-object detection, our method achieves APs = 65.71%, clearly outperforming representative models such as YOLOv11-L (63.53%), YOLOv12-L (63.98%), and RT-DETR (62.01%), demonstrating superior effectiveness in challenging small-scale scenarios. (3) Although YOLO-series methods (YOLOv11-L and YOLOv12-L) achieve higher inference speeds (32.6 FPS and 30.7 FPS, respectively), they show significantly lower overall detection accuracy and small-object performance compared to our method. (4) When compared with RT-DETR, our method not only increases mAP50–95 by approximately 10.26% but also significantly enhances small-object detection accuracy at similar inference speeds, further underscoring the efficiency of our designed modules. In summary, the proposed method effectively balances detection accuracy, especially for small-scale objects, and inference speed, providing clear advantages over both YOLO-series and DETR models.
- (6)
- Visualization of Detection Results: To further verify the practical effectiveness of the proposed method, object detection experiments were conducted on the dehazed images. As shown in Figure 10, our method is capable of accurately detecting components with significant scale variations, such as insulators (within insulator groups) and screws (support sleeve screws). This clearly demonstrates the adaptability and robustness of our method in addressing complex backgrounds and multi-scale challenges. To further validate the superiority of the proposed method, we conducted additional visualization experiments using six different detection methods. As shown in Figure 10, the letters (a) to (f) correspond to the detection results of the ATSS, DDQ, Dynamic R-CNN, YOLOv12, RT-DETR, and FF-DEIM detectors on the same image. In this experiment, we selected the detector with the best overall performance from various methods. Red arrows in the figure indicate the areas where catenary components were missing by the detectors. Based on the results, we can make the following observations: (1) the ATSS detector (a) failed to detect multiple components, including four sleeve screws, a load-bearing cable base, a sleeve double ear, and an isoelectric line; (2) the DDQ (b) and Dynamic R-CNN (c) detectors exhibited similar performance, each missing four sleeve screws and one isoelectric line; (3) RT-DETR (e) also showed some deficiency, with three sleeve screws left undetected; and (4) in contrast, both YOLOv12 (d) and FF-DEIM (f) demonstrated complete detection coverage, successfully identifying all catenary support components without any omissions. As shown in Figure 11, the proposed method achieves accurate detection of components with significant scale variations, such as insulators within insulator groups and support sleeve screws. The results demonstrate the robustness of the method under complex backgrounds and multi-scale scenarios.
4. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Duan, F.; Wang, H.; Yang, H.; Wei, C.; Zhang, C.; Song, Y.; Liu, Z. MRFM-IFCOS: An Anchor-Free Interactive Detector Based on Multireceptive Field Mamba for Detecting Catenary Support Components. IEEE Trans. Instrum. Meas. 2025, 74, 2553415. [Google Scholar] [CrossRef]
- Yang, H.; Liu, Z.; Liu, W.; Wang, H.; Zhang, Y.; Wang, H. Graph-MDETR: A Graph-Guided Mamba-DETR Network for UAV Catenary Support Components Detection in Electrified Railways. IEEE Trans. Intell. Transp. Syst. 2026, 27, 6319–6332. [Google Scholar] [CrossRef]
- Song, Y.; Liu, Z.; Ouyang, H.; Wang, H.; Lu, X. Sliding mode control with PD sliding surface for high-speed railway pantograph-catenary contact force under strong stochastic wind field. Shock. Vib. 2017, 4, 4895321. [Google Scholar]
- Yan, J.; Zhou, N.; Cheng, Y.; Zhang, F.; Wang, H.; Wang, M.; Zhang, W. Application of machine-vision-driven physics-informed neural networks in pantograph–catenary system state detection. Mech. Syst. Signal Process. 2026, 257, 114577. [Google Scholar] [CrossRef]
- Zhang, M.; Ma, L.; Wu, Y.; Shen, K.; Sun, Y.; Leung, H. High-Traversability and Precise Navigation for Mobile Robots in Constrained Environments. IEEE Sens. J. 2025, 25, 22815–22826. [Google Scholar] [CrossRef]
- Zhang, M.; Ma, L.; Wu, Y.; Shen, K.; Huang, D.; Leung, H. Tackling the Kidnapped Robot Problem via Sparse Feasible Hypothesis Sampling and Reliable Batched Multistage Inference. IEEE Trans. Instrum. Meas. 2026, 75, 7504614. [Google Scholar] [CrossRef]
- Shajeena, J.; Govindasamy, B.; Gnanasundaram, M.; Robinson Joel, M. Mobile-Le Harmonic Fusion Network for Object Recognition and Siam MoT Based Multi-Object Tracking Using Video Surveillance. Cybern. Syst. 2025, 57, 866–896. [Google Scholar] [CrossRef]
- Yang, H.; Hu, K.; Wang, H.; Hong, W.; Wang, X.; Wang, H.; Liu, Z. BCLIP-ADer: A Bayesian Prompt Contrastive Language-Image Pretraining Method for Catenary Component Anomaly Detection in Electrified Railways. IEEE Trans. Transp. Electrif. 2026, 1. [Google Scholar] [CrossRef]
- Yang, G.; Jiang, Y.; Wang, S.; Chen, K. VinsFusion-Line: Binocular Vision Inertial Navigation Real-Time SLAM System Based on Line Features. Cybern. Syst. 2025, 57, 350–375. [Google Scholar] [CrossRef]
- Song, Y.; Rønnquist, A.; Jiang, T.; Nåvik, P. Railway pantograph-catenary interaction performance in an overlap section: Modeling, validation and analysis. J. Sound Vib. 2023, 548, 117506. [Google Scholar] [CrossRef]
- Sreekala, K.; Maniraj, S.P.; Singh, A.; Pratap Singh, A.; Pyingkodi, M.; Inthiyaz, S. Enhancing Medical Diagnosis through Multimodal Image Fusion: A Novel Approach Using Modified Swin-Based Cross Attention Fusion. Cybern. Syst. 2025, 57, 765–807. [Google Scholar] [CrossRef]
- Reddy, K.; Sekhar, M.; Nelakuditi, U.R. Illustration of Image Registration-Based Novel Segmentation and Classification Model Using Multimodal Medical Images with 3D-TRRSegnet and Adaptive RAN. Cybern. Syst. 2025, 1–35. [Google Scholar] [CrossRef]
- Han, Y.; Liu, Z.; Han, Z.; Yang, H.M. Fracture detection of ear pieces of catenary support devices of high-speed railway based on SIFT feature matching. J. China Railw. Soc. 2014, 36, 31–36. [Google Scholar]
- Han, Y.; Liu, Z.; Lyu, Y.; Liu, K.; Li, C.; Zhang, W. Deep learning-based visual ensemble method for high-speed railway catenary clevis fracture detection. Neurocomputing 2020, 396, 556–568. [Google Scholar] [CrossRef]
- Yang, H.M.; Liu, Z.G.; Han, Z.W.; Han, Y. Foreign body detection between insulator pieces in electrified railway based on affine moment invariant. J. China Rail-Way Soc. 2013, 35, 30–36. [Google Scholar]
- Yang, H.; Liu, Z.; Han, Y.; Han, Z. Defective Condition detection of insulators in electrified railway based on feature matching of speeded-up robust features. Power Syst. Technol. 2013, 37, 2297–2302. [Google Scholar]
- Zhong, J.; Liu, Z.; Zhang, G.; Han, Z. Condition detection of swivel clevis pins in overhead contact system of high-speed railway. J. China Railw. Soc. 2017, 39, 65–71. [Google Scholar]
- Feng, Z.; Peng, L.; Kang, D.; Zhou, M.; Kuang, P.; Wu, M.; Su, J. TIPS: Two-level prompt selection for more stability-plasticity balance in continual learning. Pattern Recognit. 2025, 171, 112276. [Google Scholar]
- Zhang, H.; Chang, H.; Ma, B.; Wang, N.; Chen, X. Dynamic R-CNN: Towards high quality object detection via dynamic training. In Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, Scotland, 23–28 August 2020; pp. 260–275. [Google Scholar]
- Tian, X.; Xianyu, X.; Li, Z.; Chen, J.; Zhang, Y. Infrared and visible image fusion based on multi-level detail enhancement and generative adversarial network. Intell. Robot. 2024, 4, 524–543. [Google Scholar] [CrossRef]
- Liu, Z.; Liu, K.; Zhong, J.; Han, Z.; Zhang, W. A high-precision positioning approach for catenary support components with multiscale difference. IEEE Trans. Instrum. Meas. 2019, 69, 700–711. [Google Scholar]
- Liu, Z.; Lyu, Y.; Wang, L.; Han, Z. Detection approach based on an improved faster RCNN for brace sleeve screws in high-speed railways. IEEE Trans. Instrum. Meas. 2019, 69, 4395–4403. [Google Scholar]
- Hu, X.; Yang, J.; Jiang, F.; Hussain, A.; Dashtipour, K.; Gogate, M. Steel surface defect detection based on self-supervised contrastive representation learning with matching metric. Appl. Soft Comput. 2023, 145, 110578. [Google Scholar] [CrossRef]
- Ma, X.; Wu, Z.; Pan, J.; Zheng, K.; Lian, R.; Wu, W.; Zhang, W. CDMask: Change Customized Mask Architecture for Change Detection. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5619919. [Google Scholar] [CrossRef]
- Meng, C.; Huang, G.; Fu, R.; Jian, R.; Gan, Z.; Ouyang, C. CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 5–7 June 2026; pp. 1606–1615. [Google Scholar]
- Fattal, R. Dehazing using color-lines. ACM Trans. Graph. TOG 2014, 34, 1–14. [Google Scholar] [CrossRef]
- Saihood, R.T. Aerial Image Enhancement based on YCbCr Color Space. Int. J. Intell. Eng. Syst. 2021, 14, 177. [Google Scholar] [CrossRef]
- Hemalatha, S.; Acharya, U.D.; Renuka, A. Comparison of secure and high capacity color image steganography techniques in RGB and YCbCr domains. arXiv 2013, arXiv:1307.3026. [Google Scholar]
- Tufail, Z.; Khurshid, K.; Salman, A.; Nizami, I.F.; Jeon, B. Improved dark channel prior for image defogging using RGB and YCbCr color space. IEEE Access 2018, 6, 32576–32587. [Google Scholar] [CrossRef]
- Ki, S.; Sim, H.; Choi, J.S.; Kim, S.Y.; Seo, S.; Kim, S.; Kim, M. Fully end-to-end learning based conditional boundary equilibrium gan with receptive field sizes enlarged for single ultra-high resolution image dehazing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Salt Lake City, UT, USA, 18–22 June 2018; pp. 817–824. [Google Scholar]
- Belmer Gladson, V.; Kumar, S.; Rajesh, M.; Karthik, R. Image Steganography with Security Using Massive Threefold Attentional Residual GAN Optimized By Chaotic PSO Algorithm. Cybern. Syst. 2025, 1–32. [Google Scholar] [CrossRef]
- Feng, Z.; Zhou, M.; Gao, Z.; Stefanidis, A.; Su, J.; Dang, K.; Li, C. Adaptive knowledge transfer for class incremental learning. Pattern Recognit. Lett. 2024, 183, 165–171. [Google Scholar] [CrossRef]
- Liu, Z.; Zhao, W.; Jia, N.; Liu, X.; Yang, J. SANet: Scale-adaptive network for lightweight salient object detection. Intell. Robot. 2024, 4, 503–523. [Google Scholar] [CrossRef]
- Yang, C.; Yang, S.; He, Y.; Fan, L.; Gao, X.; Tang, M.; Sun, J. Research on defect detection performance of silicon carbide wafer surface based on ESN-YOLOv8 algorithm. Comput. Mater. Sci. 2026, 268, 114656. [Google Scholar] [CrossRef]
- Fan, L.; He, Y.; Mo, Y.; Cao, Y. An interpretable ensemble machine learning model for predicting carbon dioxide adsorption on magnesium oxide-based sorbents. Environ. Res. 2026, 297, 124126. [Google Scholar] [CrossRef] [PubMed]
- Chen, J.; Shao, Z.; Zhu, H.; Chen, Y.; Li, Y.; Zeng, Z.; Yang, Y.; Wu, J.; Hu, B. Sustainable interior design: A new approach to intelligent design and automated manufacturing based on Grasshopper. Comput. Ind. Eng. 2023, 183, 109509. [Google Scholar] [CrossRef]
- Wu, X.; Dong, J.; Bao, W.; Zou, B.; Wang, L.; Wang, H. Augmented intelligence of things for emergency vehicle secure trajectory prediction and task offloading. IEEE Internet Things J. 2024, 11, 36030–36043. [Google Scholar] [CrossRef]
- Zhang, J.; Song, X.; Li, Y.; Liang, D.; Zhang, Z.; Cai, J. Adaptive dual cross-attention network for multispectral object detection in autonomous driving. Expert Syst. Appl. 2026, 318, 132012. [Google Scholar] [CrossRef]
- Zhang, J.; Xiang, M.; Hu, Y.; Hao, W.; Lei, L.; Yi, K. Multivariate feature learning and associative spatial information enhancement for snow object detection in autonomous driving. Eng. Appl. Artif. Intell. 2026, 175, 114672. [Google Scholar] [CrossRef]
- Jiao, R.; Zhang, J.; Li, C.; Hu, L. Large-kernel spatially parallel feature fusion for monocular 3D perception in autonomous driving. Knowl.-Based Syst. 2026, 343, 115998. [Google Scholar] [CrossRef]
- Khosla, P.; Teterwak, P.; Wang, C.; Sarna, A.; Tian, Y.; Isola, P.; Maschinot, A.; Liu, C.; Krishnan, D. Supervised contrastive learning. Adv. Neural Inf. Process. Syst. 2020, 33, 18661–18673. [Google Scholar] [CrossRef] [PubMed]
- Xie, Z.; Zhang, Z.; Cao, Y.; Lin, Y.; Bao, J.; Yao, Z.; Dai, Q.; Hu, H. Simmim: A simple framework for masked image modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 21–24 June 2022; pp. 9653–9663. [Google Scholar]
- Li, M.; Xu, P.; Li, C.G.; Guo, J. MaskCL: Semantic mask-driven contrastive learning for unsupervised person re-identification with clothes change. arXiv 2023, arXiv:2305.13600. [Google Scholar]
- He, M.; Qin, L.; Deng, X.; Liu, K. MFI-YOLO: Multi-fault insulator detection based on an improved YOLOv8. IEEE Trans. Power Deliv. 2023, 39, 168–179. [Google Scholar] [CrossRef]
- Cui, Y.; Ren, W.; Cao, X.; Knoll, A. Image restoration via frequency selection. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 46, 1093–1108. [Google Scholar] [CrossRef] [PubMed]
- Huang, S.; Lu, Z.; Cun, X.; Yu, Y.; Zhou, X.; Shen, X. DEIM: DETR with Improved Matching for Fast Convergence. arXiv 2024, arXiv:2412.04234. [Google Scholar]
- Qiu, Y.W.; Zhang, K.; Wang, C.; Luo, W.; Li, H.; Jin, Z. Mb-taylorformer: Multi-branch efficient transformer expanded by taylor formula for image dehazing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 12802–12813. [Google Scholar]
- Fan, J.; Li, X.; Qian, J.; Li, J.; Yang, J. Non-aligned supervision for real image dehazing. arXiv 2023, arXiv:2303.04940. [Google Scholar]
- Liu, K.X.; Zhang, Y.; Li, W.; Wang, J. Image Dehazing Technique Based on DenseNet and the Denoising Self-Encoder. Processes 2024, 12, 2568. [Google Scholar] [CrossRef]
- Cui, Y.; Ren, W.; Knoll, A. Omni-kernel network for image restoration. Proc. AAAI Conf. Artif. Intell. 2024, 38, 1426–1434. [Google Scholar] [CrossRef]
- Zhang, Y.; Zhou, S.; Li, H. Depth information assisted collaborative mutual promotion network for single image dehazing. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 2846–2855. [Google Scholar]
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollár, P.; Girshick, R. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 16000–16009. [Google Scholar]
- Chen, M.; Radford, A.; Child, R.; Wu, J.; Jun, H.; Luan, D.; Sutskeve, I. Generative pretraining from pixels. In Proceedings of the International Conference on Machine Learning, PMLR, Virtual, 13–18 July 2020; pp. 1691–1703. [Google Scholar]
- Bao, H.; Dong, L.; Piao, S.; Wei, F. Beit: Bert pre-training of image transformers. arXiv 2021, arXiv:2106.08254. [Google Scholar]
- Caron, M.; Touvron, H.; Misra, I.; Jégou, H.; Mairal, J.; Bojanowski, P.; Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 9650–9660. [Google Scholar]
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path aggregation network, for instance, segmentation. In Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA; IEEE Press: Piscataway, NJ, USA, 2018; pp. 8759–8768. [Google Scholar]
- Xie, W.; Yang, H.; Shi, L.; Liu, Z. MTA-Net: A One-stage Detector Based on a Multi-scale Task-aligned Network for Catenary Support Components. IEEE Trans. Instrum. Meas. 2024, 73, 5020413. [Google Scholar] [CrossRef]
- Yang, H.; He, J.; Liu, Z.; Zhang, C. LLD-MFCOS: A Multiscale Anchor-Free Detector Based on Label Localization Distillation for Wheelset Tread Defect Detection. IEEE Trans. Instrum. Meas. 2023, 73, 5003815. [Google Scholar] [CrossRef]
- Quan, Y.; Zhang, D.; Zhang, L.; Tang, J. Centralized feature pyramid for object detection. IEEE Trans. Image Process. 2023, 32, 4341–4354. [Google Scholar] [CrossRef] [PubMed]
- Cai, Z.W.; Vasconcelos, N. Cascade R-CNN: High-quality object detection and instance segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 1483–1498. [Google Scholar] [CrossRef] [PubMed]
- Tao, W.Y.; Feng, A. ATSS-driven surface flame detection and extent evaluation using edge computing on UAVs. IEEE Access 2023, 11, 72108–72119. [Google Scholar] [CrossRef]
- Zhang, H.; Li, F.; Liu, S.; Zhang, L.; Su, H.; Zhu, J.; Shum, H.Y. DINO: DETR with Improved denoising anchor boxes for end-to-end object detection. In Proceedings of the Eleventh International Conference on Learning Representations (ICLR), Kigali, Rwanda, 1–5 May 2023. [Google Scholar] [CrossRef]
- Cai, Z.; Liu, S.; Wang, G.; Ge, Z.; Zhang, X.; Huang, D. Align-DETR: Improving DETR with simple IoU-aware BCE loss. arXiv 2023. [Google Scholar] [CrossRef]
- Jia, D.; Yuan, Y.; He, H.; Wu, X.; Yu, H.; Lin, W.; Hu, H. DETRs with hybrid matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada; IEEE Press: Piscataway, NJ, USA, 2023; pp. 19702–19712. [Google Scholar]
- Zhang, S.; Wang, X.; Wang, J.; Pang, J.; Lyu, C.; Zhang, W.; Chen, K. Dense distinct query for end-to-end object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada; IEEE Press: Piscataway, NJ, USA, 2023; pp. 7329–7338. [Google Scholar]
- Khanam, R.; Hussain, M. Yolov11: An overview of the key architectural enhancements. arXiv 2024, arXiv:2410.17725. [Google Scholar]
- Tian, Y.; Ye, Q. Doermann, DYolov12: Attention-centric real-time object detectors. arXiv 2025, arXiv:2502.12524. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]











| Methods | PSNR | SSIM | Params/M | FLOPS/G |
|---|---|---|---|---|
| MB-Talyor [47] | 19.52 | 0.605 | 7.43 | 44.05 |
| NSDNet [48] | 21.83 | 0.611 | 11.38 | 56.86 |
| DAE-Net [49] | 21.22 | 0.574 | 3.65 | 32.23 |
| OKNet [50] | 21.43 | 0.634 | 4.72 | 39.67 |
| DCMPNet [51] | 19.93 | 0.597 | 7.16 | 62.89 |
| DCFNet (Ours) | 22.56 | 0.826 | 13.32 | 53.40 |
| Module | Accuracy | Speed | ||||
|---|---|---|---|---|---|---|
| MIMCL | FFPN | Dense O2O | mAP50 | mAP50–95 | mAR | FPS |
| 0.979 | 0.652 | 0.971 | 24.1 | |||
| ✓ | 0.987 | 0.703 | 0.983 | 23.5 | ||
| ✓ | 0.982 | 0.674 | 0.974 | 23.8 | ||
| ✓ | 0.984 | 0.683 | 0.976 | 22.6 | ||
| ✓ | ✓ | 0.989 | 0.727 | 0.987 | 21.8 | |
| ✓ | ✓ | 0.990 | 0.734 | 0.988 | 22.2 | |
| ✓ | ✓ | ✓ | 0.992 | 0.755 | 0.991 | 20.9 |
| SL | mAP | AP50 | AP75 | APs | APm | APl |
|---|---|---|---|---|---|---|
| MAE [52] | 67.91 | 96.44 | 77.9 | 62.61 | 72.64 | 82.55 |
| IGPT [53] | 71.98 | 96.94 | 77.54 | 65.51 | 73.08 | 82.92 |
| SimMIM [42] | 72.18 | 96.57 | 77.34 | 64.62 | 72.42 | 82.76 |
| BEIT [54] | 74.31 | 97.33 | 78.51 | 65.68 | 74.29 | 83.54 |
| DINO [55] | 73.32 | 98.23 | 78.31 | 65.54 | 73.47 | 83.11 |
| MIMCL (ours) | 75.53 | 99.27 | 80.23 | 65.71 | 72.32 | 87.34 |
| Methods | mAP | AP50 | AP75 | APs | APm | APl |
|---|---|---|---|---|---|---|
| FPN | 60.41 | 94.28 | 67.33 | 59.24 | 70.62 | 81.39 |
| PA-FPN [56] | 61.46 | 93.21 | 67.76 | 60.96 | 69.95 | 80.07 |
| NAS-BFPN [57] | 61.73 | 94.01 | 68.16 | 60.59 | 68.61 | 80.72 |
| CR-FPN [58] | 62.47 | 93.93 | 68.84 | 60.71 | 71.03 | 81.39 |
| CFP [59] | 65.45 | 94.81 | 73.96 | 64.32 | 72.46 | 83.95 |
| FFPN (ours) | 67.43 | 95.54 | 75.26 | 65.88 | 75.08 | 83.07 |
| Methods | mAP | AP50 | AP75 | APs | APm | APl | Pa/M | FLOPs/G | FPS |
|---|---|---|---|---|---|---|---|---|---|
| Dynamic RCNN | 58.33 | 91.83 | 62.58 | 56.97 | 62.63 | 78.17 | 41.41 | 197.7 | 16.2 |
| Cascade RCNN [60] | 56.87 | 91.41 | 62.63 | 53.16 | 62.37 | 76.58 | 69.19 | 225.4 | 14.8 |
| ATSS [61] | 57.73 | 92.93 | 61.85 | 56.96 | 65.51 | 77.69 | 32.14 | 192.4 | 15.6 |
| DINO [62] | 62.87 | 94.08 | 68.93 | 61.58 | 71.58 | 81.92 | 47.56 | 265.6 | 15.5 |
| Align-DETR [63] | 64.25 | 94.58 | 71.58 | 60.83 | 70.13 | 81.54 | 47.51 | 253.8 | 15.1 |
| HDINO [64] | 63.76 | 94.10 | 70.78 | 60.12 | 69.49 | 80.83 | 68.10 | 281.4 | 16.5 |
| DDQ (DETR) [65] | 63.44 | 94.18 | 71.47 | 61.25 | 70.42 | 82.42 | 65.75 | 852.5 | 15.4 |
| YOLOv11-L [66] | 66.28 | 95.21 | 75.24 | 63.53 | 73.24 | 83.43 | 25.32 | 87.3 | 32.6 |
| YOLOv12-L [67] | 67.74 | 94.79 | 73.57 | 63.98 | 73.82 | 83.38 | 26.42 | 89.5 | 30.7 |
| RT-DETR [68] | 65.27 | 97.92 | 72.29 | 62.01 | 73.06 | 82.84 | 42.14 | 125.7 | 23.2 |
| FF-DEIM (ours) | 75.53 | 99.27 | 80.23 | 65.71 | 72.32 | 85.34 | 67.35 | 153.3 | 20.9 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, L.; Huang, J.; Qin, G.; Cao, J.; Fan, F.; Wang, H.; Yang, H. FF-DEIM: DEIM with Image Dehazing and Self-Supervised Pretraining for Catenary Support Component Detection. Sensors 2026, 26, 5000. https://doi.org/10.3390/s26155000
Zhang L, Huang J, Qin G, Cao J, Fan F, Wang H, Yang H. FF-DEIM: DEIM with Image Dehazing and Self-Supervised Pretraining for Catenary Support Component Detection. Sensors. 2026; 26(15):5000. https://doi.org/10.3390/s26155000
Chicago/Turabian StyleZhang, Lingzhi, Jinyong Huang, Guojin Qin, Jincheng Cao, Fei Fan, Hui Wang, and Haonan Yang. 2026. "FF-DEIM: DEIM with Image Dehazing and Self-Supervised Pretraining for Catenary Support Component Detection" Sensors 26, no. 15: 5000. https://doi.org/10.3390/s26155000
APA StyleZhang, L., Huang, J., Qin, G., Cao, J., Fan, F., Wang, H., & Yang, H. (2026). FF-DEIM: DEIM with Image Dehazing and Self-Supervised Pretraining for Catenary Support Component Detection. Sensors, 26(15), 5000. https://doi.org/10.3390/s26155000

