A Lightweight Multi-Scale Feature Fusion Signal Detection Model for Metro Computer Interlocking Systems
Featured Application
Abstract
1. Introduction
1.1. Railway Interlocking System
1.2. AI-Based Detection Approaches
1.3. Research Objectives and Contributions
- A detection architecture for small targets at multiple scales, YOLOv8-SFF, is proposed by integrating a P2 detection layer, SSFF modules, and TFE modules [16] on top of YOLOv8. Redundant concatenation and upsampling operations are replaced by TFE blocks, yielding both higher accuracy and a more compact network.
- A joint lightweighting strategy combining LAMP pruning and CWD knowledge distillation is applied to YOLOv8-SFF. Using the unpruned model as the teacher and the pruned model as the student, the resulting compact model retains detection accuracy while achieving an 86.2% parameter reduction, making deployment on embedded dispatching terminals feasible.
- An automatic annotation subsystem based on HSV color analysis, morphological processing, and positional reasoning is developed, which substantially reduces the manual labeling effort required for small and densely distributed interlocking signals.
- Extensive experiments on a real metro interlocking dataset demonstrate that the proposed framework provides an end-to-end solution that balances precision, speed, and deployability for practical rail transit scenarios.
2. Related Works
2.1. Object Detection in Transportation
2.2. Model Compression for Object Detection
2.3. Signal Detection in Interlocking Systems
3. Proposed YOLOv8-SFF Architecture
3.1. Small Object Detection Layer
3.2. SSFF Module
3.3. TFE Module and Structural Optimization
3.4. Loss Function
4. Model Compression via Joint Pruning and Distillation
4.1. Motivation for Lightweighting
4.2. LAMP-Based Pruning
4.3. Channel-Wise Knowledge Distillation
4.4. Joint Optimization Pipeline
5. Signal Dataset and Annotation System
5.1. Signal Classification
5.2. Automatic Annotation System
5.3. Dataset Partition and Evaluation Protocol
6. Experiments
6.1. Implementation Details
6.2. Performance Metrics
6.3. Ablation Study
6.4. Confusion-Matrix Analysis
6.5. Comparative Experiments
6.6. Pruning and Distillation Results
6.7. Detection Results
7. Discussion
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| YOLO | You Only Look Once |
| SSFF | Scale Sequence Feature Fusion |
| TFE | Triple Feature Encoding |
| LAMP | Layer-Adaptive Magnitude-based Pruning |
| CWD | Channel-wise Knowledge Distillation |
| mAP | mean Average Precision |
| FPS | Frames Per Second |
| IoU | Intersection over Union |
| EIoU | Efficient Intersection over Union |
| FPN | Feature Pyramid Network |
| HSV | Hue, Saturation, Value |
| HOG | Histogram of Oriented Gradients |
| SIFT | Scale-Invariant Feature Transform |
| R-CNN | Region-Based Convolutional Neural Network |
| SSD | Single Shot MultiBox Detector |
| RPN | Region Proposal Network |
| KL | Kullback–Leibler |
References
- Huang, L. The past, present and future of railway interlocking system. In Proceedings of the 2020 IEEE 5th International Conference on Intelligent Transportation Engineering (ICITE), Beijing, China, 11–13 September 2020; pp. 170–174. [Google Scholar]
- Qian, K.; Tian, L.; Liu, Y.; Wen, X.; Bao, J. Image robust recognition based on feature-entropy-oriented differential fusion capsule network. Appl. Intell. 2021, 51, 1108–1117. [Google Scholar] [CrossRef] [Scilit]
- Su, H.; Wen, J. Reliability and safety analysis on railway signal regional computer interlocking system. Int. J. Saf. Secur. Eng. 2014, 4, 315–328. [Google Scholar] [CrossRef] [Scilit]
- Shopa, P.; Sumitha, N.; Patra, P.S.K. Traffic sign detection and recognition using OpenCV. In Proceedings of the International Conference on Information Communication and Embedded Systems (ICICES2014), Chennai, India, 27–28 February 2014; pp. 1–6. [Google Scholar]
- Ahmed, N.; Rabbi, S.; Rahman, T.; Mia, R.; Rahman, M. Traffic sign detection and recognition model using support vector machine and histogram of oriented gradient. Int. J. Inf. Technol. Comput. Sci. 2021, 13, 61–73. [Google Scholar] [CrossRef] [Scilit]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA, 23–28 June 2014; pp. 580–587. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; pp. 21–37. [Google Scholar]
- Qian, K.; Tian, L. A topic-based multi-channel attention model under hybrid mode for image caption. Neural Comput. Appl. 2022, 34, 2207–2216. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Farhadi, A. YOLOv3: An Incremental Improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G. YOLO by Ultralytics (Version 5.7.0). GitHub. 2022. Available online: https://github.com/ultralytics/yolov5 (accessed on 1 September 2024).
- Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. YOLOv4: Optimal speed and accuracy of object detection. arXiv 2020, arXiv:2004.10934. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Qiu, J. YOLO by Ultralytics (Version 8.0.0). GitHub. 2023. Available online: https://github.com/ultralytics/ultralytics (accessed on 1 September 2024).
- Lee, J.; Park, S.; Mo, S.; Ahn, S.; Shin, J. Layer-adaptive sparsity for the magnitude-based pruning. In Proceedings of the International Conference on Learning Representations (ICLR), Virtual, 3–7 May 2021. [Google Scholar]
- Shu, C.; Liu, Y.; Gao, J.; Yan, Z.; Shen, C. Channel-wise knowledge distillation for dense prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 5311–5320. [Google Scholar]
- Kang, M.; Ting, C.M.; Ting, F.F.; Phan, R.C.W. ASF-YOLO: A novel YOLO model with attentional scale sequence fusion for cell instance segmentation. Image Vis. Comput. 2024, 147, 105057. [Google Scholar] [CrossRef] [Scilit]
- Yuan, L.; Gan, Q.; Li, K.; Fu, Q. Optimal generation of test sequences for EMU and ATP on-board equipment interface test based on deep learning and genetic algorithm. J. China Railw. Soc. 2018, 40, 7. [Google Scholar] [CrossRef]
- Yan, X. Research on Simulation Test Method of Computer Interlocking Software. Master’s Thesis, Beijing Jiaotong University, Beijing, China, 2019. [Google Scholar]
- Han, C.; Gao, G.; Zhang, Y. Real-time small traffic sign detection with revised Faster-RCNN. Multimed. Tools Appl. 2019, 78, 13263–13278. [Google Scholar] [CrossRef] [Scilit]
- Choodowicz, E.; Lisiecki, P.; Lech, P. Hybrid algorithm for the detection and recognition of railway signs. In Progress in Computer Recognition Systems 11; Springer: Cham, Switzerland, 2020; pp. 337–347. [Google Scholar]
- Dewi, C.; Chen, R.C.; Liu, Y.T.; Jiang, X.; Hartomo, K.D. YOLOv4 for advanced traffic sign recognition with synthetic training data generated by various GAN. IEEE Access 2021, 9, 97228–97242. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Guo, Z.; Wu, J.; Tian, Y.; Tang, H.; Guo, X. Real-time vehicle detection based on improved YOLOv5. Sustainability 2022, 14, 12274. [Google Scholar]
- Yang, J.; He, W.Y.; Zhang, T.L.; Zhang, C.L.; Zeng, L.; Nan, B.F. Research on subway pedestrian detection algorithms based on SSD model. IET Intell. Transp. Syst. 2020, 14, 1491–1496. [Google Scholar] [CrossRef] [Scilit]
- Lin, M.; Li, C.; Bu, X.; Sun, M.; Lin, C.; Yan, J.; Ouyang, W.; Deng, Z. DETR for crowd pedestrian detection. arXiv 2020, arXiv:2012.06785. [Google Scholar]
- Han, S.; Pool, J.; Tran, J.; Dally, W.J. Learning both weights and connections for efficient neural networks. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Montreal, QC, Canada, 7–12 December 2015; pp. 1135–1143. [Google Scholar]
- Liu, Z.; Li, J.; Shen, Z.; Huang, G.; Yan, S.; Zhang, C. Learning efficient convolutional networks through network slimming. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2736–2744. [Google Scholar]
- Li, H.; Kadav, A.; Durdanovic, I.; Samet, H.; Graf, H.P. Pruning filters for efficient ConvNets. In Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France, 24–26 April 2017. [Google Scholar]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv 2015, arXiv:1503.02531. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Li, Q. Image recognition and analysis of computer interlock interface based on OpenCV technology. In Proceedings of the 2024 IEEE 6th Advanced Information Management, Communications, Electronic and Automation Control Conference (IMCEC), Chongqing, China, 24–26 May 2024; pp. 1318–1321. [Google Scholar]
- Zhang, H.; Kuang, W. Computer interlocking host computer recognition based on OpenCV and neural network. In Proceedings of the 2024 IEEE 6th Advanced Information Management, Communications, Electronic and Automation Control Conference (IMCEC), Chongqing, China, 24–26 May 2024; pp. 1171–1174. [Google Scholar]
- Cheng, H.; He, T.; Tian, R. Research on the application of YOLOv5 in station interlocking test. In Proceedings of the Third International Conference on Image Processing and Intelligent Control (IPIC 2023), Kuala Lumpur, Malaysia, 5–7 May 2023; Volume 12782, pp. 245–250. [Google Scholar]
- Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; Savarese, S. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 658–666. [Google Scholar]
- Zheng, Z.; Wang, P.; Liu, W.; Li, J.; Ye, R.; Ren, D. Distance-IoU loss: Faster and better learning for bounding box regression. Proc. AAAI Conf. Artif. Intell. 2020, 34, 12993–13000. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.F.; Ren, W.; Zhang, Z.; Jia, Z.; Wang, L.; Tan, T. Focal and efficient IOU loss for accurate bounding box regression. Neurocomputing 2022, 506, 146–157. [Google Scholar] [CrossRef] [Scilit]









| Signal Type | Display Legend | Classification Name | Function Description |
|---|---|---|---|
| Shunting Signal | ![]() | RL | Prohibit shunting |
![]() | BL | Suspend shunting | |
| Entry/Exit Signal | ![]() | BLFS | Blocking |
![]() | RDL | Prohibit inbound | |
![]() | DRL | Outbound prohibited | |
| Turnouts | ![]() | SG | Positioning |
![]() | SY | Reverse Positioning | |
| Track Lines | ![]() | LB | Idle |
![]() | LR | Occupied | |
![]() | LSB | Incoming |
| # | P2 | SSFF | TFE | Params (M) | GFLOPs | mAP@50 (%) | P (%) | R (%) |
|---|---|---|---|---|---|---|---|---|
| 1 | – | – | – | 3.01 | 8.1 | 87.37 ± 1.39 | 91.65 ± 0.70 | 75.33 ± 3.76 |
| 2 | ✔ | – | – | 2.92 | 12.2 | 98.63 ± 0.08 | 97.70 ± 0.33 | 97.19 ± 0.36 |
| 3 | – | ✔ | – | 3.04 | 8.3 | 87.53 ± 0.63 | 91.73 ± 1.67 | 76.16 ± 2.70 |
| 4 | – | – | ✔ | 3.02 | 8.3 | 87.43 ± 0.68 | 92.77 ± 0.86 | 74.27 ± 1.17 |
| 5 | ✔ | ✔ | – | 2.46 | 11.8 | 98.64 ± 0.02 | 97.69 ± 0.40 | 97.25 ± 0.13 |
| 6 | ✔ | – | ✔ | 2.45 | 11.5 | 98.52 ± 0.04 | 97.51 ± 0.45 | 97.15 ± 0.32 |
| 7 | – | ✔ | ✔ | 3.05 | 8.5 | 86.85 ± 0.32 | 93.22 ± 0.45 | 73.76 ± 0.84 |
| 8 | ✔ | ✔ | ✔ | 2.49 | 12.0 | 98.66 ± 0.08 | 97.24 ± 0.48 | 97.39 ± 0.09 |
| # | P2 | SSFF | TFE | mAP@50 (%) | mAP@50:95 (%) | P (%) | R (%) |
|---|---|---|---|---|---|---|---|
| 1 | – | – | – | 87.20 ± 1.33 | 70.36 ± 1.31 | 91.64 ± 1.84 | 74.90 ± 3.92 |
| 2 | ✔ | – | – | 98.67 ± 0.03 | 86.86 ± 0.39 | 97.61 ± 0.33 | 97.36 ± 0.27 |
| 3 | – | ✔ | – | 87.32 ± 0.89 | 71.06 ± 1.30 | 90.84 ± 2.09 | 75.71 ± 3.12 |
| 4 | – | – | ✔ | 87.30 ± 0.58 | 69.29 ± 0.71 | 93.46 ± 0.06 | 73.81 ± 0.78 |
| 5 | ✔ | ✔ | – | 98.65 ± 0.02 | 86.70 ± 0.10 | 97.63 ± 0.40 | 97.47 ± 0.07 |
| 6 | ✔ | – | ✔ | 98.53 ± 0.06 | 85.62 ± 0.18 | 97.63 ± 0.44 | 97.20 ± 0.36 |
| 7 | – | ✔ | ✔ | 86.63 ± 0.46 | 69.68 ± 0.44 | 93.19 ± 0.78 | 72.65 ± 0.45 |
| 8 | ✔ | ✔ | ✔ | 98.67 ± 0.04 | 86.74 ± 0.10 | 97.45 ± 0.12 | 97.36 ± 0.23 |
| Model | P (%) | R (%) | mAP50 (%) | Parameters | FPS | Model Size (MB) |
|---|---|---|---|---|---|---|
| YOLOv5n | 72.8 | 82.3 | 83.7 | 1,777,447 | 42.92 | 3.8 |
| YOLOv7-tiny | 65.3 | 70.8 | 64.5 | 6,039,342 | 62.89 | 12.3 |
| YOLOv9-tiny | 90.6 | 74.2 | 86.4 | 2,620,460 | 25.38 | 6.1 |
| RT-DETR-tiny | 99.1 | 99.2 | 99.3 | 19,884,600 | 17.30 | 40.5 |
| YOLOv8-SFF | 97.2 | 97.4 | 98.7 | 2,490,488 | 60.40 | 5.1 |
| Speedup | P (%) | R (%) | mAP50 (%) | FPS | Parameters | GFLOPs | Model Size (MB) |
|---|---|---|---|---|---|---|---|
| 1.0 | 54.3 | 60.3 | 57.6 | 44.8 | 933,387 | 7.9 | 2.2 |
| 1.5 | 42.4 | 34.8 | 33.0 | 46.1 | 575,875 | 6.0 | 1.5 |
| 2.0 | 19.2 | 7.4 | 11.3 | 47.9 | 433,017 | 4.8 | 1.2 |
| 2.5 | 10.5 | 0.7 | 0.2 | 48.2 | 345,022 | 4.0 | 1.1 |
| Speedup | P (%) | R (%) | mAP50 (%) | FPS | Parameters | GFLOPs | Model Size (MB) |
|---|---|---|---|---|---|---|---|
| Baseline | 98.0 | 96.6 | 98.2 | 60.4 | 2,490,488 | 12.0 | 5.1 |
| 1.0 | 97.9 | 97.6 | 98.7 | 56.6 | 929,905 | 7.9 | 2.2 |
| 1.5 | 97.2 | 97.7 | 98.6 | 58.2 | 573,110 | 6.0 | 1.5 |
| 2.0 | 96.8 | 97.0 | 98.5 | 60.9 | 430,615 | 4.7 | 1.2 |
| 2.5 | 96.5 | 96.8 | 98.0 | 61.0 | 342,875 | 3.9 | 1.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Xu, X.; Li, Z.; Xu, S.; Song, Y.; Yu, C.; Qian, K. A Lightweight Multi-Scale Feature Fusion Signal Detection Model for Metro Computer Interlocking Systems. Appl. Sci. 2026, 16, 5452. https://doi.org/10.3390/app16115452
Xu X, Li Z, Xu S, Song Y, Yu C, Qian K. A Lightweight Multi-Scale Feature Fusion Signal Detection Model for Metro Computer Interlocking Systems. Applied Sciences. 2026; 16(11):5452. https://doi.org/10.3390/app16115452
Chicago/Turabian StyleXu, Xiaonong, Zhengyan Li, Sicheng Xu, Yaqing Song, Chong Yu, and Kui Qian. 2026. "A Lightweight Multi-Scale Feature Fusion Signal Detection Model for Metro Computer Interlocking Systems" Applied Sciences 16, no. 11: 5452. https://doi.org/10.3390/app16115452
APA StyleXu, X., Li, Z., Xu, S., Song, Y., Yu, C., & Qian, K. (2026). A Lightweight Multi-Scale Feature Fusion Signal Detection Model for Metro Computer Interlocking Systems. Applied Sciences, 16(11), 5452. https://doi.org/10.3390/app16115452











