A Lightweight Improved RT-DETR for Stereo-Vision-Based Excavator Posture Recognition
Abstract
1. Introduction
- 1.
- A posture angle recognition method for excavator working devices is proposed, and an integrated technical pipeline is established, consisting of coded marker detection, binocular vision-based 3D coordinate recovery, world coordinate transformation, and geometric solving of posture angles. By combining the layout of coded markers, the 3D point cloud information from the ZED stereo camera, and the D–H kinematic model, the proposed method enables the calculation of the posture angles of the boom, stick, and bucket.
- 2.
- To address the large number of parameters and high computational cost of RT-DETR in edge-device deployment, RepViT is introduced as a lightweight backbone network. This reduces model complexity and inference burden while improving the feature extraction capability for small targets.
- 3.
- To overcome the limitation of traditional bilinear interpolation in preserving sufficient detail during upsampling, the DySample dynamic upsampling operator is introduced, allowing the sampling positions to be learned adaptively. This enhances multi-scale feature fusion and improves small-target detection performance.
- 4.
- A convolution-and-attention fusion module (CAFM) is designed to combine the local detail modeling capability of convolution with the global context modeling capability of the attention mechanism, thereby improving the discriminability between target markers and background under complex backgrounds.
- 5.
- A multi-scale feed-forward network (MSFN) is designed to enhance multi-scale feature representation through dilated convolution, depthwise separable convolution, and a gated fusion mechanism, further improving the model’s detection accuracy and robustness for small targets in complex scenes.
2. Related Work
2.1. Object Detectors
2.2. Excavator Marker Design
3. Excavator Posture Recognition Method
3.1. Improved Real-Time Detection Transformer
Real-Time Detection Transformer Base Model
3.2. Improvements to Real-Time Detection Transformer
Lightweight Backbone Based on RepViT
3.3. DySample
3.4. Convolution and Attention Fusion Module
3.5. Multiscale Feed-Forward Network
3.6. Marker-Based Posture Recognition Method
3.6.1. Marker Installation Design
3.6.2. Kinematic Modeling Based on the D–H Method
3.6.3. Posture Angle Calculation
4. Results
4.1. Experimental Environment and Dataset
4.2. Evaluation Metrics
4.3. Model Comparison
4.4. Comparison of Backbone Designs
4.5. Comparison of Upsampling Operators
4.6. Ablation Study
4.7. Edge Deployment Evaluation on Jetson Orin NX 8GB
4.8. Pose Estimation Accuracy Analysis
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| RT-DETR | Real-Time Detection Transformer |
| DETR | Detection Transformer |
| RepViT | Reparameterized Vision Transformer |
| DySample | Dynamic Upsampler |
| CAFM | Convolution and Attention Fusion Module |
| MSFN | Multiscale Feed-Forward Network |
| FFN | Feed-Forward Network |
| CNN | Convolutional Neural Network |
| FCN | Fully Convolutional Network |
| NMS | Non-Maximum Suppression |
| AP | Average Precision |
| mAP | Mean Average Precision |
| IoU | Intersection over Union |
| D–H | Denavit–Hartenberg |
References
- Kim, J.; Lee, D.; Seo, J. Task planning strategy and path similarity analysis for an autonomous excavator. Autom. Constr. 2020, 112, 103108. [Google Scholar] [CrossRef] [Scilit]
- Fareh, R.; Baziyad, M.; Rabie, T.; Bettayeb, M. Enhancing path quality of real-time path planning algorithms for mobile robots: A sequential linear paths approach. IEEE Access 2020, 8, 167090–167104. [Google Scholar] [CrossRef] [Scilit]
- Dai, J.; Tang, J.; Huang, S.; Wang, Y. Signal-based intelligent hydraulic fault diagnosis methods: Review and prospects. Chin. J. Mech. Eng. 2019, 32, 75. [Google Scholar] [CrossRef] [Scilit]
- Baqqal, I.; Mouatassim, S.; Benabbou, R.; Benhra, J. Digital twin design for monitoring system: A case study of bucket wheel excavator. Artif. Intell. Ind. Appl. Smart Oper. Manag. 2023, 771, 156. [Google Scholar] [CrossRef] [Scilit]
- Xu, J.; Yoon, H.S. Vision-based estimation of excavator manipulator pose for automated grading control. Autom. Constr. 2019, 98, 122–131. [Google Scholar] [CrossRef] [Scilit]
- Onder, M.; Bayrak, A.; Aksoy, S. RISE-based backstepping control design for an electro-hydraulic arm system with parametric uncertainties. Int. J. Control 2022, 95, 2815–2827. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Cao, Y.; Xiang, Y. Pose estimation method for construction machine based on improved AlphaPose model. Eng. Constr. Archit. Manag. 2024, 31, 976–996. [Google Scholar] [CrossRef] [Scilit]
- Guo, Y.; Cui, H.; Li, S. Excavator joint node-based pose estimation using lightweight fully convolutional network. Autom. Constr. 2022, 141, 104435. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollar, P. Focal loss for dense object detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2980–2988. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single shot multibox detector. In European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar] [CrossRef] [Scilit]
- Bochkovskiy, A.; Wang, C.-Y.; Liao, H.-Y.M. YOLOv4: Optimal speed and accuracy of object detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef] [Scilit]
- Wang, A.; Chen, H.; Lin, Z.; Han, J.; Ding, G. RepViT: Revisiting mobile CNN from ViT perspective. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 15909–15920. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Wang, X.; Lv, W.; Bai, X.; Long, X.; Deng, K.; Dang, Q.; Han, S.; Liu, Q.; Hu, X.; et al. PP-YOLOv2: A practical object detector. arXiv 2021, arXiv:2104.10419. [Google Scholar] [CrossRef] [Scilit]
- Liu, W.; Lu, H.; Fu, H.; Cao, Z. Learning to upsample by learning to sample. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–3 October 2023; pp. 6027–6037. [Google Scholar] [CrossRef] [Scilit]
- Hu, S.; Gao, F.; Zhou, X.; Dong, J.; Du, Q. Hybrid convolutional and attention network for hyperspectral image denoising. IEEE Geosci. Remote Sens. Lett. 2024, 21, 1–5. [Google Scholar] [CrossRef] [Scilit]
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO series in 2021. arXiv 2021, arXiv:2107.08430. [Google Scholar] [CrossRef] [Scilit]
- Lin, J.; Mao, X.; Chen, Y.; Xu, L.; He, Y.; Xue, H. D2ETR: Decoder-only DETR with computationally efficient cross-scale attention. arXiv 2022, arXiv:2203.00860. [Google Scholar]
- Xu, S.; Wang, X.; Lv, W.; Chang, Q.; Cui, C.; Deng, K.; Wang, G.; Dang, Q.; Wei, S.; Du, Y.; et al. PP-YOLOE: An evolved version of YOLO. arXiv 2022, arXiv:2203.16250. [Google Scholar] [CrossRef] [Scilit]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2020; pp. 213–229. [Google Scholar] [CrossRef] [Scilit]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable transformers for end-to-end object detection. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
- Liu, S.; Li, F.; Zhang, H.; Yang, X.; Qi, X.; Su, H.; Zhu, J.; Zhang, L. DAB-DETR: Dynamic anchor boxes are better queries for DETR. In Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia, 26–30 April 2021. [Google Scholar]
- Li, F.; Zhang, H.; Liu, S.; Guo, J.; Ni, L.M.; Zhang, L. DN-DETR: Accelerate DETR training by introducing query denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 13619–13627. [Google Scholar] [CrossRef] [Scilit]
- Chen, Q.; Chen, X.; Zeng, G.; Wang, J. Group DETR: Fast training convergence with decoupled one-to-many label assignment. arXiv 2022, arXiv:2207.13085. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs beat YOLOs on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar] [CrossRef] [Scilit]
- Soni, S.; Ahirwar, S.L.; Jain, R.; Shrivastava, A.K. Simulation and static analysis on improved design of excavator boom. Int. J. Emerg. Technol. Adv. Eng. 2014, 4, 49–55. [Google Scholar]
- Yang, Q.; Li, H.; Liao, H.; Yin, F.; Cao, J.; Liu, K. Design of a Multi-DOF structure based on dynamic analysis and autonomous operation algorithm. J. Phys. Conf. Ser. 2023, 2557, 012023. [Google Scholar] [CrossRef] [Scilit]
- Feng, W.W. Design and Implementation of the Motion Attitude Detection System. Master’s Thesis, Chongqing University, Chongqing, China, 2008. [Google Scholar]
- Li, C.H.; Lu, Y. Facial expression recognition based on depth separable convolution. Comput. Eng. Des. 2021, 42, 1448–1454. [Google Scholar] [CrossRef]
- Kulyukin, V.A.; Kulyukin, A.V. Accuracy vs. Energy: An assessment of bee object inference in videos from on-hive video loggers with YOLOv3, YOLOv4-Tiny, and YOLOv7-Tiny. Sensors 2023, 23, 6791. [Google Scholar] [CrossRef] [Scilit] [PubMed]













| Marker Category | Train | Validation | Test | Total |
|---|---|---|---|---|
| Marker-a | 31,274 | 3944 | 3892 | 39,110 |
| Marker-b | 30,981 | 3812 | 3906 | 38,699 |
| Marker-c | 29,847 | 3706 | 3742 | 37,295 |
| Marker-d | 28,596 | 3517 | 3577 | 35,690 |
| Marker-e | 28,214 | 3463 | 3548 | 35,225 |
| Marker-f | 26,928 | 3305 | 3392 | 33,625 |
| Marker-g | 26,511 | 3244 | 3301 | 33,056 |
| Raw images | 36,313 | 4539 | 4539 | 45,391 |
| Video sequences | 24 | 3 | 3 | 30 |
| Model | Recall/% | mAP/% | Parameters (M) |
|---|---|---|---|
| YOLOv5m | 85.23 | 84.56 | 25.1 |
| YOLOv8m | 86.34 | 85.67 | 25.9 |
| YOLOv11m | 87.41 | 81.54 | 20.1 |
| DETR | 82.14 | 85.75 | 43.6 |
| Deformable-DETR | 81.65 | 81.56 | 35.2 |
| RT-DETR | 88.19 | 86.33 | 28.9 |
| FCN | 83.15 | 84.34 | 23.7 |
| RevCol-S | 87.24 | 89.25 | 28.2 |
| FocalNet-S | 89.23 | 91.24 | 61.3 |
| Ours | 93.57 | 94.29 | 18.8 |
| Model | APa | APb | APc | APd | APe | APf | APg | mAP/% | Params (M) |
|---|---|---|---|---|---|---|---|---|---|
| RT-DETR | 81.83 | 81.83 | 84.57 | 86.60 | 87.43 | 90.26 | 92.13 | 86.33 | 28.9 |
| RT-DETR + RepViT | 82.75 | 82.54 | 84.75 | 88.35 | 90.44 | 94.04 | 92.77 | 88.78 | 14.2 |
| RT-DETR + LSKNet | 84.79 | 81.37 | 80.16 | 85.85 | 94.25 | 93.58 | 94.26 | 87.75 | 17.7 |
| RT-DETR + EfficientViT | 82.82 | 84.16 | 85.30 | 86.96 | 86.83 | 88.75 | 89.30 | 86.30 | 11.8 |
| RT-DETR + EMO | 82.73 | 82.33 | 86.64 | 86.01 | 88.39 | 87.87 | 86.80 | 85.82 | 12.4 |
| RT-DETR + VanillaNet | 83.49 | 80.32 | 80.37 | 89.24 | 89.69 | 93.05 | 83.29 | 87.02 | 19.8 |
| RT-DETR + StarNet | 84.10 | 82.27 | 86.30 | 86.49 | 94.15 | 89.82 | 94.26 | 88.20 | 16.9 |
| Model | APa | APb | APc | APd | APe | APf | APg | mAP/% | Params (M) |
|---|---|---|---|---|---|---|---|---|---|
| RT-DETR (RepViT) | 82.75 | 82.54 | 84.75 | 88.35 | 90.44 | 94.04 | 92.77 | 88.78 | 14.2 |
| + DySample | 87.35 | 89.77 | 91.75 | 91.56 | 92.67 | 91.42 | 93.03 | 91.03 | 15.3 |
| + CARAFE | 81.34 | 88.10 | 87.81 | 89.35 | 92.63 | 93.91 | 94.89 | 89.72 | 16.2 |
| + SAPA | 87.59 | 83.20 | 90.89 | 88.13 | 90.99 | 92.27 | 92.61 | 89.38 | 15.7 |
| + Transposed Convolution | 84.46 | 87.63 | 88.48 | 88.86 | 89.54 | 90.33 | 93.10 | 90.91 | 17.8 |
| Model | APa | APb | APc | APd | APe | APf | APg | mAP/% |
|---|---|---|---|---|---|---|---|---|
| RT-DETR | 81.91 | 81.88 | 84.63 | 86.68 | 87.52 | 90.31 | 91.94 | 86.41 ± 0.22 |
| RT-DETR + RepViT | 83.44 | 83.12 | 85.58 | 89.17 | 91.46 | 94.78 | 94.47 | 88.86 ± 0.19 |
| RT-DETR + RepViT + DySample | 87.46 | 89.92 | 91.81 | 91.68 | 92.79 | 91.54 | 92.57 | 91.11 ± 0.16 |
| RT-DETR + RepViT + DySample + CAFM | 92.26 | 89.08 | 93.04 | 94.18 | 93.62 | 94.44 | 94.24 | 92.98 ± 0.14 |
| RT-DETR + RepViT + DySample + CAFM + MSFN | 94.95 | 94.16 | 93.62 | 93.58 | 95.47 | 95.02 | 93.72 | 94.36 ± 0.12 |
| Model | mAP/% | Params (M) | Params (M) | GFLOPs | GFLOPs |
|---|---|---|---|---|---|
| RT-DETR | – | 28.9 | – | 96.8 | – |
| RT-DETR + RepViT | +2.45 | 14.2 | -14.7 | 55.4 | -41.4 |
| RT-DETR + RepViT + DySample | +2.25 | 15.3 | +1.1 | 58.1 | +2.7 |
| RT-DETR + RepViT + DySample + CAFM | +1.87 | 16.9 | +1.6 | 64.7 | +6.6 |
| RT-DETR + RepViT + DySample + CAFM + MSFN | +1.38 | 18.8 | +1.9 | 72.3 | +7.6 |
| Model | Params (M) | Peak Memory (GB) | Latency (ms) | FPS |
|---|---|---|---|---|
| YOLOv5m | 25.1 | 1.24 | 28 | 36 |
| YOLOv8m | 25.9 | 1.28 | 25 | 40 |
| YOLOv11m | 20.1 | 1.16 | 22 | 45 |
| DETR | 43.6 | 2.11 | 71 | 14 |
| Deformable-DETR | 35.2 | 1.87 | 52 | 19 |
| RT-DETR | 28.9 | 1.61 | 39 | 26 |
| FCN | 23.7 | 1.36 | 34 | 29 |
| RevCol-S | 28.2 | 1.55 | 36 | 28 |
| FocalNet-S | 61.3 | 2.74 | 105 | 10 |
| Ours | 18.8 | 1.33 | 29 | 34 |
| Component | MAE (°) | RMSE (°) |
|---|---|---|
| Boom | 1.18 | 1.56 |
| Stick | 1.43 | 1.87 |
| Bucket | 2.06 | 2.71 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hou, Y.; Wu, K.; Zhang, Y.; Zhou, M.; Lu, J.; Zhang, Z. A Lightweight Improved RT-DETR for Stereo-Vision-Based Excavator Posture Recognition. Mathematics 2026, 14, 1226. https://doi.org/10.3390/math14071226
Hou Y, Wu K, Zhang Y, Zhou M, Lu J, Zhang Z. A Lightweight Improved RT-DETR for Stereo-Vision-Based Excavator Posture Recognition. Mathematics. 2026; 14(7):1226. https://doi.org/10.3390/math14071226
Chicago/Turabian StyleHou, Yunlong, Ke Wu, Yuhan Zhang, Mengying Zhou, Jiasheng Lu, and Zhao Zhang. 2026. "A Lightweight Improved RT-DETR for Stereo-Vision-Based Excavator Posture Recognition" Mathematics 14, no. 7: 1226. https://doi.org/10.3390/math14071226
APA StyleHou, Y., Wu, K., Zhang, Y., Zhou, M., Lu, J., & Zhang, Z. (2026). A Lightweight Improved RT-DETR for Stereo-Vision-Based Excavator Posture Recognition. Mathematics, 14(7), 1226. https://doi.org/10.3390/math14071226

