Improving Harvesting Efficiency and Sustainability Through Lightweight Visual Perception: COS-DETR for Resource-Constrained Tea-Harvesting Robots
Abstract
1. Introduction
- To develop COS-DETR, a lightweight visual perception framework that addresses the dual challenges of fine tea bud recognition and real-time inference on embedded platforms, by integrating the Faster CGLU block, the OmniKernel module with its internal FSAM, and SPDConv into the RT-DETR architecture.
- To design a hybrid model compression pipeline that combines layer adaptive magnitude pruning (LAMP), channel pruning, and Mimic + Linear knowledge distillation, achieving a parameter reduction of ≥20% and a computational reduction of ≥35% while keeping above .
- To evaluate the integration of the compressed detector, stereo positioning, and manipulator control through offline benchmarking and 15 static field positioning trials on an experimental prototype of a tea-harvesting platform built for this study, and to assess the feasibility and practical value of the detection, positioning, and grasping pipeline under the tested conditions. Because the platform is a research prototype rather than a production machine, this objective is limited to static positioning accuracy; continuous dynamic picking and the associated picking success rate and missed-detection rate are outside the scope of the present study.
2. Materials and Methods
2.1. Experimental Platform and Data Acquisition
2.1.1. Sample and Data Collection
2.1.2. Tea-Harvesting Robot Platform
- Visual perception module: Intel RealSense D405 (Jabil Precision Industry Co., Ltd., Guangzhou, China) stereo depth camera and fixed adjustable lighting.
- Mechanical execution module: Delta robotic arm and an end effector.
- Control and computing module: NVIDIA Jetson Orin NX (NVIDIA Corporation, Santa Clara, CA, USA).
- Power supply module: 48 V lithium battery pack and solar panels.
- Mobile operation module: Mobile chassis with adjustable width and height.
2.1.3. Training Configuration and Parameter Settings
2.2. COS-DETR: Design and Implementation
2.2.1. Input Feature Enhancement Module Based on BasicBlock Faster CGLU
2.2.2. Feature Extraction Based on OmniKernel and SPDConv
2.3. Lightweight Deployment Pipeline
2.3.1. LAMP with Channel Pruning
2.3.2. Mimic + Linear Knowledge Distillation
2.4. Evaluation Metrics
3. Results
3.1. Ablation Study
3.2. Comparative Experiment
3.3. Comparison of Different Network Models Under Challenging Field Conditions
3.4. Pruning and Distillation Experiments
3.5. Field Validation: 3D Positioning Accuracy Under Static Conditions
4. Discussion
4.1. Synergy of Multi-Scale and Frequency-Domain Features
4.2. Balancing Accuracy and Efficiency for Edge Deployment
4.3. Comparison with Recent Tea-Detection Studies
4.4. Limitations and Future Work
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Liu, Z. The development process and trend of Chinese tea comprehensive processing industry. J. Tea Sci. 2019, 39, 115–122. [Google Scholar] [CrossRef]
- Jia, J.; Wang, D.; He, L.; Wu, C.; Chen, J.; Zhang, J.; Li, Y. Research progress and prospects of mechanical technology for whole machine of high-quality tea picking robot. Trans. Chin. Soc. Agric. Mach. 2025, 56, 193–206. [Google Scholar] [CrossRef]
- Song, R.; Gao, C. Application of computer image processing technology in tea picking robot system. J. Agric. Mech. Res. 2023, 45, 177–179+183. [Google Scholar] [CrossRef]
- Tao, H.; Zhang, R.; Zhang, L.; Zhang, D.; Yi, T.; Wu, M. Tea harvest robot navigation path generation algorithm based on semantic segmentation using a visual sensor. Electronics 2025, 14, 988. [Google Scholar] [CrossRef] [Scilit]
- Jia, J.; Li, Y.; Wang, X.; Yu, T.; Chen, J.; Zhou, Y.; Wu, C. Collaborative motion planning for multi-arm tea-picking robots. Comput. Electron. Agric. 2026, 244, 111470. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Yu, J.; Chen, Y.; Yang, W.; Zhang, W.; He, Y. Real-time strawberry detection using deep neural networks on embedded system (rtsd-net): An edge AI application. Comput. Electron. Agric. 2022, 192, 106586, Corrigendum in Comput. Electron. Agric. 2023, 215, 108445. [Google Scholar] [CrossRef] [Scilit]
- Tang, Z.; Fang, L.; Sun, S.; Gong, Y.; Li, Q. ML-DETR: Multiscale-lite detection transformer for identification of mature cherry tomatoes. IEEE Trans. Instrum. Meas. 2025, 74, 2547018. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Luo, H.; Ren, L.; Huo, M.; Jiang, Y.; Kaynak, O. Data-driven design of distributed monitoring and optimization system for manufacturing systems. IEEE Trans. Ind. Inform. 2024, 20, 9455–9464. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Luo, H.; Qiao, X.; Huo, M.; Xu, X. Data-driven distributed robust monitoring and control optimization for interconnected systems. IEEE Trans. Ind. Inform. 2025, 21, 1399–1408. [Google Scholar] [CrossRef] [Scilit]
- Bac, C.W.; van Henten, E.J.; Hemming, J.; Edan, Y. Harvesting robots for high-value crops: State-of-the-art review and challenges ahead. J. Field Robot. 2014, 31, 888–911. [Google Scholar] [CrossRef] [Scilit]
- Arad, B.; Balendonck, J.; Barth, R.; Ben-Shahar, O.; Edan, Y.; Hellström, T.; Hemming, J.; Kurtser, P.; Ringdahl, O.; Tielen, T.; et al. Development of a sweet pepper harvesting robot. J. Field Robot. 2020, 37, 1027–1039. [Google Scholar] [CrossRef] [Scilit]
- Rajendran, V.; Debnath, B.; Mghames, S.; Mandil, W.; Parsa, S.; Parsons, S.; Ghalamzan-E., A. Towards autonomous selective harvesting: A review of robot perception, robot design, motion planning and control. J. Field Robot. 2024, 41, 2247–2279. [Google Scholar] [CrossRef] [Scilit]
- Yang, L. Fiber optic connector end-face defect detection based on machine vision. Opt. Fiber Technol. 2025, 91, 104158. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Xu, Q.; Liao, M.; Sun, K.; Xiang, R.; Xu, H. Feature selection based on information entropy for accurate detection of optical fiber end-face defects. Entropy 2026, 28, 462. [Google Scholar] [CrossRef] [Scilit]
- Dong, C.; Wu, W.; Han, C.; Zeng, Z.; Tang, T.; Liu, W. Plucking point and posture determination of tea buds based on deep learning. Agriculture 2025, 15, 144. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Wang, J.; Lu, G.; Fei, S.; Zheng, T.; Huang, B. Automated tea quality identification based on deep convolutional neural networks and transfer learning. J. Food Process Eng. 2023, 46, e14303. [Google Scholar] [CrossRef] [Scilit]
- Yan, C.; Chen, Z.; Li, Z.; Liu, R.; Li, Y.; Xiao, H.; Lu, P.; Xie, B. Tea sprout picking point identification based on improved DeepLabV3+. Agriculture 2022, 12, 1594. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Wang, J.; Yan, T.; Lu, X.; Lu, G.; Tang, X.; Huang, B. An instance-based deep transfer learning method for quality identification of Longjing tea from multiple geographical origins. Complex Intell. Syst. 2023, 9, 3409–3428. [Google Scholar] [CrossRef] [Scilit]
- Yang, J.; Chen, Y. Tender leaf identification for early-spring green tea based on semi-supervised learning and image processing. Agronomy 2022, 12, 1958. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Ren, Y.; Kang, S.; Yin, C.; Shi, Y.; Men, H. Identification of tea quality at different picking periods: A hyperspectral system coupled with a multibranch kernel attention network. Food Chem. 2024, 433, 137307. [Google Scholar] [CrossRef] [Scilit]
- Yang, G.; Weng, D.; Li, Z.; Wu, Y. Tomato ripeness detection model based on improved RT-DETR lightweight model. Agronomy 2026, 16, 932. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Gu, J.; Wang, M. A review on the application of computer vision and machine learning in the tea industry. Front. Sustain. Food Syst. 2023, 7, 1172543. [Google Scholar] [CrossRef] [Scilit]
- Shi, D. TransNeXt: Robust foveal visual perception for vision transformers. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; IEEE: New York, NY, USA, 2024; pp. 17773–17783. [Google Scholar] [CrossRef] [Scilit]
- Cui, Y.; Ren, W.; Knoll, A. Omni-kernel modulation for universal image restoration. IEEE Trans. Circuits Syst. Video Technol. 2024, 34, 12496–12509. [Google Scholar] [CrossRef] [Scilit]
- Sunkara, R.; Luo, T. No more strided convolutions or pooling: A new CNN building block for low-resolution images and small objects. In Machine Learning and Knowledge Discovery in Databases, Proceedings of the European Conference, ECML PKDD 2022, Grenoble, France, 19–23 September 2022; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2023; Volume 13715, pp. 443–459. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Zhang, H.; Xiao, J.; Nie, L.; Shao, J.; Liu, W.; Chua, T.S. SCA-CNN: Spatial and channel-wise attention in convolutional networks for image captioning. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; IEEE: New York, NY, USA, 2017; pp. 6298–6306. [Google Scholar] [CrossRef] [Scilit]
- Zhu, M.; Gupta, S. To prune, or not to prune: Exploring the efficacy of pruning for model compression. arXiv 2017, arXiv:1710.01878. [Google Scholar] [CrossRef] [Scilit]
- Polyak, A.; Wolf, L. Channel-level acceleration of deep face representations. IEEE Access 2015, 3, 2163–2175. [Google Scholar] [CrossRef] [Scilit]
- Aghli, N.; Ribeiro, E. Combining weight pruning and knowledge distillation for CNN compression. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Virtual Event, 19–25 June 2021; IEEE: New York, NY, USA, 2021; pp. 3185–3192. [Google Scholar] [CrossRef] [Scilit]
- Liu, D.; Zhu, Y.; Liu, Z.; Liu, Y.; Han, C.; Tian, J.; Li, R.; Yi, W. A survey of model compression techniques: Past, present, and future. Front. Robot. AI 2025, 12, 1518965. [Google Scholar] [CrossRef] [Scilit]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv 2015, arXiv:1503.02531. [Google Scholar] [CrossRef] [Scilit]
- Hirschmüller, H. Accurate and efficient stereo processing by semi-global matching and mutual information. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), San Diego, CA, USA, 20–26 June 2005; IEEE: New York, NY, USA, 2005; Volume 2, pp. 807–814. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Li, H.; Deng, X.; Liu, Y.; Wu, T.; Liu, W.; Xiao, R.; Wang, Z.; Wang, B. Improved You Only Look Once v.8 model based on deep learning: Precision detection and recognition of fresh leaves from Yunnan large-leaf tea tree. Agriculture 2024, 14, 2324. [Google Scholar] [CrossRef] [Scilit]
- Tang, X.; Tang, L.; Li, J.; Guo, X. Enhancing multilevel tea leaf recognition based on improved YOLOv8n. Front. Plant Sci. 2025, 16, 1540670. [Google Scholar] [CrossRef] [Scilit]
- Liu, Z.; Zhuo, L.; Dong, C.; Li, J.; Li, Y. TBD-Y: Automatic tea bud detection with synergistic object-spatial attention and global-local attention guided feature fusion. Smart Agric. Technol. 2025, 12, 101066. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Xi, Y.; Gu, J.; Yang, Q.; Pan, Z.; Zhang, X.; Xu, G.; Zhou, M. YOLOv8-TEA: Recognition method of tender shoots of tea based on instance segmentation algorithm. Agronomy 2025, 15, 1318. [Google Scholar] [CrossRef] [Scilit]
- Zhang, T.; Yang, Q.; Tong, X.; Hu, L.; Shao, J. A transformer-based path fusion detector network for tea bud detection. Smart Agric. Technol. 2026, 14, 102321. [Google Scholar] [CrossRef] [Scilit]


















| Parameter | Value |
|---|---|
| Learning Rate | 0.0001 |
| Momentum | 0.9 |
| Weight Decay | 0.0001 |
| Optimizer | AdamW |
| Batch Size | 4 |
| Image Size | 640 × 640 |
| Number of Training Epochs | 300 |
| Group | Add CGLU | Add OmniKernel | Add SPDConv | w/o FSAM | |
|---|---|---|---|---|---|
| 1 | × | × | × | – | 0.846 |
| 2 | ✓ | × | × | – | 0.859 |
| 3 | ✓ | ✓ | × | – | 0.868 |
| 4 | ✓ | ✓ | ✓ | × | 0.871 |
| 5 | ✓ | ✓ | ✓ | ✓ | 0.866 |
| Model | Precision (%) | Recall (%) | (%) | (%) |
|---|---|---|---|---|
| YOLOv8s | 77.92 | 62.84 | 70.29 | 54.71 |
| YOLO11s | 74.09 | 62.16 | 71.41 | 54.07 |
| YOLO12s | 68.32 | 65.54 | 70.73 | 54.23 |
| YOLOv5 | 81.5 | 78.0 | 78.1 | 41.3 |
| YOLOv8 | 83.6 | 80.7 | 81.2 | 44.5 |
| YOLOv11 | 82.3 | 79.8 | 78.3 | 42.9 |
| RT-DETR-R18 | 82.7 | 84.2 | 84.6 | 47.8 |
| RT-DETR-R50 | 83.7 | 84.7 | 85.4 | 48.2 |
| COS-DETR | 85.6 | 85.0 | 87.1 | 50.4 |
| Module | (%) |
|---|---|
| YOLOv5 | 65.0 |
| YOLOv8 | 67.0 |
| RT-DETR-R18 | 78.0 |
| COS-DETR | 83.0 |
| Pruning | Pruning Ratio | (%) | GFLOPs | Parameters | FPS | Average Latency |
|---|---|---|---|---|---|---|
| None | None | 87.1 | 54.4 | 16,019,172 | 49.6 | 12.96 |
| LAMP + Cl | 1.3 | 85.5 | 38.9 | 13,166,188 | 55.8 | 10.13 |
| LAMP + Cl | 1.5 | 84.2 | 35.3 | 12,892,196 | 58.1 | 9.50 |
| LAMP + Cl | 1.7 | 84.0 | 34.1 | 12,768,324 | 59.4 | 8.26 |
| LAMP | 1.3 | 86.1 | 41.5 | 14,077,252 | 43.3 | 15.13 |
| LAMP | 1.5 | 85.1 | 36.2 | 13,163,548 | 51.9 | 12.27 |
| LAMP | 1.7 | 84.6 | 35.5 | 12,930,188 | 54.3 | 11.01 |
| L1 | 1.3 | 86.3 | 41.4 | 14,261,764 | 43.1 | 15.25 |
| L1 | 1.5 | 83.8 | 35.6 | 12,959,028 | 43.7 | 14.98 |
| L1 | 1.7 | 81.8 | 33.2 | 12,548,732 | 49.7 | 11.01 |
| group_taylor | 1.3 | 85.3 | 40.2 | 13,741,844 | 54.2 | 11.09 |
| group_taylor | 1.5 | 82.8 | 35.6 | 12,943,396 | 48.2 | 13.24 |
| group_taylor | 1.7 | 83.0 | 35.5 | 12,943,396 | 50.0 | 13.50 |
| No. | Actual Coords (mm) | Transformation Coords (mm) | Abs. Error (mm) | Rel. Error (%) |
|---|---|---|---|---|
| 1 | 4.88 | 0.84 | ||
| 2 | 3.93 | 0.69 | ||
| 3 | 3.57 | 0.65 | ||
| 4 | 3.23 | 0.57 | ||
| 5 | 3.88 | 0.69 | ||
| 6 | 2.63 | 0.46 | ||
| 7 | 3.89 | 0.66 | ||
| 8 | 1.86 | 0.32 | ||
| 9 | 3.89 | 0.67 | ||
| 10 | 2.89 | 0.45 | ||
| 11 | 4.01 | 0.60 | ||
| 12 | 3.33 | 0.49 | ||
| 13 | 3.20 | 0.51 | ||
| 14 | 3.57 | 0.57 | ||
| 15 | 4.47 | 0.65 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Xu, H.; Liang, J.; Peng, J.; Wen, Y.; Duan, Y.; Dai, Y.; Liao, M.; Xu, Q. Improving Harvesting Efficiency and Sustainability Through Lightweight Visual Perception: COS-DETR for Resource-Constrained Tea-Harvesting Robots. Processes 2026, 14, 2963. https://doi.org/10.3390/pr14182963
Xu H, Liang J, Peng J, Wen Y, Duan Y, Dai Y, Liao M, Xu Q. Improving Harvesting Efficiency and Sustainability Through Lightweight Visual Perception: COS-DETR for Resource-Constrained Tea-Harvesting Robots. Processes. 2026; 14(18):2963. https://doi.org/10.3390/pr14182963
Chicago/Turabian StyleXu, Haonan, Jianhao Liang, Jiahao Peng, Yixiao Wen, Yanbin Duan, Yunzhong Dai, Min Liao, and Quan Xu. 2026. "Improving Harvesting Efficiency and Sustainability Through Lightweight Visual Perception: COS-DETR for Resource-Constrained Tea-Harvesting Robots" Processes 14, no. 18: 2963. https://doi.org/10.3390/pr14182963
APA StyleXu, H., Liang, J., Peng, J., Wen, Y., Duan, Y., Dai, Y., Liao, M., & Xu, Q. (2026). Improving Harvesting Efficiency and Sustainability Through Lightweight Visual Perception: COS-DETR for Resource-Constrained Tea-Harvesting Robots. Processes, 14(18), 2963. https://doi.org/10.3390/pr14182963

