GEFA-YOLO: Lightweight Weed Detection with Group-Enhanced Fusion Attention
Abstract
1. Introduction
- 1.
- Propose the GEFA attention mechanism, which enhances feature extraction capabilities through the combination of group convolution, one-dimensional convolution, and residual connections, effectively improving the comprehensive performance of the attention mechanism.
- 2.
- Propose a named group-enhanced fusion attention YOLO (GEFAY) object detection network, which integrates multiple attention mechanisms (such as SE, CBAM, mixed local channel attention (MLCA), etc.) and the proposed GEFA.
- 3.
- Experiments conducted on multiple datasets show that GEFA and GEFAY are lightweight and feasible.
- 4.
- The innovations in deployment and interaction ensure detection performance while supporting the deployment of embedded devices. It also designs a visualization output and interaction interface to solve the problem of “difficulty in implementation and operation” of traditional algorithms.
2. Methodology
2.1. Attention Mechanism
2.2. YOLOv5 Benchmark Model
2.3. Group-Enhanced Fusion Attention Mechanism
2.4. Group-Enhanced Fusion Attention Yolo
3. Experiments
3.1. Experimental Setup and Evaluation Metrics
3.2. Comparative Experiments
3.3. Verification of Cross-Scene Generalization Capability
3.4. Performance Evaluation of Object Detection Algorithms
4. Design and Implementation of Weed Detection System
4.1. Research Motivation
4.2. Design of Weed Detection System
- 1.
- Parameter Configuration Module: After the model selection is completed, the system will guide the user into the parameter configuration stage. The core of this stage lies in adjusting key hyperparameters to adapt to different detection difficulties, target densities, and equipment performance constraints.
- 2.
- Data Selection Module: Given that weed detection tasks may present multiple input forms in practice, the system has designed multiple data input methods for the detection phase to meet the different needs and operating habits of users.
- 3.
- Result Display Module: After the weed detection process is completed, the system will visualize the results obtained from the analysis, so that users can intuitively observe the distribution of weeds in the screen.
4.3. Implementation of Weed Detection System
4.3.1. Parameter Configuration Module
4.3.2. Data Selection Module
4.3.3. Result Display
4.4. Deployment of Weed Detection System Equipment
4.4.1. Introduction to Edge Devices
4.4.2. The Identification Process of Equipment Modules
| Platform | Specification |
|---|---|
| Embedded Board | Rockchip RK3588 Development Board |
| CPU | 8-core heterogeneous architecture: 4 × Cortex-A76 + 4 × Cortex-A55, up to 2.4 GHz |
| NPU | Built-in Neural Processing Unit (NPU), peak performance up to 6 TOPS |
| Camera Module | MCIMX415 high-resolution industrial camera |
| Display | 5.5-inch MIPI display, 1080p resolution |

| Deployment Scenario | Platform | Input Resolution | FPS |
|---|---|---|---|
| Online system | PC (RTX 3090) | 640 × 640 | 310.9 |
| Edge deployment | RK3588 | 640 × 640 | 29.19 |

4.5. Chapter Summary
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| GEFAY | Group-Enhanced Fusion Attention YOLO |
| GEFA | Group-Enhanced Fusion Attention |
References
- Zhou, Q.; Li, H.; Cai, Z.; Zhong, Y.; Zhong, F.; Lin, X.; Wang, L. YOLO-ACE: Enhancing YOLO with Augmented Contextual Efficiency for Precision Cotton Weed Detection. Sensors 2025, 25, 1635. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Columbus, OH, USA, 23–28 June 2014; pp. 580–587. [Google Scholar]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the European Conference on Computer Vision (ECCV), Amsterdam, The Netherlands, 11–14 October 2016; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Balasubramaniam, A.; Pasricha, S. Object Detection in Autonomous Vehicles: Status and Open Challenges. arXiv 2022, arXiv:2201.07706. [Google Scholar] [CrossRef] [Scilit]
- Juyal, A.; Sharma, S.; Matta, P. Deep Learning Methods for Object Detection in Autonomous Vehicles. In Proceedings of the 5th International Conference on Trends in Electronics and Informatics (ICOEI), Tirunelveli, Tamil Nadu, India, 3–5 June 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 751–755. [Google Scholar]
- Wei, Y.; Tran, S.; Xu, S.; Kang, B.; Springer, M. Deep Learning for Retail Product Recognition: Challenges and Techniques. Comput. Intell. Neurosci. 2020, 2020, 8875910. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sinha, A.; Banerjee, S.; Chattopadhyay, P. An Improved Deep Learning Approach for Product Recognition on Racks in Retail Stores. arXiv 2022, arXiv:2202.13081. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, H.H.; Ta, T.N.; Nguyen, N.C.; Bui, V.T.; Pham, H.M.; Nguyen, D.M. YOLO-Based Real-Time Human Detection for Smart Video Surveillance at the Edge. In Proceedings of the IEEE International Conference on Communications and Electronics (ICCE), Phu Quoc Island, Vietnam, 13–15 January 2021; IEEE: Piscataway, NJ, USA, 2021; pp. 439–444. [Google Scholar]
- Wang, Z.; Zhou, Q.; Wang, L.; Wu, Q. Tea-Bud Detection Based on a Hybrid Attention Mechanism. J. Nanjing Univ. Inf. Sci. Technol. (Natural Sci. Ed.) (In Chinese) 2025, 17, 506–514. (In Chinese) [Google Scholar] [CrossRef]
- Wang, Z.; Su, Y.; Kang, F.; Wang, L.; Lin, Y.; Wu, Q.; Li, H.; Cai, Z. PC-YOLO11s: A Lightweight and Effective Feature Extraction Method for Small Target Image Detection. Sensors 2025, 25, 348. [Google Scholar] [CrossRef] [Scilit]
- He, C.; Wan, F.; Ma, G.; Mou, X.; Zhang, K.; Wu, X.; Huang, X. Analysis of the Impact of Different Improvement Methods Based on YOLOv8 for Weed Detection. Agriculture 2024, 14, 674. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.L.; Su, W.H.; Zhang, H.Y.; Peng, Y. SE-YOLOv5x: An Optimized Model Based on Transfer Learning and Visual Attention Mechanism for Identifying and Localizing Weeds and Vegetables. Agronomy 2022, 12, 2061. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Wang, H.; Zhang, H.; Luo, T.; Wei, D.; Long, T.; Wang, Z. Weed Detection in Sesame Fields Using a YOLO Model with an Enhanced Attention Mechanism and Feature Fusion. Comput. Electron. Agric. 2022, 202, 107412. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Wang, Q.; Qiao, Y.; Zhang, X.; Lu, C.; Wang, C. Precision Weed Management for Straw-Mulched Maize Field: Advanced Weed Detection and Targeted Spraying Based on Enhanced YOLOv5s. Agriculture 2024, 14, 2134. [Google Scholar] [CrossRef] [Scilit]
- Guo, X.; Ou, Y.; Deng, K.; Fan, X.; Gao, R.; Zhou, Z. A Unmanned Aerial Vehicle-Based Image Information Acquisition Technique for the Middle and Lower Sections of Rice Plants and a Predictive Algorithm Model for Pest and Disease Detection. Agriculture 2025, 15, 790. [Google Scholar] [CrossRef] [Scilit]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Sun, C.; Zhou, X.; Zhang, M.; Qin, A. SE-VisionTransformer: Hybrid Network for Diagnosing Sugarcane Leaf Diseases Based on Attention Mechanism. Sensors 2023, 23, 8529. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Guo, Z.; Chen, X.; Li, M.; Chi, Y.; Shi, D. Construction and Validation of Peanut Leaf Spot Disease Prediction Model Based on Long Time Series Data and Deep Learning. Agronomy 2024, 14, 294. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 11534–11542. [Google Scholar]
- Qin, Z.; Zhang, P.; Wu, F.; Li, X. FCANet: Frequency Channel Attention Networks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 783–792. [Google Scholar]
- Gao, Z.; Xie, J.; Wang, Q.; Li, P. Global Second-Order Pooling Convolutional Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Beach, CA, USA, 15–20June 2019; pp. 3024–3033. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Wan, D.; Lu, R.; Shen, S.; Xu, T.; Lang, X.; Ren, Z. Mixed Local Channel Attention for Object Detection. Eng. Appl. Artif. Intell. 2023, 123, 106442. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Wang, W.; Hu, X.; Yang, J. Selective Kernel Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 510–519. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. arXiv 2017, arXiv:1706.03762. [Google Scholar]
- Fan, X.; Zhang, S.; Chen, B.; Zhou, M. Bayesian Attention Modules. Adv. Neural Inf. Process. Syst. 2020, 33, 16362–16376. [Google Scholar]
- Li, X.; Zhong, Z.; Wu, J.; Yang, Y.; Lin, Z.; Liu, H. Expectation-Maximization Attention Networks for Semantic Segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 9167–9176. [Google Scholar]
- Zhang, L.; Winn, J.; Tomioka, R. Gaussian Attention Model and Its Application to Knowledge Base Embedding and Question Answering. arXiv 2016, arXiv:1611.02266. [Google Scholar] [CrossRef] [Scilit]
- Ultralytics. YOLOv5. 2022. Available online: https://github.com/ultralytics/yolov5 (accessed on 13 February 2023).
- Liu, S.; Qi, L.; Qin, H.; Shi, J.; Jia, J. Path Aggregation Network for Instance Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 8759–8768. [Google Scholar]
- Dang, F.; Chen, D.; Lu, Y.; Li, Z. YOLOWeeds: A novel benchmark of YOLO object detectors for multi-class weed detection in cotton production systems. Comput. Electron. Agric. 2023, 205, 107655. [Google Scholar] [CrossRef] [Scilit]
- Everingham, M.; Van Gool, L.; Williams, C.K.; Winn, J.; Zisserman, A. The Pascal Visual Object Classes (VOC) Challenge. Int. J. Comput. Vis. 2010, 88, 303–338. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland, 6–12 September 2014; Springer: Berlin/Heidelberg, Germany, 2014; pp. 740–755. [Google Scholar]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate Attention for Efficient Mobile Network Design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 13713–13722. [Google Scholar]
- Chattopadhay, A.; Sarkar, A.; Howlader, P.; Balasubramanian, V.N. Grad-CAM++: Generalized Gradient-Based Visual Explanations for Deep Convolutional Networks. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, NV, USA, 12–15 March 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 839–847. [Google Scholar]
- Wang, Y.; Wang, C.; Zhang, H.; Dong, Y.; Wei, S. A SAR Dataset of Ship Detection for Deep Learning under Complex Backgrounds. Remote Sens. 2019, 11, 765. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Chen, F.; Huang, H.; Li, D.; Cheng, W. A New Steel Defect Detection Algorithm Based on Deep Learning. Comput. Intell. Neurosci. 2021, 2021, 5592878. [Google Scholar] [CrossRef] [Scilit]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 7464–7475. [Google Scholar]
- Ge, Z.; Liu, S.; Wang, F.; Li, Z.; Sun, J. YOLOX: Exceeding YOLO Series in 2021. arXiv 2021, arXiv:2107.08430. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Jiang, C. DAHD-YOLO: A New High Robustness and Real-Time Method for Smoking Detection. Sensors 2025, 25, 1433. [Google Scholar] [CrossRef] [Scilit]
- Kurniawan, H.; Arief, M.A.A.; Manggala, B.; Kim, H.; Lee, S.; Kim, M.S.; Baek, I.; Cho, B.K. Development of an Intelligent Inspection System Based on YOLOv7 for Real-Time Detection of Foreign Materials in Fresh-Cut Vegetables. Agriculture 2025, 15, 2297. [Google Scholar] [CrossRef] [Scilit]

















| Platform | Specification |
|---|---|
| CPU | Intel(R) Xeon(R) Gold 6330 CPU |
| GPU | NVIDIA GeForce RTX 3090 |
| Operating System | Ubuntu 18.04.5 LTS |
| Framework | PyTorch 1.13 |
| Attention | GFLOPS | Parameters (m) |
|---|---|---|
| NONE | 16.1 | 7.07 |
| SE/4 | 17.4 | 7.38 |
| ECA | 17.4 | 7.37 |
| CA/16 | 17.4 | 7.39 |
| CBAM/16 | 17.4 | 7.37 |
| MLCA (5 × 5) | 17.5 | 7.35 |
| MLCA (9 × 9) | 17.8 | 7.42 |
| GEFA (5 × 5) | 17.6 | 7.39 |
| GEFA (9 × 9) | 17.9 | 7.43 |
| Attention Mechanism | mAP@0.5 | mAP@0.5:0.95 | Parameters (M) | GFLOPs | GPU Speed (ms) | FPS |
|---|---|---|---|---|---|---|
| None | 94.0% | 84.6% | 7.04 | 15.9 | 6.5 | 353.8 |
| SE/4 | 94.0% | 84.0% | 7.35 | 17.2 | 9.7 | 337.6 |
| ECA | 94.6% | 85.3% | 7.36 | 17.3 | 9.1 | 282.4 |
| CA/16 | 94.3% | 85.7% | 7.36 | 17.3 | 11.1 | 299.1 |
| CBAM/16 | 94.2% | 84.4% | 7.35 | 17.3 | 10.8 | 325.7 |
| MLCA () | 94.8% | 85.3% | 7.34 | 17.3 | 10.6 | 177.7 |
| GEFA () | 95.0% | 87.0% | 7.36 | 17.5 | 9.6 | 310.9 |
| Attention Mechanism | mAP@0.5 | mAP@0.5:0.95 | Parameters (M) | GFLOPS | GPU Speed (ms) |
|---|---|---|---|---|---|
| NONE | 66.5% | 43.9% | 7.07 | 16.1 | 6.5 |
| SE/4 | 66.3% | 44.6% | 7.38 | 17.4 | 9.7 |
| ECA | 66.5% | 44.7% | 7.37 | 17.4 | 9.1 |
| CA/16 | 66.5% | 44.9% | 7.39 | 17.4 | 11.1 |
| CBAM/16 | 64.9% | 43.0% | 7.37 | 17.4 | 10.8 |
| MLCA (5 × 5) | 66.6% | 44.8% | 7.35 | 17.5 | 10.6 |
| GEFA (5 × 5) | 67.1% | 46.5% | 7.39 | 17.6 | 9.6 |
| Attention Mechanism | mAP@0.5 | mAP@0.5:0.95 | Parameters (M) | GFLOPS | GPU Speed (ms) |
|---|---|---|---|---|---|
| NONE | 56.7% | 36.5% | 7.23 | 16.6 | 6.6 |
| SE/4 | 57.3% | 37.3% | 7.54 | 17.9 | 9.9 |
| ECA | 57.7% | 37.5% | 7.53 | 17.9 | 9.3 |
| CA/16 | 58.2% | 38.1% | 7.55 | 17.9 | 11.4 |
| CBAM/16 | 57.7% | 37.6% | 7.54 | 17.9 | 10.7 |
| MLCA (5 × 5) | 58.3% | 38.0% | 7.52 | 18.0 | 10.4 |
| GEFA (5 × 5) | 58.5% | 38.2% | 7.55 | 18.1 | 9.3 |
| Model | mAP@0.5 | mAP@0.5:0.95 | Parameters (M) | GFLOPS | GPU Speed (ms) |
|---|---|---|---|---|---|
| Faster RCNN | 87.70% | - | 78.1 | 41.1 | 20.4 |
| YOLOX-s | 90.28% | 60.21% | 8.94 | 26.8 | 15.5 |
| YOLOv5-s | 98.33% | 69.32% | 6.58 | 15.8 | 6.2 |
| YOLOv7-tiny | 97.52% | 66.91% | 5.73 | 13.0 | 3.6 |
| GEFAY (Ours) | 98.43% | 71.24% | 7.33 | 17.3 | 9.4 |
| Model | mAP@0.5 | mAP@0.5:0.95 | Parameters (M) | GFLOPS | GPU Speed (ms) |
|---|---|---|---|---|---|
| Faster RCNN | 67.80% | - | 41.2 | 10.5 | 12.3 |
| YOLOX-s | 67.63% | 33.78% | 8.94 | 3.3 | 4.6 |
| YOLOv5-s | 75.03% | 37.72% | 6.70 | 15.8 | 5.2 |
| YOLOv7-tiny | 73.12% | 35.81% | 5.74 | 13.1 | 3.9 |
| GEFAY (Ours) | 75.82% | 42.01% | 7.34 | 17.5 | 9.3 |
| Platform | Specification |
|---|---|
| CPU | Intel(R) Core(TM) Ultra 5 225H (1.70 GHz) |
| GPU | Intel(R) Arc(TM) 130T GPU(16GB) |
| Operating System | Microsoft Windows 11 |
| GUI Framework | PyQt5 5.15.14 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, H.; Zhao, P.; Kang, F.; Su, Y.; Zhou, Q.; Wang, Z.; Wang, L. GEFA-YOLO: Lightweight Weed Detection with Group-Enhanced Fusion Attention. Sensors 2026, 26, 540. https://doi.org/10.3390/s26020540
Li H, Zhao P, Kang F, Su Y, Zhou Q, Wang Z, Wang L. GEFA-YOLO: Lightweight Weed Detection with Group-Enhanced Fusion Attention. Sensors. 2026; 26(2):540. https://doi.org/10.3390/s26020540
Chicago/Turabian StyleLi, Huicheng, Pushi Zhao, Feng Kang, Yuting Su, Qi Zhou, Zhou Wang, and Lijin Wang. 2026. "GEFA-YOLO: Lightweight Weed Detection with Group-Enhanced Fusion Attention" Sensors 26, no. 2: 540. https://doi.org/10.3390/s26020540
APA StyleLi, H., Zhao, P., Kang, F., Su, Y., Zhou, Q., Wang, Z., & Wang, L. (2026). GEFA-YOLO: Lightweight Weed Detection with Group-Enhanced Fusion Attention. Sensors, 26(2), 540. https://doi.org/10.3390/s26020540

