A Review of Lightweight Object Detection Technologies for Densely Occluded Scenarios in Agricultural Fields
Abstract
1. Introduction
2. Review Methodology
2.1. Search Strategy and Databases
2.2. Inclusion and Exclusion Criteria
2.3. Screening Procedure
3. Agricultural Field Occlusion Mechanisms and Feature Loss Analysis
3.1. Morphological Similarity and Class Confusion
3.2. Multi-Scale Variation and Effective Receptive Field Limitation
3.3. Dynamic Environmental Interference and Non-Rigid Deformation
4. Lightweight Backbone Networks and Feature Fusion Strategies
4.1. Lightweight Evolution of the YOLO Series
4.2. Architectural Breakthrough of Area Attention in YOLOv12
4.3. Lightweight Adaptation of End-to-End Transformers
4.4. Application of State Space Models (Mamba) in Agricultural Vision
5. Optimization and Application of Lightweight Attention Mechanisms
5.1. Global Recalibration and Multi-Dimensional Feature Enhancement
5.2. Multi-Scale Interaction and Non-Rigid Deformation Adaptation
5.3. Dynamic Routing and Sparse Computation Allocation
6. Bounding Box Regression and Multi-Dimensional Loss Function Optimization
6.1. Limitations of Traditional IoU Losses
6.2. Distance Metrics and Distribution Modeling
6.3. Hard Examples and Overlapping Region Focusing
7. Lightweight Application of Large Vision Models (LVMs) in Agricultural Perception
7.1. Downscaling Zero-Shot Segmentation Capabilities of Vision Foundation Models
7.2. Lightweight Reconstruction of Open-Vocabulary Detection
7.3. Multi-Modal Fusion and Vision-Language Large Models
8. Agricultural Perception Evaluation Systems and Frontier Datasets
8.1. Benchmarking Datasets and Field-Oriented Evaluation Frameworks
8.2. Shift Toward Multi-Dimensional Engineering Performance Evaluation
8.3. Fine-Grained Evaluation Gap in Dense Occlusion Scenarios
8.4. Standardization of Agriculture-Specific Benchmark Datasets
9. Conclusions and Outlook
9.1. Research Summary
9.2. Future Development Trends and Outlook
- (1)
- Contextual occlusion simulation and model adaptive optimization for specific agricultural field scenarios. Current general data augmentation methods struggle to realistically replicate the complex physical occlusion states of farmlands. Future efforts urgently require the development of advanced targeted data augmentation technologies, such as Context-Aware Scale-Adaptive Occlusion Simulation (CA-SAOS). Particularly in tasks like agricultural field weeding characterized by highly unstructured features, explicitly constructing complex topological relationships—such as leaf folding and severe plant intertwining—during the training phase can greatly enhance the robustness of lightweight models. This approach accelerates their practical iteration on real-world agricultural edge devices. To implement this, generative paradigms such as Diffusion Models or Generative Adversarial Networks (GANs) can be utilized to synthetically overlay photorealistic occluded leaves onto target stems. The core barrier lies in the immense computational overhead required to generate high-fidelity, agronomically accurate foliage topologies. Given the rapid maturity of generative vision, this direction may become more practical in the short term (1–2 years), depending on the availability of efficient generative pipelines and agronomically realistic simulation datasets.
- (2)
- Deep integration of multi-modal physical perception and active vision. In the depths of agricultural fields where the canopy is heavily closed, extracting features solely from optical RGB images faces insurmountable physical limits. Future agricultural robot vision systems are likely to increasingly rely on multi-modal joint perception. By introducing low-power thermal infrared, hyperspectral, and lightweight solid-state Light Detection and Ranging (LiDAR) data, information gaps will be bridged from both three-dimensional geometric and physical spectral dimensions. More importantly, perception systems will achieve “embodied intelligence” linkage with physical actuators like robotic arms, empowering agricultural machinery with “active vision” interaction capabilities—such as physically parting leaves—to obtain complete target semantic information by actively altering the physical occlusion state. Methodologically, this requires constructing cross-attention fusion mechanisms to align asynchronous heterogeneous sensor inputs in a unified coordinate space. The primary implementation challenge stems from the stringent temporal-spatial calibration alignment required under violent vehicle vibration and dust interference. Consequently, this multi-modal active interaction paradigm is expected to achieve wide field viability within the medium term (3–5 years).
- (3)
- Low-cost edge-side downscaling and distillation of Large Vision Models (LVMs). While Large Vision Models (LVMs) like SAM2 and Grounding DINO have demonstrated remarkable zero-shot boundary stripping and anti-occlusion capabilities in the cloud, their massive parameter scales remain a core obstacle to edge deployment. Future research should focus on knowledge distillation and Low-Rank Adaptation (LoRA) technologies under the cloud-edge collaborative paradigm. Exploring how to transfer part of the structural prior knowledge of LVMs to edge-side small models may improve the adaptability of lightweight networks to long-tailed diseases, sudden pest outbreaks, or previously unseen weeds. Technically, structured channel pruning paired with token-dropping attention mechanisms can be deployed to systematically compress these foundation backbones. The potential barrier involves resolving the significant accuracy collapse of open-vocabulary representations during severe architectural compression. Driven by intensive corporate and academic investment in mobile-side AI, this pathway is likely to achieve mature edge deployment within 2–3 years.
- (4)
- Construction of a closed-loop ecosystem integrating perception, cognition, and operation. Future agricultural visual perception should not remain merely at the physical localization level of outputting bounding boxes or masks; it must accelerate its evolution toward an integrated “cognition-diagnosis-decision” paradigm that deeply fuses agriculture-specific knowledge graphs (such as AgMMU). Through the coupling of multi-modal Vision-Language Models (VLMs) and Multi-modal Knowledge Graphs (MMKG), when the system identifies severely occluded weeds or lesions, it can autonomously transcend visual representation to directly map and invoke underlying agronomic expert systems, outputting decision instructions containing specific weeding plans or pesticide guidance. Integrating cognitive frameworks with perceptual models will be a key step toward robust autonomous operation in complex agricultural field environments. The realization of this ecosystem depends on deploying Graph Neural Networks (GNNs) to map vision embeddings directly onto dynamic agronomic knowledge entities. The bottleneck lies in the real-time inference latency when executing multi-turn cognitive reasoning on low-power edge microprocessors, alongside the lack of standardized cross-domain failure diagnostic corpora. As a highly integrated socio-technical challenge, it is projected as a long-term goal requiring 5+ years to reach full field autonomy.
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Mavridou, E.; Vrochidou, E.; Papakostas, G.A.; Pachidis, T.; Kaburlasos, V.G. Machine Vision Systems in Precision Agriculture for Crop Farming. J. Imaging 2019, 5, 89. [Google Scholar] [CrossRef]
- Guan, X.P.; Shi, L.Y.; Yang, W.G.; Ge, H.R.; Wei, X.H.; Ding, Y.H. Multi-Feature Fusion Recognition and Localization Method for Unmanned Harvesting of Aquatic Vegetables. Agriculture 2024, 14, 971. [Google Scholar] [CrossRef]
- Ge, C.; Zhang, G.J.; Wang, Y.J.; Shao, D.D.; Song, X.J.; Wang, Z.W. Research Status and Development Trends of Artificial Intelligence in Smart Agriculture. Agriculture 2025, 15, 2247. [Google Scholar] [CrossRef]
- Liang, Z.W.; Xu, X.Y.; Yang, D.Y.; Liu, Y.B. The Development of a Lightweight DE-YOLO Model for Detecting Impurities and Broken Rice Grains. Agriculture 2025, 15, 848. [Google Scholar] [CrossRef]
- Zhou, X.Y.; Chen, W.M.; Wei, X.H. Improved Field Obstacle Detection Algorithm Based on YOLOv8. Agriculture 2024, 14, 2263. [Google Scholar] [CrossRef]
- Zhang, T.F.; Zhou, J.H.; Liu, W.; Yue, R.C.; Yao, M.J.; Shi, J.W.; Hu, J.P. Seedling-YOLO: High-Efficiency Target Detection Algorithm for Field Broccoli Seedling Transplanting Quality Based on YOLOv7-Tiny. Agronomy 2024, 14, 931. [Google Scholar] [CrossRef]
- Lu, Y.Z.; Liu, P.F.; Tan, C. MA-YOLO: A Pest Target Detection Algorithm with Multi-Scale Fusion and Attention Mechanism. Agronomy 2025, 15, 1549. [Google Scholar] [CrossRef]
- Wang, C.D.; Chen, X.Y.; Jiao, Z.Y.; Song, S.; Ma, Z. An Improved YOLOP Lane-Line Detection Utilizing Feature Shift Aggregation for Intelligent Agricultural Machinery. Agriculture 2025, 15, 1361. [Google Scholar] [CrossRef]
- Wang, J.Z.; Gao, Z.H.; Zhang, Y.; Zhou, J.; Wu, J.Z.; Li, P.P. Real-Time Detection and Location of Potted Flowers Based on a ZED Camera and a YOLO V4-Tiny Deep Learning Algorithm. Horticulturae 2022, 8, 21. [Google Scholar] [CrossRef]
- Ji, W.; Pan, Y.; Xu, B.; Wang, J.C. A Real-Time Apple Targets Detection Method for Picking Robot Based on ShufflenetV2-YOLOX. Agriculture 2022, 12, 856. [Google Scholar] [CrossRef]
- Zhu, C.T.; Hao, S.H.; Liu, C.L.; Wang, Y.W.; Jia, X.; Xu, J.T.; Guo, S.B.; Huo, J.X.; Wang, W.M. An Efficient Computer Vision-Based Dual-Face Target Precision Variable Spraying Robotic System for Foliar Fertilisers. Agronomy 2024, 14, 2770. [Google Scholar] [CrossRef]
- Zhang, T.F.; Zhou, J.H.; Liu, W.; Yue, R.C.; Shi, J.W.; Zhou, C.J.; Hu, J.P. SN-CNN: A Lightweight and Accurate Line Extraction Algorithm for Seedling Navigation in Ridge-Planted Vegetables. Agriculture 2024, 14, 1446. [Google Scholar] [CrossRef]
- Wang, A.C.; Xu, Y.Z.; Hu, D.; Zhang, L.Y.; Li, A.; Zhu, Q.Z.; Liu, J.Z. Tomato Yield Estimation Using an Improved Lightweight YOLO11n Network and an Optimized Region Tracking-Counting Method. Agriculture 2025, 15, 1353. [Google Scholar] [CrossRef]
- Sun, J.; He, X.F.; Ge, X.; Wu, X.H.; Shen, J.F.; Song, Y.Y. Detection of Key Organs in Tomato Based on Deep Migration Learning in a Complex Background. Agriculture 2018, 8, 196. [Google Scholar] [CrossRef]
- Arthur, L.; Mahnan, S.; He, L.; Hussain, M.; Heinemann, P.; Brunharo, C. YOLOv7-CBAM and DeepSORT with pixel grid analysis for Real-Time weed localization and Intra-Row density estimation in apple orchards. Comput. Electron. Agric. 2025, 239, 111071. [Google Scholar] [CrossRef]
- Yang, C.; Zhao, B.; Mansurova, M.; Zhou, T.; Liu, Q.; Bao, J.; Zheng, D. AgriLiteNet: Lightweight Multi-Scale Tomato Pest and Disease Detection for Agricultural Robots. Horticulturae 2025, 11, 671. [Google Scholar] [CrossRef]
- Yang, Z.F.; Khan, Z.; Shen, Y.; Liu, H. GTDR-YOLOv12: Optimizing YOLO for Efficient and Accurate Weed Detection in Agriculture. Agronomy 2025, 15, 1824. [Google Scholar] [CrossRef]
- Xu, Z.J.; Liu, J.Z.; Wang, J.; Cai, L.J.; Jin, Y.C.; Zhao, S.Y.; Xie, B.B. Realtime Picking Point Decision Algorithm of Trellis Grape for High-Speed Robotic Cut-and-Catch Harvesting. Agronomy 2023, 13, 1618. [Google Scholar] [CrossRef]
- Zhou, M.L.; Wei, Z.X.; Wang, Z.L.; Sun, H.; Wang, G.B.; Yin, J.J. Design and Experimental Investigation of a Transplanting Mechanism for Super Rice Pot Seedlings. Agriculture 2023, 13, 1920. [Google Scholar] [CrossRef]
- Tao, L.; Qingchun, F.; Quan, Q.; Feng, X.; Chunjiang, Z. Occluded Apple Fruit Detection and Localization with a Frustum-Based Point-Cloud-Processing Approach for Robotic Harvesting. Remote Sens. 2022, 14, 482. [Google Scholar] [CrossRef]
- Wang, X.; Huang, Y.; Wei, S.; Xu, W.; Zhu, X.; Mu, J.; Chen, X. ELD-YOLO: A Lightweight Framework for Detecting Occluded Mandarin Fruits in Plant Research. Plants 2025, 14, 1729. [Google Scholar] [CrossRef]
- Yang, Q.; Gu, J.; Xiong, T.; Wang, Q.; Huang, J.; Xi, Y.; Shen, Z. RFA-YOLOv8: A Robust Tea Bud Detection Model with Adaptive Illumination Enhancement for Complex Orchard Environments. Agriculture 2025, 15, 1982. [Google Scholar] [CrossRef]
- Fu, H.; Guo, Z.; Feng, Q.; Xie, F.; Zuo, Y.; Li, T. MSOAR-YOLOv10: Multi-Scale Occluded Apple Detection for Enhanced Harvest Robotics. Horticulturae 2024, 10, 1246. [Google Scholar] [CrossRef]
- Peng, Y.; Wang, A.C.; Liu, J.Z.; Faheem, M. A Comparative Study of Semantic Segmentation Models for Identification of Grape with Different Varieties. Agriculture 2021, 11, 997. [Google Scholar] [CrossRef]
- Tao, K.; Wang, A.C.; Shen, Y.D.; Lu, Z.M.; Peng, F.T.; Wei, X.H. Peach Flower Density Detection Based on an Improved CNN Incorporating Attention Mechanism and Multi-Scale Feature Fusion. Horticulturae 2022, 8, 904. [Google Scholar] [CrossRef]
- Yang, X.; Zhao, W.; Wang, Y.; Yan, W.Q.; Li, Y. Lightweight and efficient deep learning models for fruit detection in orchards. Sci. Rep. 2024, 14, 26086. [Google Scholar] [CrossRef] [PubMed]
- Li, A.; Wang, C.R.; Wang, A.C.; Sun, J.P.; Gu, F.W.; Zhang, T.X. YOLO-MSRF: A Multimodal Segmentation and Refinement Framework for Tomato Fruit Detection and Segmentation with Count and Size Estimation Under Complex Illumination. Agriculture 2026, 16, 277. [Google Scholar] [CrossRef]
- Xin, X.; Sun, J.; Shi, L.; Yao, K.S.; Zhang, B. Application of hyperspectral imaging technology combined with ECA-MobileNetV3 in identifying different processing methods of Yunnan coffee beans. J. Food Compos. Anal. 2025, 143, 107625. [Google Scholar] [CrossRef]
- Sun, J.; Xin, X.; Xin, Y.; Cong, S.L. Non-destructive identification of processing methods of Yunnan coffee beans via portable near-infrared spectrometer and lightweight MobileNetV4. J. Food Compos. Anal. 2026, 149, 108724. [Google Scholar] [CrossRef]
- Zhang, F.; Chen, Z.J.; Ali, S.; Yang, N.; Fu, S.L.; Zhang, Y.K. Multi-class detection of cherry tomatoes using improved YOLOv4-Tiny. Int. J. Agric. Biol. Eng. 2023, 16, 225–231. [Google Scholar] [CrossRef]
- Xu, B.; Cui, X.; Ji, W.; Yuan, H.; Wang, J.C. Apple Grading Method Design and Implementation for Automatic Grader Based on Improved YOLOv5. Agriculture 2023, 13, 124. [Google Scholar] [CrossRef]
- Zhao, Y.Q.; Zhang, X.D.; Sun, J.J.; Yu, T.T.; Cai, Z.Y.; Zhang, Z.; Mao, H.P. Low-Cost Lettuce Height Measurement Based on Depth Vision and Lightweight Instance Segmentation Model. Agriculture 2024, 14, 1596. [Google Scholar] [CrossRef]
- Shi, Q.; Zhang, Y.Z.; Du, X.X.; Chen, T.H.; Wang, Y.F. Light-YOLO-Pepper: A Lightweight Model for Detecting Missing Seedlings. Agriculture 2026, 16, 231. [Google Scholar] [CrossRef]
- Ali, M.L.; Zhang, Z. The YOLO Framework: A Comprehensive Review of Evolution, Applications, and Benchmarks in Object Detection. Computers 2024, 13, 336. [Google Scholar] [CrossRef]
- Tai, S.; Tang, Z.; Li, B.; Wang, S.G.; Guo, X.H. Intelligent Recognition and Automated Production of Chili Peppers: A Review Addressing Varietal Diversity and Technological Requirements. Agriculture 2025, 15, 1200. [Google Scholar] [CrossRef]
- Zhang, R.X.; Zhu, H.T.; Chang, Q.L.; Mao, Q.R. A Comprehensive Review of Digital Twins Technology in Agriculture. Agriculture 2025, 15, 903. [Google Scholar] [CrossRef]
- Xia, G.; Li, X. YOLOv12-BDA: A Dynamic Multi-Scale Architecture for Small Weed Detection in Sesame Fields. Sensors 2025, 25, 6927. [Google Scholar] [CrossRef]
- Wang, Q.; Qin, W.C.; Liu, M.N.; Zhao, J.J.; Zhu, Q.Z.; Yin, Y.X. Semantic Segmentation Model-Based Boundary Line Recognition Method for Wheat Harvesting. Agriculture 2024, 14, 1846. [Google Scholar] [CrossRef]
- Zhu, W.D.; Sun, J.; Wang, S.M.; Shen, J.F.; Yang, K.F.; Zhou, X. Identifying Field Crop Diseases Using Transformer-Embedded Convolutional Neural Network. Agriculture 2022, 12, 1083. [Google Scholar] [CrossRef]
- Ji, W.; Zhai, K.L.; Xu, B.; Wu, J.W. Green Apple Detection Method Based on Multidimensional Feature Extraction Network Model and Transformer Module ☆. J. Food Prot. 2025, 88, 100397. [Google Scholar] [CrossRef]
- Wang, Z.; Liu, H.; Hu, Z.; Wang, Y. Multi-scale mamba attention and residual learning for robust agricultural image segmentation in precision agriculture. Discov. Appl. Sci. 2026, 8, 310. [Google Scholar] [CrossRef]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. Syst. Rev. 2021, 10, 89. [Google Scholar] [CrossRef]
- Pengyu, C.; Zhaojian, L.; Kaixiang, Z.; Dong, C.; Kyle, L.; Renfu, L. O2RNet: Occluder-occludee relational network for robust apple detection in clustered orchard environments. Smart Agric. Technol. 2023, 5, 100284. [Google Scholar] [CrossRef]
- Hu, T.T.; Wang, W.B.; Gu, J.A.; Xia, Z.L.; Zhang, J.; Wang, B. Research on Apple Object Detection and Localization Method Based on Improved YOLOX and RGB-D Images. Agronomy 2023, 13, 1816. [Google Scholar] [CrossRef]
- Yue, Y.; Zhao, A. Weed Discrimination at the Seedling Stage in Dryland Fields Under Maize–Soybean Rotation. Plants 2026, 15, 1114. [Google Scholar] [CrossRef]
- Cui, J.; Zhang, X.; Zhang, J.; Han, Y.; Ai, H.; Dong, C.; Liu, H. Weed identification in soybean seedling stage based on UAV images and Faster R-CNN. Comput. Electron. Agric. 2024, 227, 109533. [Google Scholar] [CrossRef]
- Chen, Z.; Wu, L.; Jia, Z.; Wang, J.; Zhou, G.; Zhang, Z. Research on Field Weed Target Detection Algorithm Based on Deep Learning. Sensors 2026, 26, 677. [Google Scholar] [CrossRef] [PubMed]
- Deng, L.; Miao, Z.H.; Zhao, X.G.; Yang, S.; Gao, Y.Y.; Zhai, C.Y.; Zhao, C.J. HAD-YOLO: An Accurate and Effective Weed Detection Model Based on Improved YOLOV5 Network. Agronomy 2025, 15, 57. [Google Scholar] [CrossRef]
- He, L.; Wu, D.; Zheng, X.; Xu, F.; Lin, S.; Wang, S.; Ni, F.; Zheng, F. RLK-YOLOv8: Multi-stage detection of strawberry fruits throughout the full growth cycle in greenhouses based on large kernel convolutions and improved YOLOv8. Front. Plant Sci. 2025, 16, 1552553. [Google Scholar] [CrossRef] [PubMed]
- Qiuchi, X.; Xiaoning, H.; Zhouxu, H.; Xingming, C.; Jintao, C.; Xiaoyu, T. Yolo-Pest: An Insect Pest Object Detection Algorithm via CAC3 Module. Sensors 2023, 23, 3221. [Google Scholar] [CrossRef]
- Wei, Z.; He, H.; Youqiang, S.; Xiaowei, W. AgriPest-YOLO: A rapid light-trap agricultural pest detection method based on deep learning. Front. Plant Sci. 2022, 13, 1079384. [Google Scholar] [CrossRef]
- Guo, Y.; Zhan, W.; Zhang, Z.; Zhang, Y.; Guo, H. FRPNet: A Lightweight Multi-Altitude Field Rice Panicle Detection and Counting Network Based on Unmanned Aerial Vehicle Images. Agronomy 2025, 15, 1396. [Google Scholar] [CrossRef]
- Wang, J.Z.; Zhang, Y.; Gu, R.R. Research Status and Prospects on Plant Canopy Structure Measurement Using Visual Sensors Based on Three-Dimensional Reconstruction. Agriculture 2020, 10, 462. [Google Scholar] [CrossRef]
- Ma, Y.; Xi, C.; Ma, T.; Sun, H.; Lu, H.; Xu, X.; Xu, C. I-YOLOv11n: A Lightweight and Efficient Small Target Detection Framework for UAV Aerial Images. Sensors 2025, 25, 4857. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Y.; Chen, Y.; Xu, X.; He, Y.; Gan, H.; Wu, N.; Wang, Z.; Sun, X.; Wang, Y.; Skobelev, P.; et al. Ta-YOLO: Overcoming target blocked challenges in greenhouse tomato detection and counting. Front. Plant Sci. 2025, 16, 1618214. [Google Scholar] [CrossRef] [PubMed]
- Yuxiang, W.; Zengling, Y.; Gert, K.; Ahmad, K.H. Correction: The impact of variable illumination on vegetation indices and evaluation of illumination correction methods on chlorophyll content estimation using UAV imagery. Plant Methods 2023, 19, 62. [Google Scholar] [CrossRef]
- Jia, W.K.; Zheng, Y.J.; Zhao, D.A.; Yin, X.; Liu, X.Y.; Du, R.C. Preprocessing method of night vision image application in apple harvesting robot. Int. J. Agric. Biol. Eng. 2018, 11, 158–163. [Google Scholar] [CrossRef]
- Gao, Z.M.; Ma, J.; Hu, W.; Wang, K.Y.; Liu, K.; Chen, J.; Wang, T.; Dong, X.Y.; Qiu, B.J. Wind-Induced Bending Characteristics of Crop Leaves and Their Potential Applications in Air-Assisted Spray Optimization. Horticulturae 2025, 11, 1002. [Google Scholar] [CrossRef]
- Hu, W.; Gao, Z.M.; Chen, J.; Dong, X.Y.; Qiu, B.J. Motion behavior of a charged droplet impacting on hydrophilic and hydrophobic leaf surfaces. J. Agric. Eng. 2025, 56, 1736. [Google Scholar] [CrossRef]
- Xu, S.; Zheng, S.; Rai, R. Dense object detection based canopy characteristics encoding for precise spraying in peach orchards. Comput. Electron. Agric. 2025, 232, 110097. [Google Scholar] [CrossRef]
- Praveen, S.; Jung, Y. CBAM-STN-TPS-YOLO: Enhancing Agricultural Object Detection through Spatially Adaptive Attention Mechanisms. arXiv 2025, arXiv:2506.07357. [Google Scholar]
- Dong, X.; Pan, J. DHS-YOLO: Enhanced Detection of Slender Wheat Seedlings Under Dynamic Illumination Conditions. Agriculture 2025, 15, 510. [Google Scholar] [CrossRef]
- Chen, J.; Fu, H.; Lin, C.; Liu, X.; Wang, L.; Lin, Y. YOLOPears: A novel benchmark of YOLO object detectors for multi-class pear surface defect detection in quality grading systems. Front. Plant Sci. 2025, 16, 1483824. [Google Scholar] [CrossRef]
- Li, S.; Guo, S.; Gao, S.; Sun, L.; Li, F.; Zhang, S. Robust detection for selective harvesting of field flat jujube: Overcoming occlusion and small-target challenges in unstructured environments. Front. Plant Sci. 2026, 17, 1795650. [Google Scholar] [CrossRef] [PubMed]
- Zhao, S.Y.; Fang, C.; Hua, T.Z.; Jiang, Y. Detecting the Maturity of Red Strawberries Using Improved YOLOv8s Model. Agriculture 2025, 15, 2263. [Google Scholar] [CrossRef]
- Weyler, J.; Magistri, F.; Marks, E.; Chong, Y.L.; Sodano, M.; Roggiolani, G.; Chebrolu, N.; Stachniss, C.; Behley, J. Phenobench: A large dataset and benchmarks for semantic image interpretation in the agricultural domain. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 9583–9594. [Google Scholar] [CrossRef]
- Olsen, A.; Konovalov, D.A.; Philippa, B.; Ridd, P.; Wood, J.C.; Johns, J.; Banks, W.; Girgenti, B.; Kenny, O.; Whinney, J. DeepWeeds: A multiclass weed species image dataset for deep learning. Sci. Rep. 2019, 9, 2058. [Google Scholar] [CrossRef]
- Duan, Y.L.; Han, W.Y.; Guo, P.; Wei, X.H. YOLOv8-GDCI: Research on the Phytophthora Blight Detection Method of Different Parts of Chili Based on Improved YOLOv8 Model. Agronomy 2024, 14, 2734. [Google Scholar] [CrossRef]
- Dang, F.; Chen, D.; Lu, Y.; Li, Z. YOLOWeeds: A novel benchmark of YOLO object detectors for multi-class weed detection in cotton production systems. Comput. Electron. Agric. 2023, 205, 107655. [Google Scholar] [CrossRef]
- Wang, W.B.; Xi, Y.D.; Gu, J.N.; Yang, Q.Y.; Pan, Z.Y.; Zhang, X.Z.; Xu, G.Y.; Zhou, M. YOLOv8-TEA: Recognition Method of Tender Shoots of Tea Based on Instance Segmentation Algorithm. Agronomy 2025, 15, 1318. [Google Scholar] [CrossRef]
- Chiu, M.T.; Xu, X.; Wei, Y.; Huang, Z.; Schwing, A.G.; Brunner, R.; Khachatrian, H.; Karapetyan, H.; Dozier, I.; Rose, G. Agriculture-vision: A large aerial image database for agricultural pattern analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 2828–2838. [Google Scholar]
- Rai, N.; Mahecha, M.; Christensen, A.; Quanbeck, J.; Howatt, K.; Ostlie, M.; Zhang, Y.; Sun, X. ImageWeeds: An Image dataset consisting of weeds in multiple formats to advance computer vision algorithms for real-time weed identification and spot spraying application. Mendeley Data 2023, 2. [Google Scholar] [CrossRef]
- Zhang, Z.; Lu, Y.Z.; Zhao, Y.Q.; Pan, Q.M.; Jin, K.; Xu, G.; Hu, Y.G. TS-YOLO: An All-Day and Lightweight Tea Canopy Shoots Detection Model. Agronomy 2023, 13, 1411. [Google Scholar] [CrossRef]
- Wang, Y.; Jia, W.; Dai, S.; Ou, M.; Dong, X.; Wang, G.; Gao, B.; Tu, D. Analytical Methods for Wind-Driven Dynamic Behavior of Pear Leaves (Pyrus pyrifolia). Agriculture 2025, 15, 886. [Google Scholar] [CrossRef]
- Zhang, Q.; Chen, Q.S.; Xu, L.Z.; Xu, X.Q.; Liang, Z.W. Wheat Lodging Direction Detection for Combine Harvesters Based on Improved K-Means and Bag of Visual Words. Agronomy 2023, 13, 2227. [Google Scholar] [CrossRef]
- Zhuang, X.B.; Li, Y.M. Segmentation and Angle Calculation of Rice Lodging during Harvesting by a Combine Harvester. Agriculture 2023, 13, 1425. [Google Scholar] [CrossRef]
- Tan, M.; Pang, R.; Le, Q.V. EfficientDet: Scalable and efficient object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 10778–10787. [Google Scholar] [CrossRef]
- Wang, Z.; Li, C.; Xu, H.; Zhu, X.; Li, H. Mamba YOLO: A Simple Baseline for Object Detection with State Space Model. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 26–27 February 2024. [Google Scholar]
- Tao, T.; Wei, X.H. STBNA-YOLOv5: An Improved YOLOv5 Network for Weed Detection in Rapeseed Field. Agriculture 2025, 15, 22. [Google Scholar] [CrossRef]
- Badgujar, C.M.; Poulose, A.; Gan, H. Agricultural object detection with You Only Look Once (YOLO) Algorithm: A bibliometric and systematic literature review. Comput. Electron. Agric. 2024, 223, 109090. [Google Scholar] [CrossRef]
- Ma, Z.; Zhang, N.; Wang, S.; Li, Y.M.; Pan, Y.; Zhang, J.Q.; Liu, C.; Gao, H.Y. A lightweight real-time potato damage detection method based on improved YOLOv8. J. Food Meas. Charact. 2025, 19, 5931–5945. [Google Scholar] [CrossRef]
- Lv, R.Y.; Hu, J.P.; Zhang, T.F.; Chen, X.X.; Liu, W. Crop-Free-Ridge Navigation Line Recognition Based on the Lightweight Structure Improvement of YOLOv8. Agriculture 2025, 15, 942. [Google Scholar] [CrossRef]
- Jegham, N.; Koh, C.Y.; Abdelatti, M.; Hendawi, A. YOLO Evolution: A Comprehensive Benchmark and Architectural Review of YOLOv12, YOLO11, and Their Previous Versions. arXiv 2024, arXiv:2411.00201. [Google Scholar]
- Wang, C.Y.; Yeh, J.; Liao, H.Y.M. YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. In Proceedings of the 18th European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 1–21. [Google Scholar]
- Rout, A.; Champatiray, C.; Mahanta, G.B.; Sahni, R.S.; Aggarwal, T.; Maji, K.; Bahubalendruni, M.R. Autonomous green vegetable growth monitoring via YOLOv9 and a vine robot with tracked mobility. Expert Syst. Appl. 2026, 321, 132165. [Google Scholar] [CrossRef]
- Guan, S.; Lin, Y.; Lin, G.; Su, P.; Huang, S.; Meng, X.; Liu, P.; Yan, J. Real-Time Detection and Counting of Wheat Spikes Based on Improved YOLOv10. Agronomy 2024, 14, 1936. [Google Scholar] [CrossRef]
- Wang, A.; Chen, H.; Liu, L.H.; Chen, K.; Lin, Z.J.; Han, J.G.; Ding, G.G. YOLOv10: Real-Time End-to-End Object Detection. In Proceedings of the 2024 38th Conference on Neural Information Processing Systems-NeurIPS, Vancouver, BC, Canada, 10–15 December 2024. [Google Scholar]
- Li, A.; Wang, C.R.; Ji, T.T.; Wang, Q.Y.; Zhang, T.X. D3-YOLOv10: Improved YOLOv10-Based Lightweight Tomato Detection Algorithm Under Facility Scenario. Agriculture 2024, 14, 2268. [Google Scholar] [CrossRef]
- Sapkota, R.; Karkee, M. Comparing YOLOv11 and YOLOv8 for instance segmentation of occluded and non-occluded immature green fruits in complex orchard environment. arXiv 2024, arXiv:2410.19869. [Google Scholar]
- Jiang, L.L.; Wang, Y.F.; Yan, H.H.; Yin, Y.Z.; Wu, C. Strawberry Fruit Deformity Detection and Symmetry Quantification Using Deep Learning and Geometric Feature Analysis. Horticulturae 2025, 11, 652. [Google Scholar] [CrossRef]
- Tian, Y.; Ye, Q.; Doermann, D. YOLOv12: Attention-Centric Real-Time Object Detectors. Adv. Neural Inf. Process. Syst. 2026, 38, 78433–78457. [Google Scholar]
- Kim, J.; Kim, G.; Yoshitoshi, R.; Tokuda, K. Real-Time Object Detection for Edge Computing-Based Agricultural Automation: A Case Study Comparing the YOLOX and YOLOv12 Architectures and Their Performance in Potato Harvesting Systems. Sensors 2025, 25, 4586. [Google Scholar] [CrossRef]
- Lv, W.; Xu, S.; Zhao, Y.; Wang, G.; Wei, J.; Cui, C.; Du, Y.; Dang, Q.; Liu, Y. DETRs Beat YOLOs on Real-time Object Detection. arXiv 2023, arXiv:2304.08069. [Google Scholar]
- Shehzadi, T.; Hashmi, K.A.; Liwicki, M.; Stricker, D.; Afzal, M.Z. Object Detection with Transformers: A Review. Sensors 2025, 25, 6025. [Google Scholar] [CrossRef]
- Hadi, S.J.; Ahmed, I.; Iqbal, A.; Alzahrani, A.S. CATR: CNN augmented transformer for object detection in remote sensing imagery. Sci. Rep. 2025, 15, 42281. [Google Scholar] [CrossRef]
- Mehdipour, S.; Mirroshandel, S.A.; Tabatabaei, S.A. Vision transformers in precision agriculture: A comprehensive survey. Intell. Syst. Appl. 2026, 29, 200617. [Google Scholar] [CrossRef]
- Zhang, X.; Mamat, S.; Liu, X.; Liu, J.; Liu, R.; Wu, G.; Zhu, P.; Li, H.; Ma, M.; Liu, X. MSRRT-DETR: A high-precision apple detection method with strong cross-domain generalization capability in complex orchard scenes. PLoS ONE 2026, 21, e0342854. [Google Scholar] [CrossRef]
- Robinson, I.; Robicheaux, P.; Popov, M.; Ramanan, D.; Peri, N. RF-DETR: Neural Architecture Search for Real-Time Detection Transformers. arXiv 2025, arXiv:2511.09554. [Google Scholar]
- Sapkota, R.; Cheppally, R.H.; Sharda, A.; Karkee, M. RF-DETR Object Detection vs YOLOv12: A Study of Transformer-based and CNN-based Architectures for Single-Class and Multi-Class Greenfruit Detection in Complex Orchard Environments Under Label Ambiguity. arXiv 2025, arXiv:2504.13099. [Google Scholar]
- Gu, A.; Dao, T. Mamba: Linear-time sequence modeling with selective state spaces. arXiv 2023, arXiv:2312.00752. [Google Scholar]
- Mamun, A.A.; Zhang, M.; Aristizabal, D.A.; Hayder, Z.; Awrangjeb, M. ConMamba: Contrastive vision Mamba for plant disease detection. Pattern Recognit. 2026, 176, 113177. [Google Scholar] [CrossRef]
- Xia, X.; Zhang, N.; Guan, Z.; Chai, X.; Ma, S.; Chai, X.; Sun, T. PAB-Mamba-YOLO: VSSM assists in YOLO for aggressive behavior detection among weaned piglets. Artif. Intell. Agric. 2025, 15, 52–66. [Google Scholar] [CrossRef]
- Bao, M.; Lyu, S.; Xu, Z.; Zhou, H.; Ren, J.; Xiang, S.; Li, X.; Cheng, G. Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook. Remote Sens. 2026, 18, 594. [Google Scholar] [CrossRef]
- You, S.C.; Li, B.H.; Chen, Y.J.; Ren, Z.Y.; Liu, Y.Y.; Wu, Q.Y.; Tao, J.H.; Zhang, Z.J.; Zhang, C.Y.; Xue, F.; et al. Rose-Mamba-YOLO: An enhanced framework for efficient and accurate greenhouse rose monitoring. Front. Plant Sci. 2025, 16. [Google Scholar] [CrossRef]
- Yuan, C.; Li, S.; Wang, K.; Liu, Q.; Li, W.; Zhao, W.; Guo, G.; Wei, L. Mamba-YOLO-ML: A State-Space Model-Based Approach for Mulberry Leaf Disease Detection. Plants 2025, 14, 2084. [Google Scholar] [CrossRef] [PubMed]
- Allmendinger, A.; Saltık, A.O.; Peteinatos, G.G.; Stein, A.; Gerhards, R. Assessing the capability of YOLO- and transformer-based object detectors for real-time weed detection. Precis. Agric. 2025, 26, 52. [Google Scholar] [CrossRef]
- Khan, Z.; Shen, Y.; Liu, H. ObjectDetection in Agriculture: A Comprehensive Review of Methods, Applications, Challenges, and Future Directions. Agriculture 2025, 15, 1351. [Google Scholar] [CrossRef]
- Guo, M.-H.; Xu, T.-X.; Liu, J.-J.; Liu, Z.-N.; Jiang, P.-T.; Mu, T.-J.; Zhang, S.-H.; Martin, R.R.; Cheng, M.-M.; Hu, S.-M. Attention mechanisms in computer vision: A survey. Comput. Vis. Media 2022, 8, 331–368. [Google Scholar] [CrossRef]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate attention for efficient mobile network design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 13713–13722. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the European conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Syed, T.N.; Zhou, J.; Lakhiar, I.A.; Marinello, F.; Gemechu, T.T.; Rottok, L.T.; Jiang, Z.Z. Enhancing Autonomous Orchard Navigation: A Real-Time Convolutional Neural Network-Based Obstacle Classification System for Distinguishing ‘Real’ and ‘Fake’ Obstacles in Agricultural Robotics. Agriculture 2025, 15, 827. [Google Scholar] [CrossRef]
- Zhao, S.Y.; Liu, J.Z.; Hua, T.Z.; Jiang, Y. Improved UNet Recognition Model for Multiple Strawberry Pests Based on Small Samples. Agronomy 2025, 15, 2252. [Google Scholar] [CrossRef]
- Zhao, S.Y.; Peng, Y.; Liu, J.Z.; Wu, S. Tomato Leaf Disease Diagnosis Based on Improved Convolution Neural Network by Attention Module. Agriculture 2021, 11, 651. [Google Scholar] [CrossRef]
- Wang, X.; Hu, C.; Wang, X.; Zha, H.; Chen, X.; Yuan, S.; Zhang, J.; Liao, J.; Ye, Z. Research on multi class pests identification and detection based on fusion attention mechanism with Mask-RCNN-CBAM. Front. Agron. 2025, 7, 1578412. [Google Scholar] [CrossRef]
- Pei, H.T.; Sun, Y.Q.; Huang, H.; Zhang, W.; Sheng, J.J.; Zhang, Z.Y. Weed Detection in Maize Fields by UAV Images Based on Crop Row Preprocessing and Improved YOLOv4. Agriculture 2022, 12, 975. [Google Scholar] [CrossRef]
- Wu, Y.; Yuan, S.; Tang, Y.; Tang, L. Application of real-time detection transformer based on convolutional block attention module and grouped convolution in maize seedling. Front. Plant Sci. 2025, 16, 1672746. [Google Scholar] [CrossRef]
- Li, S.; Chen, Z.; Xie, J.; Zhang, H.; Guo, J. PD-YOLO: A novel weed detection method based on multi-scale feature fusion. Front. Plant Sci. 2025, 16, 1506524. [Google Scholar] [CrossRef]
- Ouyang, D.; He, S.; Zhang, G.; Luo, M.; Guo, H.; Zhan, J.; Huang, Z. Efficient multi-scale attention module with cross-spatial learning. In Proceedings of the ICASSP 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes, Greece, 4–9 June 2023; pp. 1–5. [Google Scholar]
- Dai, X.; Chen, Y.; Xiao, B.; Chen, D.; Liu, M.; Yuan, L.; Zhang, L. Dynamic head: Unifying object detection heads with attentions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 7373–7382. [Google Scholar]
- Hong, W.; Ling, S.; Zhu, P.; Wang, Z.; Zhao, R.; Liu, Y.; Dong, M. Lightweight Vision–Transformer Network for Early Insect Pest Identification in Greenhouse Agricultural Environments. Insects 2026, 17, 74. [Google Scholar] [CrossRef]
- Zuo, X.; Chu, J.; Shen, J.F.; Sun, J. Multi-Granularity Feature Aggregation with Self-Attention and Spatial Reasoning for Fine-Grained Crop Disease Classification. Agriculture 2022, 12, 1499. [Google Scholar] [CrossRef]
- Jaderberg, M.; Simonyan, K.; Zisserman, A. Spatial transformer networks. Adv. Neural Inf. Process. Syst. 2015, 28. [Google Scholar] [CrossRef]
- Bookstein, F.L. Principal warps: Thin-plate splines and the decomposition of deformations. IEEE Trans. Pattern Anal. Mach. Intell. 2002, 11, 567–585. [Google Scholar] [CrossRef]
- Zhu, L.; Wang, X.; Ke, Z.; Zhang, W.; Lau, R.W. Biformer: Vision transformer with bi-level routing attention. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 18–22 June 2023; pp. 10323–10333. [Google Scholar]
- Chen, B.-J.; Bu, J.-Y.; Xia, J.-L.; Li, M.-X.; Su, W.-H. AFBF-YOLO: An Improved YOLO11n Algorithm for Detecting Bunch and Maturity of Cherry Tomatoes in Greenhouse Environments. Plants 2025, 14, 2587. [Google Scholar] [CrossRef]
- Tang, S.X.; Xia, Z.L.; Gu, J.N.; Wang, W.B.; Huang, Z.D.; Zhang, W.H. High-precision apple recognition and localization method based on RGB-D and improved SOLOv2 instance segmentation. Front. Sustain. Food Syst. 2024, 8, 1403872. [Google Scholar] [CrossRef]
- Qin, J.; Yang, X.; Zhang, T.; Bi, S. BI-TST_YOLOv5: Ground Defect Recognition Algorithm Based on Improved YOLOv5 Model. World Electr. Veh. J. 2024, 15, 102. [Google Scholar] [CrossRef]
- Zheng, T.; Chen, Z.; Bai, J.; Xie, H.; Jiang, Y.-G. Tps++: Attention-enhanced thin-plate spline for scene text recognition. arXiv 2023, arXiv:2305.05322. [Google Scholar]
- Ma, J.; Zhao, Y.K.; Fan, W.P.; Liu, J.Z. An Improved YOLOv8 Model for Lotus Seedpod Instance Segmentation in the Lotus Pond Environment. Agronomy 2024, 14, 1325. [Google Scholar] [CrossRef]
- Tong, Z.; Chen, Y.; Xu, Z.; Yu, R. Wise-IoU: Bounding box regression loss with dynamic focusing mechanism. arXiv 2023, arXiv:2301.10051. [Google Scholar]
- Rezatofighi, H.; Tsoi, N.; Gwak, J.; Sadeghian, A.; Reid, I.; Savarese, S. Generalized intersection over union: A metric and a loss for bounding box regression. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; pp. 658–666. [Google Scholar]
- Zheng, Z.; Wang, P.; Liu, W.; Li, J.; Ye, R.; Ren, D. Distance-IoU loss: Faster and better learning for bounding box regression. In Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA, 7–12 February 2020; pp. 12993–13000. [Google Scholar]
- He, J.; Erfani, S.; Ma, X.; Bailey, J.; Chi, Y.; Hua, X.-S. α-IoU: A family of power intersection over union losses for bounding box regression. Adv. Neural Inf. Process. Syst. 2021, 34, 20230–20242. [Google Scholar]
- Chu, X.; Zheng, A.; Zhang, X.; Sun, J. Detection in crowded scenes: One proposal, multiple predictions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–18 June 2020; pp. 12214–12223. [Google Scholar]
- Wang, X.; Xiao, T.; Jiang, Y.; Shao, S.; Sun, J.; Shen, C. Repulsion loss: Detecting pedestrians in a crowd. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 7774–7783. [Google Scholar]
- Wang, J.; Xu, C.; Yang, W.; Yu, L. A normalized Gaussian Wasserstein distance for tiny object detection. arXiv 2021, arXiv:2110.13389. [Google Scholar]
- Yang, X.; Yan, J.; Ming, Q.; Wang, W.; Zhang, X.; Tian, Q. Rethinking rotated object detection with gaussian wasserstein distance loss. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 11830–11841. [Google Scholar]
- Wang, G.; Chen, Y.; An, P.; Hong, H.; Hu, J.; Huang, T. UAV-YOLOv8: A small-object-detection model based on improved YOLOv8 for UAV aerial photography scenarios. Sensors 2023, 23, 7190. [Google Scholar] [CrossRef]
- Zhang, H.; Zhang, S. Focaler-iou: More focused intersection over union loss. arXiv 2024, arXiv:2401.10525. [Google Scholar]
- Zhang, H.; Xu, C.; Zhang, S. Inner-IoU: More effective intersection over union loss with auxiliary bounding box. arXiv 2023, arXiv:2311.02877. [Google Scholar]
- Zheng, S.; Jia, X.; He, M.; Zheng, Z.; Lin, T.; Weng, W. Tomato recognition method based on the YOLOv8-Tomato model in complex greenhouse environments. Agronomy 2024, 14, 1764. [Google Scholar] [CrossRef]
- Song, J.; Cheng, K.; Chen, F.; Hua, X. RDW-YOLO: A deep learning framework for scalable agricultural pest monitoring and control. Insects 2025, 16, 545. [Google Scholar] [CrossRef] [PubMed]
- Zhang, W.; Chen, K.; Wang, J.; Shi, Y.; Guo, W. Easy domain adaptation method for filling the species gap in deep learning-based fruit detection. Hortic. Res. 2021, 8. [Google Scholar] [CrossRef] [PubMed]
- Yin, S.; Xi, Y.; Zhang, X.; Sun, C.; Mao, Q. Foundation Models in Agriculture: A Comprehensive Review. Agriculture 2025, 15, 847. [Google Scholar] [CrossRef]
- Li, J.; Xu, M.; Xiang, L.; Chen, D.; Zhuang, W.; Yin, X.; Li, Z. Foundation models in smart agriculture: Basics, opportunities, and challenges. Comput. Electron. Agric. 2024, 222, 109032. [Google Scholar] [CrossRef]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.-Y. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 2–6 October 2023; pp. 4015–4026. [Google Scholar]
- Ravi, N.; Gabeur, V.; Hu, Y.-T.; Hu, R.; Ryali, C.; Ma, T.; Khedr, H.; Rädle, R.; Rolland, C.; Gustafson, L. Sam 2: Segment anything in images and videos. arXiv 2024, arXiv:2408.00714. [Google Scholar]
- Zhao, S.; Liu, J.; Wu, S. Multiple disease detection method for greenhouse-cultivated strawberry based on multiscale feature fusion Faster R_CNN. Comput. Electron. Agric. 2022, 199, 107176. [Google Scholar] [CrossRef]
- Williams, D.; Macfarlane, F.; Britten, A. Leaf only SAM: A segment anything pipeline for zero-shot automated leaf segmentation. Smart Agric. Technol. 2024, 8, 100515. [Google Scholar] [CrossRef]
- Zhang, C.; Han, D.; Qiao, Y.; Kim, J.U.; Bae, S.-H.; Lee, S.; Hong, C.S. Faster segment anything: Towards lightweight sam for mobile applications. arXiv 2023, arXiv:2306.14289. [Google Scholar] [CrossRef]
- Zhou, C.; Li, X.; Loy, C.C.; Dai, B. Edgesam: Prompt-in-the-loop distillation for on-device deployment of sam. arXiv 2023, arXiv:2312.06660. [Google Scholar]
- Oehme, L.H.; Boysen, J.; Wu, Z.; Stein, A.; Müller, J. Orchestrating segment anything models to accelerate segmentation annotation on agricultural image datasets. Front. Artif. Intell. 2026, 8, 1748468. [Google Scholar] [CrossRef] [PubMed]
- Sun, X.; Liu, J.; Shen, H.; Zhu, X.; Hu, P. On Efficient Variants of Segment Anything Model: A Survey. Int. J. Comput. Vis. 2025, 133, 7406–7436. [Google Scholar] [CrossRef]
- Wang, Y.; Fei, Z.; Li, R.; Ying, Y. Learn from foundation model: Fruit detection model without manual annotation. Pattern Recognit. 2025, 174, 112799. [Google Scholar] [CrossRef]
- Singh, R.; Puhl, R.B.; Dhakal, K.; Sornapudi, S. Few-Shot Adaptation of Grounding DINO for Agricultural Domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
- Zhang, Y.; Shao, Y.; Tang, C.; Liu, Z.; Li, Z.; Zhai, R.; Peng, H.; Song, P. E-clip: An enhanced clip-based visual language model for fruit detection and recognition. Agriculture 2025, 15, 1173. [Google Scholar] [CrossRef]
- Liu, S.; Zeng, Z.; Ren, T.; Li, F.; Zhang, H.; Yang, J.; Jiang, Q.; Li, C.; Yang, J.; Su, H. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 38–55. [Google Scholar]
- Dong, X.; Zhang, K.; Nong, Q.; Ju, M.; Tu, Y. Empowering Grounding DINO with MoE Empowering Grounding DINO with MoE: An End-to-End Framework for End Framework for Cross-Domain Few Domain Few-Shot Object Detection Shot Object Detection. ZTE Commun. 2025, 23, 77–85. [Google Scholar] [CrossRef]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. Iclr 2022, 1, 3. [Google Scholar]
- Houlsby, N.; Giurgiu, A.; Jastrzebski, S.; Morrone, B.; De Laroussilhe, Q.; Gesmundo, A.; Attariyan, M.; Gelly, S. Parameter-efficient transfer learning for NLP. In Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA, 9–15 July 2019; pp. 2790–2799. [Google Scholar]
- Jiang, X.; Liu, Y.; Debbagh, M.; Tian, Y.; Hoyos-Villegas, V.; Adamchuk, V.; Sun, S. Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions. arXiv 2025, arXiv:2509.25805. [Google Scholar]
- Bao, Q.-Z.; Yang, Y.-X.; Li, Q.; Yang, H.-C. Zero-shot instance segmentation for plant phenotyping in vertical farming with foundation models and VC-NMS. Front. Plant Sci. 2025, 16, 1536226. [Google Scholar] [CrossRef]
- Devanna, R.P.; Reina, G.; Cheein, F.A.; Milella, A. Boosting grape bunch detection in RGB-D images using zero-shot annotation with Segment Anything and GroundingDINO. Comput. Electron. Agric. 2025, 229, 109611. [Google Scholar] [CrossRef]
- Yang, Z.-X.; Li, Y.; Wang, R.-F.; Hu, P.; Su, W.-H. Deep learning in multimodal fusion for sustainable plant care: A comprehensive review. Sustainability 2025, 17, 5255. [Google Scholar] [CrossRef]
- Zhu, H.; Qin, S.; Su, M.; Lin, C.; Li, A.; Gao, J. Harnessing large vision and language models in agriculture: A review. Front. Plant Sci. 2025, 16, 1579355. [Google Scholar] [CrossRef] [PubMed]
- Sun, J.; Jiang, S.Y.; Mao, H.P.; Wu, X.H.; Li, Q.L. Classification of Black Beans Using Visible and Near Infrared Hyperspectral Imaging. Int. J. Food Prop. 2016, 19, 1687–1695. [Google Scholar] [CrossRef]
- Ban, C.; Wang, L.; Su, T.; Chi, R.; Fu, G. Fusion of monocular camera and 3D LiDAR data for navigation line extraction under corn canopy. Comput. Electron. Agric. 2025, 232, 110124. [Google Scholar] [CrossRef]
- Wei, L.L.; Yang, H.S.; Niu, Y.X.; Zhang, Y.N.; Xu, L.Z.; Chai, X.Y. Wheat biomass, yield, and straw-grain ratio estimation from multi-temporal UAV-based RGB and multispectral images. Biosyst. Eng. 2023, 234, 187–205. [Google Scholar] [CrossRef]
- Ruigrok, T.; Henten, E.J.v.; Kootstra, G. Stereo Vision for Plant Detection in Dense Scenes. Sensors 2024, 24, 1942. [Google Scholar] [CrossRef]
- Gauba, A.; Pi, I.; Man, Y.; Pang, Z.; Adve, V.S.; Wang, Y.X. AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark. Adv. Neural Inf. Process. Syst. 2025, 38. [Google Scholar] [CrossRef]
- Wang, Y.; Wang, F.; Chen, W.; Lv, B.; Liu, M.; Kong, X.; Zhao, C.; Pan, Z. A large language model for multimodal identification of crop diseases and pests. Sci. Rep. 2025, 15, 21959. [Google Scholar] [CrossRef]
- Chenzi, Z.; Xiaoyan, M.; Yu, L.; Shuaiqi, Y. MKG-CottonCapT6: A Multimodal Knowledge Graph-Enhanced Image Captioning Framework for Expert-Level Cotton Disease and Pest Diagnosis. Appl. Sci. 2026, 16, 3029. [Google Scholar]
- Joshi, A.; Guevara, D.; Earles, M. Standardizing and centralizing datasets for efficient training of agricultural deep learning models. Plant Phenomics 2023, 5, 0084. [Google Scholar] [CrossRef]
- Singh, S.; Yadav, A.; Jain, J.; Shi, H.; Johnson, J.; Desai, K. Benchmarking object detectors with coco: A new path forward. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 279–295. [Google Scholar]
- Islam, M.D.; Liu, W.; Izere, P.; Singh, P.; Yu, C.; Riggan, B.; Zhang, K.; Jhala, A.J.; Knezevic, S.; Ge, Y. Towards real-time weed detection and segmentation with lightweight CNN models on edge devices. Comput. Electron. Agric. 2025, 237, 110600. [Google Scholar] [CrossRef]
- El Jarroudi, M.; Kouadio, L.; Delfosse, P.; Bock, C.H.; Mahlein, A.-K.; Fettweis, X.; Mercatoris, B.; Adams, F.; Lenné, J.M.; Hamdioui, S. Leveraging edge artificial intelligence for sustainable agriculture. Nat. Sustain. 2024, 7, 846–854. [Google Scholar] [CrossRef]
- Li, K.; Chen, K.; Wang, H.; Hong, L.; Ye, C.; Han, J.; Chen, Y.; Zhang, W.; Xu, C.; Yeung, D.-Y. Coda: A real-world road corner case dataset for object detection in autonomous driving. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 406–423. [Google Scholar]
- Wang, L.; Yang, J.; Zhang, Y.; Wang, F.; Zheng, F. Depth-aware concealed crop detection in dense agricultural scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 17–21 June 2024; pp. 17201–17211. [Google Scholar]
- Liu, D.; Wang, P.; Zhang, Z.; Lu, Y.; Wang, B.; Hu, Y. Robust detection of dense small tea shoots across cultivars under occlusion and bud–leaf similarity for intelligent selective harvesting. Sci. Hortic. 2025, 353, 114499. [Google Scholar] [CrossRef]
- Mohammadzadeh Babr, M.; Faghihabdolahi, M.; Ristić-Durrant, D.; Michels, K. Deep learning-based occlusion handling of overlapped plants for robotic grasping. Appl. Sci. 2022, 12, 3655. [Google Scholar] [CrossRef]
- Lu, Y.; Young, S. A survey of public datasets for computer vision tasks in precision agriculture. Comput. Electron. Agric. 2020, 178, 105760. [Google Scholar] [CrossRef]






| Physical Factor | Feature Degradation Pattern | Representative Agricultural Scenarios | Typical Benchmark Datasets | Suitable Evaluation Metrics |
|---|---|---|---|---|
| Dense Occlusion |
|
|
| |
| Small-Target Detection |
|
|
| |
| Morphological Similarity |
|
|
| |
| Background Clutter |
|
|
| |
| Illumination Variation |
|
|
| |
| Non-Rigid Deformation |
|
|
|
| Versions | Precision | mAP50 | Inference Time | GFLOPs |
|---|---|---|---|---|
| YOLOv8 | 0.749 | 0.777 | 6.8 | 8.1 |
| YOLOv9 | 0.792 | 0.812 | 10 | 7.7 |
| YOLOv10 | 0.722 | 0.722 | 0.8 | 8.3 |
| YOLOv11 | 0.768 | 0.757 | 0.6 | 6.4 |
| YOLOv12 | 0.883 | 0.802 | 4.6 | 6.4 |
| Model Architecture Category | Representative Models | Core Structural Innovations | Official Repository/Code | Advantages in Agricultural Field Dense Occlusion Scenarios |
|---|---|---|---|---|
| Single-stage Convolutional Networks (CNNs) | YOLOv10, YOLOv11 | NMS-free, C3k2, C2PSA, Consistent Dual Assignment | https://github.com/THU-MIG/yolov10 (accessed on 24 April 2026), https://github.com/ultralytics/ultralytics (accessed on 24 April 2026) | Extremely high computational efficiency; effectively avoids the erroneous deletion of overlapping clustered fruits by NMS; cross-stage attention enhances local feature representation [13,79,80,82,83,84,85,86,87,88,89,90]. |
| Single-stage Attention Networks | YOLOv12 | Area Attention (A2), R-ELAN, FlashAttention | https://github.com/sunsmarterjie/yolov12 (accessed on 24 April 2026) | While maintaining CNN-like low latency, it significantly expands the receptive field through regional division mechanisms, facilitating the inferential mining of weak features and occluded targets [17,34,91,92]. |
| Real-time Transformers | RT-DETR, RF-DETR | Efficient Hybrid Encoder, Contrastive Denoising Training, Decoupled Fusion | https://github.com/roboflow/rf-detr (accessed on 27 April 2026) | Completely eliminates reliance on heuristic NMS; built-in global self-attention can effectively overcome information gaps caused by local rigid occlusion [39,93,94,95,96,97,98,99]. |
| State Space Models (SSMs) | ROSE-MAMBA-YOLO | Hardware-aware linear sequence modeling, Visual selective memory | No official public repository identified | Achieves global context association with extremely low memory overhead and linear computational complexity, effectively handling extreme scale variations and severe multi-layer geometric overlapping [100,101,102,103,104,105]. |
| Attention Mechanism Category | Typical Representatives | Core Computational Logic | Addressed Agricultural Vision Pain Points | Limitations and Computational Cost |
|---|---|---|---|---|
| Global Recalibration | SE, CBAM | Relies on global pooling to compute channel/spatial weight matrices | Effectively suppresses large-area macroscopic background noise such as soil and sky [110,111,112,113,114,115,116]. | Single aggregation method, making it difficult to finely delineate the mutual occlusion boundaries of dense targets; extremely low computational overhead [108,117]. |
| Multi-scale and Deformation Perception | EMA, CSLA, TPS-Attention | Relies on global pooling to compute channel/spatial weight matrices | Fuses shallow edges and deep semantics to reconstruct fractured features; actively adapts to non-rigid leaf deformations caused by wind [61,118,119,120,121,122,123,124]. | Requires constructing complex parallel branches, resulting in cumbersome structural design; slightly increases inference latency [119,129]. |
| Dynamic Sparse Routing | BiFormer (BRA), Area Attention | Combines content-aware coarse-grained screening with intra-region fine-grained self-attention | Effectively suppresses interference from large-area low-value backgrounds, enabling adaptive focusing of computational power on dense, clustered fruits or tiny lesion regions [91,125,126]. | High engineering implementation complexity, requiring highly optimized underlying operator support (e.g., FlashAttention) [128]. |
| Loss Function Series | Core Mathematical Mechanism | Engineering Advantages in Agricultural Dense Occlusion Scenarios |
|---|---|---|
| NWD | Models bounding boxes as 2D Gaussian distributions and calculates the Wasserstein distance between distributions [137]. | Reduces scale sensitivity; provides smooth gradients under zero overlap or slight offsets, serving as a useful option for tiny crop detection from UAV high-altitude perspectives [138]. |
| WIoUv3 | Based on a dynamic non-monotonic focusing mechanism, it dynamically evaluates the “outlier degree” of samples [139]. | Intelligently isolates “low-quality noise samples” with extreme severe occlusion or annotation errors, preventing the model from collapsing due to force-fitting noise, and improving generalizability [143]. |
| Focaler-IoU | Reconstructs the numerical space of IoU based on linear interval mapping [140]. | Dynamically amplifies the gradient weights of hard examples (e.g., similar colors, deep occlusion), enhancing the model’s bounding box localization capability under complex illumination and cluttered backgrounds [140]. |
| Inner-IoU | Introduces an auxiliary inner bounding box with a scaling factor to focus on the core of the overlap [141]. | Circumvents the edge blur defects caused by occlusion, prioritizes aligning the clear internal regions of the target, may improve convergence behavior in dense clustered fruit environments, and possesses multi-dimensional synergistic optimization capabilities [142]. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yan, M.; Sun, Z.; Xu, Y.; Gong, C.; Kang, C. A Review of Lightweight Object Detection Technologies for Densely Occluded Scenarios in Agricultural Fields. Agronomy 2026, 16, 1059. https://doi.org/10.3390/agronomy16111059
Yan M, Sun Z, Xu Y, Gong C, Kang C. A Review of Lightweight Object Detection Technologies for Densely Occluded Scenarios in Agricultural Fields. Agronomy. 2026; 16(11):1059. https://doi.org/10.3390/agronomy16111059
Chicago/Turabian StyleYan, Mingzhi, Zeyu Sun, Yijun Xu, Chen Gong, and Can Kang. 2026. "A Review of Lightweight Object Detection Technologies for Densely Occluded Scenarios in Agricultural Fields" Agronomy 16, no. 11: 1059. https://doi.org/10.3390/agronomy16111059
APA StyleYan, M., Sun, Z., Xu, Y., Gong, C., & Kang, C. (2026). A Review of Lightweight Object Detection Technologies for Densely Occluded Scenarios in Agricultural Fields. Agronomy, 16(11), 1059. https://doi.org/10.3390/agronomy16111059

