YieldNet: A Lightweight YOLOv8n Enhancement for Immature Green Tomato Detection in UAV Images: Real-Time Edge Demonstration Toward Pre-Harvest Yield Estimation
Abstract
1. Introduction
- We construct a dedicated greenhouse immature green tomato dataset captured by low-altitude UAVs, providing a realistic and challenging benchmark for early-stage fruit detection.
- We modify YOLOv8n by replacing its backbone with ShuffleNetV2, integrating ECA modules in the neck, and adopting PIoU v2 loss. Under the present protocol, the final model reports higher validation metrics than YOLOv8n, with 0.3 M more parameters and 0.1 G fewer FLOPs.
- We demonstrate YieldNet on an RK3588-based Orange Pi 5 Max, where the live UVC pipeline provides countable detections for future pre-harvest yield-estimation studies.
2. Materials and Methods
2.1. Dataset and Data Collection
- GreenTomato-UAV Dataset (Self-Built): Images were collected in a greenhouse at the China Nongjiyuan Agricultural Machinery Test Station (shown in Figure 1), 577M+7X2, Changping District, Beijing, China, 102206, on 19 July 2025, between 16:30 and 18:00. An F450 quadcopter equipped with a 640 × 480-pixel camera was flown at 0.5–0.8 m above the canopy (this height range corresponds to the area where tomato fruits are most densely concentrated in this greenhouse’s tomato cultivation). The 1.5 h interval was the total collection session, not a recorded continuous flight; per-sortie flight times and battery-discharge data were not logged. The UAV traversed between planting beds following the serpentine pattern illustrated in Figure 2 for image acquisition. Approximately 80,000 raw frames were captured. Scenes included soil, stakes, and weeds; most targets were green, unripe tomatoes exhibiting frequent overlapping, leaf occlusion, dense planting, and motion blur caused by UAV movement. Figure 3 illustrates those key challenges encountered in the dataset. An open-source annotation tool, Labelme, was employed, with five annotators independently labeling objects using rectangular bounding boxes. After manual review and cleaning, 600 images were retained (100 without targets). The target-free images were retained as negative samples.
- Tomato-Recog Public Dataset: A publicly available tomato dataset consisting of 1986 images covering green, turning, and red tomatoes was used for an external multi-ripeness benchmark. We divided all images into 1390 training images and 596 validation images at a 7:3 ratio; the latter is termed the Tomato-Recog public validation set below. The dataset includes official YOLO-format annotations and is fully open-source.
2.2. Incremental Data Augmentation
- Proportional Sampling: In each augmentation round, the number of images to be augmented is calculated as a fixed proportion (20%) of the original dataset size, while images are randomly sampled from the current dataset (i.e., the master set including all previously augmented images). This procedure permits cumulative combinations of transformations while expanding the training set.
- Augmentation Operations: Each sampled image undergoes three successive augmentation steps:
- Motion Blur: Kernel size of 15, with the blur angle randomly chosen from .
- Affine Rotation: Rotation angle randomly selected from the range .
- Lighting Adjustment: Brightness and contrast scaling factors randomly set within .
- Dataset Management: For geometric transformations, bounding-box coordinates were updated with the images. Filenames were appended with the augmentation-round suffix (e.g., _r1, _r2), and each augmented set was merged into the mother set before the next round.
- Cumulative Effect: Following three augmentation rounds, the overall dataset size grew by about 60% relative to the original, providing additional training examples for motion blur, rotation, and illumination variation. Color jittering and random cropping were not evaluated in this study. Mosaic augmentation was enabled during training and disabled for the final 10 epochs (close_mosaic=10).
2.3. Improvement in YieldNet Network
2.3.1. Improvement in Backbone Network
2.3.2. Improvement in Neck Network
2.3.3. Improvement in Loss Function
2.4. Experimental Environment and Model Evaluation Indicators
2.4.1. Experimental Environment
- Input resolution: pixels;
- Training epochs: 150;
- Batch size: 8;
- Optimizer: SGD (initial learning rate 0.01, momentum 0.937);
- Single-class mode (single_cls=True) for tomato detection.
2.4.2. Model Evaluation Indicators
- F1 Score:tallied at IoU ≥ 0.5.
- mAP@50:Derived from the precision-recall curve area per category:then averaged across C classes at IoU = 0.5:
- mAP@50-95:Averaged over IoU ∈ [0.50:0.05:0.95]:where t indexes the ten thresholds. This metric imposes stricter localization constraints than mAP@50.
2.4.3. Supplementary Stability Evaluation
3. Results
3.1. Comparison of Model Before and After Improvement
3.2. Ablation Experiment
3.3. Comparison Between Different Models
3.4. Training Visualization Analysis
3.5. Edge-Device Deployment and Performance
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Kitinoja, L.; AlHassan, H.Y. Identification of Appropriate Postharvest Technologies for Small Scale Horticultural Farmers and Marketers in Sub-Saharan Africa and South Asia—Part 1. Postharvest Losses and Quality Assessments. Acta Hortic. 2012, 934, 31–40. [Google Scholar] [CrossRef] [Scilit]
- Izdori, F.J.; Mkwambisi, D.; Karuaihe, S.T.; Papargyropoulou, E. Multi-stakeholder collaboration framework for post-harvest loss reduction: The case of tomato value chain in Iringa and Morogoro regional in Tanzania. Agric. Food Econ. 2025, 13, 6. [Google Scholar] [CrossRef] [Scilit]
- Shawon, S.M.; Ema, F.B.; Mahi, A.K.; Niha, F.L.; Zubair, H.T. Crop yield prediction using machine learning: An extensive and systematic literature review. Smart Agric. Technol. 2025, 10, 100718. [Google Scholar] [CrossRef] [Scilit]
- Tsouros, D.C.; Bibi, S.; Sarigiannidis, P.G. A Review on UAV-Based Applications for Precision Agriculture. Information 2019, 10, 349. [Google Scholar] [CrossRef] [Scilit]
- Radoglou-Grammatikis, P.; Sarigiannidis, P.; Lagkas, T.; Moscholios, I. A compilation of UAV applications for precision agriculture. Comput. Netw. 2020, 172, 107148. [Google Scholar] [CrossRef] [Scilit]
- Hu, P.; Zhang, R.; Yang, J.; Chen, L. Development Status and Key Technologies of Plant Protection UAVs in China: A Review. Drones 2022, 6, 354. [Google Scholar] [CrossRef] [Scilit]
- Rejeb, A.; Abdollahi, A.; Rejeb, K.; Treiblmaier, H. Drones in agriculture: A review and bibliometric analysis. Comput. Electron. Agric. 2022, 198, 107017. [Google Scholar] [CrossRef] [Scilit]
- Zhou, H.; Huang, F.; Lou, W.; Gu, Q.; Ye, Z.; Hu, H.; Zhang, X. Yield prediction through UAV-based multispectral imaging and deep learning in rice breeding trials. Agric. Syst. 2025, 223, 104214. [Google Scholar] [CrossRef] [Scilit]
- Shahi, T.B.; Xu, C.-Y.; Neupane, A.; Guo, W. Machine learning methods for precision agriculture with UAV imagery: A review. Electron. Res. Arch. 2022, 30, 4277–4317. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar] [CrossRef] [Scilit]
- Redmon, J.; Farhadi, A. YOLOv3: An Incremental Improvement. arXiv 2018, arXiv:1804.02767. [Google Scholar] [CrossRef] [Scilit]
- Sohan, M.; Ram, T.S.; Reddy, C.V.R. A Review on YOLOv8 and Its Advancements. In Data Intelligence and Cognitive Informatics; Springer: Singapore, 2024; pp. 529–545. [Google Scholar] [CrossRef] [Scilit]
- Egi, Y.; Hajyzadeh, M.; Eyceyurt, E. Drone-Computer Communication Based Tomato Generative Organ Counting Model Using YOLO V5 and Deep-Sort. Agriculture 2022, 12, 1290. [Google Scholar] [CrossRef] [Scilit]
- Wang, A.; Xu, Y.; Hu, D.; Zhang, L.; Li, A.; Zhu, Q.; Liu, J. Tomato Yield Estimation Using an Improved Lightweight YOLO11n Network and an Optimized Region Tracking-Counting Method. Agriculture 2025, 15, 1353. [Google Scholar] [CrossRef] [Scilit]
- Chai, S.; Wen, M.; Li, P.; Zeng, Z.; Tian, Y. DCFA-YOLO: A Dual-Channel Cross-Feature-Fusion Attention YOLO Network for Cherry Tomato Bunch Detection. Agriculture 2025, 15, 271. [Google Scholar] [CrossRef] [Scilit]
- Ji, W.; Pan, Y.; Xu, B.; Wang, J. A Real-Time Apple Targets Detection Method for Picking Robot Based on ShufflenetV2-YOLOX. Agriculture 2022, 12, 856. [Google Scholar] [CrossRef] [Scilit]
- Zeng, T.; Li, S.; Song, Q.; Zhong, F.; Wei, X. Lightweight tomato real-time detection method based on improved YOLO and mobile deployment. Comput. Electron. Agric. 2023, 205, 107625. [Google Scholar] [CrossRef] [Scilit]
- Chen, W.; Liu, M.; Zhao, C.; Li, X.; Wang, Y. MTD-YOLO: Multi-task deep convolutional neural network for cherry tomato fruit bunch maturity detection. Comput. Electron. Agric. 2024, 216, 108533. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Luo, T.; Lai, Y.; Liu, Y.; Kang, W. EdgeFormer-YOLO: A Lightweight Multi-Attention Framework for Real-Time Red-Fruit Detection in Complex Orchard Environments. Mathematics 2025, 13, 3790. [Google Scholar] [CrossRef] [Scilit]
- Ma, N.; Zhang, X.; Zheng, H.-T.; Sun, J. ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design. In Proceedings of the Computer Vision—ECCV 2018, Munich, Germany, 8–14 September 2018; pp. 122–138. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 11531–11539. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.; Wang, K.; Li, Q.; Zhao, F.; Zhao, K.; Ma, H. Powerful-IoU: More straightforward and faster bounding box regression loss with a nonmonotonic focusing mechanism. Neural Netw. 2024, 170, 276–284. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, X.; Zhou, X.; Lin, M.; Sun, J. ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices. In Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 6848–6856. [Google Scholar] [CrossRef] [Scilit]
- Dong, Q.; Sun, L.; Han, T.; Cai, M.; Gao, C. PestLite: A Novel YOLO-Based Deep Learning Technique for Crop Pest Detection. Agriculture 2024, 14, 228. [Google Scholar] [CrossRef] [Scilit]
- Kotthapalli, M.; Ravipati, D.; Bhatia, R. YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges. arXiv 2025, arXiv:2508.02067. [Google Scholar] [CrossRef] [Scilit]














| Component | Configuration |
|---|---|
| Operating System | Ubuntu 20.04.6 LTS (Kernel 5.15.0-139-generic) |
| CPU | 13th Gen Intel Core i9-13900HX @ 2.4 GHz (Turbo Boost 5.4 GHz) |
| GPU | NVIDIA GeForce RTX 4060 Laptop GPU 8 GB GDDR6 |
| GPU Accelerator | CUDA 12.1, cuDNN 8.7 |
| Deep Learning Framework | PyTorch 2.4.1 + cu121 |
| Hardware Configuration | 15 GB RAM, 320 GB NVMe SSD |
| Compiler | Python 3.8.20 (Anaconda environment) |
| Programming Language | Python 3.8.20 |
| Model | mAP@50 | mAP@50-95 | Precision | Recall | F1 |
|---|---|---|---|---|---|
| YOLOv8n baseline | 0.8887 ± 0.0072 | 0.4583 ± 0.0032 | 0.8727 ± 0.0230 | 0.8253 ± 0.0186 | 0.8480 ± 0.0028 |
| YieldNet | 0.9097 ± 0.0057 | 0.4730 ± 0.0036 | 0.9057 ± 0.0081 | 0.8483 ± 0.0086 | 0.8760 ± 0.0043 |
| Model | Class | mAP@50 | mAP@50-95 | Precision | Recall | F1 | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|---|---|
| YOLOv8n baseline | All | 0.926 | 0.666 | 0.921 | 0.868 | 0.894 | 3.00 | 8.1 |
| Ripe | 0.931 | 0.709 | 0.959 | 0.882 | 0.919 | 3.00 | — | |
| Semiripe | 0.906 | 0.652 | 0.919 | 0.820 | 0.867 | 3.00 | — | |
| Unripe | 0.942 | 0.637 | 0.884 | 0.901 | 0.892 | 3.00 | — | |
| YieldNet (Ours) | All | 0.970 | 0.792 | 0.972 | 0.921 | 0.946 | 3.30 | 8.0 |
| Ripe | 0.990 | 0.846 | 1.000 | 0.960 | 0.980 | 3.30 | — | |
| Semiripe | 0.951 | 0.765 | 0.953 | 0.883 | 0.917 | 3.30 | — | |
| Unripe | 0.969 | 0.763 | 0.964 | 0.921 | 0.942 | 3.30 | — |
| Model | Class | mAP@50 | mAP@50-95 | Precision | Recall | F1 | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|---|---|
| YOLOv8n baseline | All | 0.884 | 0.457 | 0.872 | 0.820 | 0.845 | 3.00 | 8.1 |
| YieldNet (Ours) | All | 0.915 | 0.473 | 0.905 | 0.855 | 0.879 | 3.30 | 8.0 |
| Model | mAP@50 | mAP@50-95 | Precision | Recall | F1 | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|---|
| YOLOv8n baseline (no PIoU) | 0.884 | 0.457 | 0.872 | 0.820 | 0.845 | 3.00 | 8.1 |
| +ShuffleNetV2 (no PIoU) | 0.912 | 0.460 | 0.904 | 0.834 | 0.867 | 3.30 | 8.0 |
| +ECA (no PIoU) | 0.905 | 0.459 | 0.893 | 0.827 | 0.858 | 3.00 | 8.1 |
| +ShuffleNetV2-ECA (no PIoU) | 0.912 | 0.475 * | 0.904 | 0.839 | 0.870 | 3.30 | 8.0 |
| +PIoU v2 | 0.896 | 0.456 | 0.862 | 0.842 | 0.851 | 3.00 | 8.1 |
| +ShuffleNetV2 + PIoU v2 | 0.908 | 0.462 | 0.896 | 0.871 * | 0.883 * | 3.30 | 8.0 |
| +ECA + PIoU v2 | 0.896 | 0.469 | 0.888 | 0.814 | 0.849 | 3.00 | 8.1 |
| YieldNet (ShuffleNetV2-ECA + PIoU v2) | 0.915 * | 0.473 | 0.905 * | 0.855 | 0.879 | 3.30 | 8.0 |
| Model | mAP@50 | mAP@50-95 | Precision | Recall | F1 | Params (M) | GFLOPs |
|---|---|---|---|---|---|---|---|
| Tomato-Recog public validation set (596 images, 2228 instances) | |||||||
| YOLOv8n (baseline) | 0.926 | 0.666 | 0.921 | 0.868 | 0.894 | 3.00 | 8.1 |
| YieldNet (Ours) | 0.970 * | 0.792 * | 0.972 * | 0.921 * | 0.946 * | 3.30 | 8.0 |
| YOLOv8s | 0.958 | 0.747 | 0.956 | 0.898 | 0.926 | 11.10 | 28.4 |
| YOLOv5n | 0.915 | 0.641 | 0.872 | 0.865 | 0.868 | 2.50 | 7.1 |
| YOLOv5s | 0.960 | 0.721 | 0.952 | 0.907 | 0.929 | 9.10 | 23.8 |
| YOLOv11n | 0.920 | 0.636 | 0.901 | 0.856 | 0.878 | 2.60 | 6.3 |
| YOLOv11s | 0.951 | 0.722 | 0.963 | 0.868 | 0.913 | 9.40 | 21.3 |
| RT-DETR-n | 0.869 | 0.633 | 0.885 | 0.844 | 0.864 | 15.50 | 37.4 |
| GreenTomato-UAV validation set (180 images, 590 instances) | |||||||
| YOLOv8n (baseline) | 0.884 | 0.457 | 0.872 | 0.820 | 0.845 | 3.00 | 8.1 |
| YieldNet (Ours) | 0.915 * | 0.473 | 0.905 | 0.855 | 0.879 * | 3.30 | 8.0 |
| YOLOv8s | 0.901 | 0.473 | 0.894 | 0.861 * | 0.877 | 11.10 | 28.4 |
| YOLOv5n | 0.898 | 0.456 | 0.849 | 0.853 | 0.851 | 2.50 | 7.1 |
| YOLOv5s | 0.900 | 0.474 * | 0.910 * | 0.824 | 0.865 | 9.10 | 23.8 |
| YOLOv11n | 0.897 | 0.453 | 0.882 | 0.834 | 0.857 | 2.60 | 6.3 |
| YOLOv11s | 0.901 | 0.468 | 0.885 | 0.835 | 0.859 | 9.40 | 21.3 |
| RT-DETR-n | 0.862 | 0.430 | 0.824 | 0.816 | 0.820 | 15.50 | 37.4 |
| Category | Item | Value |
|---|---|---|
| Hardware | Device | Orange Pi 5 Max |
| SoC and memory | Rockchip RK3588; 8 GB RAM | |
| Software | Operating system | Ubuntu 24.04.1 LTS ARM64 |
| NPU runtime | RKNN Lite/librknnrt 2.3.2 | |
| Model | Format | Non-quantized YieldNet RKNN |
| Input | ; batch size 1 | |
| Live pipeline | Camera | Insta360 Ace Pro 2 UVC |
| Stream | MJPEG; 30 FPS | |
| Processing | Centre 25% ROI; three asynchronous NPU cores |
| Model | Params (M) | GFLOPs | RKNN (MB) | p50 (ms) | FPS | Memory (MiB) |
|---|---|---|---|---|---|---|
| YOLOv8n | 3.00 | 8.1 | 7.99 | 52.79 | 47.00 | 314.0 |
| YieldNet (Ours) | 3.30 | 8.0 | 13.58 | 69.65 | 37.46 | 332.1 |
| YOLOv8s | 11.10 | 28.4 | 24.66 | 98.00 | 26.42 | 450.4 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Yu, C.; Li, L.; Huang, B. YieldNet: A Lightweight YOLOv8n Enhancement for Immature Green Tomato Detection in UAV Images: Real-Time Edge Demonstration Toward Pre-Harvest Yield Estimation. AgriEngineering 2026, 8, 311. https://doi.org/10.3390/agriengineering8080311
Yu C, Li L, Huang B. YieldNet: A Lightweight YOLOv8n Enhancement for Immature Green Tomato Detection in UAV Images: Real-Time Edge Demonstration Toward Pre-Harvest Yield Estimation. AgriEngineering. 2026; 8(8):311. https://doi.org/10.3390/agriengineering8080311
Chicago/Turabian StyleYu, Chenyu, Lu Li, and Bolin Huang. 2026. "YieldNet: A Lightweight YOLOv8n Enhancement for Immature Green Tomato Detection in UAV Images: Real-Time Edge Demonstration Toward Pre-Harvest Yield Estimation" AgriEngineering 8, no. 8: 311. https://doi.org/10.3390/agriengineering8080311
APA StyleYu, C., Li, L., & Huang, B. (2026). YieldNet: A Lightweight YOLOv8n Enhancement for Immature Green Tomato Detection in UAV Images: Real-Time Edge Demonstration Toward Pre-Harvest Yield Estimation. AgriEngineering, 8(8), 311. https://doi.org/10.3390/agriengineering8080311

