Next Article in Journal
Comparative Economic Evaluation of Greenhouse Pepper Cultivation Under Conventional and Organic Management
Previous Article in Journal
Integrated Nutrient Management Enhances Root Growth, Nutrient Use Efficiency, and Ratooning Ability in Rice Under Acidic Paddy Soils
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Cucumber Robotic Continuous Harvesting: Enhanced YOLOv8n Detection and Dynamic Bézier Curve-Assisted Collision-Free Path Generation

College of Optical, Mechanical and Electrical Engineering, Zhejiang A&F University, Hangzhou 311300, China
*
Author to whom correspondence should be addressed.
Agriculture 2026, 16(8), 888; https://doi.org/10.3390/agriculture16080888
Submission received: 1 March 2026 / Revised: 6 April 2026 / Accepted: 14 April 2026 / Published: 16 April 2026
(This article belongs to the Section Agricultural Technology)

Abstract

To address the inefficiency of the long single-fruit grasping cycle in traditional fruit harvesting robots, this study proposes a collision-free continuous harvesting solution for cucumber cultivation scenarios, coupled with a customized robotic system equipped with a continuous harvesting end-effector. In terms of visual perception, the YOLOv8n model is enhanced by integrating the GhostNet lightweight architecture, the Context-Guided Fusion Module (CGFM), and the MPDIoU loss function. Ablation experiments confirm the optimal model configuration, and the optimized model achieves a reduced model size of 5.3 MB and computational load of 6.6 GFLOPs while improving the mean average precision (mAP@50) by 2.5%, which facilitates low-cost deployment. For path planning, an Enhanced Bézier Continuous Picking (EBCP) algorithm is developed by combining 3D Gaussian kernel modeling and cubic Bézier curves to generate collision-free continuous trajectories. Simulation and practical experiments demonstrate that the path length of the proposed continuous picking method is only 31.1% that of the traditional path, with a theoretical collision-free rate of 96.69% and an actual collision-free rate of 92.24%. The feasibility and effectiveness of the proposed system are fully verified, providing a technical reference for the continuous operation of fruit harvesting robots.

1. Introduction

Cucumbers, as an important fruit and vegetable cash crop, are widely cultivated, with abundant market supply and high economic benefits. At the same time, they have a short growth cycle, delicious taste and rich nutritional value, and are widely favored by consumers. However, cucumbers are prone to rapid fruit expansion and need to be harvested in a timely manner. Delayed picking not only affects the commercial quality of the fruit but also inhibits the ripening of subsequent fruits, and the optimal harvest window is often merely one day [1]. At present, China’s fruit and vegetable sector is still dominated by manual harvesting. With an aging population and increasingly scarce agricultural labor, manual harvesting accounts for 33% to 55% of the total production cost of cucumbers, directly reducing the market competitiveness of the crop [2]. Therefore, the integration of artificial intelligence and automated equipment is an imperative for the development of intelligent harvesting technology [3].
Object detection, as the core task of computer vision, provides key technical support for the accurate identification of cucumber fruits. The mainstream widely used deep learning algorithms today include the two-stage R-CNN series [4,5], the single-stage YOLO series [6,7,8] and the SSD algorithm [9]. With their unique network structures and detection mechanisms, these algorithms can effectively locate and recognize target fruits, providing an important technical solution for agricultural automated monitoring. Chang et al. (2024) [10] proposed an enhanced YOLOv5s model to identify tomatoes with a complex background with 94.2% recognition accuracy. Pan et al. (2022) [11] proposed a sugarcane seedling detection method based on an enhanced rapid R-CNN structure. ResNet50 is used as a backbone network for deep feature extraction to generate high-resolution, semantically rich feature representations. The introduction of an SN-block attention mechanism module enabled the network to better focus on key feature channels associated with seedlings while suppressing background interference, with an average detection accuracy of 93.67%. Li M. et al. (2024) [12] proposed YOLOv8 model for chestnut fruit recognition. By introducing Pconv and a weighted bidirectional feature pyramid network, and modifying the bounding box loss function with a dynamic non-monotonic focus mechanism, they improved the recall rate by 1.5% and the average precision by 1.8%. Xu et al. (2024) [13] designed a lightweight YOLO-RFEW detection model for melon fruit ripeness identification in complex greenhouse environments with an average accuracy of 90.82% in a single frame processing time of 1.5 milliseconds. Luo et al. (2024) [14] proposed an enhanced strawberry recognition method based on the YOLOv8n model to address the challenge of missed detection and misclassification of small strawberries due to their complex background. The improved algorithm reduced model size by 59.7%. Yane Li et al. (2023) [15] combined the Mask R-CNN with an attention mechanism to develop a model for enhanced grape cluster segmentation and ripeness detection with an average accuracy of 94.4%. Zhong et al. (2024) [16] designed a lightweight detection model, Light-YOLO, which introduced residual structures into the neck network and incorporates an attention mechanism with an average accuracy of 96.1%. Wang et al. (2025) [17] developed a two-stage localization method for the picking points of wine grapes by enhancing the YOLO_GSConv and YOLACT_CBAM models, achieving an mAP@50 of 93.9%.
Currently, enhancing the operational efficiency of harvesting robots constitutes a core objective within agricultural robotics research. Yajun Li et al. (2024) [18] proposed an improved HER-SAC algorithm to guide end-effectors during tomato grasping, achieving an average harvesting time of 11.42 s. D. Wang et al. (2023) [19] employed an RRT algorithm based on quintic polynomial interpolation as the motion planner for a tomato-picking robot, achieving an 88% success rate but requiring 20 s per single fruit harvest. Ye et al. (2021) [20] proposed a Bidirectional Rapidly exploring Random Tree (Bi-RRT) algorithm for collision-free path planning in a lychee-picking robot, attaining a 100% success rate with an average path generation time of 4.24 s. T. Li et al. (2023) [21] designed a multi-agent reinforcement learning algorithm based on the Markov decision processes for multi-arm collaborative task allocation. Applied to a four-arm picking robot, it achieved a 79.31% picking success rate with an average single-fruit picking time of 6.7 s. H. Zhang et al. (2024) [22] transformed the picking task allocation problem in trellised pear orchards into a two-dimensional Traveling Salesman Problem. The picking sequence optimized via the simulated annealing algorithm improved harvesting efficiency by 20%. Xiong et al. (2023) [23] proposed a path planning method for citrus-picking robotic arms integrating artificial potential field theory with deep reinforcement learning. By incorporating a long short-term memory structure to process time-series data, they achieved a 16.20% reduction in path length compared with the RRT-Connect algorithm, alongside a 9.67% increase in path planning success rate. Hou et al. (2025) [24] employed a picking priority planning strategy, determining optimal harvesting sequences based on the relationship between cherry tomato clusters and mature fruits. This approach achieved an 82.1% harvesting success rate with an average harvesting time of 9.8 s per fruit cluster. Riboli et al. (2025) [25] proposed a novel approach for collision-free motion planning of dual-gantry robot systems for pick-and-place applications, which delivers higher efficiency than most current mainstream industrial practices and achieves an average cycle time reduction of approximately 22%.
This research suggests a novel continuous harvesting solution to overcome the efficiency bottlenecks seen in the traditional “single-fruit grasping-transport to collection bin-return” operating model that is commonly used by harvesting robots today. Restructuring the harvesting process is at the heart of this strategy. The robot uses a customized end-effector to work constantly between fruit clusters rather than go back to the collection bin after each pick.
Through an integrated conveyor channel in the end-effector, harvested fruits slide straight into the collection bin, significantly cutting down on idle time and the robotic arm’s travel distance between picking sites and the collection bin. This continuous procedure significantly minimizes the ineffective travel of the robotic arm compared to prevalent intermittent harvesting methods, while also reducing the time required per individual fruit picked. Lastly, to guarantee that the robot can safely and effectively perform continuous harvesting tasks within dense fruit clusters, we developed an enhanced visual identification algorithm and a collision-free path planning algorithm.
This paper focuses on automated continuous cucumber harvesting requirements and presents a comprehensive robotic operation system with three main components:
(a)
A dedicated robotic body for continuous cucumber harvesting with an end-effector mechanical structure designed according to continuous harvesting principles;
(b)
An enhanced YOLOv8n object detection algorithm for visual perception, which includes a lightweight GhostNet architecture, a CGFM feature fusion module, and an MPDIoU loss function. The optimal model configuration was determined through multiple rounds of ablation experiments. This model accurately identifies cucumbers in complex dense environments, providing precise picking positions for the robot; it has the advantages of fewer parameters and low computational complexity, enabling practical engineering deployment.
(c)
A collision-free trajectory generation technique based on continuous harvesting pathways is proposed for path planning. This method combines a cubic Bézier curve obstacle avoidance technique with fruit-picking sequence planning. This approach first identifies local fruit clusters via the DBSCAN clustering algorithm, and then uses a 3D Gaussian kernel function to calculate the distance-weighted index j for each cluster. After allocating pathways among fruits in clusters, it optimizes trajectories by avoiding obstacles. Enhanced Bézier Continuous Picking, or EBCP, is the name of this path generation technique.
The improved object detection algorithm and the collision-free continuous harvesting path planning algorithm are systematically validated in the experimental section.

2. Materials and Methods

2.1. Harvesting Pathways and Strategies

2.1.1. Continuous Harvesting Pathway Model

This experiment is based on the collision-free assumption in order to compare the trajectory length and time efficiency of the continuous harvesting technique with those of the conventional method. A number of cucumber fruits (n1, n2, n3, n4…) are randomly arranged along the vine, and the collection box is fixed at a set position in 3D space. For the conventional harvesting technique depicted in Figure 1, the robotic arm must transport each successfully grasped fruit to the collection box to deposit it, and only then move to the next harvesting position. Equation (1) is used to calculate the total travel distance of the robotic arm during this process. This conventional operation mode leads to a large number of repeated empty-load trips and inefficient motion. As shown in Figure 2, the continuous harvesting method entails going one after the other between fruit clusters. Harvested fruits can be collected in real time and conveyed to the collection box via the dedicated end-effector and integrated conveyor channel. As a result, repeated return trips to the collection box are no longer necessary. Equation (2) is used to calculate the total travel distance of the robotic arm. When the moving speed of the robotic arm is consistent, the continuous harvesting approach significantly reduces the total path length and time required for the robotic arm to complete the entire harvesting task, as shown by Equations (3) and (4). This fundamentally improves the overall operating efficiency of the harvesting system.
L s u m 1 = 2 ( L 1 + L 2 + L 3 + + L n )
L s u m 2 = L 1 + L n + L 12 + L 23 + + L ( n 1 ) n
T 1 = 2 k n i = 1 L i
T 2 = k L 1 + L n + n 1 i = 1 L i ( i + 1 )
where L i is the distance from picking point i to the collection box (located at the origin of the robotic arm), L i ( i + 1 ) is the Euclidean distance between picking point i and the subsequent picking point i + 1, L s u m 1 is the total travel distance of the conventional intermittent pick-and-place harvesting method and L s u m 2 is the total travel distance of the continuous picking method. T 1 is the total time of the conventional pick-and-place method, T 2 is the total time of the continuous picking method, and k is the motion time coefficient of the robotic arm, which is defined as k = 1 v , where v is the constant moving speed of the robotic arm.

2.1.2. Cucumber Harvesting Robot

Figure 3 shows a robot that harvests cucumbers. Its main parts include (1) a robotic arm that was created in a lab and has six degrees of freedom and a 1.2 m reach. It is powered by six servo motors and a gimbal mechanism. (2) It also includes a mobile platform that can move at a speed of 0.5 m per second. (3) A computer (FA506IV) is also included. (4) An Intel RealSense D435 depth camera—the depth measurement range of which is 0.1m~10m, the maximum depth frame rate is 90fps, the field of view (H × V) is 87° × 58°, and the depth accuracy is ≤1% within 1m distance—also features, and (5) a continuous harvesting end-effector that can pick cucumbers up to 40 cm in length and 3 cm in diameter is also included. This study uses a specially designed end-effector for continuous cucumber harvesting, as seen in Figure 4. A cutting mechanism, a grasping mechanism, and a collection mechanism make up its main parts.
The Intel RealSense D435 depth camera (Intel Corporation, Santa Clara, CA, USA) is adopted as the core visual perception unit of the harvesting robot. Based on active stereo vision technology, this camera is equipped with a built-in infrared laser projector and a binocular infrared imaging module. It acquires depth information of the scene by calculating the disparity of the projected infrared speckle patterns and can simultaneously output RGB color images and pixel-aligned depth images, which means it fully meets the functional requirements of target detection, spatial localization, and obstacle detection for the harvesting robot. This camera fulfilled three core functions in the field harvest experiments, as detailed below: (1) It acquired RGB images of cucumber plants in the field, which were fed into the improved YOLOv8-GCM model for target detection and classification, to obtain the two-dimensional (2D) pixel coordinates of cucumber fruits and their corresponding harvestability categories; (2) Based on the pixel-aligned depth images, it converted the 2D pixel coordinates of the target fruits into three-dimensional (3D) spatial coordinates in the robot base coordinate system, so as to provide accurate position information of picking points for the path planning of the robotic manipulator; (3) It collected depth information of the harvest scene to construct a 3D point cloud map of the robot’s workspace, which provided environmental perception data for obstacle detection, collision risk assessment, and obstacle avoidance path generation of the proposed EBCP algorithm. For the collected depth data, the following processing flow was adopted: (1) Bilateral filtering was used to remove the noise of the depth image while retaining the depth information of the fruit edge; (2) ROI depth extraction was performed based on the detection bounding box, only the depth data in the target cucumber area was retained, and the invalid background information was filtered out; (3) RANSAC algorithm was used to fit the depth data, eliminate outliers, and calculate the optimal 3D coordinates of the cucumber picking point; (4) Hand-eye calibration was completed by the nine-point calibration method to realize the coordinate conversion between the camera coordinate system and the robotic arm base coordinate system, and the re-projection error after calibration was ≤0.5 mm, which ensured the positioning accuracy of the picking point.

2.1.3. Planning Collision-Free Harvesting Sequences for Cucumbers

Picking sequence planning is a core step in the automated harvesting process of cucumber harvesting robots. An optimal picking sequence can effectively reduce fruit damage and improve operational efficiency. Figure 5 illustrates the collision-free continuous picking method proposed in this paper, which adopts the origin of the robotic arm as the origin of the coordinate system. The detailed picking strategy is described as follows:
Step 1: DBSCAN [26], as a density-based clustering algorithm, automatically forms cluster structures based on the spatial proximity of data points. This algorithm can partition densely clustered yet disordered target fruits into independent clusters. Individual, isolated cucumbers are classified into separate clusters. In this paper, we employ the DBSCAN algorithm to partition cucumber picking points into clusters suitable for continuous harvesting.
Step 2: To determine the harvesting sequence of cucumber clusters, we employ a Gaussian kernel function to assign distance weights to clusters within the robotic arm’s operational workspace:
G ( x , y , z ) = 1 ( 2 π ) 3 2 σ 3 e x 2 + y 2 + z 2 2 σ 2
In the equation, G ( x , y , z ) denotes the Gaussian kernel function, ( x , y , z ) represents three-dimensional coordinates, and σ 2 denotes the variance. The closer a cucumber cluster is to the origin of the robotic arm, the higher its weight. Conversely, clusters farther away from the origin receive lower weights. Each cucumber cluster is then ranked in descending order according to its weight.
Step 3: Winner-Take-All (WTA) mechanism selection. The WTA mechanism is a classic competitive selection method widely used in neuroscience and machine learning. Its core principle is to automatically select the most salient target from multiple input signals. This mechanism mimics the information processing mode of the visual nervous system, in which attention prioritizes the salient locations or features with the highest activation in the visual field. In this paper, we calculate the Euclidean distance between the end-effector of the robotic arm and the target cucumber in 3D space. A shorter distance indicates higher target salience, which aligns with the competitive principle of the WTA mechanism. This selection rule guides the robot to prioritize grasping the most prominent and most accessible individual cucumber within the cluster.
Step 4: Path adjustment and obstacle avoidance. When considering the collision bounding volume of the cucumber fruit model during the actual harvesting operation, we sequentially apply the following adjustments to the path planned in Step 3: (a), If the nearest target picking point is a collision-free safe point, the target picking point is set as this nearest point; (b), If the nearest target picking point is unsafe and a collision-free safe point exists within the cluster, the target picking point is set as the nearest safe point; (c), If the nearest target picking point is unsafe and no collision-free safe point exists within the cluster, the target picking point remains the nearest point, and obstacle avoidance is implemented via the Bézier curve method according to Equation (6). (d), (a) to (c) are repeated until all target points are fully harvested.
B ( t ) = P 0 ( 1 t ) 3 + 3 P 1 t ( 1 t ) 2 + 3 P 2 t 2 ( 1 t ) + P 3 t 3
where t is a parameter ranging from 0 to 1, used to control the position along the curve.
We perform collision detection along the straight path intended for the end-effector movement. If a collision is detected, a Bézier curve is generated to bypass the obstacle; otherwise, the straight path is retained. When target clusters are adjacent or target points are excessively dense, the Bézier curve—constructed based on the normal offset from the nearest obstacle—may encroach into the collision detection regions of other targets. Although control points P1 and P2 are added to avoid obstacles (as shown in Figure 6), the resulting curve still eventually enters the collision zone of P4. To solve this problem, this study proposes an improved Bézier-based obstacle avoidance method. The method explores multiple feasible avoidance directions, considers various avoidance options in three-dimensional space, and selects the path with the lowest collision probability by employing multiple sets of independent control points. Meanwhile, the curve is optimized by dynamically adjusting the offset according to the environmental conditions, taking both collision constraints and the length of the generated path into consideration. This design aims to minimize the occurrence of collisions while maintaining a short path length. The proposed method enhances obstacle avoidance capability in scenes with multiple obstacles.

2.2. Machine Vision Inspection Models

2.2.1. Data Collection and Construction

Cucumber data was collected in Yuhang District, Hangzhou City, Zhejiang Province, with photographs taken between 13:00 and 17:00 on 19 June 2025 and between 09:00 and 12:00 on 2 July 2025. All images were captured at a distance of 0.2–1 m from the cucumber fruit surface, acquired from different plots at the sampling sites, and covered a comprehensive range of imaging angles and natural field illumination conditions. The acquisition device was a Huawei P30 with a resolution of 2736 × 2736 pixels and a focal length of 27mm. Images were saved in JPG format. Following manual review to exclude images lacking target cucumbers, a total of 1106 natural images were collected. As shown in Figure 7, These were divided into training, validation, and test sets in a 7:2:1 ratio. Given the limited sample size and to prevent overfitting in the neural network model, a series of data augmentation techniques were applied. These included cropping, rotation, mirroring, translation, noise addition, and brightness adjustment to expand the dataset. After data augmentation, a total of 2212 valid images were obtained, which were divided into three mutually independent subsets with no data overlap: a training set containing 1548 images, a validation set with 442 images, and a test set comprising 222 images. This approach enhanced sample diversity, thereby improving the convolutional neural network’s generalization capability and robustness.
To ensure that the dataset constructed in this study fully matches the actual in-field scenarios of commercial continuous cucumber harvesting, cucumber samples with maturity meeting commercial harvesting criteria were labeled as harvest-ready cucumbers during the annotation process, while immature young fruits, malformed fruits, diseased fruits, and pest-damaged fruits were labeled as non-harvest-ready cucumbers. All sample annotations were completed using LabelImg (v1.8.6), with annotation files saved in the PASCAL VOC format compatible with the YOLO family of models. During the annotation workflow, minimum bounding rectangle (MBR) bounding boxes were drawn with the intact harvestable main body of cucumber fruits as the reference benchmark. For target fruits partially occluded by stems and foliage yet meeting the pre-specified inclusion criteria, bounding boxes were delineated to maximally cover the complete visible contour of the fruits, ensuring high consistency between the bounding boxes and the actual morphological features of the target fruits to avoid compromising the model’s localization accuracy due to annotation deviations. To guarantee the reliability and reproducibility of the dataset, the annotation task was independently performed by two trained researchers. Upon completion of the initial annotation, a double-blind cross-validation procedure was implemented: the two annotators conducted a one-by-one cross-check on the inclusion compliance of all samples and the accuracy of bounding box annotations. Annotation discrepancies were resolved through mutual consultation between the two annotators, followed by standardized recalibration, which not only meets the reliability requirements for agricultural vision dataset construction but also ensures the full reproducibility of the dataset establishment workflow.

2.2.2. Improve the Overall Structure of the YOLO v8 Model

The architecture of the improved YOLOv8-GCM network is illustrated in Figure 8, which is mainly composed of three core components: the backbone network, the neck network, and the head network. The key improvements are detailed as follows: (1) To reduce the model size and the number of parameters, the C2f module in the backbone feature extraction network is reconstructed. Specifically, the original Bottleneck module inside the C2f module is replaced with the Ghost Bottleneck, and four original convolutional modules in the network are substituted with Ghost Conv modules. This design can effectively reduce the computational complexity of the model while avoiding the loss of feature information. (2) To enhance the network’s perception capability for large-scale targets, the proposed CGFM feature fusion module is introduced to replace the original concatenation (Concat) module at the P5 stage. (3) To mitigate the risk of model overfitting induced by the interference of annotation noise, the MPDIoU loss function is adopted to replace the original CIoU loss function as the bounding box regression loss.

2.2.3. Optimization of the Backbone Network

Multiple convolutional modules are embedded in traditional convolutional neural network (CNN) models, which results in a high number of parameters and computational demands. These models generate a large number of redundant feature maps during feature extraction and forward propagation, which hinders their lightweight implementation and makes them unsuitable for deployment on embedded platforms. Although existing mainstream CNNs use small-sized convolutional kernels to construct lightweight and efficient models, the remaining 1 × 1 convolutional layers still occupy a large amount of memory and consume significant FLOPs. Han et al. (2020) [27] designed a lightweight network architecture named GhostNet to address this challenge. The core mechanism of the Ghost module consists of two steps: first, intrinsic feature maps are generated from input features via conventional convolution; second, additional ghost feature maps are generated from the intrinsic features via cheap linear operations. The final output feature map is then obtained by concatenating the generated ghost features with the original intrinsic features. GhostNet achieves excellent inference speed with fewer parameters and lower computational overhead, benefiting from this dual-branch design. Figure 9 illustrates the convolution process of the Ghost module, where Φk represents the linear transformation operation corresponding to the k-th feature channel.
To extract image features, YOLOv8 adopts standard convolutional layers, which have a large number of parameters and generate redundant feature information, even though they can capture detailed features. In this study, we introduce the linear transformation mechanism of the Ghost module to reduce the model complexity. Using a small number of conventional convolutions, this module first generates intrinsic feature maps. It then applies cheap linear operations to generate supplementary “ghost feature maps”, significantly reducing the consumption of computational resources. We design a lightweight network architecture based on the feature generation mechanism of the Ghost module: to reduce the number of model parameters, we replace part of the standard convolutional layers in the YOLOv8 backbone network with GhostConv modules. We construct the Ghost-C2f module by reconstructing the standard C2f module, replacing its original Bottleneck module with a Ghost Bottleneck (as shown in Figure 10). This module design retains the feature fusion capability of the original C2f module, while combining the computational efficiency of the linear transformation of the Ghost module. By fusing multi-path features, the proposed Ghost-C2f module enhances the feature representation capability of the model while reducing model complexity.

2.2.4. CGFM Feature Fusion Module

In the cucumber fruit detection task under complex field environments, the model must accurately distinguish target fruits from trellises and vines, despite severe occlusion by stems and leaves, as well as interference from the similar color distribution between fruits and the background. By default, the original YOLOv8 model adopts the channel concatenation (Concat) operation to fuse feature maps of different scales. However, the fusion of high-level semantic features and low-level detail features is insufficient, because this operation only performs channel-wise concatenation without a feature selection mechanism, and thus cannot distinguish task-critical channels from irrelevant background interference information. Meanwhile, the P5 feature layer, generated by 32 × downsampling, loses most of the fine-grained geometric contour and texture features of cucumber fruits. Although the P5 feature layer has a large receptive field, its fixed-size convolutional kernels cannot dynamically model contextual information and adapt to irregular occlusions, which leads to feature confusion between target fruits and backgrounds with similar color distributions [28]. This paper proposes the Context-Guided Fusion Module (CGFM) to replace the concatenation operation in layer P5, which is inspired by the adaptive calibration of feature channels via channel attention in SENet [29] and the mutual enhancement of registration and fusion through commonality mining and contrastive learning in C2RF [30]. As illustrated in Figure 11, the proposed CGFM achieves occlusion awareness and monochromatic decoupling through a three-stage technical framework encompassing three core sequential components: channel alignment, cross-attention fusion, and bidirectional feature enhancement.
First, to ensure dimensional compatibility for subsequent operations, the acquired detail features X 0 R C 0 × H × W and semantic features X 1 R C 1 × H × W undergo 1 × 1 convolutions to adjust their channels to C. Following this, the two feature maps are feature-stitched and aligned to generate a 2C-channel feature.
Second, to capture global information across each channel, the input feature maps undergo global average pooling, reducing the dimensionality of each channel’s features to a single global vector of size 2C × 1 × 1. The subsequent network comprises two fully connected layers (FC), employing ReLU and Sigmoid activation functions. To mitigate model complexity overhead, the number of channels is reduced by an appropriate reduction ratio r (r = 16). Here, the FC layer takes all neuron outputs from the preceding layer as input, computing the weighted sum for each neuron via the weight matrix. This compels the network to integrate features from different regions. After processing through the subsequent network, the number of output weight values is split and maintained to match the channel count of the input feature map, resulting in a dimension of C × 1 × 1.
Finally, the resulting weight values are multiplied element-wise with the original feature map [31] and cross-multiplied element-wise with another feature map. The two generated feature sets are then concatenated via a Concat operation, yielding a 2C × H × W fused feature output.
X 0 a d j = Conv 1 × 1 ( X 0 ) if   C 0 C 1 X 0 otherwise
s = σ ( W 2 δ ( W 1 z ) )
where σ is the Sigmoid function, W 1 , W 2 are two 1 × 1 convolutions sharing parameters, with W 1 R C / r × C and W 2 R C × C / r , δ is the ReLU activation function; z is the feature vector undergoing global average pooling.

2.2.5. MPDIoU Loss Function

The original YOLOv8 model uses CIoU (Complete Intersection over Union) as its default bounding box regression loss function. By simultaneously optimizing geometric parameters such as the overlap area, centroid distance, and aspect ratio between the predicted and ground-truth bounding boxes, this loss function can effectively improve localization accuracy and achieves excellent performance in general object detection tasks. Nevertheless, the CIoU penalty term may degenerate to zero when the predicted bounding box has the same aspect ratio as the ground-truth box but completely different width and height values, making it impossible for the model to optimize mismatched bounding boxes [32]. In cucumber detection scenarios, target fruits have a similar color to the background, are sparsely distributed, and are frequently occluded by stems and leaves. This frequently leads to inaccuracies in bounding box dimensions or boundary deviations during the annotation process. Directly increasing the loss weight of low-quality samples will aggravate the interference of annotation noise, which will lead to model overfitting and ultimately reduce detection reliability.
To address the above issues, this study adopts the MPDIoU loss function to construct a novel bounding box similarity metric by directly minimizing the distance between the top-left and bottom-right corner points of the predicted bounding box and the ground-truth bounding box. This method holistically takes into account all critical geometric factors, including overlapping and non-overlapping regions, center point distance, and aspect ratio deviation. In addition, the MPDIoU loss function significantly reduces computational complexity by directly deriving other required parameters from corner coordinates, thus improving both detection accuracy and inference speed.
d 1 2 = ( x 1 B x 1 A ) 2 + ( y 1 B y 1 A ) 2
d 2 2 = ( x 2 B x 2 A ) 2 + ( y 2 B y 2 A ) 2
M P D I o U = A B A B d 1 2 ω 2 + h 2 d 2 2 ω 2 + h 2
where A and B are two arbitrary convex shapes; ω and h are the width and height of the image, respectively; d 1 and d 2 are the Euclidean distances between the top-left and bottom-right corners of A and B, respectively; ( x 1 A , y 1 A ) and ( x 2 A , y 2 A ) are the coordinates of the top-left and bottom-right points of A, respectively; and ( x 1 B , y 1 B ) and ( x 2 B , y 2 B ) are the coordinates of the top-left and bottom-right points of B, respectively.

2.2.6. Model Training and Evaluation Metrics

The experimental hardware configuration comprised a GeForce RTX 3060 12G GPU graphics card, an Intel® Core™ i5-12600KF CPU with a base frequency of 3.19 GHz, and 16 GB of RAM. The software configuration utilized the Windows 11 operating system, PyTorch 1.13.1, and CUDA 11.7.
Model training parameters were set as follows: image dimensions 640 × 640, learning rate 0.01, training iterations 300, SGD momentum 0.937, decay rate 0.0005, and batch size 16.
To comprehensively validate model performance, this study employs key evaluation metrics including precision, recall, mean average precision (mAP), number of parameters, floating-point operations per second (FLOPs), and inference speed. This ensures a multidimensional assessment encompassing prediction accuracy, resource consumption, and operational efficiency.

3. Experiments and Results

3.1. Machine Vision Inspection Results

3.1.1. Loss Function Performance Validation

The model’s overall performance in object detection tasks is directly influenced by the loss function selection. The efficiency of various loss functions varies in cucumber identification tasks due to their notable variations in convergence characteristics, computing complexity, and optimization goals. This work uses the MPDIoU loss function for training in order to improve the YOLOv8 model’s object localization capabilities. On a specially created cucumber dataset, it contrasts this function with more complex IoU variations, such as CIoU, DIoU, GIOU, EIoU, and SIoU. Figure 12 displays the outcomes of the experiment.
The six loss functions’ loss curves on the validation set display the following patterns, as shown in Figure 12: The loss values drop quickly during the first 50 training epochs. As the number of iterations rises, the rate of loss reduction progressively decreases and the curve stabilizes, suggesting that the model is moving closer to a low-loss state. With the smallest fluctuation amplitude and the fastest loss reduction rate among them, MPDIoU offers the best convergence stability and performance.
After the model training was completed, this study evaluated the performance of six bounding box regression loss functions using the test set (Table 1). The experimental data indicated that EIoU had the weakest overall performance, with its mAP@50 and mAP@50:95 lagging behind MPDIoU by 3.8% and 4.7%, respectively, suggesting that this loss function is not suitable for cucumber detection tasks. In contrast, MPDIoU demonstrated significant advantages: compared with CIoU, it improved mAP@50, precision (P), and recall (R) by 0.6%, 0.9%, and 0.6% respectively; compared with DIoU, mAP@50 and R increased by 1.2% and 2.1% respectively; compared with GIoU, MPDIoU led in mAP@50, P, and R by 0.8%, 0.8%, and 1.2% respectively; and compared with SIoU, it was 2.4%, 2.6%, and 3.4% higher respectively.

3.1.2. Ablation Experiment

This study used a self-built cucumber dataset under a unified experimental environment, with YOLOv8 as the baseline model, to validate the independent and synergistic improvements in detection performance achieved by Ghost optimization, the CGFM feature fusion module, and the MPDIoU bounding box loss function within the backbone network. Table 2 displays the outcomes of the experiment.
As shown in Table 2, after optimizing the backbone network of the YOLOv8 model with GhostConv and C2f Ghost modules in Experiment 2, the mAP decreased by 0.5%. However, the computational load and model size were reduced by 1.6 GFLOPs and 1.2 MB, respectively, achieving a substantial reduction in model size. Experiment 3, employing the CGFM, saw a 0.2MB increase in model size. Nevertheless, it demonstrated notable improvements in accuracy, recall, and mAP, with mAP rising by 1.2% compared to the baseline model. This indicates its superior ability to fuse information from features at different scales. Experiment 4 introduced the MPDIoU loss function to refine the similarity metric for bounding box regression, enabling more precise matching between predicted and ground-truth boxes. This resulted in a 1.2% improvement in mAP performance. Experiment 5 was constructed by incorporating the proposed CGFM into the framework of Experiment 2. This design not only compensated for the accuracy degradation caused by the Ghost-based lightweight modification, but also achieved a 1.9% improvement in mAP relative to the baseline. Meanwhile, it retained a low computational burden of 6.6 GFLOPs and a compact model size of 5.3 MB, realizing a preliminary trade-off between model lightweighting and detection accuracy. In contrast, neither Experiment 6 (Ghost + MPDIoU) nor Experiment 7 (CGFM + MPDIoU) could simultaneously achieve optimal performance in both lightweight design and detection accuracy. This finding validates the irreplaceable role of the proposed CGFM in compensating for the feature information loss induced by model lightweighting. Experiment 8 synergistically integrates the Ghost module, CGFM, and MPDIoU loss function. This achieves optimal precision, recall, and mean average precision for cucumber detection in occlusion scenarios, representing improvements of 1.9%, 2.8%, and 2.5%, respectively, over the baseline model. Concurrently, computational complexity and model size are reduced by 18.5% and 15.9%, meeting the requirements for embedded deployment tasks.

3.1.3. Comparison of Detection Capabilities Across Different Algorithms

A self-constructed cucumber detection dataset under diverse backgrounds was used in this study to thoroughly assess the performance benefits of the enhanced model. The model was thoroughly tested against several popular object detection techniques under regular training rounds. YOLOv5n, YOLOv7-tiny, YOLOv8n, YOLOv10n and YOLOv11n were among the more sophisticated versions of the YOLO series that were included in the comparison models, along with exemplary two-stage detection frameworks like Faster R-CNN and single-stage detectors like SSD. Table 3 has comprehensive performance comparison data.
As shown in Table 3, the YOLOv8-GCM model outperforms other models in terms of precision, recall, and mAP. Concurrently, YOLOv8-GCM maintains low computational complexity and model size, at 6.6 GFLOPs and 5.3 MB, respectively. Compared to the two-stage object detection algorithm Faster R-CNN, the proposed improved model YOLOv8-GCM demonstrates significant enhancements across multiple performance metrics: precision improves by 13.5%, recall increases by 5%, and mean average precision (mAP) rises by 13.7%. Concurrently, the model weight is substantially reduced by 95.2%, markedly enhancing its practicality and deployment efficiency. When compared with the single-stage model SSD, YOLOv8-GCM achieves improvements of 11.1% and 5.8% in recall and mAP, respectively. Compared to previous lightweight models in the YOLO series—YOLOv5n, YOLOv7-tiny, and YOLOv8n—YOLOv8-GCM demonstrates recall gains of 5.6%, 4.5%, and 2.8%, respectively, alongside mAP improvements of 5.6%, 3.9%, and 2.5%. Compared with state-of-the-art lightweight YOLOv10n and YOLOv11n, our proposed YOLOv8-GCM improves Precision by 2% and 1.6%, Recall by 3% and 2.2%, and mAP@50 by 2.3% and 2.1%, respectively. With an inference time of only 1.7 ms (faster than all compared models), YOLOv8-GCM also shows prominent advantages in model size and training time, making it ideal for outdoor embedded deployment, especially for cucumber detection in complex field backgrounds.

3.1.4. Visual Analysis of Detection Results

Comparative validation will be performed using both the YOLOv8n model and the modified YOLOv8-GCM model on a sample of images in order to more clearly illustrate the detection performance of the improved method given here. Figure 13 displays the detection findings, with red boxes signifying false negatives and yellow boxes signifying false positives. When vines are hidden by foliage and tendrils, the baseline model misidentifies them as cucumbers and produces false negatives, as seen in Figure 13. The original model further misclassifies a single cucumber as several cucumbers in complex backdrops. On the other hand, the enhanced algorithm successfully addresses these problems and produces favorable outcomes on datasets with intricate backgrounds and comparable color schemes.

3.2. Experimental Results of Collision-Free Continuous Harvesting of Cucumbers

3.2.1. Simulation Experiments

In simulation experiments, this paper conducted tests on five scenarios involving 10, 14, 18, 22, and 26 cucumbers, respectively, evaluating four algorithms: EBCP (Enhanced Bezier Continuous picking), RRT-CP, CP(Continuous picking), RP(Round-trip picking), AYDY [33] and the K-G method [34]. RRT-CP is an RRT-based algorithm optimized for continuous picking, CP is a sequential harvesting algorithm adhering solely to the winner-takes-all principle, RP employs a conventional harvesting approach prioritizing the closest target fruit, AYDY was originally a collision-free path planning algorithm for green pepper harvesting and the K-G method was originally developed as a picking sequence planning algorithm for tomato harvesting scenarios. The simulation program was developed in Python 3.8.20.
The simulated robotic arm traversed all cucumbers under ideal conditions with no environmental obstacles, considering only potential collisions with other fruits during movement. Results are recorded in Table 4. As the number of cucumbers increased from 10 to 22, the path values for AYDY and RP rose rapidly. In contrast, EBCP, RRT-CP, CP and K-G method demonstrated superior path length performance, exhibiting slower path growth with increasing fruit quantity.
In simulated experiments with 10 and 14 cucumbers, the EBCP and AYDY algorithms experienced virtually no collisions due to the limited number and sparse distribution of cucumbers. Compared to RRT-CP, EBCP improved traversal paths by 1.3% and 13.3%, while collision-free picking rates increased by 10% and 14.29%. Compared to CP, EBCP improved traversal paths by 3.99% and 31.67%, while collision-free picking rates increased by 10% and 14.29%. Compared to RP, EBCP reduced traversal paths by 67.8% and 64.7%, respectively, while its collision-free rate increased by 20% and 21.43%. Against AYDY, EBCP shortened paths by 67.8% and 64.7%, respectively, with comparable collision frequencies. Compared to the K-G Method, EBCP’s traversal paths were marginally longer by 0.3% in the 10-cucumber group and shortened by 4.8% in the 14-cucumber group, while collision-free picking rates increased by 10% and 7.15%, respectively.
In simulation experiments with 18, 22, and 26 cucumbers, as distribution density within the picking space increased, EBCP and AYDY experienced minor collisions, while RRT-CP, CP, RP and the K-G method exhibited significantly higher collision rates. The traversal paths of EBCP and RRT-CP are similar, but the collision-free picking rates increased by 11.11%, 18.18%, and 11.53%, respectively. Compared to CP, EBCP’s traversal paths were marginally longer, yet collision-free picking rates improved by 5.55%, 18.18%, and 19.23%. Against RP, EBCP reduced traversal paths by 67.37%, 71.70%, and 70.24%, respectively; while its collision-free picking rates improved by 16.66%, 36.39%, and 42.30%. Compared to AYDY, EBCP traversed paths 67.37%, 71.70%, and 70.24% shorter, with collision-free picking rates increasing by 5.55%, 9.09%, and 3.84%, respectively. Compared to the K-G Method, EBCP’s traversal paths were marginally longer by 2.1% and 2.9% in the 18 and 22 cucumber groups, while collision-free picking rates increased by 11.11% and 13.39%, respectively.
On average, EBCP exhibits the highest collision-free picking rate and the lowest collision frequency. Its traversal path length is marginally higher than RRT-CP and CP, comparable to the K-G Method, but significantly lower than RP and AYDY. This demonstrates that EBCP not only adapts to varying numbers of target fruits but also substantially reduces both path length and collision frequency. However, collisions still occur when two or more fruits are in close proximity.

3.2.2. Practical Experiments

To validate the effectiveness of the EBCP algorithm in field applications, we conducted field experiments on a cucumber harvesting robot platform, in which the fruit arrangement in the experimental scene was consistent with the simulated fruit layout described in Section 3.2.1. The calibration of the repetitive positioning accuracy of the robotic manipulator was completed prior to the formal experiments. The experimental results for the simulated trajectory are shown in Figure 14, with specific data presented in Table 5. Analysis indicates that across five test scenarios, the EBCP algorithm achieved an average collision-free picking rate of 92.24%. This performance of the proposed EBCP algorithm represents improvements of 13.1%, 16.02%, 28.22% and 10.02% over the RRT-CP, CP, RP and K-G method, respectively, alongside a 0.9% enhancement relative to the AYDY method.
During the field harvest experiments (Figure 15a), it was observed that during the back-and-forth movement between the fruit and the collection box, the end-effector repeatedly collided with non-target fruits. This is likely the primary reason for the higher collision frequency in the Round-trip picking method. Additionally, no missed detection, false detection or erroneous picking occurred during formal harvesting experiments. Gripping and shearing failure from target cucumbers’ close proximity to stems and trellis supports (Figure 15b) and end-effector collisions caused by stem and foliage overlap (Figure 15c) were the primary causes of harvesting failure. However, minor stem collisions typically did not adversely affect the overall harvesting process.

4. Conclusions

The goal of this study is to harvest cucumbers continuously and automatically while avoiding obstacles throughout the harvesting process. This study proposes a collision-free continuous harvesting approach that combines intelligent path planning and machine vision, and this approach has been verified by simulation and real-world field experiments. The main findings of this study are as follows: (1) To identify and locate cucumber fruits, we propose an enhanced YOLOv8-GCM detection method. With a model size of 5.3 MB, this method can be deployed on embedded platforms in field scenarios, and it achieves a mean average precision (mAP@50) of 84.2%. (2) For continuous collision-free cucumber harvesting, we propose an Enhanced Bézier Continuous Picking (EBCP) algorithm. To generate continuous harvesting paths, this approach combines DBSCAN clustering, a 3D Gaussian kernel function for distance weighting, and the winner-takes-all principle with cubic Bézier curves. (3) The performance of the proposed EBCP algorithm is verified via simulation and physical field experiments. The proposed method achieves an average simulation collision-free rate of 96.69% while significantly reducing the total travel distance of the robotic arm. The actual field collision-free harvesting rate of 92.24% is verified via on-site prototype experiments.
The proposed method significantly improves the efficiency of cucumber harvesting and provides an important technical reference for the automated harvesting of fruit crops. Harvesting failures are mainly caused by collisions occurring when cucumbers are too close to stems or trellises, or when they are occluded by foliage during the actual field harvesting process. To reduce collision risks and improve the harvesting success rate, future research will focus on developing collision-free continuous harvesting algorithms for cucumbers under dense proximity and severe occlusion conditions.

Author Contributions

C.Z.: methodology, software, validation, formal analysis, investigation, writing—original draft, writing—review and editing, visualization. M.Q.: resource supervision, project administration, funding acquisition. H.W.: methodology, formal analysis. W.L.: methodology, formal analysis. H.Z.: methodology. L.Z.: formal analysis. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Talent Start-up Project of the University’s Scientific. Research Fund and “Pioneer” and “Leading Goose” R&D Program of Zhejiang under Grant 2023C02049.

Institutional Review Board Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author. The data are not publicly available due to privacy restrictions.

Acknowledgments

The authors would like to thank all the anonymous reviewers and editors for their useful comments and suggestions that greatly improved this paper.

Conflicts of Interest

The authors declare that they have no conflicts of interest or personal relationships related to the work in this paper.

References

  1. Yang, Q.-H.; Qi, L.-Y.; Bao, G.-J.; Gao, F.; Zhang, L.-B. Cucumber Image Segmentation Algorithm Based on Rough Set Theory. N. Z. J. Agric. Res. 2007, 50, 989–996. [Google Scholar] [CrossRef]
  2. El-Ansary, D.O. Smart Farming and Orchard Management: Insights and Innovations. Curr. Food Sci. Technol. Rep. 2025, 3, 10. [Google Scholar] [CrossRef]
  3. Paul, A.; Machavaram, R.; Ambuj; Kumar, D.; Nagar, H. Smart Solutions for Capsicum Harvesting: Unleashing the Power of YOLO for Detection, Segmentation, Growth Stage Classification, Counting, and Real-Time Mobile Identification. Comput. Electron. Agric. 2024, 219, 108832. [Google Scholar] [CrossRef]
  4. Wang, L.; Hou, Y.; He, J. Target recognition and detection of Camellia oleifera fruit in natural scene based on Mask-RCNN. J. Chin. Agric. Mech. 2022, 43, 148–154, 189. [Google Scholar] [CrossRef]
  5. Zhang, W.; Wang, Y.; Shen, G.; Li, C.; Li, M.; Guo, Y. Tobacco Leaf Segmentation Based on Improved MASK RCNN Algorithm and SAM Model. IEEE Access 2023, 11, 103102–103114. [Google Scholar] [CrossRef]
  6. Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You Only Look Once: Unified, Real-Time Object Detection. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 779–788. [Google Scholar]
  7. Wang, C.-Y.; Bochkovskiy, A.; Liao, H.-Y.M. YOLOv7: Trainable Bag-of-Freebies Sets New State-of-the-Art for Real-Time Object Detectors. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 7464–7475. [Google Scholar]
  8. Wang, C.-Y.; Yeh, I.-H.; Mark Liao, H.-Y. YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. In Proceedings of the Computer Vision—ECCV 2024; Springer: Cham, Switzerland, 2025; pp. 1–21. [Google Scholar]
  9. Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Proceedings of the Computer Vision—ECCV 2016; Leibe, B., Matas, J., Sebe, N., Welling, M., Eds.; Springer International Publishing: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
  10. Chang, W.; Tan, Y.; Zhou, L.; Yang, Q. Tomato Ripening Detection in Natural Environment Based on Improved YOLOv5s. Acta Agric. Univ. Jiangxiensis 2024, 46, 1025–1036. [Google Scholar] [CrossRef]
  11. Pan, Y.; Zhu, N.; Ding, L.; Li, X.; Goh, H.-H.; Han, C.; Zhang, M. Identification and Counting of Sugarcane Seedlings in the Field Using Improved Faster R-CNN. Remote Sens. 2022, 14, 5846. [Google Scholar] [CrossRef]
  12. Li, M.; Xiao, Y.; Zong, W.; Song, B. Detecting chestnuts using improved lightweight YOLOv8. Trans. Chin. Soc. Agric. Eng. 2024, 40, 201–209. [Google Scholar]
  13. Xu, D.; Ren, R.; Zhao, H.; Zhang, S. Intelligent Detection of Muskmelon Ripeness in Greenhouse Environment Based on YOLO-RFEW. Agronomy 2024, 14, 1091. [Google Scholar] [CrossRef]
  14. Luo, Q.; Wu, C.; Wu, G.; Li, W. A Small Target Strawberry Recognition Method Based on Improved YOLOv8n Model. IEEE Access 2024, 12, 14987–14995. [Google Scholar] [CrossRef]
  15. Li, Y.; Wang, Y.; Xu, D.; Zhang, J.; Wen, J. An Improved Mask RCNN Model for Segmentation of ‘Kyoho’ (Vitis labruscana) Grape Bunch and Detection of Its Maturity Level. Agriculture 2023, 13, 914. [Google Scholar] [CrossRef]
  16. Zhong, Z.; Yun, L.; Cheng, F.; Chen, Z.; Zhang, C. Light-YOLO: A Lightweight and Efficient YOLO-Based Deep Learning Model for Mango Detection. Agriculture 2024, 14, 140. [Google Scholar] [CrossRef]
  17. Wang, Q.; Xu, F.; Chen, Q.; Mi, Z.; Fan, Y.; Su, B. Detection and Location of Wine Grape (Cabernet Sauvignon) Picking Points by Using a Dual-Stage Deep Learning Method. Comput. Electron. Agric. 2025, 237, 110637. [Google Scholar] [CrossRef]
  18. Li, Y.; Feng, Q.; Zhang, Y.; Peng, C.; Ma, Y.; Liu, C.; Ru, M.; Sun, J.; Zhao, C. Peduncle Collision-Free Grasping Based on Deep Reinforcement Learning for Tomato Harvesting Robot. Comput. Electron. Agric. 2024, 216, 108488. [Google Scholar] [CrossRef]
  19. Wang, D.; Dong, Y.; Lian, J.; Gu, D. Adaptive End-Effector Pose Control for Tomato Harvesting Robots. J. Field Robot. 2023, 40, 535–551. [Google Scholar] [CrossRef]
  20. Ye, L.; Duan, J.; Yang, Z.; Zou, X.; Chen, M.; Zhang, S. Collision-Free Motion Planning for the Litchi-Picking Robot. Comput. Electron. Agric. 2021, 185, 106151. [Google Scholar] [CrossRef]
  21. Li, T.; Xie, F.; Zhao, Z.; Zhao, H.; Guo, X.; Feng, Q. A Multi-Arm Robot System for Efficient Apple Harvesting: Perception, Task Plan and Control. Comput. Electron. Agric. 2023, 211, 107979. [Google Scholar] [CrossRef]
  22. Zhang, H.; Li, X.; Wang, L.; Liu, D.; Wang, S. Construction and Optimization of a Collaborative Harvesting System for Multiple Robotic Arms and an End-Picker in a Trellised Pear Orchard Environment. Agronomy 2024, 14, 80. [Google Scholar] [CrossRef]
  23. Xiong, C.; Xiong, J.; Yang, Z.; Hu, W. Path planning method for citrus picking manipulator based on deep reinforcement learning. J. South China Agric. Univ. 2023, 44, 473–483. [Google Scholar] [CrossRef]
  24. Hou, G.; Chen, H.; Niu, R.; Li, T.; Ma, Y.; Zhang, Y. Research on Multi-Layer Model Attitude Recognition and Picking Strategy of Small Tomato Picking Robot. Comput. Electron. Agric. 2025, 232, 110125. [Google Scholar] [CrossRef]
  25. Riboli, M.; Jaccard, M.; Silvestri, M.; Manconi, E.; Aimi, A.; Garziera, R. Collision-Free Motion Generation for Dual-Gantry Robotic Systems. Robot. Auton. Syst. 2025, 192, 105031. [Google Scholar] [CrossRef]
  26. Kramer, O.; Danielsiek, H. DBSCAN-Based Multi-Objective Niching to Approximate Equivalent Pareto-Subsets. In Proceedings of the 12th Annual Conference on Genetic and Evolutionary Computation; Association for Computing Machinery: New York, NY, USA, 2010; pp. 503–510. [Google Scholar]
  27. Han, K.; Wang, Y.; Tian, Q.; Guo, J.; Xu, C.; Xu, C. GhostNet: More Features from Cheap Operations. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 14–19 June 2020. [Google Scholar]
  28. Wang, J.; Gao, J.; Zhang, B. A Small Object Detection Model in Aerial Images Based on CPDD-YOLOv8. Sci. Rep. 2025, 15, 770. [Google Scholar] [CrossRef]
  29. Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
  30. Tang, L.; Yan, Q.; Xiang, X.; Fang, L.; Ma, J. C2RF: Bridging Multi-Modal Image Registration and Fusion via Commonality Mining and Contrastive Learning. Int. J. Comput. Vis. 2025, 133, 5262–5280. [Google Scholar] [CrossRef]
  31. Ma, X.; Dai, X.; Bai, Y.; Wang, Y.; Fu, Y. Rewrite the Stars. In Proceedings of the CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 5694–5703. [Google Scholar]
  32. Ma, S.; Xu, Y. MPDIoU: A Loss for Efficient and Accurate Bounding Box Regression. arXiv 2023, arXiv:2307.07662. [Google Scholar] [CrossRef]
  33. Ning, Z.; Luo, L.; Ding, X.; Dong, Z.; Yang, B.; Cai, J.; Chen, W.; Lu, Q. Recognition of Sweet Peppers and Planning the Robotic Picking Sequence in High-Density Orchards. Comput. Electron. Agric. 2022, 196, 106878. [Google Scholar] [CrossRef]
  34. Wang, K.; Yang, S.; Liu, X.; Song, J.; Li, Y.; Xie, F. Research on Target Recognition, Localization, and Picking Sequence Planning for Tomato-Picking Robots in Greenhouse Environments. Smart Agric. Technol. 2025, 12, 101628. [Google Scholar] [CrossRef]
Figure 1. Round-trip picking.
Figure 1. Round-trip picking.
Agriculture 16 00888 g001
Figure 2. Continuous picking.
Figure 2. Continuous picking.
Agriculture 16 00888 g002
Figure 3. Cucumber harvesting robot.
Figure 3. Cucumber harvesting robot.
Agriculture 16 00888 g003
Figure 4. Picking end-effector.
Figure 4. Picking end-effector.
Agriculture 16 00888 g004
Figure 5. Continuous harvesting strategy.
Figure 5. Continuous harvesting strategy.
Agriculture 16 00888 g005
Figure 6. Two Bezier curves: (a) Bezier curve. (b) Enhanced Bezier curve.
Figure 6. Two Bezier curves: (a) Bezier curve. (b) Enhanced Bezier curve.
Agriculture 16 00888 g006
Figure 7. Different image enhancement methods: (a) original image, (b) cropping, (c) vertical mirror projection, (d) brightness change, (e) translation, (f) rotation, and (g) Gaussian noise injection.
Figure 7. Different image enhancement methods: (a) original image, (b) cropping, (c) vertical mirror projection, (d) brightness change, (e) translation, (f) rotation, and (g) Gaussian noise injection.
Agriculture 16 00888 g007
Figure 8. Structure of YOLOv8-GCM.
Figure 8. Structure of YOLOv8-GCM.
Agriculture 16 00888 g008
Figure 9. Ghost Conv module.
Figure 9. Ghost Conv module.
Agriculture 16 00888 g009
Figure 10. C2f Ghost module.
Figure 10. C2f Ghost module.
Agriculture 16 00888 g010
Figure 11. CGFM.
Figure 11. CGFM.
Agriculture 16 00888 g011
Figure 12. Comparison of validation losses of different loss functions in improving the model.
Figure 12. Comparison of validation losses of different loss functions in improving the model.
Agriculture 16 00888 g012
Figure 13. Comparison of detection effects.
Figure 13. Comparison of detection effects.
Agriculture 16 00888 g013
Figure 14. The EBCP method was employed for planning cucumber harvesting sequences, with experiments conducted for scenarios involving different fruit quantities: (a) 10, (b) 14, (c) 18, (d) 22, and (e) 26 cucumbers. Dots of the same color denote belonging to the same harvesting cluster.
Figure 14. The EBCP method was employed for planning cucumber harvesting sequences, with experiments conducted for scenarios involving different fruit quantities: (a) 10, (b) 14, (c) 18, (d) 22, and (e) 26 cucumbers. Dots of the same color denote belonging to the same harvesting cluster.
Agriculture 16 00888 g014
Figure 15. Cucumber picking experiment in outdoor environment. (a) Cucumber picking robot experiment. (b) Fruit proximity to stem and trellis supports. (c) Stem and foliage overlap.
Figure 15. Cucumber picking experiment in outdoor environment. (a) Cucumber picking robot experiment. (b) Fruit proximity to stem and trellis supports. (c) Stem and foliage overlap.
Agriculture 16 00888 g015
Table 1. Comparison results of different loss functions.
Table 1. Comparison results of different loss functions.
LossP/%R/%mAP@50/%mAP@50:95/%
CIoU86.877.983.657.8
DIoU87.976.48356.2
GIoU86.977.383.455.4
EIoU8575.380.455
SIoU85.175.181.855.8
MPDIoU87.778.584.259.7
Table 2. Ablation experiment.
Table 2. Ablation experiment.
SchemeGHOSTCGFMMPDIoUP/%R/%FLOPs/GmAP@50/%Model Size/MB
1×××85.875.78.181.76.3
2××82.873.56.581.25.1
3××86.577.98.282.96.5
4××86.677.38.182.96.3
5×86.877.96.683.65.3
6×84.773.26.581.85.1
7×87.674.08.283.46.5
887.778.56.684.25.3
√ indicates that the corresponding module is adopted in the scheme; × indicates that the corresponding module is not adopted in the scheme.
Table 3. Comparison of detection results of different models.
Table 3. Comparison of detection results of different models.
ModelP/%R/%mAP@50/%FLOPs/GModel Size/MBInference Time/msTraining
Time/min
Faster R-CNN73.2 73.570.5140.1108.218.9385
SSD87.1 67.478.475.270.68.7182
YOLOv5n85.0 72.978.64.23.92.275
YOLOv7-tiny87.6 74.080.313.012.32.595
YOLOv8n85.875.781.78.16.31.978
YOLOv10n85.775.581.98.46.72.080
YOLOv11n86.176.382.16.55.51.872
YOLOv8-GCM87.7 78.584.26.65.31.769
Table 4. Simulation results for planning a robotic picking sequence for different numbers of cucumbers.
Table 4. Simulation results for planning a robotic picking sequence for different numbers of cucumbers.
Number of
Cucumbers
AlgorithmValue of
Traversal Path (m)
Number of
Collisions
Collision-Free Picking Rate (%)
10EBCP10.160100
RRT-CP10.02190
CP9.77190
RP31.52280
AYDY31.520100
K-G Method10.13190
14EBCP15.55192.86
RRT-CP13.72378.57
CP11.81378.57
RP44.11471.43
AYDY44.110100
K-G Method16.34285.71
18EBCP18.96194.44
RRT-CP17.66383.33
CP13.01288.89
RP58.10477.78
AYDY58.10288.89
K-G Method18.57383.33
22EBCP19.400100
RRT-CP18.21481.82
CP13.67481.82
RP68.55863.61
AYDY68.55290.91
K-G Method18.86382.61
26EBCP24.17196.15
RRT-CP22.69484.62
CP17.83676.92
RP81.231253.85
AYDY81.23292.31
K-G Method24.51484.62
averageEBCP17.650.696.69
RRT-CP16.463.083.67
CP13.223.283.24
RP56.70669.33
AYDY56.701.294.42
K-G Method17.682.685.25
Table 5. Experimental data on cucumber harvesting in outdoor environments.
Table 5. Experimental data on cucumber harvesting in outdoor environments.
Number of
Cucumbers
AlgorithmSimulate
Collision-Free Pickup Rates (%)
Actual Number of CollisionsActual Collision-Free Pickup Rate (%)
10EBCP1000100
RRT-CP1000100
CP90190
RP90280
AYDY1000100
K-G Method100190
14EBCP92.9192.9
RRT-CP71.4471.4
CP71.4471.4
RP64.3564.3
AYDY92.9192.9
K-G Method78.6378.6
18EBCP94.4288.9
RRT-CP88.9477.8
CP83.3477.8
RP77.8666.7
AYDY94.4288.9
K-G Method88.3388.3
22EBCP95.5290.9
RRT-CP77.3577.3
CP77.3672.7
RP68.2959.1
AYDY90.9386.4
K-G Method81.8577.3
26EBCP88.5388.5
RRT-CP73.1869.2
CP73.1869.2
RP57.71350
AYDY88.5388.5
K-G Method88.5676.9
averageEBCP94.261.692.24
RRT-CP82.144.279.14
CP79.024.676.22
RP71.6764.02
AYDY93.341.891.34
K-G Method87.443.682.22
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhao, C.; Wang, H.; Li, W.; Zheng, H.; Zhou, L.; Qian, M. Cucumber Robotic Continuous Harvesting: Enhanced YOLOv8n Detection and Dynamic Bézier Curve-Assisted Collision-Free Path Generation. Agriculture 2026, 16, 888. https://doi.org/10.3390/agriculture16080888

AMA Style

Zhao C, Wang H, Li W, Zheng H, Zhou L, Qian M. Cucumber Robotic Continuous Harvesting: Enhanced YOLOv8n Detection and Dynamic Bézier Curve-Assisted Collision-Free Path Generation. Agriculture. 2026; 16(8):888. https://doi.org/10.3390/agriculture16080888

Chicago/Turabian Style

Zhao, Chengheng, Huan Wang, Wenhao Li, Hengyi Zheng, Le Zhou, and Mengbo Qian. 2026. "Cucumber Robotic Continuous Harvesting: Enhanced YOLOv8n Detection and Dynamic Bézier Curve-Assisted Collision-Free Path Generation" Agriculture 16, no. 8: 888. https://doi.org/10.3390/agriculture16080888

APA Style

Zhao, C., Wang, H., Li, W., Zheng, H., Zhou, L., & Qian, M. (2026). Cucumber Robotic Continuous Harvesting: Enhanced YOLOv8n Detection and Dynamic Bézier Curve-Assisted Collision-Free Path Generation. Agriculture, 16(8), 888. https://doi.org/10.3390/agriculture16080888

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop