Next Article in Journal
Parameter Optimization of Companion Particles and Anti-Blocking Mechanism for Sesame Seed Metering in a Binary Sesame–Companion Particle System
Previous Article in Journal
Effects of Integrated Tillage-Surface-Cover Systems on Soil Hydrothermal Conditions, Cotton Growth, and Yield in Arid Xinjiang, China
Previous Article in Special Issue
Different Perspectives on the Same Target: Field and Laboratory Spectroscopy for Estimating Nitrogen Content in Sugarcane Leaves
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

System-Level Smart Robotic Harvesting for High-Value Greenhouse Crops: A Review

1
School of Agricultural Engineering, Jiangsu University, Zhenjiang 212013, China
2
School of Medicine, Jiangsu University, Zhenjiang 212013, China
*
Author to whom correspondence should be addressed.
Agronomy 2026, 16(18), 1795; https://doi.org/10.3390/agronomy16181795 (registering DOI)
Submission received: 15 August 2026 / Revised: 9 September 2026 / Accepted: 11 September 2026 / Published: 13 September 2026
(This article belongs to the Special Issue Smart Farming: Advancing Techniques for High-Value Crops)

Abstract

High-value greenhouse crops require quality-sensitive selective harvesting under labor shortages and variable crop conditions. This review synthesizes 188 unique primary studies and examines robotic harvesting as a complete task chain rather than a set of isolated sensing or manipulation modules. Embodied intelligence is used as an analytical lens to connect perception, harvestability assessment, decision making, manipulation, feedback, failure propagation, and recovery. A greenhouse-cucumber case study illustrates how errors propagate across these stages. The synthesis shows that the dominant bottleneck is system integration: incomplete observability, short-lived scene states, contact uncertainty, and weak outcome verification or recovery cause cumulative losses in complete-task success, effective throughput, and product quality. Most reported systems remain at closed-loop automation or early embodied interaction, with limited evidence of interaction-driven learning or cross-context generalization. We therefore frame evaluation around six production-relevant dimensions: complete-task success, effective cycle time and retry cost, product quality, failure detection and recovery, human intervention, and continuous operation. Near-term priorities are reliable, recoverable operation in defined crop–facility systems and standardized reporting; adaptation across cultivars and greenhouse configurations is a medium-term objective, whereas embodied learning and broad cross-platform transfer remain longer-term research directions. This system-level perspective reframes smart robotic harvesting around information continuity, fault containment, and deployable greenhouse performance.

1. Introduction

1.1. Research Background and Practical Demand

Protected agriculture uses controlled structures and intensive cultivation to support year-round production, but high-value crops such as tomatoes, cucumbers, sweet peppers, and strawberries still rely heavily on selective manual harvesting. Robot viability therefore depends not only on technical feasibility but also on complete-task success, cycle time, seasonal utilization, maintenance, energy demand, crop damage, and the amount of human supervision that remains after automation [1,2,3]. Commercial evaluation should therefore consider capital and retrofit cost, maintenance and service burden, energy use, effective labor substitution, product losses, and return on investment rather than the proportion of picking actions automated. Because this review does not perform a dedicated economic analysis, these variables are treated as deployment requirements to be demonstrated in greenhouse trials rather than as established economic outcomes.
Greenhouses are semi-structured rather than fixed industrial workspaces. Row and aisle geometry may be regular, but fruit position, maturity, occlusion, illumination, and local clearance change continuously, while fruits, peduncles, leaves, and branches deform when contacted [4,5,6]. A harvesting robot must therefore do more than detect a mature fruit: it must determine whether the target is harvestable, choose an approach and detachment action, execute safely, and verify the outcome.
Research has accordingly progressed from component-level studies of detection, localization, manipulators, and end effectors toward integrated systems that combine navigation or stationing, target sequencing, motion planning, local servoing, detachment, collection, and feedback [7]. However, strong module-level performance does not guarantee reliable harvesting because perception, calibration, planning, contact, and verification errors can accumulate across the task chain, while long-duration deployment adds safety, maintenance, and operating constraints [8].
This review therefore treats greenhouse harvesting robots as mobile manipulation systems that repeatedly update information and actions through crop–robot interaction. The central analytical question is how environmental and crop conditions shape perception, decision making, manipulation, feedback, failure propagation, and recovery across the complete harvesting process.

1.2. Conceptual Definition and Scope of Review

Embodied intelligence is used here as an analytical framework rather than a label applied to every harvesting robot. The review distinguishes greenhouse harvesting robots, embodied robots, embodied intelligence, and compliant interaction so that conventional AI-assisted or closed-loop systems are not automatically reclassified as embodied-intelligence systems.
A greenhouse harvesting robot selectively harvests individual fruits or fruit clusters that meet maturity, size, or marketability requirements in protected cultivation [9]. In this review, picking or detachment denotes the local end-effector action, whereas harvesting denotes the complete workflow, from target detection and maturity assessment through localization, harvestability assessment, task sequencing, approach, detachment, transfer, and outcome verification.
An embodied robot is an agent whose physical body contributes to perception and manipulation through its morphology, sensors, end effector, materials, actuation, and computation. For this review, embodied intelligence is the capacity to adapt task behavior through interaction with the greenhouse environment within a perception–decision–action–feedback cycle [10,11]. Deep learning detection or repeated execution of preset trajectories alone is therefore not treated as evidence of embodied intelligence.
Reported capabilities are interpreted at three levels: closed-loop automation, embodied interaction, and embodied learning. Closed-loop automation corrects predefined actions using visual, force, or operational feedback; embodied interaction additionally acquires information by changing viewpoints or physically interacting with the crop; and embodied learning uses accumulated interaction experience to improve adaptation or transfer. Most greenhouse harvesting studies currently fall within the first two levels, while online learning and cross-scenario generalization remain exploratory.
Compliant interaction denotes the ability to absorb positioning error, limit impact, and adapt to biological tissues through structure, materials, sensing, and control [12,13,14]. Passive compliance can arise from flexible or underactuated mechanisms, while active regulation uses vision, force, tactile, pressure, or current feedback. Softness alone is insufficient: the system must also recognize contact, retention, detachment, and release states and adjust actions accordingly.
The primary scope is selective harvesting of high-value greenhouse fruits and vegetables, especially tomatoes, cucumbers, sweet peppers, and strawberries. Supporting studies from orchards or fields are included only when they provide explicitly transferable evidence for a greenhouse task stage or constraint. The review covers platform positioning, perception and 3D scene representation, harvestability assessment, task and motion planning, manipulation, collection, feedback, failure diagnosis, and recovery; non-selective harvesting, transport-only systems, and studies without a clear manipulation link are excluded.

1.3. Related Reviews and Analytical Gap

Existing reviews cover integrated agricultural robots, perception, manipulators, end effectors, crop-specific harvesting, and embodied intelligence [15,16,17,18,19,20,21,22,23,24,25]. Their organizing units are usually a technology family, component, crop, or general robotics theme; they provide valuable module-level syntheses but devote less attention to cross-stage uncertainty, failure propagation, recovery, and production-relevant outcomes under greenhouse-specific constraints [26,27,28,29].
Accordingly, the analytical gap is system-level rather than terminological. Module-centered surveys ask which sensing, planning, or gripper methods perform well within a stage; crop-centered surveys organize methods by commodity; and broad embodied-intelligence reviews discuss interaction and adaptation across general robotics. This review instead uses the complete harvesting episode as the unit of analysis and follows how information, action effects, and faults cross stage boundaries from target discovery to outcome verification and recovery. Embodied intelligence is therefore an interpretive lens, not the novelty claim itself.
The review makes four linked contributions: (1) a complete harvesting-chain framework from target discovery and harvestability assessment to detachment, transfer, and outcome verification; (2) explicit analysis of uncertainty and failure propagation across perception, localization, approach, contact, detachment, and collection; (3) the integration of feedback, diagnosis, recovery, human intervention, and continuous operation into system-level evaluation; and (4) a three-level interpretation—closed-loop automation, embodied interaction, and embodied learning—that positions current evidence without reclassifying conventional harvesting robots as embodied-intelligence systems.

1.4. Literature Search Strategy and Study Selection

This study is a comprehensive narrative and technical review of selective robotic harvesting in protected cultivation. A structured search and screening procedure was adopted, with selected PRISMA 2020 and PRISMA-S items used solely as reporting aids to improve transparency and reproducibility [30,31]. Because the evidence spans heterogeneous algorithms, mechanisms, subsystems, and integrated robots, no meta-analysis or single formal risk-of-bias instrument was applied. Instead, we used a six-domain descriptive rubric: validation environment, experimental scale, test duration, denominator-based outcome reporting, reproducibility, and system completeness. Each domain was scored from 0 to 2 based on information reported in the full text; missing or ambiguous information received 0 rather than being inferred. The resulting 0–12 score provides a qualitative indication of confidence in cross-study interpretation rather than ranking crops, platforms, or algorithms.
The scoring anchors were defined as follows: For validation environment, 0 indicated simulation or image-only evaluation, 1 a laboratory or highly controlled greenhouse test, and 2 an operational greenhouse or field/farm setting under variable conditions. For experimental scale, 0 indicated a single object or demonstration, 1 limited batches or plants, and 2 repeated trials across multiple plants, rows, or targets. For test duration, 0 indicated a single run or snapshot, 1 one session or short series, and 2 repeated sessions across days or harvest cycles. For denominator-based outcome reporting, 0 indicated no denominator, 1 a partial or ambiguous denominator, and 2 explicit denominators for targets, attempts, failures, and retries where relevant. For reproducibility, 0 indicated insufficient hardware, data, or procedural detail, 1 partial detail, and 2 a sufficiently specified setup and protocol for independent reconstruction. For system completeness, 0 indicated an isolated module without a harvesting link, 1 a subsystem with a defined interface to another task stage or an integrated chain without explicit outcome verification, and 2 an integrated chain including manipulation and outcome verification. Scores of 0–4, 5–8, and 9–12 were interpreted as low, moderate, and high evidence strength, respectively. The same fixed anchors were applied to all 188 unique primary studies before aggregate counts were calculated.
Although fixed scoring anchors were applied consistently across the evidence corpus, domain-level scoring necessarily involves some interpretation of the information reported in individual studies and may therefore retain a degree of reviewer subjectivity. The scoring was conducted by a single reviewer, and no independent duplicate scoring was performed. To reduce discretionary judgment, missing or ambiguous information was assigned the lowest score for the corresponding domain, and the same predefined anchors were applied uniformly to all studies. The resulting scores should therefore be interpreted as a structured indication of evidence strength rather than as an objective or bias-free measurement.
Among the 188 unique primary studies, 55 (29.3%) were rated as high strength, 133 (70.7%) as moderate, and none as low; the median score was 8/12. These categories describe validation and reporting strength under the fixed engineering rubric rather than formal risk of bias. The corpus is therefore dominated by moderate-strength evidence, while high-strength evidence remains a minority, reflecting recurring limitations in test duration, denominator reporting, reproducibility, or complete-system validation. This pattern limits direct performance comparison and motivates the system-level emphasis of this review.
Evidence strength and evidence origin are treated as separate dimensions. Greenhouse/protected-cultivation evidence is distinguished from laboratory or controlled validation and from non-greenhouse studies used only for transferable mechanism-level inference. Where non-greenhouse evidence supports a conclusion, its role is qualified explicitly rather than used to transfer numerical performance. These context distinctions do not alter the 0–12 evidence-strength score and are intended to prevent results from heterogeneous validation settings from being read as a common ranking.
The principal bibliographic sources were four scholarly databases and search platforms: Web of Science Core Collection, IEEE Xplore, ScienceDirect, and SpringerLink. A supplementary reference-based route was also used: 200 seed records from the manuscript reference set were screened with Google Scholar-assisted bibliographic verification. This supplementary route was not treated as an independent reproducible database search. The four formal database searches were completed on 30 July 2026 and were limited to publications from 1 January 2000 to that date. The supplementary reference-based route was not subject to the same publication-year filter; it contributed one eligible 1999 primary study [32], which was retained because it satisfied the review’s technical inclusion criteria. The structured search combined greenhouse/protected-cultivation terms with harvesting terms and either robotic-system or intelligent-technology terms. Database-specific search fields, exact query strings, search dates and coverage, recorded language and document-type restrictions, result counts, and export formats are reported in Supplementary Table S1. Because the complete Boolean expression was not used as a single ScienceDirect search, the Group 3 and Group 4 subsearches were executed separately, returning 64 and 55 records, respectively; 49 exact-DOI overlaps were subsequently reconciled during deduplication. Exact DOI matching was the only formal pre-screen deduplication step in the frozen screening workflow; no secondary title/author/year/version reconciliation was performed before screening. Following reviewer feedback, a post-revision bibliographic reconciliation by title, author, and year was conducted to identify any residual duplicates missed by DOI-only matching.
Title/abstract screening, full-text eligibility assessment, and data extraction were all conducted by a single reviewer (J.W.). No independent duplicate screening or data extraction was performed; the potential for reviewer-dependent selection or extraction error is therefore acknowledged explicitly in Section 6.4.
Studies were included when they investigated a robot or harvest-oriented subsystem for selective harvesting, addressed a defined task stage, and provided experimental or implementation evidence sufficient for technical interpretation. Non-greenhouse evidence was retained only when four conditions were jointly satisfied: (i) the same harvesting-task function was involved; (ii) the biological or environmental constraint was functionally analogous; (iii) the relevant sensor, robot, or control interface and output could be interpreted for greenhouse harvesting; and (iv) the validation conditions supported mechanism-level inference. Such studies were treated as transferable evidence for mechanisms or design requirements, not as a basis for transferring numerical performance rankings. Failure to satisfy these conditions made the evidence context-specific. Studies unrelated to robotic harvesting, transport-only work, generic algorithms without a harvesting link, duplicate reports, reviews without original technical evidence, and technically insufficient reports were excluded.
The searches yielded 769 records: 569 from the four principal scholarly sources and 200 supplementary reference-based seed records screened with Google Scholar-assisted bibliographic verification. After removing 101 records by exact DOI matching, 668 records proceeded to title/abstract screening; 199 full texts were assessed, 10 were excluded, and 189 reports were retained after full-text assessment (154 database-identified and 35 supplementary). The post-revision title/author/year reconciliation identified one residual duplicate report—“Strategies for Selecting Best Approach Direction for a Sweet-Pepper Harvesting Robot” [33]—because the database record contained DOI 10.1007/978-3-319-64107-2_41 whereas the supplementary citation record lacked a DOI. The final evidence corpus therefore comprised 188 unique primary studies for study-level synthesis and 0–12 evidence scoring. Of these 188 primary studies, 53 are cited directly in the main manuscript, while all 188 are documented in Supplementary Table S2 together with bibliographic metadata and the study-level evidence assessment. Additional supporting references—including reviews, greenhouse-management and crop-quality studies, selected mechanism-level technical sources, and prospective robot-learning literature—were used for interpretation but were not included in the study-level evidence-quality scoring. Figure 1 summarizes the screening flow and post-revision reconciliation.

1.5. Research Framework and Section Organization

Within this review, smart farming defines the application context, selective robotic harvesting is the research object, and embodied intelligence provides the analytical perspective. Rather than organizing the literature around isolated modules, the review treats the complete harvesting task as the unit of analysis and asks how greenhouse geometry, illumination, occlusion, crop morphology, and biological compliance shape observation, reachability, manipulation, and feedback.
Transferability is assessed separately, from evidence strength through crop compatibility, facility compatibility, and platform compatibility. Findings are treated as directly transferable only when task function, operating constraints, interfaces, and validation conditions are comparable; otherwise, they are considered conditionally transferable or context-specific. Performance values are not generalized across crops, facilities, or robot platforms without cross-context validation.
System evaluation is organized around six production-relevant categories: complete-task success, effective cycle time and retry cost, product quality, failure detection and recovery, human intervention, and continuous-operation capability. These measures complement module-level metrics and make it possible to trace how an error at one stage changes the probability and cost of later stages.
Outcome verification and exception handling are therefore part of the harvesting chain rather than auxiliary functions. Vision, force, tactile, pressure, actuator, and device-state information can trigger correction, replanning, target deferral, or recovery. The embodied-intelligence perspective is used to interpret the degree of closed-loop integration, interaction, and adaptation demonstrated by these behaviors.
A greenhouse-cucumber case study is used to make these dependencies concrete because cucumber harvesting combines strong visual similarity to foliage, slender and often occluded peduncles, restricted high-wire workspaces, and contact-sensitive detachment. The case follows target recognition, local observation, harvestability assessment, ordering, approach, retention, cutting, transfer, and verification to show how failures propagate and how recovery can be organized.
The review proceeds from environment and crop constraints (Section 2) to system requirements (Section 3), implementation architectures and subsystems (Section 4), and the cucumber task-chain case study (Section 5). Section 6 synthesizes the principal system-level findings, bottlenecks, and research priorities for reliable, adaptive, and eventually more generalizable harvesting.

2. Greenhouse Harvesting Environments and Crop Characteristics

2.1. Greenhouse Types and Cultivation Systems

Greenhouse facilities vary in covering, span, floor conditions, illumination, environmental control, and load capacity. Construction age, investment level, and regional climate add further heterogeneity. A scoping review covering studies published through 2023 included 257 of the 330 relevant publications across diverse structures, crops, and robotic tasks [34]. Performance from one facility therefore cannot be assumed to transfer directly to another; facility characteristics should be treated as prior constraints on robot mobility, sensing, manipulation, and evaluation.
Cultivation practices determine both crop geometry and usable workspace. Tomatoes, cucumbers, and sweet peppers are commonly trained into tall row canopies, whereas strawberries are often grown on elevated gutters or benches. These systems constrain rows, working heights, and material-flow routes but not the exact pose or visibility of individual fruits. Growth, fruit loading, and differences in pruning continually change the local workspace. Greenhouses are therefore semi-structured environments: facility geometry is comparatively regular, whereas crop geometry remains variable at the robot’s operating scale. Typical systems are shown in Figure 2.
Crop management should be treated as part of the harvesting system rather than as background. Cultivar selection, training systems, pruning and canopy architecture, plant spacing, fruit load, and harvest scheduling alter the target geometry, organ visibility and occlusion, manipulator accessibility, tissue state, detachment conditions, and time available for robotic operation. Comparative studies should therefore report these variables together with robot performance, and agronomic co-design should be evaluated against any added labor or effects on yield and quality [35,36]. At the facility level, recent intelligent greenhouse-control research also shows that climate and energy-management strategies can materially affect resource efficiency, reinforcing the need to integrate harvesting robotics with broader greenhouse operation rather than treat it as an isolated automation layer [37].
Improving robotic accessibility through changes in crop architecture or training does not necessarily maximize agronomic or operational performance. More open canopies or wider access corridors may improve visibility, reachability, and collision avoidance but may also require lower planting density or alter light interception, microclimate, pruning demand, yield, and marketable quality. Conversely, crop configurations optimized primarily for yield or fruit quality may increase occlusion and manipulation difficulty. Robotic accessibility, yield, marketable quality, added crop-management labor, and facility constraints should therefore be evaluated jointly as co-design objectives [35,36].
Existing greenhouse-robotics reviews likewise show strong dependence on crop and validation context, with many studies still relying on modeling, laboratory testing, or limited greenhouse trials [36]. Cultivation systems influence target distribution, locomotion mode, sensor placement, manipulator workspace, and detachment motion; changes in these conditions may therefore require recalibration or physical redesign rather than simple software transfer [38].
Greenhouse type also affects deployment cost, maintenance, and commercial viability. A European Mediterranean greenhouse case study identified spraying and harvesting as priority robotic applications, while technology readiness, regulation, and profitability still constrained deployment [39]. Robot specifications should therefore state the target crop, cultivation regime, covering structure, floor or rail conditions, hygiene procedures, and crop-replacement cycle. These boundary conditions determine whether laboratory performance can translate into stable production capacity [40,41].
Integrated harvesting robots are most readily deployed in high-technology greenhouses, where standardized rows, level or rail-guided mobility, reliable utilities and data links, controlled climate, standardized management practices, and maintenance support reduce system uncertainty. Medium- and low-technology protected systems impose different constraints—uneven floors, variable aisles, natural ventilation, limited power and data infrastructure, heterogeneous crop management, and tighter investment and service capacity—and may favor lighter, modular, shared, or semi-autonomous architectures. Claims of commercial readiness should therefore specify the target facility class, required retrofits, operator input, management assumptions, and maintenance support [34,36,37,38,39,40,41].
Among the configurations reviewed in Section 4.4, free-ranging wheeled platforms combined with compact or task-specific manipulators, together with operator-supervised semi-autonomous operation, appear most compatible with medium-technology commercial greenhouses because they can use existing aisles and crop rows without requiring major reconstruction. Their principal practical limitations include repeatable stationing, slip and turning in narrow aisles, payload stability, and continued human intervention. Fixed-rail, gantry, and highly integrated multi-arm systems may provide greater workspace control or throughput but generally require more standardized layouts, supporting infrastructure, and coordination. Medium-technology deployment should therefore be viewed as a modular, retrofit-minimizing pathway rather than as evidence of immediate commercial readiness [34,36,38,39].
Overall, greenhouses exhibit a degree of regularity and manageability at the macroscopic layout level, but the facility type, cultivation regime, and crop growth state jointly create substantial scenario variability. Robot design should therefore shift from pursuing universally applicable prototypes toward facility-specific co-design of the robotic platform and cultivation system, with operating boundaries, environmental interfaces, and maintenance requirements defined before deployment.

2.2. Greenhouse Spatial Structure and Access Aisles

Spatial constraints occur at three nested scales: the aisle used by the platform, the manipulator’s reachable space, and the local clearance needed for contact and detachment. Although row topology is regular, aisles may contain rails, channels, posts, irrigation lines, support strings, crates, workers, and protruding foliage. Effective width changes with crop growth and management, while wet floors can increase wheel slip, odometric drift, and braking risk. Platform, arm, and end-effector dimensions must therefore be assessed together rather than against aisle width alone.
At the mobility scale, some greenhouses provide heating-pipe rails for guidance and load support, but chassis geometry must match the rail gauge and end-of-row conditions. Free-ranging platforms must handle narrow passing, turning, uneven floors, and temporary obstacles. Navigation and guidance accounted for 31% of studies in one greenhouse-robotics review, spanning monorail, pipe-rail, and free-ranging solutions [36]. Platform dimensions, turning radius, stability, and localization should consequently be matched to the facility rather than inferred from regular aisle topology.
Manipulator design should be matched to crop geometry and approach requirements rather than maximizing degrees of freedom. Narrow canopy gaps demand sufficient pose adjustment while minimizing arm volume, collision risk, control complexity, and cost; task-specific optimization studies show that fewer degrees of freedom can be adequate when target height and end-effector pose requirements are well constrained [42].
Across sweet-pepper studies, simplifying canopy occlusion and improving fruit exposure consistently increased complete harvesting success, despite differences in robots and evaluation protocols [43,44]. The common implication is more important than any single percentage: local manipulation space, fruit visibility, and canopy management directly affect reachability and end-effector access. Robotic adaptation should therefore be complemented by robot-aware crop training and pruning.
Overall, greenhouse aisles have relatively regular topology, but their effective trafficable dimensions and local manipulation spaces change continuously with crop growth and production management. Rigid infrastructure and compliant plants jointly form dense obstacles, coupling chassis trafficability, manipulator reachability, and end-effector approach pose; these factors must therefore be co-optimized at the system level.

2.3. Greenhouse Illumination and Visual Conditions

Greenhouse coverings reduce wind and rain but do not produce constant illumination. Solar angle, clouds, aging covers, shade screens, supplemental lights, frames, wet leaves, and fruit surfaces create rapid changes in exposure, reflection, and local contrast. Camera or manipulator motion can also alter color and depth quality. Vision systems therefore need robustness to illumination variation, complex backgrounds, target overlap, and crop differences; the camera type and mounting directly affect recognition, localization, and visual servoing [45,46,47,48].
High humidity further increases uncertainty in visual imaging. Spraying, crop transpiration, and temperature differences can produce airborne mist or lens condensation. Water droplets and pesticide residues on leaves may cause scattering, halos, and local blur, while dust and condensation can compromise structured-light, stereo-matching, and time-of-flight depth measurements. A dataset developed for cherry-tomato harvesting includes normal light, bright light, low light, water mist, severe occlusion, and fruit overlap, capturing variation in the appearance and visibility of the same target class under different local conditions [49,50].
Sunlight, shadow, foliage occlusion, and fruit overlap are also represented in the tomato dataset shown in Figure 3 [51]. The samples illustrate how illumination, clustering, and organ relationships jointly shape fruit appearance. Correct-recognition rates were similar in sunlight and shadow, whereas false positives and missed detections rose with occlusion.
Scenario-specific cherry-tomato experiments show that illumination and humidity affect perception differently: ordinary brightness variation can often be mitigated through training and exposure control, whereas water mist and condensation directly reduce contrast and depth quality [49]. The system-level implication is that environmental robustness must be assessed under the conditions that feed subsequent localization and manipulation rather than inferred from a single image-level metric.
High fruit-detection scores do not imply that a target is ready for manipulation. Harvesting also requires depth, organ connectivity, detachment site visibility, safe approach direction, and uncertainty. RGB, stereo, structured-light, ToF, and wrist-mounted sensing offer complementary strengths and limitations under greenhouse illumination, reflection, low texture, and occlusion [45,52,53,54,55,56,57]. Perception should therefore be judged by whether it supplies a stable, actionable scene state for harvestability assessment, planning, and local control rather than by image-level accuracy alone.
Practical harvesting systems therefore generally combine active illumination and shading, automatic exposure, high-dynamic-range imaging, RGB-D perception, multisensor fusion, local verification at the end effector, and multiview observation while using continuous sensing and online correction to improve the reliability of target-state estimation [45,49,53,54,55]. The objective of greenhouse visual perception should not be limited to improving detection accuracy in individual images. Equal attention should be paid to stability under changing illumination and humidity, recovery of organ relationships under occlusion, and whether perception outputs provide effective support for harvestability assessment, motion planning, and contact-rich manipulation [56,57].

2.4. Crop Morphology, Spatial Distribution, and Biophysical Interaction Characteristics

2.4.1. Crop Morphology and Spatial Distribution of Targets

Greenhouse harvest targets form a hierarchy from the plant and branch or cluster to the fruit and detachment organ. A robot therefore needs a relational crop state rather than an isolated fruit centroid. Fruit geometry affects end-effector aperture and contact, peduncle geometry constrains detachment, and relative positions within a cluster affect viewpoints, approach paths, and harvesting order. At the crop scale, reachability is consequently expressed through three linked conditions: whether a target is visible, approachable, and detachable.
Standardized cultivation constrains approximate target distribution but does not remove interfruit variability. Table 1 summarizes representative cherry-tomato geometry and mass [58]. The large cluster span and approximately 12 fruits per cluster create a dense, multidepth target set that requires layered observation and target-order planning.
These measurements can inform end-effector dimensions, but planning must still accommodate cluster-scale variation, maturity, orientation, and hidden attachment structures. Target representation should therefore combine fruit geometry with organ connectivity and approachability; location or maturity alone cannot establish safe harvestability [59,60,61,62,63,64,65,66].

2.4.2. Occlusion, Overlap, and Structural Uncertainty

Unlike the illumination changes discussed in Section 2.3, structural occlusion does not merely degrade image quality; it directly removes information required to complete an action. Insufficient illumination or glare can often be mitigated by supplemental lighting, exposure adjustment, and data augmentation. By contrast, a peduncle fully hidden by a leaf, an attachment point behind a fruit, or a rear target covered by foreground fruit cannot be recovered from the current viewpoint. Occlusion in greenhouse canopies includes external occlusion of fruits by leaves and branches, mutual overlap among fruits, self-occlusion of the peduncle by the fruit, and occlusion of safe cutting regions by main stems and lateral branches. These forms often occur simultaneously, creating a continuous error chain across target detection, organ-relationship inference, and manipulator approach planning.
Fruit visibility does not imply visibility of the detachment site. Sweet-pepper peduncles are generally located at the top of the fruit, may be markedly curved, and can even lie against the fruit surface. Moreover, peduncles, leaves, branches, and immature green fruits have similar color characteristics [63]. In the scene shown in Figure 4, mature sweet peppers are relatively easy to observe, whereas the peduncle orientation, attachment location, and relationships with surrounding foliage may remain occluded. Target recognition can identify fruit location, but robotic harvesting must additionally infer the fruit-plant connection, a feasible approach direction, and the precise detachment site [64,65,66].
Sweet-pepper peduncle studies illustrate the gap between fruit detection and executable harvesting: structural constraints improve rejection of implausible peduncle candidates, yet small, curved, and occluded attachment organs remain substantially harder to localize than fruits themselves [63]. The system-level bottleneck is therefore not only class recognition but recovery of the fruit–peduncle–stem relationship needed for safe approach and detachment.
Structural uncertainty also has a temporal dimension. Ventilation, plant motion, and manipulator approach alter leaf positions. Removing one fruit may release loads carried by a peduncle or branch, changing the poses and occlusion relationships of adjacent fruits; even slight contact between the end effector and the canopy may rapidly invalidate a previously acquired three-dimensional model. A greenhouse harvesting scene should therefore not be treated as a static environment that remains unchanged after a single mapping operation but as a local state that must be continually updated through observation and manipulation. The occlusion problem thus extends from information visibility to object response after contact [55,67].

2.4.3. Compliant Deformation, Fragility, and Detachment Characteristics

Fruits, peduncles, branches, leaves, and stems differ in compliance and mechanical response. Gripping, pulling, twisting, or cutting loads propagate through the plant, causing cluster motion, foliage displacement, and possible collisions. Contact can also shift the target and convert localization error into contact-location or load-direction error. Detachment is therefore a biomechanical interaction governed by tissue strength, load direction, contact location, and action duration, rather than a purely geometric move to a coordinate [68,69,70].
For robotic harvesting, biologically relevant traits should be treated as an operating envelope rather than represented by a mean fruit geometry. Fruit-size variation and mass determine whether the end effector can enclose and retain the target, while maturity progression changes the color, firmness, tissue strength, and detachment response. The peduncle length, diameter, curvature, orientation, and visibility determine whether the detachment site can be observed and reached without contacting the main stem or neighboring fruit. Detachment force and the presence or absence of a distinct abscission zone govern the required action and the load transmitted to the plant. Market-quality requirements, including allowable bruising, calyx or stem retention, and maturity class, define success more narrowly than fruit removal alone. Evaluation should report trait ranges and stratify complete-task success, damage, and marketable yield by cultivar and maturity stage [58,63,68,69,70]. A non-greenhouse orange-harvesting study further associated greater occlusion and pose deviation with higher fruit damage and higher illuminance with lower damage; these study-specific relationships are used here only as transferable mechanism-level evidence and are not generalized numerically to greenhouse crops [71].
Table 2 compares four cherry-tomato detachment modes and shows that action choice affects both removal and the subsequent crop state [58]. Methods that reduce cluster disturbance can preserve the validity of later observations, whereas actions that transmit larger loads may change neighboring target poses and occlusion.
No detachment mode is universally superior: the preferred action depends on crop mechanics, visibility of the attachment structure, required market condition, end-effector design, and available control authority [58].
Soft end effectors can accommodate size and pose variation by increasing contact area and reducing local pressure, but compliance alone does not imply damage-free handling. Insufficient stiffness may weaken retention, transmit detachment force unreliably, or allow the cutting site to move. Preset pressure or displacement also cannot confirm grasping, detachment, or slip [72]. Effective designs therefore treat compliance as a controlled system property, combining mechanical accommodation with visual, force, tactile, or pressure feedback. Once contact begins, robot action changes the crop scene; this transition from observation to state-changing interaction is an important source of embodiment [73]. Representative soft-gripper configurations are shown in Figure 5.

2.4.4. Preharvest Environment and Postharvest Quality

Robot harvestability is also conditioned by the physiological state in which fruit reaches harvest. Temperature and light, together with relative humidity, water availability and irrigation, and nutrient supply, influence quality formation and shelf-life potential in species-specific ways [74]. These conditions affect firmness, cuticle integrity, water status, pigmentation, composition, and susceptibility to bruising or decay. Fruit may therefore satisfy geometric picking criteria yet tolerate contact, suction, or detachment poorly. Evaluation should record relevant preharvest conditions and crop management and assess marketable quality after robotic handling, using crop-appropriate measures such as visible damage, firmness, water loss, color, calyx or stem retention, decay incidence, and storage life. A strategy that raises detachment success at the expense of postharvest quality cannot be considered commercially effective.
Overall, greenhouse harvesting conditions form a coupled system: facility and cultivation define operating boundaries; canopy structure and illumination determine observability and reachability; and crop morphology, fragility, and deformation govern safe interaction. These dependencies create a continuous pathway from perception uncertainty to approach error, contact outcome, scene change, and recovery, motivating the system requirements developed in Section 3.

3. Technical Requirements for Greenhouse Harvesting Robots

To distinguish reported evidence from author synthesis, Section 3, Section 4 and Section 5 use two forms of statement. Descriptions of implemented capabilities are tied to cited studies, whereas prescriptive constructs introduced as “this review proposes” or “this review recommends” are synthesis principles for system design and evaluation. The latter should not be read as features already validated together in a single harvesting platform.

3.1. Requirements for Multimodal Perception and Scene Understanding

Based on the environmental uncertainties identified in Section 2, this review proposes an actionable scene state that includes target identity, maturity, three-dimensional pose, organ relationships, free space, detachment site, approach region, and estimate confidence. Perception supports manipulation only when these outputs describe connectivity and conditions for safe approach [45,75,76,77,78,79]. Section 4 then examines how existing hardware and software implementations approximate these requirements.
This scene state must span the different scales associated with global search, the target neighborhood, and the periods before and after contact. Long-range observations delimit the operating area, target distribution, and robot pose, whereas close-range observations supplement information about organ connectivity, local clearance, and relative pose; the two must be integrated within a common temporal and coordinate framework. Manipulator motion, foliage movement, and fruit removal continuously alter prior observations. Target identity, pose, and occlusion relationships must therefore be updated as actions unfold, rather than treating a one-time estimate as indefinitely valid.
The appearance, geometry, and contact state each have distinct information blind spots. The system must exploit the complementarity of multimodal information rather than merely adding sensors. Fusion should address, at a minimum, calibration and synchronization, missing data, conflicting observations, confidence propagation, and information freshness while adjusting modality weights according to task stage: prior to contact, emphasis is placed on target and spatial relationships; after contact, the relative motion, load, and holding state become more important [80,81]. The fundamental requirement for multimodal perception is therefore reliable observability of critical states, not a predetermined fixed sensor suite [48,56,57].
Rapid, non-destructive optical sensing can be used to acquire biological information from protected-crop vegetables, but sensor compatibility, environmental variation, and model stability continue to affect practical deployment [82]. Greenhouse tomato image datasets are often limited by sample scarcity and class imbalance; data augmentation based on the geometric locations of salient features can improve the training-data distribution [83].
Multimodal information must ultimately be organized into a unified representation for manipulation. Using greenhouse cherry tomatoes as an example, Figure 6 illustrates the correspondence among RGB images, point-cloud normal information, and target-detection results. The object layer records class, maturity, position, pose, and detachment site; the relational layer describes organ ownership, occlusion order, and traversable space; the robot-state layer provides the states of the platform and manipulator; and the confidence layer specifies estimation error and validity duration. This representation should support the current action while retaining target identity and the provenance of state estimates for subsequent updates, allowing local verification to correct the scene model rather than reconstruct it in its entirety.
When the peduncle is invisible, depth data are missing, or local structure is insufficient, the system must actively modify its observation conditions. A new viewpoint should be selected jointly according to expected information gain, motion cost, collision risk, and task urgency; limited movements of the camera, manipulator, or platform can then be used to recover critical structure. The purpose of active perception is not to increase the number of observations but to determine whether missing information affects safe manipulation and to resolve, with the minimum necessary action, the uncertainty to which the decision is most sensitive [45,80,85,86].
Continuous harvesting further requires scene understanding to be relational, temporal, and explicitly uncertainty-aware. The system should continuously update the fruit-peduncle association, target identity, occlusion order, and free space while distinguishing among three states: the target is genuinely absent, temporarily invisible because of occlusion, or already harvested. Verification should be triggered when observations expire or confidence is insufficient. A scene state should therefore describe not only current observations but also their reliability, the information that remains to be confirmed, and the actions that may change the state [75,80].
Supporting non-greenhouse examples are used here only to illustrate transferable perception functions, not greenhouse performance. Field studies on broccoli and strawberries extend detection toward growth-pose estimation, picking-point localization, and peduncle-inclination estimation [87,88,89]. Together, these studies illustrate the progression from class recognition to manipulation-oriented state estimation; their numerical performance is not generalized to greenhouse harvesting.
Perception outputs should therefore be action-oriented, continuously updatable, and accompanied by explicit confidence estimates. In addition to recognition accuracy, evaluation should cover three-dimensional localization error, visibility of critical organs, correctness of relational inference, scene-update latency, information completeness, and quality of uncertainty calibration. These metrics must be interpreted according to the validation level. Offline or laboratory tests report image-level performance on curated datasets or controlled scenes, whereas commercial harvesting performance must be established through integrated greenhouse trials in which perception guides approach, detachment, and collection under production conditions. High precision, recall, or mAP cannot account by itself for false detections that trigger wasted motions, missed marketable fruit, pose errors, fruit damage, or failed operations. Unless these downstream outcomes and all eligible targets are retained in the evaluation, results should be described as module-level perception performance rather than evidence of commercial harvesting capability.
Production-level reporting should make the unit of success explicit. At minimum, studies should report marketable fruit collected per hour (including travel, unloading, retries, and downtime), complete-task success and unharvested-fruit rate per eligible target, damage or downgrade rate, human interventions per hour, energy use per fruit or kilogram, and operational availability with fault and recovery time. These measures should be calculated over consecutive harvests with all eligible targets counted; otherwise, a fast result on a small or preselected sample can be mistaken for production capacity.

3.2. Requirements for Harvestability Assessment, Task Decision Making, and Motion Planning

After acquiring the scene state, the robot must still determine whether manipulating the current target is appropriate and feasible. Harvestability is not a static label determined by maturity alone but a task state formed jointly by the target’s biological condition, visibility of critical organs, spatial reachability, end-effector pose margin, collision and damage risks, and perceptual confidence. Its purpose is to translate target discovery into a decision about the next action and to define safety boundaries for task-level decision making.
Building on the idea that perception objectives should change with task demands [90], this review proposes that harvestability be represented as a set of convertible graded states rather than a binary classification. Targets with sufficient information and satisfactory operating conditions may proceed directly to harvesting; unclear critical structures should trigger additional observation; inappropriate robot position or approach direction should prompt platform or manipulator adjustment; temporarily high-risk targets may be deferred; and persistently unreachable targets or those likely to be damaged should be skipped. Each state should include the evidence supporting the judgment, its validity period, and conditions for reassessment, enabling targets to transition among states as viewpoint, harvesting order, and canopy configuration change.
Task-level decision making then selects among observation, movement, harvesting, deferral, and withdrawal while maintaining dynamic priorities under multi-target conditions. Ranking should not be based solely on target distance but should jointly consider expected motion and manipulation costs, perceptual uncertainty, collision and damage risks, the effect of target removal on subsequent visibility, and system constraints such as energy and collection capacity [33,91,92,93,94]. Online decisions should be updated in response to new observations and completed actions [95,96], while target-density estimation can inform harvesting strategy [97]. These decision requirements are summarized in Table 3.
Motion planning converts task-level decisions into pre-manipulation poses, approach trajectories, and retreat paths. In addition to rigid infrastructure, the planner must distinguish compliant foliage that may tolerate limited contact and account for joint limits, pose margins, end-effector geometry, cables, and postharvest transfer. Geometric reachability alone is insufficient; a target is executable only when approach, interaction risk, and retreat are simultaneously feasible.
Because the scene changes during approach and manipulation, replanning should be triggered only when new observations invalidate existing assumptions. The motion layer should return the cause of infeasibility—insufficient information, unsuitable robot position, or blocked local geometry—so that the task layer can request re-observation, repositioning, a new approach, or target deferral. Evaluation should report invalid-approach rate, collision-free arrival, update latency, task completion, and risk events in addition to success rate and cycle time.

3.3. Requirements for Compliant Manipulation and Closed-Loop Interaction

Once the target, approach direction, and pre-manipulation pose have been established, the task transitions from non-contact motion to physical interaction. Position and orientation errors now translate directly into errors in contact location, action direction, and load, while inter-object variability changes the actual response to gripping, cutting, twisting, or pulling. Compliant manipulation should therefore not be understood as executing a prescribed detachment motion; rather, it requires control of the contact process as object state continuously changes, with risk maintained within recoverable bounds.
Compliance must provide both error accommodation and controllability. The structure or controller should absorb small pose errors, enlarge the effective contact region, and limit local pressure; however, the holding and detachment stages must still provide sufficient support and force transmission to prevent excessive compliance from causing slip, incomplete cutting, or distortion of the intended action. The safe operating envelope should be adjusted according to crop type, maturity, detachment method, and current manipulation stage, rather than applying a single fixed stiffness or load threshold to all objects.
Action-mode selection should likewise be based on effects on the complete task. In addition to whether the target can be detached, the system should compare visibility of the detachment site, required local workspace, end-effector pose and holding requirements, and disturbance to the plant, neighboring fruit, and stability of subsequent scenes [58]. At the requirements level, these differences can be summarized by a common principle: the action mode must match current observability, robot capabilities, allowable loads, and the objective of continuous harvesting while providing alternative strategies or a termination path for targets that cannot be handled reliably. The corresponding stage-specific feedback requirements are summarized in Table 4.
Stage-specific feedback should combine vision with force, tactile, pressure, and actuator signals. Approach feedback corrects relative pose and clearance; contact feedback detects collision or abnormal load; holding feedback verifies retention; and post-detachment feedback confirms release and collection. No single modality should determine completion independently, particularly when signals differ in sampling rate, latency, and noise [98,99,100,101,102,103,104].
Compliant manipulation therefore requires feedback to drive explicit state transitions and bounded corrective actions throughout approach, contact, holding, detachment, and transfer. When displacement, slip, or incomplete detachment is detected, the robot should adjust its pose or action parameters, verify the target again, or retreat safely. The operation should terminate when force, collision risk, or elapsed time exceeds predefined limits, thereby linking physical interaction with outcome verification and subsequent task decisions.

3.4. Requirements for Continuous Operation, Autonomous Recovery, and System Coordination

Successful detachment and transfer of one fruit indicate only the feasibility of a single operation; productive capacity depends on maintaining consistent state across targets and over extended periods. Continuous operation requires reliable handoffs of target, environment, robot, and logistics information across detection, assessment, approach, detachment, collection, and confirmation while preventing unresolved errors or outdated models from propagating into the next task.
For continuous operation, this review proposes state management at three linked levels. At the target level, the system records objects pending harvest, harvested, deferred, skipped, or awaiting verification while maintaining unique identities. At the environment level, it updates local relationships and free space after platform motion, fruit removal, and foliage disturbance. At the robot level, it maintains the states of the platform, manipulator, end effector, energy supply, communications, and collection device. These are synthesis requirements derived from recurring interface failures rather than a claim that one reviewed platform already implements the complete structure.
A practical requirement for continuous operation is not the elimination of every anomaly but the ability to localize and interpret it. Missed targets, localization drift, blocked paths, unstable holding, detachment failure, dropped fruit, collection blockage, and equipment faults should be assigned to distinct classes with their stage, target, last reliable pose, and trigger recorded. Recovery can then change the failure conditions rather than repeat the same action indiscriminately [105,106,107,108,109,110]. The resulting anomaly hierarchy and recovery requirements are summarized in Table 5.
Recovery should be hierarchical and should alter the conditions that caused the failure. Local perception or pose errors may be addressed through verification and realignment; invalid task or path assumptions require safe retreat, replanning, or target reordering; and equipment or safety faults require a controlled stop. Each recovery class should define attempt limits, time and force budgets, and explicit exit conditions [103,104,105,109].
A retry is justified only when a new viewpoint, approach direction, manipulation parameter, or target order changes the failure conditions. After recovery, target, robot, and collection states must be reconciled before the next operation. Shared target identities and interlocks should prevent manipulation before platform stabilization, transfer before detachment confirmation, and task switching before collection or end-effector reset [94,111,112,113].
Overall, this section concerns the organization of individual operations into a sustainable production process. Its principal requirements are maintaining state consistency, making anomalies diagnosable, supporting recovery actions that change the failure conditions, and limiting fault propagation through module interlocks, resource management, and logistics coordination. This completes the transition from environmental constraints to capability requirements. Section 4 examines how these requirements are implemented through robot architectures, functional modules, interfaces, and closed-loop system designs.

4. System Architecture of Greenhouse Harvesting Robots

4.1. Overall Architecture and Closed-Loop Mechanisms

Section 3 defined the information, decisions, control responses, and verification states required for reliable harvesting. This section maps those requirements to physical components, software services, and interfaces. Its focus is implementation: how mobility, observation, manipulation, collection, and recovery share a common task state and coordinate execution [114,115,116,117,118,119,120,121,122,123,124].
From a system-implementation perspective, greenhouse harvesting robots principally comprise four parts: physical execution, perception and state construction, task and motion control, and logistics and safety assurance. The mobile platform, height-adjustment mechanism, manipulator, end effector, and collection device directly perform locomotion and manipulation tasks. Vision, pose, joint, and contact sensors acquire environmental and robot states. Task management, scene updating, harvestability assessment, and motion control convert observations into actions, whereas energy supply, communications, container management, and safety interlocks support continuous operation. These parts exchange information through target identity, coordinate relationships, task stage, and equipment state; none can function independently of the overall process.
Control and computing units are generally deployed in accordance with the physical structure of the robot. Platform-side sensors and controllers are responsible for navigation, stationing, and chassis safety; the manipulator controller executes joint trajectories and returns motion states; wrist-mounted or eye-in-hand sensors supplement information about the target neighborhood; and the end-effector control unit performs holding, detachment, and contact-event assessment. The onboard computer undertakes scene understanding, task scheduling, and motion planning and issues commands to individual devices through real-time controllers or programmable logic controllers (PLCs). This division of responsibilities reduces interference between high-level computation and real-time device control and facilitates localization of the stage at which a fault occurs. Figure 7 illustrates the general configuration of a selective harvesting robot and the relationships among its functional modules.
Operationally, the harvesting loop converts a shared scene state into observation, locomotion, approach, detachment, transfer, and verification actions and then writes the outcome back to the task state. Integrated grape and greenhouse harvesting systems demonstrate this perception-to-action chain, but the general requirement is that target identity, pose, action stage, and outcome remain traceable across modules rather than being passed as isolated coordinates [125].
Hierarchical control is therefore common: the task layer manages targets and anomalies, the motion layer coordinates platform and manipulator actions, and device controllers execute end-effector and safety functions. Systems such as SWEEPER distribute these responsibilities across onboard computing and real-time controllers, illustrating the need for explicit interfaces between high-level planning and device-level execution [126].
Table 6 summarizes the principal implementations and interfaces of the system by functional level. Compared with enumerating components for each prototype, this organization more clearly specifies the states that each level must provide, the constraints it must receive, and the means by which anomalies are identified by higher levels and prevented from propagating.
Table 6. Levels and interfaces in the overall architecture of a greenhouse harvesting robot.
Table 6. Levels and interfaces in the overall architecture of a greenhouse harvesting robot.
System LevelPrincipal ImplementationKey Interfaces and Outputs
Execution layerMobile platform, height-adjustment mechanism, manipulator, end effector, and collection deviceDevice pose, action stage, load state, fruit-retention state, and bin-entry state
State layerPlatform-mounted and fixed-view cameras, eye-in-hand sensors, joint and contact sensors, and scene-update moduleTarget identity, three-dimensional pose, organ relationships, robot state, confidence, and information validity period
Task and motion layerTask management, harvestability assessment, target ranking, motion planning, trajectory control, and local correctionAction type, pre-manipulation pose, permitted approach direction, trajectory, termination conditions, and recovery entry point
Safety layerContainer and conveyance management, energy and communication monitoring, device interlocks, watchdogs, and emergency stopsCapacity and bin-entry states, energy and communication states, fault codes, safe stop, and maintenance request
Source: Synthesized from References [114,115,126].
Architecture also determines the system-level time budget. Logistics, platform stationing, manipulation, and verification can each contribute materially to cycle time, so coordination requires shared timestamps, coordinate frames, target identities, completion states, and fault codes together with an independent safety layer [126].
The architecture summarized here is a conceptual synthesis rather than a configuration already validated in one system. Its value lies in defining interfaces through which observations, actions, outcomes, and faults remain traceable across the execution, state, task, and safety layers.

4.2. Perception, Scene Understanding, and Decision-Making Functions

In implementation, perception, scene understanding, and decision making form a software chain that converts distributed sensor observations into executable commands. Platform-, arm-, and end-effector-mounted observations are transformed into the shared state defined in Section 3; the task manager then invokes observation, planning, and execution services through explicit interfaces.
Sensor placement determines each sensor’s role in the functional chain. Platform-mounted RGB-D cameras, LiDAR, and localization sensors support perception of aisles, infrastructure, obstacles, and global pose. Fixed-view or shoulder-mounted cameras perform target search and determine the initial manipulator workspace. Eye-in-hand cameras update the detachment site, local clearance, and relative pose in the target neighborhood. Encoders, drive current, force, tactile, and pressure sensors supplement robot and contact states. Sensor configurations should be derived from task variables, thereby avoiding redundant measurements of the same information while critical organ states or operation outcomes remain unobserved [45,75,80,127].
Before entering the scene state, heterogeneous data must undergo calibration, synchronization, and coordinate management. The platform updates its global position and work area at a relatively low frequency; the target neighborhood is refreshed on an event basis before and after manipulator motion; and joint and contact information is sampled continuously within the local control cycle. Each state record should carry its source, timestamp, and validity period to prevent outdated point clouds, old joint angles, or unsynchronized images from being interpreted as spatial relationships from the same instant. Long-duration systems should also include online extrinsic-calibration checks and drift warnings, rather than relying exclusively on one-time calibration before deployment.
The scene-understanding module may combine an object table, relational graph, robot state, and task state. The object table stores fruit identity, maturity, three-dimensional pose, and detachment site. The relational graph expresses fruit-peduncle associations, occlusion order, obstacle boundaries, and traversable space. The robot state records the platform, manipulator, end effector, and container, whereas the task state identifies the current stage, completed actions, confidence, and failure history. A unified representation allows target detection, pose estimation, reachability analysis, and outcome verification to operate on the same object, reducing mismatches caused when individual modules maintain separate identifiers and local maps.
Perception algorithms contribute to the scene state only after their outputs are associated with targets, coordinate frames, confidence, and task relevance. Recent work on difficult fruit detection, occlusion-aware localization, and morphology-aware segmentation illustrates the range of methods available [128,129,130], but their value to harvesting depends on whether they supply the relational and geometric information needed for decision making rather than on standalone image metrics.
State updating encompasses both multi-frame target association and local correction induced by actions. After platform restationing, manipulator approach, target removal, or foliage disturbance, only the affected objects and relationships are updated. When critical structures are invisible or uncertainty exceeds a threshold, the task manager invokes an active-observation service to generate limited viewpoint adjustments by the camera, manipulator, or platform. Active perception is thereby implemented as a task action constrained by motion cost and safety, rather than as an auxiliary algorithm independent of the decision chain [80]. Based on this mechanism for state construction and active updating, Figure 8 presents the conceptual framework for perception, scene representation, and decision making proposed in this study.
Within this framework, sensor data enter the state service after temporal and spatial alignment. The decision module sequentially invokes harvestability gating, task selection, and motion-constraint generation, and the execution module returns observation results, equipment states, and action outcomes. In the proposed framework, the shared state serves as the single authoritative source of task state for all modules, reducing inconsistencies caused when detection, planning, and control maintain separate target lists.
In implementation, the harvestability gate should function as a state filter upstream of the task scheduler. It reads target maturity, visibility of critical organs, reachability, pose margin, local clearance, and uncertainty and outputs states such as direct manipulation, additional observation, robot adjustment, deferral, or skipping while retaining the basis for the judgment and the corresponding scene version. When the environment changes, the gating result is recomputed for the affected object rather than treating harvestability as a permanent label.
Task decision making generally comprises three levels: global, local, and execution constraints. The global level manages crop rows, platform stations, unloading, and work areas. The local level generates an order from target relationships and harvestability states while considering how target removal will affect subsequent visibility and clearance. The execution-constraint level specifies the pre-manipulation pose, approach direction, allowable end-effector pose range, prohibited region, and timeout condition. All three levels share the same target identity and scene version, enabling the platform, manipulator, and end effector to coordinate their actions around the same task object [95].
Decision results should be transferred to the execution layer through a structured interface rather than as a target coordinate alone. At a minimum, the interface should include the target identifier, action type, priority, reference coordinate frame, pose tolerance, permitted approach region, obstacle constraints, confidence, and termination conditions. If motion planning finds that these constraints cannot be satisfied, it should return a specific reason for infeasibility, such as insufficient information, inappropriate robot position, or restricted local space. The task manager can then request additional observation, reposition the platform, change the approach direction, or downgrade the target.
Representative systems increasingly connect maturity, detection, organ relationships, pose, workspace constraints, and multiview observations directly to task selection and motion generation [93,106,107,131]. Across these implementations, the common requirement is semantic continuity from perception to action: target identity, uncertainty, and scene changes must remain accessible to the task manager. Representative implementations are summarized in Table 7.
Across the representative systems, the common requirement is semantic continuity from perception to action. Coordinates alone cannot preserve organ relationships, target identity, reachability, or state changes; a shared state allows local observations to revise global estimates and converts planning failures into interpretable task adjustments [93,106,107,131].
Online updates should be triggered by events such as target removal, localization change, platform motion, failed approach, or resource-state change. Learning-based outputs may support detection and ranking, but they must remain subject to confidence checks, kinematic reachability, collision constraints, device limits, and safety logic. This organization preserves intermodule traceability while avoiding unnecessary full-scene replanning [95,132].

4.3. End-Effector and Fruit Collection System

The end-effector and collection subsystems form the final physical segment of the harvesting chain. The end effector converts the task-layer target state into approach, retention, and detachment actions; the collection subsystem then receives the fruit and reports transfer, capacity, blockage, or loss. Their shared interface is the verified task outcome returned to the task manager, so they should be designed and evaluated together rather than as independent devices.
Holding may use gripping, enveloping, suction, or hybrid mechanisms, whereas detachment may use cutting, breaking, pulling, twisting, or combined actions [133,134,135,136]. The relevant system question is not which mechanism is universally best but whether crop morphology, attachment mechanics, local clearance, market-quality requirements, and observable feedback support reliable progression from contact to confirmed detachment.
Geometry and product-quality constraints jointly define the operating envelope. Fruit size and curvature shape contact geometry; surface properties and allowable pressure constrain retention; and peduncle or abscission geometry constrains the detachment tool and force direction. Cutting can limit plant disturbance when the detachment site is visible and accessible, whereas twisting or pulling can relax cutting-point precision but demand stronger retention and tighter control of crop disturbance [58,72,133,137].
Compliance is useful only when coupled to state estimation. Flexible or underactuated structures can absorb small pose errors during initial contact, but stable holding and detachment still require sufficient stiffness or regulated pressure to prevent slip and uncertain force transmission [72,138,139,140,141]. Figure 9 therefore summarizes the information flow among task state, end-effector action, collection, and outcome verification rather than classifying hardware in isolation.
Across implementations, local vision, force, tactile, pressure, or actuator signals are most useful when interpreted by task stage: approach, holding, detachment, transfer, and release have different completion and anomaly criteria [98,126,142,143,144,145,146]. Integrated end effectors can reduce coordinate transformations by combining local sensing, holding, detachment, and fruit reception, but increased wrist mass, collision envelope, calibration effort, and maintenance burden must be traded against these benefits.
Detachment modes can be represented as parameterized procedures with visibility requirements, pre-manipulation pose, holding conditions, load and speed bounds, completion criteria, and safe exits [58]. This allows the task manager to select an action according to the current target state and to invoke re-observation, adjustment, or retreat when feedback shows that the assumed conditions no longer hold. The corresponding end-effector and collection-system functions are summarized in Table 8.
Outcome verification should cross-check visual evidence with the force, pressure, tactile, or actuator state and should trigger parameter adjustment or safe retreat rather than indiscriminate force escalation [98]. Product quality is part of the same outcome: crop-appropriate measures of bruising or tissue damage, calyx or stem retention, firmness, water loss, microbial susceptibility or decay, and storage life should accompany detachment and transfer rates. Cross-cultivar or cross-species transfer should not be assumed because fruit morphology, attachment strength, harvest requirements, and market standards can require revalidation of contact materials, force limits, tool geometry, and detachment procedures [58,72,74,133,137,138,139,140,141].
Collection architecture trades reduced manipulator travel against wrist payload, transfer-path length, impact, and blockage risk. Bin entry, capacity, blockage, fruit loss, and unloading completion must therefore be visible to the task manager and coordinated with platform and manipulator motion [68,69,148].

4.4. Mobile Platform and Manipulator

The mobile platform, height-adjustment mechanism, and manipulator should be treated as one hierarchical positioning system: the platform provides coarse stationing, the lift places the arm near the target height, and the manipulator performs local reach and pose adjustment. This division reduces operation near workspace limits and links facility geometry directly to manipulator dexterity.
Platform choice is facility-specific. Pipe-rail or monorail systems offer repeatable guidance but depend on rail geometry and headland transfer; free-ranging platforms improve flexibility but must manage narrow aisles, slip, turning, and repeatable stationing; and gantry or over-row mechanisms expand coverage at the cost of infrastructure and robot size [149,150,151,152,153,154,155]. Table 9 summarizes these trade-offs.
Stationing accuracy and structural settling determine the local manipulation envelope. Eye-in-hand servoing and contact operations should begin only after the platform and lift are stable, with end-effector length, payload, container load, residual vibration, and calibration drift included in the pose budget.
Manipulator configuration should match target distributions and required approach directions rather than maximize degrees of freedom. Industrial six-axis arms provide pose margin and mature interfaces but increase collision envelope and platform load; task-specific or Cartesian mechanisms can reduce mass and control complexity when crop geometry is sufficiently regular [42,126,156].
Table 9. Suitability of mobile platform and manipulator configurations for greenhouse harvesting.
Table 9. Suitability of mobile platform and manipulator configurations for greenhouse harvesting.
ConfigurationApplicable ScenarioPrincipal CharacteristicKey Constraint
Pipe-rail platform, lift, and industrial manipulatorGreenhouses with standardized infrastructure and a wide range of target heightsStable longitudinal guidance and a broad range of pose adjustmentLimited by rail gauge and headland conditions; relatively high robot mass and large collision envelope
Free-ranging wheeled platform and industrial manipulatorWork areas without fixed rails that require flexible repositioningFlexible deployment and the ability to bypass temporary obstaclesHigh requirements for stationing accuracy, slip prevention, turning, and platform stability
Task-specific manipulator with platform-assisted positioningCrops with relatively regular target distributions and approach directionsCompact structure, fewer degrees of freedom, and lower control complexityLimited cross-crop adaptability and a need for precise matching to the cultivation system
Dual-arm or multi-arm mobile platformHigh target density with conditions suitable for parallel operationExtended spatial coverage and potential for parallel operationMore complex target assignment, inter-arm collision avoidance, shared logistics, and platform loading
Gantry or over-row mechanism with manipulatorElevated or regular row layouts that permit supporting facility modificationsLarge coverage, with motion directions aligned to the cultivation structureLimited by robot dimensions, modification cost, and cross-row transfer
Source: Compiled from References [42,126,131,147,157].
Multi-arm systems can increase coverage only when target density and fruit logistics support parallel work; otherwise, target assignment, inter-arm collision avoidance, shared collection, and platform loading can offset the expected efficiency gain [95,157,158,159].
Mechanical integration should jointly consider center of gravity, structural deflection, calibration stability, end-effector payload, and changing collection load. When a target lies outside the effective workspace or local clearance becomes invalid, control should return to platform repositioning or task deferral rather than forcing the manipulator toward kinematic limits.
Orchard or field navigation studies are used here only as transferable mechanism-level evidence when sensor inputs, vehicle-motion constraints, obstacle classes, validation conditions, and pose or traversability outputs match greenhouse stationing needs; their numerical performance is not generalized to greenhouse harvesting [160,161,162,163,164].

4.5. Interaction Feedback, Autonomous Recovery, and System Integration

During execution, crop motion, platform vibration, calibration drift, and contact deformation can make the current robot-crop state diverge from the planned model. The implementation therefore routes visual, force, tactile, pressure, and device-status events to local control, task management, or safety services. This subsection examines the state-machine, feedback, and recovery mechanisms used to realize the Section 3 requirements.
A practical controller represents approach, contact, holding, detachment, transfer, and release as explicit states with observations, completion conditions, timeouts, and anomaly exits. Relative pose and clearance guide approach; pressure, load, and slip guide holding; and post-action vision and device signals verify detachment, retention, and bin entry. Because a signal can have different meanings at different stages, thresholds are bound to the active state.
Local feedback loops correct approach and contact using fixed-view or eye-in-hand vision together with force, pressure, tactile, or device signals [98,165,166]. Higher-level state management determines when these loops start, terminate, or roll back, so that unresolved anomalies do not propagate into later stages. Explicit states for scanning, approach, manipulation, verification, and error handling also provide a defined safe point from which recovery can restart.
The Error State in Figure 10 illustrates the system-level handling of anomalies. Recovery logic determines whether a fault is reversible and whether action is needed at the local, task, or system level. Temporary occlusion, small pose errors, or insufficient contact may trigger observation, slower motion, adjustment, or a short retreat. If a path becomes invalid or the target moves substantially, the robot should return to a safe pose and replan. Repeated failure, persistent unreachability, equipment faults, or safety-boundary activation should trigger deferral, skipping, or human takeover. Recovery policies should specify attempt and time budgets, force and displacement limits, and a safe fallback path. These implementation levels are summarized in Table 10.
Reported experiments indicate that retries help only when they alter the failure conditions. Force sensing can support local stopping or correction, whereas a target enclosed by obstacles may remain inaccessible [98,105]. Before repetition, recovery should therefore change the viewpoint, approach, action parameters, or target order.
System-level recovery requires modules to share times, coordinates, target identities, scene versions, and fault semantics. State machines or behavior trees can coordinate module transitions, while independent watchdogs and interlocks enforce safety. A digital twin may extend this shared state across planning and recovery, but model latency, drift, and missing observations mean that it cannot replace real-time device-level assessment [107].
Together, these mechanisms allow local anomalies to be interpreted, contained, and recorded rather than propagated through the harvesting chain.
Commercial deployment adds requirements beyond recovery of a single fruit. Systems should specify compatibility with greenhouse infrastructure and utilities, cleanability and sanitation, maintenance access to wear parts and calibration points, operator training and operating procedures, and the logistics needed for fruit handling. Long-duration trials should report uptime, stopping failures, repair and recovery time, intervention frequency, and performance drift across repeated shifts and changing crop conditions. Without such evidence, commercial readiness remains unverified even when short trials show high technical performance.

5. Illustrative Case Study: Greenhouse Cucumber Harvesting

This section uses cucumber harvesting as an illustrative stress test of the task-chain framework, not as a representative benchmark for all greenhouse crops. The cucumber was selected because foliage-like color, slender and often occluded peduncles, pendulous geometry, restricted high-wire workspaces, and contact-sensitive detachment expose linked uncertainties across perception, planning, manipulation, and outcome verification. The evidential role is therefore illustrative: cross-crop transfer of numerical performance is not assumed.

5.1. Scene Perception and Harvest-Target Representation

5.1.1. Greenhouse Cucumber Harvesting Scenario and System Configuration

High-wire cucumber cultivation provides a concrete sequence for applying the task-chain framework. Target information is progressively transformed from long-range detection and 3D localization into local verification, path entry, peduncle detachment, fruit transfer, and outcome verification.
High-wire cultivation gives cucumber distribution a degree of regularity along crop rows and over the vertical range, yet the position and pose of individual fruits still vary substantially. The elongated fruits hang naturally, their color resembles that of the foliage, and their slender peduncles are frequently interwoven with the main stem, tendrils, and support strings. Regular rows help constrain the robot’s search and stationing regions but do not directly provide a gripping location, peduncle pose, or safe cutting direction. This case study therefore does not reduce the target to a single fruit-center coordinate; instead, it progressively enriches the target state by incorporating fruit geometry, peduncle connectivity, surrounding plant tissues, and locally operable space.
A recent continuous cucumber-harvesting platform integrates RGB-D perception, a manipulator, a task-specific end effector, mobile support, computation, and fruit collection within one workflow [167]. It provides a useful system example because detection and localization feed path generation, manipulator motion, detachment, and collection rather than being evaluated as isolated functions.
Figure 11 illustrates that continuous harvesting depends on robot-level coordination among perception, manipulator motion, end-effector operation, and fruit collection, rather than on single-target recognition alone. Global perception provides candidate targets and positions, the planning module organizes inter-target paths, and changes in position or obstacle relationships determine whether later trajectories require adjustment.

5.1.2. Cucumber Detection and Three-Dimensional Localization

Global perception must detect partially occluded cucumbers, estimate their 3D state, and maintain target identity across camera or platform motion. Representative systems use RGB-D detection to construct candidate lists and generate continuous-harvesting paths [167,168]. For manipulation, the target record should extend beyond the detection-box center to include fruit orientation, visible extent, depth quality, occlusion, and uncertainty so that the robot can determine a pre-approach pose and whether local observation is still required.
Experiments in complex backgrounds further show that determining the presence of a fruit is generally easier than estimating its orientation and connectivity [169]. Occlusion, overexposure, and chromatic similarity to the background have a relatively limited effect on centroid position but can substantially amplify errors in the major-axis direction. Detection confidence therefore cannot serve directly as a reliability measure for the manipulation pose. When the fruit axis or its upper connection is unclear, the system should schedule additional observation rather than generate a cutting pose immediately.
Recovering a three-dimensional state from two-dimensional results additionally requires handling color-depth registration, depth holes, foreground leaves, and errors in hand–eye extrinsic calibration. A more robust procedure first uses detection or segmentation to constrain the sampling region, removes boundary pixels and depth discontinuities, estimates the fruit surface, centerline, and local uncertainty, and then transforms the results into a unified coordinate frame. If the fruit apex is occluded, morphological priors should guide the next observation direction rather than substitute for actual peduncle localization. These depth-based localization, RGB–depth registration, stereo-reconstruction, and pose-estimation methods provide the three-dimensional state of individual observations [170,171,172,173]. Cross-frame target association must also prevent the same fruit from being recorded repeatedly after platform or camera motion [67,107].
After global processing, each candidate target should be represented by a record containing target identity, three-dimensional position, fruit-axis direction, visible proportion, depth quality, and estimated uncertainty. This information supports multi-target ordering while identifying the local information that remains to be acquired. When the fruit position is reliable but the peduncle state is unclear, the robot should first perform close-range verification rather than derive the end-effector cutting action directly from global detection.

5.1.3. Local Observation and Visual-Servo Correction

As the manipulator enters the target neighborhood, the perceptual problem shifts from global fruit localization to the relative pose of the peduncle, fruit, main stem, and end effector. Eye-in-hand sensing and dynamic visual servoing can update this local state as calibration error, platform stationing, and crop motion accumulate [168,174,175]. Global perception should therefore guide the robot into an observable region, while final alignment uses the most recent local state.
Because cucumbers and peduncles can oscillate during ventilation or approach, a single local estimate may become invalid before contact. Dynamic peduncle-tracking and pose-estimation methods demonstrate the value of maintaining observability throughout the servo cycle rather than optimizing a single frame [175]. Table 11 summarizes representative perception and manipulation-target data.
These studies show that target representation becomes more task-specific as the robot approaches: global sensing establishes the identity, position, orientation, and visibility, whereas local sensing resolves the peduncle pose, relative motion, and safe approach region. Because calibration, motion control, and canopy disturbance still affect execution, perception should be evaluated by its contribution to successful approach and manipulation rather than by image-level accuracy alone [168,169,175].

5.2. Harvestable-Target Assessment and Harvest-Task Decision Making

5.2.1. Target Harvestability Assessment

Harvestability assessment determines, from the current target information, whether the robot should execute harvesting as its next action. The gating information can be divided into four categories. Fruit size, maturity, and marketability constitute biological conditions. Target identity, peduncle visibility, and pose confidence constitute perceptual conditions. Inverse kinematics, joint margin, and end-effector pose constitute robot conditions. Whether the main stem, neighboring fruits, or greenhouse infrastructure enters the approach or retreat envelope constitutes interaction-safety conditions. Together, these categories determine whether to harvest directly, obtain additional observations, adjust the robot, defer the target, or skip it; mature fruits should not be placed indiscriminately in the manipulator queue.
These conditions should not be compressed into a static binary label. A mature and visible fruit may remain temporarily unharvestable because the detachment site is unobservable, the manipulator approach pose is constrained, or the retreat path is blocked. Conversely, a partially occluded target may become harvestable after a viewpoint change, platform adjustment, or limited obstacle avoidance. The assessment should retain uncertainty in pose and structure estimates and distinguish the main stem and greenhouse infrastructure, which must not be contacted [176], from leaves that may permit limited disturbance [177]. Otherwise, an excessively conservative collision model may exclude targets that could be handled through additional observation or slight avoidance [177]. Direct harvesting should therefore be triggered only when critical conditions satisfy the safety requirements; targets with insufficient evidence or improvable conditions should enter intermediate states such as additional observation or robot adjustment.
Agronomic conditions also alter the assessment outcome. Arima and Kondo [32] showed that inclined-trellis training can improve fruit exposure and facilitate robotic harvesting. Peduncle exposure and the safe distance from the main stem must still be evaluated with respect to the specific trellis geometry and plant pose. The set of harvestable targets is therefore constrained by robot capability but can also be improved before operation through trellis design, leaf and stem management, and cultivar selection.

5.2.2. Multi-Target Harvesting-Order Generation

After multiple targets pass the harvestability gate, the task layer must determine the visitation order at the current platform station. The items to be ordered are not all fruits visible in the image but rather candidate targets whose identities, three-dimensional positions, and basic manipulation conditions have been verified. Park et al. used a global camera to construct a candidate set and generated an initial harvesting order before manipulator execution, linking target selection to platform station and the current manipulator state [168]. This order constrains subsequent motion planning but should not be treated as a fixed list that remains valid throughout continuous operation.
Target ordering should consider not only geometric distance among targets but also manipulator-motion cost, expected manipulation time, collision and fruit damage risks, perceptual uncertainty, and recovery cost after failure. The relative priorities of these factors should change with the task state. When platform repositioning is costly, target coverage within a single station should be increased. In dense canopies, outer targets or those with more reliable states can be processed first to release space for subsequent observation and entry. When peduncle or detachment site information is insufficient, the target should be transferred to additional observation rather than executed early merely because it is nearby.
Continuous-harvesting planners can cluster spatially related targets, order visits, connect them with efficient transfer paths, and locally adjust segments that conflict with occupied space [167]. These methods reduce unnecessary return motion between fruits, but they address only part of the task chain: peduncle alignment, retention, detachment, and verification still require local sensing and end-effector control.
The harvesting order must also be updated as field conditions change. Fruit removal, foliage rebound, and manipulator contact may alter the positions, visibility, and reachable space of remaining targets. After each target is completed, the three-dimensional target set should therefore be updated and the affected target clusters and paths rechecked. The existing order may continue only when target identity, local space, and path conditions remain valid; otherwise, the system should reorder locally or return to additional observation. This keeps the harvestable-target set, visitation order, and approach control consistent, preventing an outdated plan from being executed after the scene has changed. Figure 12 illustrates the resulting multi-target ordering and path-adjustment process.

5.2.3. Manipulator Approach Paths and Online Adjustment

The foregoing harvestability assessment determines whether a target enters the current task set, multi-target ordering determines the visitation sequence among harvestable targets, and approach planning converts the currently selected target into executable manipulator motion. These functions occupy different levels. Only targets that pass the harvestability gate participate in ordering; the result identifies the current target and subsequent visitation direction for the manipulator but does not replace local collision avoidance and precise alignment within the target neighborhood. A failed local approach therefore does not necessarily invalidate the target’s maturity or long-term harvestability. The system should further determine whether the failure arose from insufficient information, an inappropriate robot position, or genuinely infeasible current spatial conditions.
For each selected target, manipulation typically progresses from a pre-manipulation pose through coarse approach, local observation and alignment, tool entry, and safe retreat. Global or inter-target planning brings the manipulator into the target neighborhood, while eye-in-hand sensing corrects residual localization and calibration error before contact [167,168]. This separation keeps efficient transfer and precise final alignment within distinct but connected control levels.
Earlier cucumber systems showed that static-model search can produce collision-free paths, but changing foliage, fruit motion, and failed approaches require online correction and rollback [178,179]. Later systems increasingly combine coarse planning with local observation and state-based recovery rather than treating the initial path as permanently valid.
Approach anomalies should be handled according to cause. Loss of the target or rising pose uncertainty should trigger additional observation; poor stationing or blocked clearance should trigger retreat, repositioning, or replanning; persistent infeasibility near the main stem or greenhouse structure should defer or skip the target. After successful harvesting, only scene elements affected by fruit removal and foliage rebound need to be updated. This links harvestability assessment, ordering, local planning, and recovery through a common scene state. The corresponding target states, responses, and reassessment conditions are summarized in Table 12.

5.3. Compliant Execution, State Feedback, and System Performance

5.3.1. End-Effector Localization and Peduncle Cutting

After the manipulator enters the target neighborhood, the principal error sources shift from global localization to hand–eye calibration, target oscillation, trajectory tracking, and contact deformation. Because cucumber peduncles are slender and the tool window is limited, the final approach segment should use continuous eye-in-hand observation at low speed while jointly constraining relative-pose error, end-effector pose, and the tool-safety envelope. This strategy prevents the manipulator from continuing solely on the basis of an initial coordinate and reduces compression or pulling of the fruit before detachment is complete.
Cucumber end effectors combine fruit retention, peduncle guidance or alignment, cutting, and transfer within a confined canopy. Representative designs integrate local vision with suction or gripping, mechanical guidance, cutting, and collection so that vision need not localize an exact cutting point in isolation; instead, sensing and mechanism design jointly constrain the peduncle into a safe tool region [167,168]. This illustrates a broader system principle: mechanical design can reduce perceptual precision requirements, but the resulting interaction still requires state feedback. Figure 13 shows a representative end-effector structure and harvesting sequence for greenhouse cucumbers.
End-effector actions should not be driven by fixed delays alone but should progress sequentially through holding, peduncle guidance or alignment, cutting, and transfer according to the current state. With a gripping strategy, the system should confirm that the fruit has entered the gripping range, the gripper has closed, and retention is stable before permitting the shearing mechanism to act. With a suction strategy, both vacuum retention and peduncle guidance must be confirmed. Cutting should be triggered only after the peduncle enters the tool window and neighboring tissues remain outside the hazardous region. After cutting, the system must determine whether the connection has been released and confirm that the fruit remains retained or has entered the collection channel. Compliant execution thus arises from the combination of continuous observation, mechanical accommodation of localization error, and stage-based control, rather than from compliant materials alone.
Differences between component tests and field results show that the ability of a blade to cut a peduncle does not imply that the complete robot can detach fruit reliably. Field performance also depends on reliable peduncle detection, an appropriate end-effector entry pose, confinement of gripping or guidance forces to the target tissues, continued retention after detachment, and successful passage into the collection system. End-effector evaluation should therefore report entry success, retention success, erroneous and missed cuts, fruit damage, detachment success, and transfer success separately, rather than substituting offline cutting capability for complete operational performance.

5.3.2. Fruit Transfer and Harvest-State Verification

Peduncle cutting does not by itself complete the task. A cucumber may fall at the instant of cutting because of inadequate retention, remain lodged in the collection channel, or remain connected to the plant because detachment is incomplete. The system must distinguish at least among detached and retained, detached but lost, incompletely detached, delivered to the container, and uncertain outcomes. The assessed state must be returned to the target list and task scheduler; completion of the blade motion must not be treated directly as harvesting success.
State verification can combine images acquired before and after the action, vacuum pressure or gripper displacement, actuator current, relative fruit motion, and container-side detection. Disappearance of the connection region, synchronous motion of the fruit with the end effector, and a stable retention signal provide evidence that the fruit has been detached and retained. Photoelectric, weight, or vision events at the transfer inlet or container confirm that the fruit has left the end effector and entered the collection system. Gravity transfer or channel collection reduces manipulator travel while carrying fruit but introduces additional states such as channel blockage, fruit impact damage, and full container capacity. The logistics system should therefore be incorporated into the same task-state machine.
When the outcome is uncertain, recovery should change the conditions that caused the failure rather than repeat the same action. A new observation, lower speed, different approach, or platform repositioning may recover reversible errors, whereas persistent missed cuts, retention failure, severe oscillation, or rising collision risk should trigger safe retreat and target deferral [180]. Retry limits should therefore balance expected success gain against damage risk and time cost.
Feedback should continue through retreat and scene updating. Dynamic visual servoing improves local observation and approach for oscillating targets but cannot eliminate calibration error and robot-execution error completely [175]. Post-action images can determine whether the peduncle remains connected, whether the fruit moves with the end effector, and whether the retreat corridor is occupied by rebounding foliage. These changes must be written back to the target state to prevent repeated handling of detached fruit and to reassess the visibility, reachable space, and harvesting order of neighboring targets.

5.3.3. System Performance and Task-Chain Implications

Complete-system performance is determined jointly by target detection, canopy entry, local alignment, fruit retention, peduncle detachment, postharvest transfer, and anomaly recovery. Table 13 reorganizes representative cucumber-system results around validation context, complete-task success, cycle time or retry cost, and reported autonomy or human intervention. Missing values are retained as not reported rather than inferred. Because cultivars, protocols, sample sizes, and success definitions differ, the table is intended to expose reporting gaps and stage interfaces rather than rank robots directly [167,168,178,180].
Stage-level results reveal different system bottlenecks. Early systems validated the complete harvesting chain but incurred high retry costs [178,180]; in the hierarchical system, serial detection, entry, and cutting rates explain the lower complete-system success [168]. Continuous-path planning reduced unproductive travel, but path length and collision-free rate do not measure retention, detachment, or bin-entry success [167]. System evaluation should therefore report conditional stage transitions, failure locations, recovery costs, and effective time per successfully collected fruit.
The cucumber case illustrates that embodiment arises from coordination among crop structure, sensing, mechanical guidance, and control. Local perception and task-specific end effectors can reduce uncertainty, while trellis design and plant training improve visibility and access [32,167,168]. Nevertheless, chromatic similarity, peduncle oscillation, constrained canopy entry, and incomplete post-cut verification remain important limitations [181,182,183].
However, the cucumber case should not be read as direct evidence for other greenhouse crops. Cucumber fruits are elongated and pendulous and are detached at a slender peduncle within high-wire canopies, whereas tomato fruits often occur in clusters, sweet peppers have distinct calyx and side-stem geometries, and strawberry fruits are borne on short, exposed peduncles above gutters or benches with different support and maturity constraints [43,44,59,105,106,110]. These differences can change the dominant perception errors, approach directions, end-effector design, damage risks, and attainable throughput. The transferable contribution is therefore the task-chain analysis and its evaluation logic; target representations, harvestability criteria, contact limits, and performance values must be revalidated for each crop, cultivar, training system, and greenhouse.

6. Conclusions and Outlook for Smart Robotic Harvesting of High-Value Greenhouse Crops

6.1. System-Level Findings

Across the 188 unique primary studies, the clearest system-level finding is that complete-system reliability depends on continuity across perception, decision making, manipulation, feedback, and recovery rather than on the peak result of any single module. Greenhouse harvesting fails when missing organ information, short-lived scene models, contact uncertainty, or unverified transfer propagate downstream. The most informative unit of analysis is therefore the complete harvest of a marketable fruit under a defined crop–facility–platform configuration. Marketable success also links robot mechanics to crop physiology: preharvest conditions shape susceptibility to handling damage and shelf-life potential, so detachment efficiency and throughput should be interpreted together with postharvest quality.
The embodied-intelligence lens helps position this evidence without relabeling conventional systems. Most reported platforms implement closed-loop automation and some embodied interaction through eye-in-hand observation, compliant contact, or outcome-dependent correction; comparatively little evidence shows accumulated interaction experience being used for online learning, cross-cultivar adaptation, or transfer. The practical contribution of this framework is to connect maturity of adaptation with production-relevant outcomes rather than to treat embodied intelligence as a binary label.

6.2. Major Bottlenecks and Benchmarking Priorities

Four bottlenecks recur across the task chain: incomplete observability of peduncles, free space, and attachment structures; short-lived scene validity after crop motion or fruit removal; uncertain contact outcomes during retention and detachment; and weak diagnosis and recovery when a stage fails. These bottlenecks interact multiplicatively: a locally small error can increase cycle time, damage risk, or human intervention at later stages, so recovery must change the failure conditions rather than simply repeat the same action. Sensing robustness can also deteriorate with illumination and background changes [184].
Evaluation is a second system-level bottleneck. Definitions of harvestable targets, complete success, damage, cycle time, retry cost, autonomy, and human intervention remain inconsistent. Although large-scale evaluation has been reported for some harvesting systems [185], long-duration greenhouse trials across seasons, cultivars, and robot embodiments remain scarce. Current evidence supports technical feasibility more strongly than broad commercial readiness. Economic feasibility remains insufficiently demonstrated because investment and retrofit cost, maintenance, energy use, seasonal utilization, net labor substitution, product losses, and return on investment are rarely measured together in sustained commercial greenhouse trials. Readiness should therefore be stated only for a defined crop–facility–platform combination and supported by repeated operation reporting marketable throughput, intervention and downtime, maintenance and recovery burden, quality preservation, and operating economics. Table 14 gives a minimum reporting set proposed by this review; it is intended as a benchmarking recommendation, not as a claim that all existing studies already report these variables.
For future commercial-scale studies, the minimum economic reporting set should include capital and greenhouse retrofit costs; maintenance, consumables, energy, and operator or supervision costs; net labor hours saved or reallocated per kilogram of marketable output; productive operating hours as a proportion of the available harvest window; and payback period or return on investment under clearly stated assumptions regarding utilization, service life, labor cost, discount rate, and product value. These metrics should be reported together with marketable throughput, fruit quality, intervention frequency, downtime, and total operating hours [2,3,39].

6.3. Research Priorities by Evidence Maturity

Research priorities should follow a staged sequence of reliability, adaptation, and generalization. Near-term work should establish stable perception, robust manipulation, explicit outcome verification and recovery, acceptable cycle time, low damage, and sustained operation for a clearly defined crop and cultivation system. Medium-term work should reduce the data, calibration, and engineering effort required to adapt the same system across cultivars, seasons, and greenhouse configurations. Broader transfer across crops or robot bodies should be pursued only after these reliability and adaptation requirements are demonstrated [186,187,188].
Validation should extend beyond a single cultivar and production cycle to multiple cultivars and successive production cycles and, where feasible, to additional crop species, seasons, and crop-growth stages. Such longitudinal and cross-context trials are needed to determine whether the proposed strategies remain effective under changes in canopy morphology, fruit distribution, illumination, humidity, and management conditions [186,187,188].
Datasets and benchmarks should move from isolated images toward temporal records of complete harvesting interactions, including global and eye-in-hand observations, robot state, contact signals, actions, failure and recovery events, crop and facility context, and postharvest quality [186,189,190,191]. Common definitions for harvestable-target coverage, complete-task success, damage, recovery cost, intervention frequency, and effective throughput would allow for system-level comparison without collapsing heterogeneous protocols into a single ranking.
Agronomic and robotic co-design is an immediate opportunity. Cultivar choice, training, pruning, deleafing, plant spacing, fruit-load management, and harvest scheduling can improve visibility and access, while smart greenhouse management determines the climate, energy, and operational conditions within which the robot must work [32,35,36,37]. Robot-oriented crop ideotypes and management strategies could further prioritize exposed and consistently positioned fruiting zones, synchronized maturity within practical harvest windows, standardized canopy architecture, and predictable peduncle or abscission traits for low-damage detachment. These interventions should be evaluated jointly with yield, marketable quality, disease risk, additional labor, and resource use rather than treated as free reductions in environmental complexity.
Evidence maturity should be distinguished explicitly. Experimentally established directions include RGB/RGB-D perception, local visual servoing, task-specific end effectors, stage-based feedback, and limited state-machine recovery in greenhouse prototypes. Emerging or enabling directions include active perception, multimodal contact interpretation, imitation/diffusion policies, learning-assisted adaptation, and structured human–robot collaboration; their validation for greenhouse harvesting remains limited or context-specific [192,193,194,195]. Embodied learning, foundation models, vision–language–action control, and broad cross-robot or cross-crop transfer remain frontier research questions rather than demonstrated greenhouse harvesting capabilities [196,197,198,199,200,201,202].
Figure 14 summarizes the qualitative positioning used in this review: most current evidence lies in closed-loop automation or early embodied interaction, while embodied learning remains exploratory, and every level should be judged using the same six production-relevant dimensions. Overall, the most immediate route to deployable smart robotic harvesting is reliable integration of crop-aware perception, task-conditioned decision-making, compliant manipulation, outcome feedback, and bounded recovery under documented greenhouse conditions.

6.4. Limitations of This Review

This review has several limitations. Experimental protocols, success definitions, sample sizes, and crop/facility conditions are heterogeneous; no meta-analysis or universal risk-of-bias instrument was applied, and the six-domain evidence score remains a descriptive engineering judgment. Title/abstract screening, full-text eligibility assessment, and data extraction were conducted by a single reviewer and were not independently duplicated, which increases susceptibility to selection and extraction error. The formal pre-screen deduplication relied only on exact DOI matching. Post-revision reconciliation by title, author, and year identified one residual duplicate among the 189 retained reports because one citation-route record lacked a DOI; this was corrected before the final study-level synthesis, yielding 188 unique primary studies, but it illustrates the limitation of DOI-only deduplication. The supplementary route was limited to 200 seed records drawn from the manuscript reference set and used Google Scholar-assisted bibliographic verification rather than an independently reproducible database query. Selected orchard or field studies were included only as transferable mechanism-level evidence, but transferability judgments remain interpretation-dependent. Publication and reporting bias are also plausible: successful prototypes and favorable module results are more visible than failed trials, frequent human interventions, maintenance burden, downtime, or performance degradation during sustained operation. Accordingly, absence of a reported failure should not be interpreted as evidence of reliability. Future reviews would benefit from prospective protocols, independently duplicated screening and extraction, independent verification of evidence scoring, and standardized system-level reporting.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/agronomy16181795/s1, Supplementary Table S1 reports the database- and route-specific search fields, exact query strings, search dates and coverage, recorded language and document-type restrictions, result counts, export formats, and flow-audit notes. Supplementary Table S2 provides the reconciled 188-study primary evidence corpus and six-domain evidence-scoring tracker, including bibliographic metadata, crop, study type and task stage, validation context, total evidence score, evidence class, and supporting extraction notes.

Author Contributions

J.W. conceived the project; conducted the literature searches, title/abstract screening, full-text eligibility assessment, and data extraction; wrote the manuscript; and prepared the figures. J.W., C.X., Y.C., L.S. and Z.T. revised the manuscript. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Jiangsu University College Student Innovation Training Program (X202610299803) and the 25th Batch of the College Student Scientific Research Project of Jiangsu University (25B087).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

The authors express their sincere gratitude for the valuable technical support and resources that contributed to this research.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Tong, X.; Zhang, X.; Fensholt, R.; Jensen, P.R.D.; Li, S.; Larsen, M.N.; Reiner, F.; Tian, F.; Brandt, M. Global Area Boom for Greenhouse Cultivation Revealed by Satellite Mapping. Nat. Food 2024, 5, 513–523. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Marinoudi, V.; Sørensen, C.G.; Pearson, S.; Bochtis, D. Robotics and Labour in Agriculture. A Context Consideration. Biosyst. Eng. 2019, 184, 111–121. [Google Scholar] [CrossRef] [Scilit]
  3. Lowenberg-DeBoer, J.; Huang, I.Y.; Grigoriadis, V.; Blackmore, S. Economics of Robots and Automation in Field Crop Production. Precis. Agric. 2020, 21, 278–299. [Google Scholar] [CrossRef] [Scilit]
  4. Shamshiri, R.R.; Weltzien, C.; Hameed, I.A.; Yule, I.J.; Grift, T.E.; Balasundram, S.K.; Pitonakova, L.; Ahmad, D.; Chowdhary, G. Research and development in agricultural robotics: A perspective of digital farming. Int. J. Agric. Biol. Eng. 2018, 11, 1–14. [Google Scholar] [CrossRef]
  5. Bechar, A.; Vigneault, C. Agricultural Robots for Field Operations: Concepts and Components. Biosyst. Eng. 2016, 149, 94–111. [Google Scholar] [CrossRef] [Scilit]
  6. Fountas, S.; Mylonas, N.; Malounas, I.; Rodias, E.; Santos, C.H.; Pekkeriet, E. Agricultural Robotics for Field Operations. Sensors 2020, 20, 2672. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Rong, J.; Hu, L.; Zhou, H.; Dai, G.; Yuan, T.; Wang, P. A Selective Harvesting Robot for Cherry Tomatoes: Design, Development, Field Evaluation Analysis. J. Field Robot. 2024, 41, 2564–2582. [Google Scholar] [CrossRef] [Scilit]
  8. Rose, D.C.; Lyon, J.; de Boon, A.; Hanheide, M.; Pearson, S. Responsible Development of Autonomous Robotics in Agriculture. Nat. Food 2021, 2, 306–309. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  9. van Henten, E.J.; Bac, C.W.; Hemming, J.; Edan, Y. Robotics in Protected Cultivation. IFAC Proc. Vol. 2013, 46, 170–177. [Google Scholar] [CrossRef] [Scilit]
  10. Pfeifer, R.; Lungarella, M.; Iida, F. Self-Organization, Embodiment, and Biologically Inspired Robotics. Science 2007, 318, 1088–1093. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Sun, F.; Chen, R.; Ji, T.; Luo, Y.; Zhou, H.; Liu, H. A Comprehensive Survey on Embodied Intelligence: Advancements, Challenges, and Future Perspectives. CAAI Artif. Intell. Res. 2024, 3, 9150042. [Google Scholar] [CrossRef] [Scilit]
  12. Zhao, Z.; Wu, Q.; Wang, J.; Zhang, B.; Zhong, C.; Zhilenkov, A.A. Exploring Embodied Intelligence in Soft Robotics: A Review. Biomimetics 2024, 9, 248. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Navas, E.; Shamshiri, R.R.; Dworak, V.; Weltzien, C.; Fernández, R. Soft Gripper for Small Fruits Harvesting and Pick and Place Operations. Front. Robot. AI 2024, 10, 1330496. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Wang, Q.; Zhang, Y.; Liu, W.; Li, Q.; Zhang, J.; Knoll, A.; Zhou, M.; Jiang, H.; Ying, Y. Toward Damage-Less Robotic Fragile Fruit Grasping: A Closed-Loop Force Control Method for Pneumatic-Driven Soft Gripper. Soft Robot. 2024. online first. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  15. Zhang, J.; Kang, N.; Qu, Q.; Zhou, L.; Zhang, H. Automatic Fruit Picking Technology: A Comprehensive Review of Research Advances. Artif. Intell. Rev. 2024, 57, 54. [Google Scholar] [CrossRef] [Scilit]
  16. Botta, A.; Cavallone, P.; Baglieri, L.; Colucci, G.; Tagliavini, L.; Quaglia, G. A Review of Robots, Perception, and Tasks in Precision Agriculture. Appl. Mech. 2022, 3, 830–854. [Google Scholar] [CrossRef] [Scilit]
  17. Ge, C.; Zhang, G.; Wang, Y.; Shao, D.; Song, X.; Wang, Z. Research Status and Development Trends of Artificial Intelligence in Smart Agriculture. Agriculture 2025, 15, 2247. [Google Scholar] [CrossRef] [Scilit]
  18. Zhu, Y.; Zhang, S.; Tang, S.; Gao, Q. Research Progress and Applications of Artificial Intelligence in Agricultural Equipment. Agriculture 2025, 15, 1703. [Google Scholar] [CrossRef] [Scilit]
  19. Jin, Y.; Liu, J.; Xu, Z.; Yuan, S.; Li, P.; Wang, J. Development Status and Trend of Agricultural Robot Technology. Int. J. Agric. Biol. Eng. 2021, 14, 1–19. [Google Scholar] [CrossRef] [Scilit]
  20. Yang, Y.; Han, Y.; Li, S.; Yang, Y.; Zhang, M.; Li, H. Vision Based Fruit Recognition and Positioning Technology for Harvesting Robots. Comput. Electron. Agric. 2023, 213, 108258. [Google Scholar] [CrossRef] [Scilit]
  21. Jin, T.; Han, X. Robotic Arms in Precision Agriculture: A Comprehensive Review of the Technologies, Applications, Challenges, and Future Prospects. Comput. Electron. Agric. 2024, 221, 108938. [Google Scholar] [CrossRef] [Scilit]
  22. Vrochidou, E.; Tsakalidou, V.N.; Kalathas, I.; Gkrimpizis, T.; Pachidis, T.; Kaburlasos, V.G. An Overview of End Effectors in Agricultural Robotic Harvesting Systems. Agriculture 2022, 12, 1240. [Google Scholar] [CrossRef] [Scilit]
  23. Zhang, B.; Xie, Y.; Zhou, J.; Wang, K.; Zhang, Z. State-of-the-Art Robotic Grippers, Grasping and Control Strategies, as Well as Their Applications in Agricultural Robots: A Review. Comput. Electron. Agric. 2020, 177, 105694. [Google Scholar] [CrossRef] [Scilit]
  24. Xie, D.; Chen, L.; Liu, L.; Chen, L.; Wang, H. Actuators and Sensors for Application in Agricultural Robots: A Review. Machines 2022, 10, 913. [Google Scholar] [CrossRef] [Scilit]
  25. Shi, X.; Wang, S.; Zhang, B.; Zhang, Z.; Wang, S.; Ding, X.; Wang, S.; Qi, P.; Yang, H. Advances in Berry Harvesting Robots. Horticulturae 2025, 11, 1042. [Google Scholar] [CrossRef] [Scilit]
  26. Wang, W.; Li, C.; Xi, Y.; Gu, J.; Zhang, X.; Zhou, M.; Peng, Y. Research Progress and Development Trend of Visual Detection Methods for Selective Fruit Harvesting Robots. Agronomy 2025, 15, 1926. [Google Scholar] [CrossRef] [Scilit]
  27. Gao, W.; Liu, J.; Deng, J.; Jiang, Y.; Jin, Y. Research Status and Trends in Universal Robotic Picking End-Effectors for Various Fruits. Agronomy 2025, 15, 2283. [Google Scholar] [CrossRef] [Scilit]
  28. Khan, Z.; Shen, Y.; Liu, H. Object Detection in Agriculture: A Comprehensive Review of Methods, Applications, Challenges, and Future Directions. Agriculture 2025, 15, 1351. [Google Scholar] [CrossRef] [Scilit]
  29. Ma, J.; Li, M.; Fan, W.; Liu, J. State-of-the-Art Techniques for Fruit Maturity Detection. Agronomy 2024, 14, 2783. [Google Scholar] [CrossRef] [Scilit]
  30. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 Statement: An Updated Guideline for Reporting Systematic Reviews. Syst. Rev. 2021, 10, 89. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Rethlefsen, M.L.; Kirtley, S.; Waffenschmidt, S.; Ayala, A.P.; Moher, D.; Page, M.J.; Koffel, J.B.; Blunt, H.; Brigham, T.; Chang, S.; et al. PRISMA-S: An Extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Syst. Rev. 2021, 10, 39. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Arima, S.; Kondo, N. Cucumber Harvesting Robot and Plant Training System. J. Robot. Mechatron. 1999, 11, 208–212. [Google Scholar] [CrossRef] [Scilit]
  33. Ringdahl, O.; Kurtser, P.; Edan, Y. Strategies for Selecting Best Approach Direction for a Sweet-Pepper Harvesting Robot. In Proceedings of the Towards Autonomous Robotic Systems; Gao, Y., Fallah, S., Jin, Y., Lekakou, C., Eds.; Springer International Publishing: Cham, Switzerland, 2017; Volume 10454, pp. 516–525. [Google Scholar]
  34. Sánchez-Molina, J.A.; Rodríguez, F.; Moreno, J.C.; Sánchez-Hermosilla, J.; Giménez, A. Robotics in Greenhouses. Scoping Review. Comput. Electron. Agric. 2024, 219, 108750. [Google Scholar] [CrossRef] [Scilit]
  35. Kootstra, G.; Wang, X.; Blok, P.M.; Hemming, J.; van Henten, E. Selective Harvesting Robotics: Current Research, Trends, and Future Directions. Curr. Robot. Rep. 2021, 2, 95–104. [Google Scholar] [CrossRef] [Scilit]
  36. Bagagiolo, G.; Matranga, G.; Cavallo, E.; Pampuro, N. Greenhouse Robots: Ultimate Solutions to Improve Automation in Protected Cropping Systems—A Review. Sustainability 2022, 14, 6436. [Google Scholar] [CrossRef] [Scilit]
  37. Dudnyk, A.; Pasichnyk, N.; Yakymenko, I.; Lendiel, T.; Witaszek, K.; Durczak, K.; Czekała, W. Smart Resource Management and Energy-Efficient Regimes for Greenhouse Vegetable Production. Energies 2025, 18, 4690. [Google Scholar] [CrossRef] [Scilit]
  38. Hayashi, S.; Ganno, K.; Ishii, Y.; Tanaka, I. Robotic Harvesting System for Eggplants. Jpn. Agric. Res. Q. JARQ 2002, 36, 163–168. [Google Scholar] [CrossRef] [Scilit]
  39. Moreno, J.C.; Rodriguez, F.; Sánchez-Hermosilla, J.; Gimenez, A.; Sánchez-Molina, J. Feasibility Analysis of Robots in Greenhouses. A Case Study in European Mediterranean Countries. Smart Agric. Technol. 2024, 9, 100638. [Google Scholar] [CrossRef] [Scilit]
  40. van Herck, L.; Kurtser, P.; Wittemans, L.; Edan, Y. Crop Design for Improved Robotic Harvesting: A Case Study of Sweet Pepper Harvesting. Biosyst. Eng. 2020, 192, 294–308. [Google Scholar] [CrossRef] [Scilit]
  41. Bloch, V.; Degani, A.; Bechar, A. A Methodology of Orchard Architecture Design for an Optimal Harvesting Robot. Biosyst. Eng. 2018, 166, 126–137. [Google Scholar] [CrossRef] [Scilit]
  42. Van Henten, E.J.; Van’t Slot, D.A.; Hol, C.W.J.; Van Willigenburg, L.G. Optimal Manipulator Design for a Cucumber Harvesting Robot. Comput. Electron. Agric. 2009, 65, 247–257. [Google Scholar] [CrossRef] [Scilit]
  43. Bac, C.W.; Hemming, J.; Van Tuijl, B.A.J.; Barth, R.; Wais, E.; Van Henten, E.J. Performance Evaluation of a Harvesting Robot for Sweet Pepper. J. Field Robot. 2017, 34, 1123–1139. [Google Scholar] [CrossRef] [Scilit]
  44. Lehnert, C.; McCool, C.; Sa, I.; Perez, T. Performance Improvements of a Sweet Pepper Harvesting Robot in Protected Cropping Environments. J. Field Robot. 2020, 37, 1197–1223. [Google Scholar] [CrossRef] [Scilit]
  45. Huang, Y.; Xu, S.; Chen, H.; Li, G.; Dong, H.; Yu, J.; Zhang, X.; Chen, R. A Review of Visual Perception Technology for Intelligent Fruit Harvesting Robots. Front. Plant Sci. 2025, 16, 1646871. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  46. Jia, W.; Zheng, Y.; Zhao, D.; Yin, X.; Liu, X.; Du, R. Preprocessing Method of Night Vision Image Application in Apple Harvesting Robot. Int. J. Agric. Biol. Eng. 2018, 11, 158–163. [Google Scholar] [CrossRef] [Scilit]
  47. Yang, Q.; Gu, J.; Xiong, T.; Wang, Q.; Huang, J.; Xi, Y.; Shen, Z. RFA-YOLOv8: A Robust Tea Bud Detection Model with Adaptive Illumination Enhancement for Complex Orchard Environments. Agriculture 2025, 15, 1982. [Google Scholar] [CrossRef] [Scilit]
  48. Li, A.; Wang, C.; Wang, A.; Sun, J.; Gu, F.; Zhang, T. YOLO-MSRF: A Multimodal Segmentation and Refinement Framework for Tomato Fruit Detection and Segmentation with Count and Size Estimation Under Complex Illumination. Agriculture 2026, 16, 277. [Google Scholar] [CrossRef] [Scilit]
  49. Gao, J.; Zhang, J.; Zhang, F.; Gao, J. LACTA: A Lightweight and Accurate Algorithm for Cherry Tomato Detection in Unstructured Environments. Expert Syst. Appl. 2024, 238, 122073. [Google Scholar] [CrossRef] [Scilit]
  50. Zhang, F.; Chen, Z.; Ali, S.; Yang, N.; Fu, S.; Zhang, Y. Multi-Class Detection of Cherry Tomatoes Using Improved YOLOv4-Tiny. Int. J. Agric. Biol. Eng. 2023, 16, 225–231. [Google Scholar] [CrossRef] [Scilit]
  51. Liu, G.; Hou, Z.; Liu, H.; Liu, J.; Zhao, W.; Li, K. TomatoDet: Anchor-Free Detector for Tomato Detection. Front. Plant Sci. 2022, 13, 942875. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  52. Liu, X.; Jia, W.; Ruan, C.; Zhao, D.; Gu, Y.; Chen, W. The Recognition of Apple Fruits in Plastic Bags Based on Block Classification. Precis. Agric. 2018, 19, 735–749. [Google Scholar] [CrossRef] [Scilit]
  53. Hemming, J.; Ruizendaal, J.; Hofstee, J.W.; van Henten, E.J. Fruit Detectability Analysis for Different Camera Positions in Sweet-Pepper. Sensors 2014, 14, 6032–6044. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  54. Arad, B.; Kurtser, P.; Barnea, E.; Harel, B.; Edan, Y.; Ben-Shahar, O. Controlled Lighting and Illumination-Independent Target Detection for Real-Time Cost-Efficient Applications. The Case Study of Sweet Pepper Robotic Harvesting. Sensors 2019, 19, 1390. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Burusa, A.K.; Van Henten, E.J.; Kootstra, G. Attention-Driven next-Best-View Planning for Efficient Reconstruction of Plants and Targeted Plant Parts. Biosyst. Eng. 2024, 246, 248–262. [Google Scholar] [CrossRef] [Scilit]
  56. Hu, T.; Wang, W.; Gu, J.; Xia, Z.; Zhang, J.; Wang, B. Research on Apple Object Detection and Localization Method Based on Improved YOLOX and RGB-D Images. Agronomy 2023, 13, 1816. [Google Scholar] [CrossRef] [Scilit]
  57. Tang, S.; Xia, Z.; Gu, J.; Wang, W.; Huang, Z.; Zhang, W. High-Precision Apple Recognition and Localization Method Based on RGB-D and Improved SOLOv2 Instance Segmentation. Front. Sustain. Food Syst. 2024, 8, 1403872. [Google Scholar] [CrossRef] [Scilit]
  58. Gao, J.; Zhang, F.; Zhang, J.; Guo, H.; Gao, J. Picking Patterns Evaluation for Cherry Tomato Robotic Harvesting End-Effector Design. Biosyst. Eng. 2024, 239, 1–12. [Google Scholar] [CrossRef] [Scilit]
  59. Ge, Y.; Xiong, Y.; Tenorio, G.L.; From, P.J. Fruit Localization and Environment Perception for Strawberry Harvesting Robots. IEEE Access 2019, 7, 147642–147652. [Google Scholar] [CrossRef] [Scilit]
  60. Jiang, L.; Wang, Y.; Yan, H.; Yin, Y.; Wu, C. Strawberry Fruit Deformity Detection and Symmetry Quantification Using Deep Learning and Geometric Feature Analysis. Horticulturae 2025, 11, 652. [Google Scholar] [CrossRef] [Scilit]
  61. Zhao, S.; Fang, C.; Hua, T.; Jiang, Y. Detecting the Maturity of Red Strawberries Using Improved YOLOv8s Model. Agriculture 2025, 15, 2263. [Google Scholar] [CrossRef] [Scilit]
  62. Peng, Y.; Sun, J.; Wu, Z.; Shi, L.; Ji, X.; Jia, Y.; Xie, Y. Apple Maturity Quantification Based on Fine-Grained Coloration Analysis Using Deep Learning and Computer Vision. Food Meas. 2025, 19, 9637–9653. [Google Scholar] [CrossRef] [Scilit]
  63. Lehnert, C.; McCool, C.; Perez, T. In-Field Peduncle Detection of Sweet Peppers for Robotic Harvesting: A Comparative Study. arXiv 2017, arXiv:1709.10275. [Google Scholar] [CrossRef] [Scilit]
  64. McCool, C.; Sa, I.; Dayoub, F.; Lehnert, C.; Perez, T.; Upcroft, B. Visual Detection of Occluded Crop: For Automated Harvesting. In Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16–21 May 2016; IEEE: New York, NY, USA, 2016; pp. 2506–2512. [Google Scholar]
  65. Lehnert, C.; Sa, I.; McCool, C.; Upcroft, B.; Perez, T. Sweet Pepper Pose Detection and Grasping for Automated Crop Harvesting. In Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden, 16–21 May 2016; IEEE: New York, NY, USA, 2016; pp. 2428–2434. [Google Scholar]
  66. Sun, X.; Wang, L.; Zheng, Y.; Zhu, H.; Sui, Y.; Ma, Z.; Guo, R.; Hu, W.; Zhang, T.; Yan, P.; et al. Accurate Apple Fruit Stalk Cutting Technology Based on Improved YOLOv8 with Dual Cameras. Appl. Eng. Agric. 2025, 41, 97–107. [Google Scholar] [CrossRef] [Scilit]
  67. Rapado-Rincón, D.; Van Henten, E.J.; Kootstra, G. Development and Evaluation of Automated Localisation and Reconstruction of All Fruits on Tomato Plants in a Greenhouse Based on Multi-View Perception and 3D Multi-Object Tracking. Biosyst. Eng. 2023, 231, 78–91. [Google Scholar] [CrossRef] [Scilit]
  68. Liu, J.; Yuan, Y.; Gao, Y.; Tang, S.; Li, Z. Virtual Model of Grip-and-Cut Picking for Simulation of Vibration and Falling of Grape Clusters. Trans. ASABE 2019, 62, 603–614. [Google Scholar] [CrossRef] [Scilit]
  69. Xu, B.; Liu, J.; Jin, Y.; Yang, K.; Zhao, S.; Peng, Y. Vibration-Collision Coupling Modeling in Grape Clusters for Non-Damage Harvesting Operations. Agriculture 2025, 15, 154. [Google Scholar] [CrossRef] [Scilit]
  70. Xu, Y.; Zhang, X.; Sun, X.; Wang, J.; Liu, J.; Li, Z.; Guo, Q.; Li, P. Tensile Mechanical Properties of Greenhouse Cucumber Cane. Int. J. Agric. Biol. Eng. 2016, 9, 1–8. [Google Scholar] [CrossRef]
  71. Zeeshan, S.; Aized, T.; Riaz, F. Analysis of the impact of damage rate on the performance of orange fruit harvesting robot. Ain Shams Eng. J. 2025, 16, 103735. [Google Scholar] [CrossRef] [Scilit]
  72. Navas, E.; Fernandez, R.; Sepúlveda, D.; Armada, M.; Gonzalez-de-Santos, P. Soft Grippers for Automatic Crop Harvesting: A Review. Sensors 2021, 21, 2689. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  73. Ji, W.; Qian, Z.; Xu, B.; Tang, W.; Li, J.; Zhao, D. Grasping Damage Analysis of Apple by End-effector in Harvesting Robot. J. Food Process Eng. 2017, 40, e12589. [Google Scholar] [CrossRef] [Scilit]
  74. Fanourakis, D.; Makraki, T.; Spyrou, G.P.; Karavidas, I.; Tsaniklidis, G.; Ntatsi, G. Environmental Drivers of Fruit Quality and Shelf Life in Greenhouse Vegetables: Species-Specific Insights. Agronomy 2026, 16, 48. [Google Scholar] [CrossRef] [Scilit]
  75. Tang, Y.; Chen, M.; Wang, C.; Luo, L.; Li, J.; Lian, G.; Zou, X. Recognition and Localization Methods for Vision-Based Fruit Picking Robots: A Review. Front. Plant Sci. 2020, 11, 510. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  76. Ji, W.; Pan, Y.; Xu, B.; Wang, J. A Real-Time Apple Targets Detection Method for Picking Robot Based on ShufflenetV2-YOLOX. Agriculture 2022, 12, 856. [Google Scholar] [CrossRef] [Scilit]
  77. Sun, J.; He, X.; Ge, X.; Wu, X.; Shen, J.; Song, Y. Detection of Key Organs in Tomato Based on Deep Migration Learning in a Complex Background. Agriculture 2018, 8, 196. [Google Scholar] [CrossRef] [Scilit]
  78. Ji, W.; Gao, X.; Xu, B.; Pan, Y.; Zhang, Z.; Zhao, D. Apple Target Recognition Method in Complex Environment Based on Improved YOLOv4. J. Food Process Eng. 2021, 44, e13866. [Google Scholar] [CrossRef] [Scilit]
  79. Li, A.; Wang, C.; Ji, T.; Wang, Q.; Zhang, T. D3-YOLOv10: Improved YOLOv10-Based Lightweight Tomato Detection Algorithm Under Facility Scenario. Agriculture 2024, 14, 2268. [Google Scholar] [CrossRef] [Scilit]
  80. Magalhães, S.A.; Moreira, A.P.; Santos, F.N.D.; Dias, J. Active Perception Fruit Harvesting Robots—A Systematic Review. J. Intell. Robot. Syst. 2022, 105, 14. [Google Scholar] [CrossRef] [Scilit]
  81. Kang, H.; Wang, X.; Chen, C. Accurate Fruit Localisation Using High Resolution LiDAR-Camera Fusion and Instance Segmentation. Comput. Electron. Agric. 2022, 203, 107450. [Google Scholar] [CrossRef] [Scilit]
  82. Zhang, X.; Leng, Z.; Wang, X.; Tian, S.; Zhang, Y.; Han, X.; Li, Z. Analysis of the Current Situation and Trends of Optical Sensing Technology Application for Facility Vegetable Life Information Detection. Agronomy 2025, 15, 2229. [Google Scholar] [CrossRef] [Scilit]
  83. Lu, P.; Zheng, W.; Lv, X.; Xu, J.; Zhang, S.; Li, Y.; Zhangzhong, L. An Extended Method Based on the Geometric Position of Salient Image Features: Solving the Dataset Imbalance Problem in Greenhouse Tomato Growing Scenarios. Agriculture 2024, 14, 1893. [Google Scholar] [CrossRef] [Scilit]
  84. Cai, Y.; Cui, B.; Deng, H.; Zeng, Z.; Wang, Q.; Lu, D.; Cui, Y.; Tian, Y. Cherry Tomato Detection for Harvesting Using Multimodal Perception and an Improved YOLOv7-Tiny Neural Network. Agronomy 2024, 14, 2320. [Google Scholar] [CrossRef] [Scilit]
  85. Lehnert, C.; Tsai, D.; Eriksson, A.; McCool, C. 3D Move to See: Multi-Perspective Visual Servoing towards the next Best View within Unstructured and Occluded Environments. In Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 3–8 November 2019; IEEE: New York, NY, USA, 2019; pp. 3890–3897. [Google Scholar]
  86. Zapotezny-Anderson, P.; Lehnert, C. Towards Active Robotic Vision in Agriculture: A Deep Learning Approach to Visual Servoing in Occluded and Unstructured Protected Cropping Environments. IFAC-PapersOnLine 2019, 52, 120–125. [Google Scholar] [CrossRef] [Scilit]
  87. Zuo, Z.; Gao, S.; Peng, H.; Xue, Y.; Han, L.; Ma, G.; Mao, H. Lightweight Detection of Broccoli Heads in Complex Field Environments Based on LBDC-YOLO. Agronomy 2024, 14, 2359. [Google Scholar] [CrossRef] [Scilit]
  88. Xie, H.; Zhang, Z.; Zhang, K.; Yang, L.; Zhang, D.; Yu, Y. Research on the Visual Location Method for Strawberry Picking Points under Complex Conditions Based on Composite Models. J. Sci. Food Agric. 2024, 104, 8566–8579. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. He, Z.; Yuan, F.; Zhou, Y.; Cui, B.; He, Y.; Liu, Y. Stereo Vision Based Broccoli Recognition and Attitude Estimation Method for Field Harvesting. Artif. Intell. Agric. 2025, 15, 526–536. [Google Scholar] [CrossRef] [Scilit]
  90. Vitzrabin, E.; Edan, Y. Changing Task Objectives for Improved Sweet Pepper Detection for Robotic Harvesting. IEEE Robot. Autom. Lett. 2016, 1, 578–584. [Google Scholar] [CrossRef] [Scilit]
  91. Kurtser, P.; Edan, Y. Planning the Sequence of Tasks for Harvesting Robots. Robot. Auton. Syst. 2020, 131, 103591. [Google Scholar] [CrossRef] [Scilit]
  92. Chen, B.; Gong, L.; Yu, C.; Du, X.; Chen, J.; Xie, S.; Le, X.; Li, Y.; Liu, C. Workspace Decomposition Based Path Planning for Fruit-Picking Robot in Complex Greenhouse Environment. Comput. Electron. Agric. 2023, 215, 108353. [Google Scholar] [CrossRef] [Scilit]
  93. Dai, N.; Fang, J.; Yuan, J.; Liu, X. 3MSP2: Sequential Picking Planning for Multi-Fruit Congregated Tomato Harvesting in Multi-Clusters Environment Based on Multi-Views. Comput. Electron. Agric. 2024, 225, 109303. [Google Scholar] [CrossRef] [Scilit]
  94. Zion, B.; Mann, M.; Levin, D.; Shilo, A.; Rubinstein, D.; Shmulevich, I. Harvest-Order Planning for a Multiarm Robotic Harvester. Comput. Electron. Agric. 2014, 103, 75–81. [Google Scholar] [CrossRef] [Scilit]
  95. Xie, F.; Guo, Z.; Li, T.; Feng, Q.; Zhao, C. Dynamic Task Planning for Multi-Arm Harvesting Robots Under Multiple Constraints Using Deep Reinforcement Learning. Horticulturae 2025, 11, 88. [Google Scholar] [CrossRef] [Scilit]
  96. Ji, W.; Zhang, T.; Xu, B.; He, G. Apple Recognition and Picking Sequence Planning for Harvesting Robot in a Complex Environment. J. Agric. Eng. 2024, 55, 1549. [Google Scholar] [CrossRef] [Scilit]
  97. Jiang, L.; Wang, Y.; Wu, C.; Wu, H. Fruit Distribution Density Estimation in YOLO-Detected Strawberry Images: A Kernel Density and Nearest Neighbor Analysis Approach. Agriculture 2024, 14, 1848. [Google Scholar] [CrossRef] [Scilit]
  98. Shi, Y.; Jin, S.; Zhao, Y.; Huo, Y.; Liu, L.; Cui, Y. Lightweight Force-Sensing Tomato Picking Robotic Arm with a “Global-Local” Visual Servo. Comput. Electron. Agric. 2023, 204, 107549. [Google Scholar] [CrossRef] [Scilit]
  99. Zhang, T.; Huang, Z.; You, W.; Lin, J.; Tang, X.; Huang, H. An Autonomous Fruit and Vegetable Harvester with a Low-Cost Gripper Using a 3D Sensor. Sensors 2019, 20, 93. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. Gamba Camacho, J.D.; From, P.J.; Leite, A.C. A Visual Servoing Approach for Robotic Fruit Harvesting in the Presence of Parametric Uncertainties. In Proceedings of the XXII Congresso Brasileiro de Automática, João Pessoa, Paraíba, Brazil, 9–12 September 2018; Sociedade Brasileira de Automática (SBA): Campinas, Brazil, 2020; Volume 1, p. CBA2018. [Google Scholar] [CrossRef] [Scilit]
  101. Chen, K.; Li, T.; Yan, T.; Xie, F.; Feng, Q.; Zhu, Q.; Zhao, C. A Soft Gripper Design for Apple Harvesting with Force Feedback and Fruit Slip Detection. Agriculture 2022, 12, 1802. [Google Scholar] [CrossRef] [Scilit]
  102. Zhang, H.; Ji, W.; Xu, B.; Yu, X. Optimizing Contact Force on an Apple Picking Robot End-Effector. Agriculture 2024, 14, 996. [Google Scholar] [CrossRef] [Scilit]
  103. Yu, X.; Ji, W.; Zhang, H.; Ruan, C.; Xu, B.; Wu, K. Grasping Force Optimization and DDPG Impedance Control for Apple Picking Robot End-Effector. Agriculture 2025, 15, 1018. [Google Scholar] [CrossRef] [Scilit]
  104. Shi, H.; Xu, G.; Lu, W.; Ding, Q.; Chen, X. An Electric Gripper for Picking Brown Mushrooms with Flexible Force and In Situ Measurement. Agriculture 2024, 14, 1181. [Google Scholar] [CrossRef] [Scilit]
  105. Xiong, Y.; Ge, Y.; Grimstad, L.; From, P.J. An Autonomous Strawberry-Harvesting Robot: Design, Development, Integration, and Field Evaluation. J. Field Robot. 2020, 37, 202–224. [Google Scholar] [CrossRef] [Scilit]
  106. Kim, J.; Pyo, H.; Jang, I.; Kang, J.; Ju, B.; Ko, K. Tomato Harvesting Robotic System Based on Deep-ToMaToS: Deep Learning Network Using Transformation Loss for 6D Pose Estimation of Maturity Classified Tomatoes with Side-Stem. Comput. Electron. Agric. 2022, 201, 107300. [Google Scholar] [CrossRef] [Scilit]
  107. Lang, Y.; Zhang, Y.; Sun, T.; Chai, X.; Zhang, N. Digital Twin-Driven System for Efficient Tomato Harvesting in Greenhouses. Comput. Electron. Agric. 2025, 236, 110451. [Google Scholar] [CrossRef] [Scilit]
  108. Williams, H.A.M.; Jones, M.H.; Nejati, M.; Seabright, M.J.; Bell, J.; Penhall, N.D.; Barnett, J.J.; Duke, M.D.; Scarfe, A.J.; Ahn, H.S.; et al. Robotic Kiwifruit Harvesting Using Machine Vision, Convolutional Neural Networks, and Robotic Arms. Biosyst. Eng. 2019, 181, 140–156. [Google Scholar] [CrossRef] [Scilit]
  109. Birrell, S.; Hughes, J.; Cai, J.Y.; Iida, F. A Field-Tested Robotic Harvesting System for Iceberg Lettuce. J. Field Robot. 2020, 37, 225–245. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  110. Lehnert, C.; English, A.; McCool, C.; Tow, A.W.; Perez, T. Autonomous Sweet Pepper Harvesting for Protected Cropping Systems. IEEE Robot. Autom. Lett. 2017, 2, 872–879. [Google Scholar] [CrossRef] [Scilit]
  111. Li, T.; Xie, F.; Zhao, Z.; Zhao, H.; Guo, X.; Feng, Q. A Multi-Arm Robot System for Efficient Apple Harvesting: Perception, Task Plan and Control. Comput. Electron. Agric. 2023, 211, 107979. [Google Scholar] [CrossRef] [Scilit]
  112. Ceres, R.; Pons, J.L.; Jiménez, A.R.; Martín, J.M.; Calderón, L. Design and Implementation of an Aided Fruit-harvesting Robot (Agribot). Ind. Robot Int. J. 1998, 25, 337–346. [Google Scholar] [CrossRef] [Scilit]
  113. Ling, X.; Zhao, Y.; Gong, L.; Liu, C.; Wang, T. Dual-Arm Cooperation and Implementing for Robotic Harvesting Tomato Using Binocular Vision. Robot. Auton. Syst. 2019, 114, 134–143. [Google Scholar] [CrossRef] [Scilit]
  114. Droukas, L.; Doulgeri, Z.; Tsakiridis, N.L.; Triantafyllou, D.; Kleitsiotis, I.; Mariolis, I.; Giakoumis, D.; Tzovaras, D.; Kateris, D.; Bochtis, D. A Survey of Robotic Harvesting Systems and Enabling Technologies. J. Intell. Robot. Syst. 2023, 107, 21. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  115. Rajendran, V.; Debnath, B.; Mghames, S.; Mandil, W.; Parsa, S.; Parsons, S.; Ghalamzan-E., A. Towards Autonomous Selective Harvesting: A Review of Robot Perception, Robot Design, Motion Planning and Control. J. Field Robot. 2024, 41, 2247–2279. [Google Scholar] [CrossRef] [Scilit]
  116. Baeten, J.; Donné, K.; Boedrij, S.; Beckers, W.; Claesen, E. Autonomous Fruit Picking Machine: A Robotic Apple Harvester. In Field and Service Robotics; Laugier, C., Siegwart, R., Eds.; Springer Tracts in Advanced Robotics; Springer: Berlin/Heidelberg, Germany, 2008; Volume 42, pp. 531–539. [Google Scholar]
  117. Hayashi, S.; Yamamoto, S.; Tsubota, S.; Ochiai, Y.; Kobayashi, K.; Kamata, J.; Kurita, M.; Inazumi, H.; Peter, R. Automation Technologies for Strawberry Harvesting and Packing Operations in Japan. J. Berry Res. 2014, 4, 19–27. [Google Scholar] [CrossRef] [Scilit]
  118. Hayashi, S.; Shigematsu, K.; Yamamoto, S.; Kobayashi, K.; Kohno, Y.; Kamata, J.; Kurita, M. Evaluation of a Strawberry-Harvesting Robot in a Field Test. Biosyst. Eng. 2010, 105, 160–171. [Google Scholar] [CrossRef] [Scilit]
  119. Agostini, A.; Alenyà, G.; Fischbach, A.; Scharr, H.; Wörgötter, F.; Torras, C. A Cognitive Architecture for Automatic Gardening. Comput. Electron. Agric. 2017, 138, 69–79. [Google Scholar] [CrossRef] [Scilit]
  120. Huang, M.; Jiang, X.; He, L.; Choi, D.; Pecchia, J.; Li, Y. Development of a Robotic Harvesting Mechanism for Button Mushrooms. Trans. ASABE 2021, 64, 565–575. [Google Scholar] [CrossRef] [Scilit]
  121. Yu, Y.; Xie, H.; Zhang, K.; Wang, Y.; Li, Y.; Zhou, J.; Xu, L. Design, Development, Integration, and Field Evaluation of a Ridge-Planting Strawberry Harvesting Robot. Agriculture 2024, 14, 2126. [Google Scholar] [CrossRef] [Scilit]
  122. Liu, F.; Ji, W. Design of an Apple Harvesting Robot Based on Hybrid Pneumatic-Electric Drive System. Agriculture 2026, 16, 619. [Google Scholar] [CrossRef] [Scilit]
  123. Guan, X.; Shi, L.; Ge, H.; Ding, Y.; Nie, S. Development, Design, and Improvement of an Intelligent Harvesting System for Aquatic Vegetable Brasenia schreberi. Agronomy 2025, 15, 1451. [Google Scholar] [CrossRef] [Scilit]
  124. Wang, W.; Yang, S.; Zhang, X.; Xia, X. Research on the Smart Broad Bean Harvesting System and the Self-Adaptive Control Method Based on CPS Technologies. Agronomy 2024, 14, 1405. [Google Scholar] [CrossRef] [Scilit]
  125. Xu, Z.; Liu, J.; Wang, J.; Cai, L.; Jin, Y.; Zhao, S.; Xie, B. Realtime Picking Point Decision Algorithm of Trellis Grape for High-Speed Robotic Cut-and-Catch Harvesting. Agronomy 2023, 13, 1618. [Google Scholar] [CrossRef] [Scilit]
  126. Arad, B.; Balendonck, J.; Barth, R.; Ben-Shahar, O.; Edan, Y.; Hellström, T.; Hemming, J.; Kurtser, P.; Ringdahl, O.; Tielen, T.; et al. Development of a Sweet Pepper Harvesting Robot. J. Field Robot. 2020, 37, 1027–1039. [Google Scholar] [CrossRef] [Scilit]
  127. Font, D.; Pallejà, T.; Tresanchez, M.; Runcan, D.; Moreno, J.; Martínez, D.; Teixidó, M.; Palacín, J. A proposal for automatic fruit harvesting by combining a low cost stereovision camera and a robotic arm. Sensors 2014, 14, 11557–11579. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  128. Ji, W.; Zhai, K.; Xu, B.; Wu, J. Green Apple Detection Method Based on Multidimensional Feature Extraction Network Model and Transformer Module. J. Food Prot. 2025, 88, 100397. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  129. Jiang, H.; Liu, J.; Lei, X.; Xu, B.; Jin, Y. Multi-Stage Fusion of Dual Attention Mask R-CNN and Geometric Filtering for Fast and Accurate Localization of Occluded Apples. Artif. Intell. Agric. 2026, 16, 187–205. [Google Scholar] [CrossRef] [Scilit]
  130. Peng, Y.; Sun, J.; Wu, Z.; Gao, J.; Shi, L.; Shi, Z. A Vision-Based Information Processing Framework for Vineyard Grape Picking Using Two-Stage Segmentation and Morphological Perception. Horticulturae 2025, 11, 1039. [Google Scholar] [CrossRef] [Scilit]
  131. Li, X.; Ma, N.; Han, Y.; Yang, S.; Zheng, S. AHPPEBot: Autonomous Robot for Tomato Harvesting Based on Phenotyping and Pose Estimation. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: New York, NY, USA, 2024; pp. 18150–18156. [Google Scholar]
  132. Liu, J.; Liang, J.; Zhao, S.; Jiang, Y.; Wang, J.; Jin, Y. Design of a Virtual Multi-Interaction Operation System for Hand-Eye Coordination of Grape Harvesting Robots. Agronomy 2023, 13, 829. [Google Scholar] [CrossRef] [Scilit]
  133. Ao, J.; Ji, W.; Yu, X.; Ruan, C.; Xu, B. End-Effectors for Fruit and Vegetable Harvesting Robots: A Review of Key Technologies, Challenges, and Future Prospects. Agronomy 2025, 15, 2650. [Google Scholar] [CrossRef] [Scilit]
  134. Ji, W.; He, G.; Xu, B.; Zhang, H.; Yu, X. A New Picking Pattern of a Flexible Three-Fingered End-Effector for Apple Harvesting Robot. Agriculture 2024, 14, 102. [Google Scholar] [CrossRef] [Scilit]
  135. Zhang, F.; Chen, Z.; Wang, Y.; Bao, R.; Chen, X.; Fu, S.; Tian, M.; Zhang, Y. Research on Flexible End-Effectors with Humanoid Grasp Function for Small Spherical Fruit Picking. Agriculture 2023, 13, 123. [Google Scholar] [CrossRef] [Scilit]
  136. Zuo, Z.; Xue, Y.; Gao, S.; Zhang, S.; Dai, Q.; Ma, G.; Mao, H. Design and Evaluation of a Novel Actuated End Effector for Selective Broccoli Harvesting in Dense Planting Conditions. Agriculture 2025, 15, 1537. [Google Scholar] [CrossRef] [Scilit]
  137. Navas, E.; Fernandez, R.; Sepúlveda, D.; Armada, M.; Gonzalez-de-Santos, P. A Design Criterion Based on Shear Energy Consumption for Robotic Harvesting Tools. Agronomy 2020, 10, 734. [Google Scholar] [CrossRef] [Scilit]
  138. Navas, E.; Fernández, R.; Armada, M.; Gonzalez-de-Santos, P. Diaphragm-Type Pneumatic-Driven Soft Grippers for Precision Harvesting. Agronomy 2021, 11, 1727. [Google Scholar] [CrossRef] [Scilit]
  139. Pi, J.; Liu, J.; Zhou, K.; Qian, M. An Octopus-Inspired Bionic Flexible Gripper for Apple Grasping. Agriculture 2021, 11, 1014. [Google Scholar] [CrossRef] [Scilit]
  140. Zhou, K.; Xia, L.; Liu, J.; Qian, M.; Pi, J. Design of a Flexible End-Effector Based on Characteristics of Tomatoes. Int. J. Agric. Biol. Eng. 2022, 15, 13–24. [Google Scholar] [CrossRef] [Scilit]
  141. Shi, H.; Xu, G.; Xie, Y.; Lu, W.; Ding, Q.; Chen, X. Cable-Driven Underactuated Flexible Gripper for Brown Mushroom Picking. Agriculture 2025, 15, 832. [Google Scholar] [CrossRef] [Scilit]
  142. Wang, G.; Yu, Y.; Feng, Q. Design of End-Effector for Tomato Robotic Harvesting. IFAC-PapersOnLine 2016, 49, 190–193. [Google Scholar] [CrossRef] [Scilit]
  143. Gao, J.; Zhang, F.; Zhang, J.; Yuan, T.; Yin, J.; Guo, H.; Yang, C. Development and Evaluation of a Pneumatic Finger-like End-Effector for Cherry Tomato Harvesting Robot in Greenhouse. Comput. Electron. Agric. 2022, 197, 106879. [Google Scholar] [CrossRef] [Scilit]
  144. Yaguchi, H.; Nagahama, K.; Hasegawa, T.; Inaba, M. Development of an Autonomous Tomato Harvesting Robot with Rotational Plucking Gripper. In Proceedings of the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, South Korea, 9–14 October 2016; IEEE: New York, NY, USA, 2016; pp. 652–657. [Google Scholar] [CrossRef] [Scilit]
  145. Feng, Q.; Zou, W.; Fan, P.; Zhang, C.; Wang, X. Design and test of robotic harvesting system for cherry tomato. Int. J. Agric. Biol. Eng. 2018, 11, 96–100. [Google Scholar] [CrossRef] [Scilit]
  146. Yang, Q.; Zhong, X.; Qu, G.; Liu, L.; Yang, X.; Hu, X.; Min, A. A Three-Finger End-Effector with Retractable Suction for Cluster Tomato Harvesting: Design and Dynamic Performance Evaluation. Appl. Eng. Agric. 2025, 41, 603–616. [Google Scholar] [CrossRef] [Scilit]
  147. Xiong, Y.; Peng, C.; Grimstad, L.; From, P.J.; Isler, V. Development and Field Evaluation of a Strawberry Harvesting Robot with a Cable-Driven Gripper. Comput. Electron. Agric. 2019, 157, 392–402. [Google Scholar] [CrossRef] [Scilit]
  148. Faheem, M.; Liu, J.; Chang, G.; Ahmad, I.; Peng, Y. Hanging Force Analysis for Realizing Low Vibration of Grape Clusters during Speedy Robotic Post-Harvest Handling. Int. J. Agric. Biol. Eng. 2021, 14, 62–71. [Google Scholar] [CrossRef] [Scilit]
  149. Foglia, M.M.; Reina, G. Agricultural Robot for Radicchio Harvesting. J. Field Robot. 2006, 23, 363–377. [Google Scholar] [CrossRef] [Scilit]
  150. Tanigaki, K.; Fujiura, T.; Akase, A.; Imagawa, J. Cherry-Harvesting Robot. Comput. Electron. Agric. 2008, 63, 65–72. [Google Scholar] [CrossRef] [Scilit]
  151. Ma, Z.; Yang, S.; Li, J.; Qi, J. Research on SLAM Localization Algorithm for Orchard Dynamic Vision Based on YOLOD-SLAM2. Agriculture 2024, 14, 1622. [Google Scholar] [CrossRef] [Scilit]
  152. Li, M.; Gao, H.; Zhao, M.; Mao, H. Development and Experimentation of a Real-Time Greenhouse Positioning System Based on IUKF-UWB. Agriculture 2024, 14, 1479. [Google Scholar] [CrossRef] [Scilit]
  153. Yu, Y.; Li, Z.; Dai, B.; Pan, J.; Xu, L. High-Precision Mapping and Real-Time Localization for Agricultural Machinery Sheds and Farm Access Roads Environments. Agriculture 2025, 15, 2248. [Google Scholar] [CrossRef] [Scilit]
  154. Guan, X.; Ge, H.; Nie, S.; Ding, Y. Research on Agricultural Autonomous Positioning and Navigation System Based on LIO-SAM and Apriltag Fusion. Agronomy 2025, 15, 2731. [Google Scholar] [CrossRef] [Scilit]
  155. Ma, Z.; Wang, X.; Chen, X.; Hu, B.; Li, J. Advances in Crop Row Detection for Agricultural Robots: Methods, Performance Indicators, and Scene Adaptability. Agriculture 2025, 15, 2151. [Google Scholar] [CrossRef] [Scilit]
  156. Wang, L.; Zhao, B.; Fan, J.; Hu, X.; Wei, S.; Li, Y.; Zhou, Q.; Wei, C. Development of a tomato harvesting robot used in greenhouse. Int. J. Agric. Biol. Eng. 2017, 10, 140–149. [Google Scholar] [CrossRef] [Scilit]
  157. Zhao, Y.; Gong, L.; Liu, C.; Huang, Y. Dual-Arm Robot Design and Testing for Harvesting Tomato in Greenhouse. IFAC-PapersOnLine 2016, 49, 161–165. [Google Scholar] [CrossRef] [Scilit]
  158. Sepúlveda, D.; Fernández, R.; Navas, E.; Armada, M.; González-De-Santos, P. Robotic Aubergine Harvesting Using Dual-Arm Manipulation. IEEE Access 2020, 8, 121889–121904. [Google Scholar] [CrossRef] [Scilit]
  159. Lei, X.; Liu, J.; Jiang, H.; Xu, B.; Jin, Y.; Gao, J. Design and Testing of a Four-Arm Multi-Joint Apple Harvesting Robot Based on Singularity Analysis. Agronomy 2025, 15, 1446. [Google Scholar] [CrossRef] [Scilit]
  160. Wu, H.; Wang, X.; Chen, X.; Zhang, Y.; Zhang, Y. Review on Key Technologies for Autonomous Navigation in Field Agricultural Machinery. Agriculture 2025, 15, 1297. [Google Scholar] [CrossRef] [Scilit]
  161. Chen, Z.; Yin, J.; Farhan, S.M.; Liu, L.; Zhang, D.; Zhou, M.; Cheng, J. A Comprehensive Review of Obstacle Avoidance for Autonomous Agricultural Machinery in Multi-Operational Environment. Artif. Intell. Agric. 2026, 16, 139–163. [Google Scholar] [CrossRef] [Scilit]
  162. Syed, T.N.; Zhou, J.; Lakhiar, I.A.; Marinello, F.; Gemechu, T.T.; Rottok, L.T.; Jiang, Z. Enhancing Autonomous Orchard Navigation: A Real-Time Convolutional Neural Network-Based Obstacle Classification System for Distinguishing ‘Real’ and ‘Fake’ Obstacles in Agricultural Robotics. Agriculture 2025, 15, 827. [Google Scholar] [CrossRef] [Scilit]
  163. Jiang, Q.; Shen, Y.; Liu, H.; Khan, Z.; Sun, H.; Huang, Y. A Hybrid Path Planning Algorithm for Orchard Robots Based on an Improved D* Lite Algorithm. Agriculture 2025, 15, 1698. [Google Scholar] [CrossRef] [Scilit]
  164. Shen, Y.; Shen, Y.; Zhang, Y.; Huo, C.; Shen, Z.; Su, W.; Liu, H. Research Progress on Path Planning and Tracking Control Methods for Orchard Mobile Robots in Complex Scenarios. Agriculture 2025, 15, 1917. [Google Scholar] [CrossRef] [Scilit]
  165. Zhao, Y.; Gong, L.; Huang, Y.; Liu, C. A Review of Key Techniques of Vision-Based Control for Harvesting Robot. Comput. Electron. Agric. 2016, 127, 311–323. [Google Scholar] [CrossRef] [Scilit]
  166. Barth, R.; Hemming, J.; Van Henten, E.J. Design of an Eye-in-Hand Sensing and Servo Control Framework for Harvesting Robotics in Dense Vegetation. Biosyst. Eng. 2016, 146, 71–84. [Google Scholar] [CrossRef] [Scilit]
  167. Zhao, C.; Wang, H.; Li, W.; Zheng, H.; Zhou, L.; Qian, M. Cucumber Robotic Continuous Harvesting: Enhanced YOLOv8n Detection and Dynamic Bézier Curve-Assisted Collision-Free Path Generation. Agriculture 2026, 16, 888. [Google Scholar] [CrossRef] [Scilit]
  168. Park, Y.; Seol, J.; Pak, J.; Jo, Y.; Kim, C.; Son, H.I. Human-Centered Approach for an Efficient Cucumber Harvesting Robot System: Harvest Ordering, Visual Servoing, and End-Effector. Comput. Electron. Agric. 2023, 212, 108116. [Google Scholar] [CrossRef] [Scilit]
  169. Fernández, R.; Montes, H.; Surdilovic, J.; Surdilovic, D.; Gonzalez-De-Santos, P.; Armada, M. Automatic Detection of Field-Grown Cucumbers for Robotic Harvesting. IEEE Access 2018, 6, 35512–35527. [Google Scholar] [CrossRef] [Scilit]
  170. Tian, Y.; Duan, H.; Luo, R.; Zhang, Y.; Jia, W.; Lian, J.; Zheng, Y.; Ruan, C.; Li, C. Fast Recognition and Location of Target Fruit Based on Depth Information. IEEE Access 2019, 7, 170553–170563. [Google Scholar] [CrossRef] [Scilit]
  171. Salinas, C.; Fernández, R.; Montes, H.; Armada, M. A New Approach for Combining Time-of-Flight and RGB Cameras Based on Depth-Dependent Planar Projective Transformations. Sensors 2015, 15, 24615–24643. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  172. Xiang, R.; Jiang, H.; Ying, Y. Recognition of Clustered Tomatoes Based on Binocular Stereo Vision. Comput. Electron. Agric. 2014, 106, 75–90. [Google Scholar] [CrossRef] [Scilit]
  173. Yin, W.; Wen, H.; Ning, Z.; Ye, J.; Dong, Z.; Luo, L. Fruit detection and pose estimation for grape cluster-harvesting robot using binocular imagery based on deep neural networks. Front. Robot. AI 2021, 8, 626989. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  174. Xiong, J.; He, Z.; Lin, R.; Liu, Z.; Bu, R.; Yang, Z.; Peng, H.; Zou, X. Visual Positioning Technology of Picking Robots for Dynamic Litchi Clusters with Disturbance. Comput. Electron. Agric. 2018, 151, 226–237. [Google Scholar] [CrossRef] [Scilit]
  175. Park, Y.; Kim, C.; Son, H.I. Fast and Stable Pedicel Detection for Robust Visual Servoing to Harvest Shaking Fruits. Comput. Electron. Agric. 2024, 220, 108863. [Google Scholar] [CrossRef] [Scilit]
  176. Bac, C.W.; Hemming, J.; van Henten, E.J. Stem Localization of Sweet-Pepper Plants Using the Support Wire as a Visual Cue. Comput. Electron. Agric. 2014, 105, 111–120. [Google Scholar] [CrossRef] [Scilit]
  177. Nemlekar, H.; Liu, Z.; Kothawade, S.; Niyaz, S.; Raghavan, B.; Nikolaidis, S. Robotic Lime Picking by Considering Leaves as Permeable Obstacles. In Proceedings of the 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Prague, Czech Republic, 27 September–1 October 2021; IEEE: New York, NY, USA, 2021; pp. 3278–3284. [Google Scholar]
  178. Van Henten, E.J.; Hemming, J.; Van Tuijl, B.A.J.; Kornet, J.G.; Meuleman, J.; Bontsema, J.; Van Os, E.A. An Autonomous Robot for Harvesting Cucumbers in Greenhouses. Auton. Robots 2002, 13, 241–258. [Google Scholar] [CrossRef] [Scilit]
  179. Van Henten, E.J.; Hemming, J.; Van Tuijl, B.A.J.; Kornet, J.G.; Bontsema, J. Collision-Free Motion Planning for a Cucumber Picking Robot. Biosyst. Eng. 2003, 86, 135–144. [Google Scholar] [CrossRef] [Scilit]
  180. Van Henten, E.J.; Van Tuijl, B.A.J.; Hemming, J.; Kornet, J.G.; Bontsema, J.; Van Os, E.A. Field Test of an Autonomous Cucumber Picking Robot. Biosyst. Eng. 2003, 86, 305–313. [Google Scholar] [CrossRef] [Scilit]
  181. Ahmed, A.; Zhang, Z.; Manzoor, S.H.; Abdelhamid, M.A.; Gul, N.; Ahmed, R.; Mhamed, M.; Sun, R.; Hao, C.; Huo, W.; et al. Cucumber Picking Robots: Technological Progress, Challenges, and Future Directions. Smart Agric. Technol. 2026, 13, 101813. [Google Scholar] [CrossRef] [Scilit]
  182. Zhou, H.; Wang, X.; Au, W.; Kang, H.; Chen, C. Intelligent Robots for Fruit Harvesting: Recent Developments and Future Challenges. Precis. Agric. 2022, 23, 1856–1907. [Google Scholar] [CrossRef] [Scilit]
  183. Bac, C.W.; Van Henten, E.J.; Hemming, J.; Edan, Y. Harvesting Robots for High-value Crops: State-of-the-art Review and Challenges Ahead. J. Field Robot. 2014, 31, 888–911. [Google Scholar] [CrossRef] [Scilit]
  184. Beldek, C.; Cunningham, J.; Aydin, M.; Sariyildiz, E.; Phung, S.L.; Alici, G. Sensing-Based Robustness Challenges in Agricultural Robotic Harvesting. In Proceedings of the 2025 IEEE International Conference on Mechatronics (ICM), Wollongong, Australia, 28 February–2 March 2025; IEEE: New York, NY, USA, 2025; pp. 1–6. [Google Scholar]
  185. Williams, H.; Ting, C.; Nejati, M.; Jones, M.H.; Penhall, N.; Lim, J.Y.; Seabright, M.; Bell, J.; Ahn, H.S.; Scarfe, A.; et al. Improvements to and Large-Scale Evaluation of a Robotic Kiwifruit Harvester. J. Field Robot. 2020, 37, 187–201. [Google Scholar] [CrossRef] [Scilit]
  186. Wei, P.; Cao, S.; Liu, J.; Liu, Z.; Sun, W.; Kong, F. Embodied Intelligent Agricultural Robots: Key Technologies, Application Analysis, Challenges and Prospects. Smart Agric. 2025, 7, 141–158. [Google Scholar] [CrossRef]
  187. Liu, B.; Zhu, Y.; Gao, C.; Feng, Y.; Liu, Q.; Zhu, Y.; Stone, P. LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning. Adv. Neural Inf. Process. Syst. 2023, 36, 44776–44791. [Google Scholar] [CrossRef] [Scilit]
  188. Bousmalis, K.; Vezzani, G.; Rao, D.; Devin, C.; Lee, A.X.; Bauza, M.; Davchev, T.; Zhou, Y.; Gupta, A.; Raju, A.; et al. RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation. arXiv 2023, arXiv:2306.11706. [Google Scholar] [CrossRef] [Scilit]
  189. Zhao, W.; Queralta, J.P.; Westerlund, T. Sim-to-Real Transfer in Deep Reinforcement Learning for Robotics: A Survey. In Proceedings of the 2020 IEEE Symposium Series on Computational Intelligence (SSCI), Canberra, ACT, Australia, 1–4 December 2020; IEEE: New York, NY, USA, 2020; pp. 737–744. [Google Scholar]
  190. Walke, H.R.; Black, K.; Zhao, T.Z.; Vuong, Q.; Zheng, C.; Hansen-Estruch, P.; He, A.W.; Myers, V.; Kim, M.J.; Du, M.; et al. BridgeData V2: A Dataset for Robot Learning at Scale. In Proceedings of the 7th Conference on Robot Learning; Tan, J., Toussaint, M., Darvish, K., Eds.; Proceedings of Machine Learning Research (PMLR): Cambridge, MA, USA, 2023; Volume 229, pp. 1723–1736. [Google Scholar]
  191. Williams, E.; Polydoros, A. Zero-Shot Sim-to-Real Reinforcement Learning for Fruit Harvesting. In Proceedings of the 2025 IEEE 21st International Conference on Automation Science and Engineering (CASE), Los Angeles, CA, USA, 17–21 August 2025; IEEE: New York, NY, USA, 2025; pp. 1423–1428. [Google Scholar]
  192. Mahmoudi, S.; Davar, A.; Sohrabipour, P.; Bist, R.B.B.; Tao, Y.; Wang, D. Leveraging Imitation Learning in Agricultural Robotics: A Comprehensive Survey and Comparative Analysis. Front. Robot. AI 2024, 11, 1441312. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  193. Chi, C.; Xu, Z.; Feng, S.; Cousineau, E.; Du, Y.; Burchfiel, B.; Tedrake, R.; Song, S. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion. Int. J. Robot. Res. 2025, 44, 1684–1704. [Google Scholar] [CrossRef] [Scilit]
  194. Bechar, A.; Edan, Y. Human-robot Collaboration for Improved Target Recognition of Agricultural Robots. Ind. Robot 2003, 30, 432–436. [Google Scholar] [CrossRef] [Scilit]
  195. Zhao, J.; Fan, S.; Zhang, B.; Wang, A.; Zhang, L.; Zhu, Q. Research Status and Development Trends of Deep Reinforcement Learning in the Intelligent Transformation of Agricultural Machinery. Agriculture 2025, 15, 1223. [Google Scholar] [CrossRef] [Scilit]
  196. Firoozi, R.; Tucker, J.; Tian, S.; Majumdar, A.; Sun, J.; Liu, W.; Zhu, Y.; Song, S.; Kapoor, A.; Hausman, K.; et al. Foundation Models in Robotics: Applications, Challenges, and the Future. Int. J. Robot. Res. 2025, 44, 701–739. [Google Scholar] [CrossRef] [Scilit]
  197. Driess, D.; Xia, F.; Sajjadi, M.S.M.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. PaLM-E: An Embodied Multimodal Language Model. In Proceedings of the 40th International Conference on Machine Learning; Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J., Eds.; Proceedings of Machine Learning Research (PMLR): Cambridge, MA, USA, 2023; Volume 202, pp. 8469–8488. [Google Scholar]
  198. Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control. In Proceedings of the 7th Conference on Robot Learning; Tan, J., Toussaint, M., Darvish, K., Eds.; Proceedings of Machine Learning Research (PMLR): Cambridge, MA, USA, 2023; Volume 229, pp. 2165–2183. [Google Scholar]
  199. Ghosh, D.; Walke, H.R.; Pertsch, K.; Black, K.; Mees, O.; Dasari, S.; Hejna, J.; Kreiman, T.; Xu, C.; Luo, J.; et al. Octo: An open-source generalist robot policy. In Proceedings of the Robotics: Science and Systems XX, Delft, The Netherlands, 15–19 July 2024. [Google Scholar] [CrossRef] [Scilit]
  200. Kim, M.J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.P.; Sanketi, P.R.; Vuong, Q.; et al. OpenVLA: An Open-Source Vision-Language-Action Model. In Proceedings of the 8th Conference on Robot Learning; Agrawal, P., Kroemer, O., Burgard, W., Eds.; Proceedings of Machine Learning Research (PMLR): Cambridge, MA, USA, 2025; Volume 270, pp. 2679–2713. [Google Scholar]
  201. Yin, S.; Xi, Y.; Zhang, X.; Sun, C.; Mao, Q. Foundation Models in Agriculture: A Comprehensive Review. Agriculture 2025, 15, 847. [Google Scholar] [CrossRef] [Scilit]
  202. O’Neill, A.; Rehman, A.; Maddukuri, A.; Gupta, A.; Padalkar, A.; Lee, A.; Pooley, A.; Gupta, A.; Mandlekar, A.; Jain, A.; et al. Open X-Embodiment: Robotic learning datasets and RT-X models. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 6892–6903. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Literature identification, screening, and post-revision duplicate reconciliation.
Figure 1. Literature identification, screening, and post-revision duplicate reconciliation.
Agronomy 16 01795 g001
Figure 2. Typical cultivation systems for greenhouse crops: (a) V-shaped training of sweet pepper; (b) high-wire training of tomato; and (c) raised-bed cultivation of strawberry. Reproduced from Kootstra et al. [35] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 2. Typical cultivation systems for greenhouse crops: (a) V-shaped training of sweet pepper; (b) high-wire training of tomato; and (c) raised-bed cultivation of strawberry. Reproduced from Kootstra et al. [35] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g002
Figure 3. Representative tomato samples under different growing conditions: (A) a single tomato; (B) a tomato cluster; (C) occlusion; (D) overlap; (E) shading; and (F) sunlight. Reproduced from Liu et al. [51] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 3. Representative tomato samples under different growing conditions: (A) a single tomato; (B) a tomato cluster; (C) occlusion; (D) overlap; (E) shading; and (F) sunlight. Reproduced from Liu et al. [51] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g003
Figure 4. Complex spatial relationships among greenhouse sweet-pepper fruits, peduncles, and surrounding foliage. Photograph by Yan Krukau, Pexels; used under the Pexels License (accessed 22 July 2026).
Figure 4. Complex spatial relationships among greenhouse sweet-pepper fruits, peduncles, and surrounding foliage. Photograph by Yan Krukau, Pexels; used under the Pexels License (accessed 22 July 2026).
Agronomy 16 01795 g004
Figure 5. Hypothetical applications of different soft-gripper configurations in fruit and vegetable harvesting: (a) continuum enveloping gripper; (b) suction gripper; (c) corrugated soft gripper; (d) multi-finger soft gripper; (e) ring gripper; and (f) tendon-driven gripper. Adapted from Navas et al. [72] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 5. Hypothetical applications of different soft-gripper configurations in fruit and vegetable harvesting: (a) continuum enveloping gripper; (b) suction gripper; (c) corrugated soft gripper; (d) multi-finger soft gripper; (e) ring gripper; and (f) tendon-driven gripper. Adapted from Navas et al. [72] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g005
Figure 6. RGB-D multimodal information and target-detection results in a greenhouse cherry-tomato scene: (a) RGB image; (b) multimodal image overlaid with point-cloud normals; (c) detection results from the baseline RGB-D model; and (d) detection results from the improved RGB-D model. Orange arrows indicate additional detections of distant or partially occluded fruits. Adapted from Cai et al. [84] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 6. RGB-D multimodal information and target-detection results in a greenhouse cherry-tomato scene: (a) RGB image; (b) multimodal image overlaid with point-cloud normals; (c) detection results from the baseline RGB-D model; and (d) detection results from the improved RGB-D model. Orange arrows indicate additional detections of distant or partially occluded fruits. Adapted from Cai et al. [84] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g006
Figure 7. General system configuration of a selective harvesting robot. Reproduced from Rajendran et al. [115] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 7. General system configuration of a selective harvesting robot. Reproduced from Rajendran et al. [115] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g007
Figure 8. Proposed conceptual framework for perception, scene representation, and decision making in a greenhouse harvesting robot. Perceptual data are calibrated, temporally synchronized, and updated to form the task-scene state S_t. The decision layer sequentially performs harvestability gating, task selection, and execution-constraint generation, whereas active observation, platform adjustment, manipulator approach, and end-effector operation in turn alter the scene and its observability. Synthesized from References [80,93,106,107,131]. Colors distinguish the four functional blocks, while arrows indicate information/action flow, the active-perception loop, and execution feedback. Proposed by this review.
Figure 8. Proposed conceptual framework for perception, scene representation, and decision making in a greenhouse harvesting robot. Perceptual data are calibrated, temporally synchronized, and updated to form the task-scene state S_t. The decision layer sequentially performs harvestability gating, task selection, and execution-constraint generation, whereas active observation, platform adjustment, manipulator approach, and end-effector operation in turn alter the scene and its observability. Synthesized from References [80,93,106,107,131]. Colors distinguish the four functional blocks, while arrows indicate information/action flow, the active-perception loop, and execution feedback. Proposed by this review.
Agronomy 16 01795 g008
Figure 9. System-level interfaces among task and scene state, end-effector operation, fruit collection and transfer, outcome verification, and failure recovery during robotic harvesting. Solid arrows indicate normal task progression, dashed arrows indicate feedback from outcome verification to upstream task stages, and dotted arrows represent failure and recovery pathways. The framework emphasizes cross-stage state updating and fault recovery rather than hardware taxonomy. Proposed by this review.
Figure 9. System-level interfaces among task and scene state, end-effector operation, fruit collection and transfer, outcome verification, and failure recovery during robotic harvesting. Solid arrows indicate normal task progression, dashed arrows indicate feedback from outcome verification to upstream task stages, and dotted arrows represent failure and recovery pathways. The framework emphasizes cross-stage state updating and fault recovery rather than hardware taxonomy. Proposed by this review.
Agronomy 16 01795 g009
Figure 10. Application-control states and transitions for an eye-in-hand perception and visual-servo framework. Adapted from Barth et al. [166] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 10. Application-control states and transitions for an eye-in-hand perception and visual-servo framework. Adapted from Barth et al. [166] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g010
Figure 11. Continuous greenhouse cucumber-harvesting robot and its principal components. Reproduced from Zhao et al. [167] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 11. Continuous greenhouse cucumber-harvesting robot and its principal components. Reproduced from Zhao et al. [167] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g011
Figure 12. Multi-target continuous-harvesting order and path-adjustment method for greenhouse cucumbers. Circles represent target fruits; different colors distinguish target/task clusters, while arrows indicate the successive clustering, distance-weighting, sequence-planning, and path-adjustment steps. Adapted from Zhao et al. [167] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 12. Multi-target continuous-harvesting order and path-adjustment method for greenhouse cucumbers. Circles represent target fruits; different colors distinguish target/task clusters, while arrows indicate the successive clustering, distance-weighting, sequence-planning, and path-adjustment steps. Adapted from Zhao et al. [167] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g012
Figure 13. Structure and operation of a continuous-harvesting end effector for greenhouse cucumbers: (a) end-effector structure; (b) cucumber-harvesting trial. Adapted from Zhao et al. [167] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Figure 13. Structure and operation of a continuous-harvesting end effector for greenhouse cucumbers: (a) end-effector structure; (b) cucumber-harvesting trial. Adapted from Zhao et al. [167] under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Agronomy 16 01795 g013
Figure 14. Qualitative research landscape across three levels of embodied intelligence and six production-relevant evaluation dimensions in greenhouse robotic harvesting. (A) Representative capabilities and research directions identified in the reviewed literature, distinguishing experimentally validated closed-loop capabilities, emerging embodied-interaction mechanisms, and prospective embodied-learning directions. (B) Qualitative synthesis of the evidence status for six production-relevant evaluation dimensions. The color coding represents comparative evidence maturity and research gaps rather than a quantitative classification or scoring of all 188 unique primary studies. Proposed by this review.
Figure 14. Qualitative research landscape across three levels of embodied intelligence and six production-relevant evaluation dimensions in greenhouse robotic harvesting. (A) Representative capabilities and research directions identified in the reviewed literature, distinguishing experimentally validated closed-loop capabilities, emerging embodied-interaction mechanisms, and prospective embodied-learning directions. (B) Qualitative synthesis of the evidence status for six production-relevant evaluation dimensions. The color coding represents comparative evidence maturity and research gaps rather than a quantitative classification or scoring of all 188 unique primary studies. Proposed by this review.
Agronomy 16 01795 g014
Table 1. Typical morphological parameters of cherry-tomato clusters and individual fruits and their implications for robotic operation.
Table 1. Typical morphological parameters of cherry-tomato clusters and individual fruits and their implications for robotic operation.
ObjectParameterMeasured ValuePrimary Implications for Robotic Operation
Fruit clusterLength/width/height69.34 ± 4.97/30.09 ± 1.78/194.86 ± 14.15 mmThe fruit cluster spans a large vertical range, requiring layered observation and planning of the harvesting order for individual fruits.
Fruit clusterMain rachis length/number of fruits45.78 ± 7.54 mm/approximately 12The attachment structure and number of fruits jointly determine occlusion relationships and feasible approach directions within the cluster.
Individual fruitLong axis/short axis/fruit height29.54 ± 1.72/27.65 ± 2.06/27.65 ± 2.00 mmThese dimensions inform gripper aperture, enveloping scale, and the spatial resolution required for visual localization.
Individual fruitPeduncle length/mass17.31 ± 1.12 mm/13.35 ± 2.64 gThe small detachment structure increases localization-accuracy requirements, while fruit mass affects holding force and transfer dynamics.
Individual fruit populationDominant size and mass ranges>90% of diameters: 27–33 mm; >95% of masses: 10–18 gThe end effector can be designed around the dominant population range, but compliance must be retained for outliers and pose variation.
Source: Reference [58].
Table 2. Comparison of interaction characteristics among four cherry-tomato detachment modes.
Table 2. Comparison of interaction characteristics among four cherry-tomato detachment modes.
Detachment ModePeak Force (N)Cluster Disturbance Distance (mm)Peak Angle
(°)
Calyx Retention (%)Implications for Robotic Operation
Press and snap12.325.8 ± 2.970.8 ± 6.296.7Low disturbance and good marketability but requires precise identification of the small abscission zone and bending direction.
Pulling15.2175.9 ± 20.2Not applicable80.0Simple action, but both the required peak force and plant disturbance are high, making it unsuitable for continuous harvesting in dense clusters.
Combined pull and twist14.6134.0 ± 18.9132.4 ± 8.976.7More complex action and control; still causes substantial disturbance and has relatively poor overall compatibility.
Twisting13.422.3 ± 4.31158.6 ± 122.376.7Minimal plant disturbance but requires a large and continuous end-effector rotation range.
Source: Reference [58]. Maximum peak force is the largest of the three finger measurements; the large peak angle in the twisting mode results from multiple consecutive rotations.
Table 3. Mapping of requirements for harvestability decision making and motion planning.
Table 3. Mapping of requirements for harvestability decision making and motion planning.
Decision ScenarioCore AssessmentRequired Output
Insufficient critical informationAre current observations sufficient for safe manipulation, and can additional observation substantially reduce uncertainty?Object to verify, recommended viewpoint, minimum confidence, and return conditions
Target visible but operating conditions inadequateCan reachability be restored by repositioning the platform or manipulator or by changing the approach direction?Repositioning method, pre-manipulation pose, permitted directions, and prohibited regions
Multiple coupled targetsWould changing the order improve subsequent visibility, local clearance, and total cost?Dynamic priority, expected cost and risk, and triggers for reordering
Source: Synthesized from References [91,92,93].
Table 4. Feedback and regulation requirements across stages of compliant manipulation.
Table 4. Feedback and regulation requirements across stages of compliant manipulation.
Manipulation StagePrimary Uncertainty or RiskRequired Feedback and Regulation
ApproachResidual pose error and changes in local clearanceVisual or distance-based verification; adjust speed, pose, and approach direction
Initial contactIncorrect contact location, collision, or sudden load increaseDetect force, tactile, or actuator events; stop or perform a short retreat
Stable holdingInsufficient retention, excessive load, or target slipFuse pressure, load, and relative-motion signals; adjust retention or terminate
Detachment and transferIncomplete detachment, dropped fruit, or incomplete collectionCross-validate vision and state before and after action; limited retry, safe placement, or task rollback
Source: Synthesized from References [58,98]; load and displacement limits for different crops should be determined through object-specific measurements and online feedback.
Table 5. Hierarchy of anomalies and recovery requirements in continuous operation.
Table 5. Hierarchy of anomalies and recovery requirements in continuous operation.
Failure LevelTypical ConditionRecovery Requirement
Local perception or poseTemporary occlusion, reduced confidence, or small target displacementAdditional observation, local realignment, reduced speed, or short retreat
Task, path, or manipulationInvalid approach direction, blocked path, or incomplete retention or detachmentReturn to a safe pose, replan, change order, or perform a limited retry
System, equipment, or safetyEnd-effector or collection fault, communication or energy failure, or risk-limit violationStop related actions, enter a safe state, retain logs, and request maintenance or human takeover
Source: Synthesized from References [105,106,107,111]; recovery strategies should be matched to failure level, risk, and time budget.
Table 7. Functional implementation in representative perception–scene representation–decision systems.
Table 7. Functional implementation in representative perception–scene representation–decision systems.
System or StudyState RepresentationDecision/PlanningImplementation Implication
Deep-ToMaToS [106]Maturity, target detection, and six-dimensional pose are unified within the same target objectTarget state is transferred directly to harvesting-action controlPerception outputs should correspond to the task variables required for execution
3MSP2 [93]Shoulder-mounted and eye-in-hand observations maintain multi-fruit-cluster relationships and candidate posesHarvesting order and manipulation poses are generated jointly and updated after target removalMultiview states should support joint decision making over order and pose
AHPPEBot [131]Cluster-fruit relationships, maturity, volume, and peduncle keypoints jointly represent the targetTarget selection and path planning are performed in conjunction with the manipulator workspaceTarget structure and robot capability must be validated within the same state
Digital twin [107]Dynamic scanning constructs a greenhouse-scale scene of targets, plants, and the robotGlobal and local tasks, platform position, and manipulator actions share a common scene versionA unified state can connect platform scheduling with local manipulation
General state serviceObject table, relational graph, robot state, task stage, and uncertaintyEvent-triggered updates provide interfaces to observation, planning, and execution servicesModules should share a unique target identity and scene version
Source: References [93,106,107,131].
Table 8. Functional chain of the end -effector and fruit-collection subsystems.
Table 8. Functional chain of the end -effector and fruit-collection subsystems.
Functional StageCommon ImplementationKey State InterfacePrincipal Engineering Constraint
Target holdingGripping, enveloping, suction, or combined mechanismsContact location, pressure/load, slip, and retention confirmationFruit dimensions, surface damage, localization error, and retention stability
Fruit detachmentCutting, breaking, pulling, twisting, or composite actionsDetachment site, action parameters, connection-release state, and termination conditionsLocal clearance, plant disturbance, rotational stroke, and blade safety
End-effector perception and controlEye-in-hand vision, distance, force, tactile, pressure, and actuator statesStage, confidence, anomaly events, and local-correction outcomeSensor synchronization, noise, calibration, and end-effector volume
Fruit reception and temporary storageEnd-effector container, tracking receptacle, chute, or pneumatic tubeBin entry, capacity, blockage, and dropped-fruit statesWrist payload, fruit impact, stacking, and channel blockage
Unloading and logisticsManipulator transfer, independent fruit container, and automatic unloading interfaceUnloading request, completion confirmation, and container resetAdditional trajectory, time budget, localization, and system interlocks
Source: Compiled from References [72,126,133,142,143,147].
Table 10. Implementation levels for interaction feedback, autonomous recovery, and system integration.
Table 10. Implementation levels for interaction feedback, autonomous recovery, and system integration.
LevelPrimary State InputControl or Management MechanismAnomaly Handling
Motion feedbackImage error, relative pose, joint state, and collision distanceVisual servoing, speed adjustment, and local trajectory correctionPause, retreat, additional observation, or replanning
Contact feedbackForce, tactile, pressure, actuator state, and target motionStage-specific thresholds, retention regulation, and safe stoppingRelease load, adjust pose, perform a limited retry, or terminate
Task controlTarget identity, scene version, action stage, and completion eventFinite-state machine, behavior tree, or hierarchical task schedulingDefer or skip the target, roll back the task, and verify recovery
System safetyCommunication, energy, container, and equipment states and standardized fault codesWatchdogs, interlocks, emergency stops, logging, and human-takeover interfaceIsolate the fault, enter a safe state, and retain traceable records
Source: Compiled from References [98,105,107,166].
Table 11. Representative data for scene perception and manipulation-target representation in cucumber harvesting.
Table 11. Representative data for scene perception and manipulation-target representation in cucumber harvesting.
System/ScenarioPerceptual and Representational ObjectRepresentative DataImplication for Manipulation
Global RGB-D [168]Cucumber targets, depth, and global position1920 × 1080; approximately 5000 images and 30 video sequences; candidate-target confidence threshold > 0.95Construct candidate targets and harvesting order; a detection box cannot substitute for peduncle and approach poses
Complex-background detection [169]Visible region, centroid, and major-axis direction45 scenes; P/R/F1 = 85.65%/90.10%/87.8%; centroid errors of 6/5 px; orientation MAE of 10.1 degreesOrientation is more sensitive to occlusion, overexposure, and chromatic similarity; the manipulation representation must retain uncertainty
Eye-in-hand local perception [168]Fruit contour, peduncle region, and local three-dimensional pose640 × 480; 16–23 FPS; depth-based background removal and FPFH featuresUse continuous local observation to replace single-frame long-range localization and correct the peduncle-end-effector relative pose
6DRVS peduncle perception [175]Peduncle segmentation, tracking, and six-dimensional pose under oscillation100 cucumber samples; 640 × 480; peduncle detection at 15–37 FPS; with 6DRVS, perception and approach success rates were 90.00% and 82.22%, respectively, versus 70.00% and 77.14% without 6DRVSVideo stabilization, frequency-domain filtering, and six-dimensional pose estimation convert dynamic observations into online approach constraints
Source: Compiled from References [168,169,175]. Values originate from different experiments and are shown to illustrate changes in task-state requirements; they should not be interpreted as a direct performance ranking.
Table 12. Cucumber-target states, response actions, and reassessment conditions.
Table 12. Cucumber-target states, response actions, and reassessment conditions.
Cucumber-Target StateKey EvidenceDecision ActionPlanning/Control OutputReassessment Trigger
Fruit and peduncle reliable; sufficient local clearanceStable identity and reliable pose; main stem and neighboring fruits outside the tool envelopeExecute the current targetPre-manipulation pose, approach direction, speed, termination conditions, and retreat pathTarget oscillation, clearance change, or excessive approach error
Fruit reliable; peduncle or upper connection unclearMissing depth, occlusion, or high axial uncertaintyPause and obtain additional local observationsRecommended viewpoint, observation distance, and minimum information requirementPeduncle visibility reaches the threshold; defer after the observation budget is exhausted
Target passes gating; current entry path congestedOuter targets cause occlusion, platform position is poor, or a local obstacle occupies the corridorReturn to the pre-manipulation pose, reorder, or repositionPrioritize outer targets; adjust platform or manipulator; generate a new pre-manipulation pose and local replanTarget removal, completed platform repositioning, or restoration of local-path feasibility
High main-stem or infrastructure risk; repeated entry failureHazardous tissue enters the tool window, inverse-kinematic margin is insufficient, or failure conditions remain unchangedMark temporarily unharvestable; defer or skipFailure cause, reassessment conditions, and safe retreat pathSubstantive change in viewpoint, agronomic state, or robot position
Source: Compiled from References [32,167,168,175,176,177,178,179].
Table 13. Representative complete-system and stage-level data for the cucumber-harvesting task chain.
Table 13. Representative complete-system and stage-level data for the cucumber-harvesting task chain.
System/StudyValidation Context and Autonomy/Human InterventionKey Reported PerformanceSystem-Level Implication
Early autonomous cucumber robots [178,180]Greenhouse/protected-cultivation trials; autonomous single-fruit operation reported; intervention frequency NRComplete-task success approximately 80% in one trial and 74.4% in another; cycle approximately 45 s/fruit, or 65.2 s per successful fruit and 124 s/fruit when all attempts were includedComplete-chain feasibility was demonstrated, but retries and return motion strongly reduced effective
throughput
Hierarchical perception and integrated end effector [168]Greenhouse/protected cultivation; hierarchical perception, local servoing, suction retention, cutting, and gravity transfer; intervention frequency NRComplete-system success 56.6%; stage rates 83.8% detection, 78.4% entry, and 86.2% cutting; mean cycle 56.0 s/fruitSerial stage losses accumulate; canopy entry and local alignment are critical interfaces
Continuous-harvesting system [167]Greenhouse/protected cultivation; multi-target continuous-path operation; intervention frequency NRComplete-task harvesting success and full cycle time NR; collision-free rate 92.24%; path length 31.1% of the return-path baselineNonproductive inter-target motion was reduced, but complete outcome and intervention metrics remain necessary
Source: Data compiled from References [167,168,178,180]. NR, not reported. Values originate from heterogeneous cultivars, protocols, sample sizes, and performance definitions and cannot be compared directly as a ranking.
Table 14. Minimum system-level reporting set proposed by this review for greenhouse robotic harvesting.
Table 14. Minimum system-level reporting set proposed by this review for greenhouse robotic harvesting.
Evaluation DomainMinimum Reporting SetOperational Definition and Context
Complete-task successTotal candidate targets; eligible/harvestable targets; successful marketable fruits; explicit failure/exclusion countsReport success per eligible target with crop/cultivar, maturity criterion, scene preparation, and sample size
Effective cycle time and retry costTotal operating time; successful fruits; retries; recovery time; travel/unloading/waiting included or excluded explicitlyReport time per successfully collected fruit including observation, approach, retries, and recovery
Product qualityHarvested fruits assessed; damaged/downgraded fruits; crop-appropriate quality endpointsReport damage/downgrade with maturity, contact method, calyx/stem requirements, and postharvest assessment interval
Failure detection and recoveryFailures by stage; retries; recovered failures; unrecovered stops; recovery timeUse an explicit failure taxonomy, retry limit, and exit/safety conditions; report recovery success and cost
Human intervention/autonomyInterventions and takeover events; intervention duration; manual tasks remainingReport interventions per hour or eligible target and define the autonomous operating scope, including reset, fruit handling, container exchange, recovery, and supervision
Economic viabilityCapital and retrofit costs; maintenance, consumables, energy, and operator costs; net labor hours saved or reallocated; seasonal utilization; payback period or return on investmentReport costs per kilogram of marketable output or per operating season and state assumptions for utilization, service life, labor cost, discount rate, and product value
Continuous operationOperating duration; uptime/downtime by cause; marketable output; maintenance and energy useReport effective fruit/h or kg/h and operational availability over consecutive harvests, shifts, or days under documented greenhouse conditions
Note: Report all metrics with explicit denominators and validation duration. Numerical values from different crops, facilities, protocols, or robot platforms should not be interpreted as direct rankings unless test conditions are equivalent.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, J.; Xi, C.; Chen, Y.; Su, L.; Tang, Z. System-Level Smart Robotic Harvesting for High-Value Greenhouse Crops: A Review. Agronomy 2026, 16, 1795. https://doi.org/10.3390/agronomy16181795

AMA Style

Wang J, Xi C, Chen Y, Su L, Tang Z. System-Level Smart Robotic Harvesting for High-Value Greenhouse Crops: A Review. Agronomy. 2026; 16(18):1795. https://doi.org/10.3390/agronomy16181795

Chicago/Turabian Style

Wang, Junyi, Chenyu Xi, Yiming Chen, Ling Su, and Zhong Tang. 2026. "System-Level Smart Robotic Harvesting for High-Value Greenhouse Crops: A Review" Agronomy 16, no. 18: 1795. https://doi.org/10.3390/agronomy16181795

APA Style

Wang, J., Xi, C., Chen, Y., Su, L., & Tang, Z. (2026). System-Level Smart Robotic Harvesting for High-Value Greenhouse Crops: A Review. Agronomy, 16(18), 1795. https://doi.org/10.3390/agronomy16181795

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop