GPC-Frame: Bridging Scene Perception and Harvesting Cognition for Active Disocclusion and Harvesting-Point Localization in Trellised Table-Grape Vineyards
Abstract
1. Introduction
1.1. Background
1.2. Related Work
1.3. Research Hypothesis and Objectives
1.4. Main Contributions
2. Materials and Methods
2.1. Experimental Sites and Vineyard Characteristics
2.2. Robotic Platform and Image Acquisition
2.3. Dataset Construction and Annotation
2.4. Overview of GPC-Frame
2.5. Perception Layer: Multi-Target Scene Segmentation
2.5.1. Backbone for Irregular Morphological Feature Extraction
2.5.2. Dynamic Feature Fusion for Significant Scale Disparities
2.5.3. Transformer-Based Decoder for Multi-Target Segmentation
2.6. Cognition Layer: Reasoning and Localization
2.6.1. Grape Cluster Structure and Occlusion Reasoning
2.6.2. D/H-Point Decision-Making and Localization
2.7. Framework Loss Function and Two-Stage Optimization Strategy
2.7.1. Perception Layer Optimization
2.7.2. Cognition Layer Optimization
2.8. Training Parameters and Evaluation Metrics
2.8.1. Implementation and Training Parameters
2.8.2. Evaluation Metrics
2.9. Implementation of Comparative Methods
3. Results
3.1. Perception–Cognition Performance Evaluation
3.1.1. Multi-Target Scene Segmentation Performance
3.1.2. Grape Cluster Structure and Occlusion Reasoning with D/H-Point Localization
3.1.3. Performance Under Different Lighting Conditions
3.2. Comparison with Representative Grape-Harvesting Localization Methods
3.3. Field Robotic Validation
3.3.1. Field Perception–Cognition Performance Across Time Windows
3.3.2. Autonomous Harvesting Operations
4. Discussion
4.1. From Point Localization to State-Aware Harvesting Decisions
4.2. Illumination Effects and Failure Propagation from Perception to Cognition
4.3. Robotic Execution Constraints After Perception–Cognition
4.4. Implications for Harvesting Cognition in Horticultural Fruit Crops
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Wang, C.; Luo, L. Application of smart technology and equipment in horticulture. Horticulturae 2024, 10, 676. [Google Scholar] [CrossRef] [Scilit]
- Shi, X.; Wang, S.; Zhang, B.; Zhang, Z.; Wang, S.; Ding, X.; Wang, S.; Qi, P.; Yang, H. Advances in berry harvesting robots. Horticulturae 2025, 11, 1042. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Pan, W.; Zou, T.; Li, C.; Han, Q.; Wang, H.; Yang, J.; Zou, X. A review of perception technologies for berry fruit-picking robots: Advantages, disadvantages, challenges, and prospects. Agriculture 2024, 14, 1346. [Google Scholar] [CrossRef] [Scilit]
- Lin, T.; Sun, F.; Li, X.; Guo, X.; Ying, J.; Wu, H.; Li, H. A review of key technologies and recent advances in intelligent fruit-picking robots. Horticulturae 2026, 12, 158. [Google Scholar] [CrossRef] [Scilit]
- Seol, J.; Park, Y.; Pak, J.; Jo, Y.; Lee, G.; Kim, Y.; Ju, C.; Hong, A.; Son, H.I. Human-centered robotic system for agricultural applications: Design, development, and field evaluation. Agriculture 2024, 14, 1985. [Google Scholar] [CrossRef] [Scilit]
- Fu, H.; Li, T.; Feng, Q.; Chen, L. Push-or-avoid: Deep reinforcement learning of obstacle-aware harvesting for orchard robots. Agriculture 2026, 16, 670. [Google Scholar] [CrossRef] [Scilit]
- Behroozi-Khazaei, N.; Maleki, M.R. A robust algorithm based on color features for grape cluster segmentation. Comput. Electron. Agric. 2017, 142, 41–49. [Google Scholar] [CrossRef] [Scilit]
- Coll-Ribes, G.; Torres-Rodríguez, I.J.; Grau, A.; Guerra, E.; Sanfeliu, A. Accurate detection and depth estimation of table grapes and peduncles for robot harvesting, combining monocular depth estimation and CNN methods. Comput. Electron. Agric. 2023, 215, 108362. [Google Scholar] [CrossRef] [Scilit]
- Weng, W.; Lai, Z.; Cui, Z.; Chen, Z.; Chen, H.; Lin, T.; Wang, J.; Zheng, S.; Chen, G. GCD-YOLO: A deep learning network for accurate tomato fruit stalks identification in unstructured environments. Smart Agric. Technol. 2025, 12, 101465. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Ma, A.; Huang, L.; Su, Y.; Li, W.; Zhang, H.; Wang, Z. GA-YOLO: A lightweight YOLO model for dense and occluded grape target detection. Horticulturae 2023, 9, 443. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Chen, H.; Xu, F.; Lin, M.; Zhang, D.; Zhang, L. Real-time detection of mature table grapes using ESP-YOLO network on embedded platforms. Biosyst. Eng. 2024, 246, 122–134. [Google Scholar] [CrossRef] [Scilit]
- Wang, Q.; Xu, F.; Chen, Q.; Mi, Z.; Fan, Y.; Su, B. Detection and location of wine grape (Cabernet Sauvignon) picking points by using a dual-stage deep learning method. Comput. Electron. Agric. 2025, 237, 110637. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Peng, D.; Qi, H.; Zhou, L.; Zhang, C. Detection and Instance Segmentation of Grape Clusters in Orchard Environments Using an Improved Mask R-CNN Model. Agriculture 2024, 14, 918. [Google Scholar] [CrossRef] [Scilit]
- Mehdipour, S.; Mirroshandel, S.A.; Tabatabaei, S.A. Vision transformers in precision agriculture: A comprehensive survey. arXiv 2025, arXiv:2504.21706. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Zhang, Z.; Luo, L.; Zhu, W.; Chen, J.; Wang, W. SwinGD: A robust grape bunch detection model based on Swin Transformer in complex vineyard environment. Horticulturae 2021, 7, 492. [Google Scholar] [CrossRef] [Scilit]
- Devanna, R.P.; Reina, G.; Auat Cheein, F.; Milella, A. Boosting grape bunch detection in RGB-D images using zero-shot annotation with Segment Anything and GroundingDINO. Comput. Electron. Agric. 2025, 229, 109611. [Google Scholar] [CrossRef] [Scilit]
- Peng, Y.; Sun, J.; Wu, Z.; Gao, J.; Shi, L.; Shi, Z. A vision-based information processing framework for vineyard grape picking using two-stage segmentation and morphological perception. Horticulturae 2025, 11, 1039. [Google Scholar] [CrossRef] [Scilit]
- Shen, L.; Su, J.; Huang, R.; Quan, W.; Song, Y.; Fang, Y.; Su, B. Fusing attention mechanism with Mask R-CNN for instance segmentation of grape cluster in the field. Front. Plant Sci. 2022, 13, 934450. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jiang, T.; Li, Y.; Feng, H.; Wu, J.; Sun, W.; Ruan, Y. Research on a trellis grape stem recognition method based on YOLOv8n-GP. Agriculture 2024, 14, 1449. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Yu, X.; Zhang, D.; Yang, Y.; Qiu, Y.; Pang, L.; Wang, H. TinySeg: A deep learning model for small target segmentation of grape pedicels with multi-attention and multi-scale feature fusion. Comput. Electron. Agric. 2025, 237, 110726. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Zhang, Z.; Luo, L.; Wei, H.; Wang, W.; Chen, M.; Luo, S. DualSeg: Fusing transformer and CNN structure for image segmentation in complex vineyard environment. Comput. Electron. Agric. 2023, 206, 107682. [Google Scholar] [CrossRef] [Scilit]
- Li, P.; Wen, M.; Zeng, Z.; Tian, Y. Cherry tomato bunch and picking point detection for robotic harvesting using an RGB-D sensor and a StarBL-YOLO network. Horticulturae 2025, 11, 949. [Google Scholar] [CrossRef] [Scilit]
- Zhou, X.; Zou, X.; Meng, H.; Wu, F.; Chen, S.; Luo, X. Multi-task perception and three-dimensional picking point localization method for grapes based on structural constraints and geometric analysis. Comput. Electron. Agric. 2025, 238, 110814. [Google Scholar] [CrossRef] [Scilit]
- Lin, X.; Wang, J.; Wang, J.; Wei, H.; Chen, M.; Luo, L. Picking point localization method based on semantic reasoning for complex picking scenarios in vineyards. Artif. Intell. Agric. 2025, 15, 744–756. [Google Scholar] [CrossRef] [Scilit]
- Hussain, M.; He, L.; Schupp, J.R.; Lyons, D. Green Fruit-Stem Pairing and Clustering for Machine Vision System in Robotic Thinning of Apples. J. Field Robot. 2025, 42, 1463–1490. [Google Scholar] [CrossRef] [Scilit]
- Du, W.; Jia, Z.; Sui, S.; Liu, P. Table grape inflorescence detection and clamping point localisation based on channel pruned YOLOV7-TP. Biosyst. Eng. 2023, 235, 100–115. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Y.; Sui, S.; Du, W.; Li, X.; Liu, P. Picking point localization method of table grape picking robot based on you only look once version 8 nano. Eng. Appl. Artif. Intell. 2025, 146, 110266. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Zhang, G.; Cao, H.; Hu, K.; Wang, Q.; Deng, Y.; Gao, J.; Tang, Y. Geometry-aware 3D point cloud learning for precise cutting-point detection in unstructured field environments. J. Field Robot. 2025, 42, 3063–3076. [Google Scholar] [CrossRef] [Scilit]
- Zhang, T.; Wu, F.; Wang, M.; Chen, Z.; Li, L.; Zou, X. Grape-bunch identification and location of picking points on occluded fruit axis based on YOLOv5-GAP. Horticulturae 2023, 9, 498. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Ma, A.; Huang, L.; Li, H.; Zhang, H.; Huang, Y.; Zhu, T. Efficient and lightweight grape and picking point synchronous detection model based on key point detection. Comput. Electron. Agric. 2024, 217, 108612. [Google Scholar] [CrossRef] [Scilit]
- Lu, J.; Cao, Z.; Wang, J.; Wang, Z.; Zhao, J.; Zhang, M. A picking point localization method for table grapes based on PGSS-YOLOv11s and morphological strategies. Agriculture 2025, 15, 1622. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Lin, X.; Luo, L.; Chen, M.; Wei, H.; Xu, L.; Luo, S. Cognition of grape cluster picking point based on visual knowledge distillation in complex vineyard environment. Comput. Electron. Agric. 2024, 225, 109216. [Google Scholar] [CrossRef] [Scilit]
- Yi, T.; Zhang, D.; Luo, L.; Wang, Y.; Liu, B. View planning for grape harvesting based on self-supervised deep reinforcement learning under occlusion. Comput. Electron. Agric. 2025, 239, 110913. [Google Scholar] [CrossRef] [Scilit]
- Luo, L.; Liu, B.; Chen, M.; Wang, J.; Wei, H.; Lu, Q.; Luo, S. DRL-enhanced 3D detection of occluded stems for robotic grape harvesting. Comput. Electron. Agric. 2025, 229, 109736. [Google Scholar] [CrossRef] [Scilit]
- Xiong, Y.; Ge, Y.; From, P.J. An obstacle separation method for robotic picking of fruits in clusters. Comput. Electron. Agric. 2020, 175, 105397. [Google Scholar] [CrossRef] [Scilit]
- Xiong, Y.; Ge, Y.; From, P.J. An improved obstacle separation method using deep learning for object detection and tracking in a hybrid visual control loop for fruit picking in clusters. Comput. Electron. Agric. 2021, 191, 106508. [Google Scholar] [CrossRef] [Scilit]
- He, Z.; Liu, Z.; Zhou, Z.; Karkee, M.; Zhang, Q. Improving picking efficiency under occlusion: Design, development, and field evaluation of an innovative robotic strawberry harvester. Comput. Electron. Agric. 2025, 237, 110684. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Dai, J.; Chen, Z.; Huang, Z.; Li, Z.; Zhu, X.; Hu, X.; Lu, T.; Lu, L.; Li, H.; et al. InternImage: Exploring Large-Scale Vision Foundation Models with Deformable Convolutions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 14408–14419. [Google Scholar] [CrossRef] [Scilit]
- Dai, X.; Chen, Y.; Xiao, B.; Chen, D.; Liu, M.; Yuan, L.; Zhang, L. Dynamic Head: Unifying Object Detection Heads with Attentions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 7373–7382. [Google Scholar] [CrossRef] [Scilit]
- Li, F.; Zhang, H.; Xu, H.; Liu, S.; Zhang, L.; Ni, L.M.; Shum, H.-Y. Mask DINO: Towards a Unified Transformer-Based Framework for Object Detection and Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 3041–3050. [Google Scholar] [CrossRef] [Scilit]
- Yu, R.; Li, Y.; Liang, H.; Chen, Z. GeoExplainer: Interpreting graph convolutional networks with geometric masking. Neurocomputing 2024, 605, 128393. [Google Scholar] [CrossRef] [Scilit]
- Dai, M.; Cheng, W.; Liu, J.J.; Yang, S.; Cai, W.; Sun, Y.; Yang, W. DeRIS: Decoupling perception and cognition for enhanced referring image segmentation through loopback synergy. arXiv 2025, arXiv:2507.01738. [Google Scholar] [CrossRef] [Scilit]
- Nordström, M.; Maki, A.; Hult, H. The impact label noise and choice of threshold has on cross-entropy and soft-Dice in image segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 10–17 June 2025; pp. 20820–20829. [Google Scholar]
- Xu, M. Understanding graph embedding methods and their applications. SIAM Rev. 2021, 63, 825–853. [Google Scholar] [CrossRef] [Scilit]
- Parr, B.; Legg, M.; Alam, F. Analysis of Depth Cameras for Proximal Sensing of Grapes. Sensors 2022, 22, 4179. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ben Hazem, Z. A fuzzy-TD3 hybrid reinforcement learning framework for robust trajectory tracking of the Mitsubishi RV-2AJ robotic arm. Sci. Rep. 2026, 16, 12269. [Google Scholar] [CrossRef] [Scilit] [PubMed]











| Category | Instance Number | Precision/% | Recall/% | F1-Score/% | AP/% | IoU/% |
|---|---|---|---|---|---|---|
| Grape clusters | 1080 | 96.72 | 95.58 | 96.15 | 96.24 | 92.58 |
| Cluster stems | 522 | 90.45 | 87.92 | 89.17 | 89.65 | 80.46 |
| Petioles | 2237 | 87.15 | 84.68 | 85.90 | 86.32 | 75.28 |
| Leaves | 4141 | 94.39 | 92.85 | 93.61 | 93.92 | 87.99 |
| Canes | 2150 | 91.05 | 89.05 | 90.04 | 90.51 | 81.88 |
| Mean | N/A | 91.95 | 90.02 | 90.97 | 91.33 | 83.64 |
| Grape Cluster Status | Grape Cluster Number | Structural Reasoning Accuracy/% | Occlusion Reasoning Accuracy/% | D/H-Point Localization Success/% | Reasoning and Localization Time/s |
|---|---|---|---|---|---|
| Non-occluded | 294 | 98.64 | 98.30 | 96.60 | 0.04 |
| Leaf/petiole-occluded | 381 | 96.59 | 96.06 | 95.01 | 0.05 |
| Overlap-occluded | 228 | 94.74 | 94.30 | 93.86 | 0.04 |
| Cane-occluded | 177 | 95.48 | 94.92 | N/A | 0.03 |
| Mean | N/A | 96.36 | 95.90 | 95.16 | 0.04 |
| Lighting Condition | Grape Cluster Number | mIoU/% | Structural Reasoning Accuracy/% | D-Point Localization Success/% | H-Point Localization Success/% | Perception Time/s | Cognition Time/s |
|---|---|---|---|---|---|---|---|
| Sunny | 368 | 81.88 | 95.38 | 93.70 | 94.32 | 0.21 | 0.04 |
| Cloudy | 364 | 84.35 | 96.98 | 95.33 | 95.86 | 0.20 | 0.05 |
| Night (Fixed LED) | 348 | 86.23 | 97.70 | 96.80 | 96.79 | 0.20 | 0.04 |
| Mean | N/A | 84.15 | 96.69 | 95.28 | 95.66 | 0.20 | 0.04 |
| Method | Core Technique | Occluded D-Point Localization | Non-Occluded H-Point Localization/% | Lead of GPC-Frame | Inference Time/s |
|---|---|---|---|---|---|
| Chen et al. [30] | YOLOv8-pose + One keypoint | N/A | 81.63 | +14.97 | 0.08 |
| Jiang et al. [19] | YOLOv8-pose + Three keypoints | N/A | 84.35 | +12.25 | 0.06 |
| Lu et al. [31] | YOLOv11-seg + Morphological label | N/A | 87.41 | +9.19 | 0.13 |
| Zhou et al. [23] | YOLACT + Box constraints | N/A | 90.48 | +6.12 | 0.16 |
| Lin et al. [24] | SegFormerB2 + Geometric rules | N/A | 92.52 | +4.08 | 0.37 |
| GPC-Frame | Perception layer + Cognition layer | Yes | 96.60 | N/A | 0.24 |
| Time Periods | Time Window | Lighting Condition | Grape Cluster Number | mIoU/% | Structural Reasoning Accuracy/% | D/H-Point Localization Success/% | Perception–Cognition Time/s |
|---|---|---|---|---|---|---|---|
| Early morning | 06:00–08:00 | Natural + Fixed LED | 41 | 81.25 | 87.80 | 85.37 | 0.24 |
| Morning | 08:00–11:00 | Natural (Side light) | 45 | 78.18 | 84.44 | 80.00 | 0.24 |
| Midday | 11:00–14:00 | Natural (Top light) | 42 | 75.45 | 78.57 | 73.81 | 0.25 |
| Afternoon | 14:00–18:00 | Natural (Side light) | 44 | 77.50 | 81.82 | 77.27 | 0.24 |
| Evening | 18:00–20:00 | Natural + Fixed LED | 43 | 80.53 | 88.37 | 86.05 | 0.24 |
| Night | 20:00–06:00 | Fixed LED | 46 | 82.48 | 93.48 | 91.30 | 0.25 |
| Mean | N/A | N/A | N/A | 79.23 | 85.75 | 82.30 | 0.24 |
| Scene Type | Mode | Grape Cluster Number | D-Point Localization Success/% | Harvesting Success/% | Mean Harvesting Time per Cluster/s |
|---|---|---|---|---|---|
| Non-occluded | Direct harvesting | 37 | N/A | 89.19 | 23 |
| Occluded | Side-view/disocclusion harvesting | 141 | 86.52 | 85.11 | 75 |
| Mean | N/A | N/A | N/A | 87.15 | 49 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ning, Z.; Li, J.; Du, P.; Wang, Y.; Zhu, L.; Zhuang, Y. GPC-Frame: Bridging Scene Perception and Harvesting Cognition for Active Disocclusion and Harvesting-Point Localization in Trellised Table-Grape Vineyards. Horticulturae 2026, 12, 869. https://doi.org/10.3390/horticulturae12070869
Ning Z, Li J, Du P, Wang Y, Zhu L, Zhuang Y. GPC-Frame: Bridging Scene Perception and Harvesting Cognition for Active Disocclusion and Harvesting-Point Localization in Trellised Table-Grape Vineyards. Horticulturae. 2026; 12(7):869. https://doi.org/10.3390/horticulturae12070869
Chicago/Turabian StyleNing, Zhengtong, Jian Li, Pengfei Du, Yangwei Wang, Liangkuan Zhu, and Yu Zhuang. 2026. "GPC-Frame: Bridging Scene Perception and Harvesting Cognition for Active Disocclusion and Harvesting-Point Localization in Trellised Table-Grape Vineyards" Horticulturae 12, no. 7: 869. https://doi.org/10.3390/horticulturae12070869
APA StyleNing, Z., Li, J., Du, P., Wang, Y., Zhu, L., & Zhuang, Y. (2026). GPC-Frame: Bridging Scene Perception and Harvesting Cognition for Active Disocclusion and Harvesting-Point Localization in Trellised Table-Grape Vineyards. Horticulturae, 12(7), 869. https://doi.org/10.3390/horticulturae12070869

