A Robust Visual Grasping Method for Robots in Cluttered and Stacked Scenes
Abstract
1. Introduction
2. Methods
2.1. Principles and Applications of SAM
2.2. Principles and Applications of FoundationPose
2.3. SAM-FoundationPose Closed-Loop Fusion Framework
2.3.1. Overall Framework Design
- (1)
- If (where denotes the preset high-precision convergence threshold), the currently estimated pose is regarded as a highly optimized and reliable solution within the search space. The system terminates the closed-loop optimization and directly outputs the final pose , which is used to drive the trajectory planning of the robot arm end-effector and perform actual physical visual grasping.
- (2)
- If , it indicates that the current pose hypothesis has fallen into a local minimum due to severe clutter or occlusion. The system immediately triggers the iterative optimization control stage. The imperfect 3D pose is fed into a forward projection operator to synthesize a refined geometric guided mask :where and denote the rotation matrix and translation vector of the current rigid-body pose Tk, respectively. represents the vertex coordinates of the 3D object model, denotes the 3D point cloud transformed into the camera coordinate system, and K is the camera intrinsic matrix. denotes the perspective dehomogenization projection function, while is the rasterization operator that converts the projected pixel set into a binary topological mask.
2.3.2. Dynamic Prompt-Based SAM Segmentation Module
2.3.3. FoundationPose Pose Estimation Module
2.3.4. Multi-Dimensional Confidence Assessment Module
2.3.5. Iterative Optimization Control
3. Experiments and Analysis
3.1. Platform Setup and Experimental Preparation
3.2. Object Pose Estimation in Cluttered but Non-Stacked Scenes
3.2.1. Robustness Analysis
3.2.2. Positioning Accuracy Analysis
3.3. Object Pose Estimation in Cluttered and Stacked Scenes
3.4. Comparative Experiments with Different Estimation Methods
3.4.1. Selection of Comparative Algorithms and Evaluation Metrics
3.4.2. Result Comparison and Analysis
3.4.3. Visual Comparison of Results Across Different Methods
3.5. Core Module Ablation Experiments
3.5.1. Ablation Variant Design
- (1)
- Variant A (FoundationPose): Only the basic FoundationPose network is used without introducing any additional modules. This variant serves as the performance baseline, reflecting the inherent performance upper bound of the basic method under complex stacking backgrounds and providing a unified reference standard for all subsequent comparisons.
- (2)
- Variant B (FoundationPose + SAM): Based on Variant A, the SAM segmentation module is introduced solely at the input end to provide an initial foreground mask, while maintaining an open-loop process without subsequent iterative optimization. This variant is used to verify the suppression effect of the SAM segmentation module on background interference.
- (3)
- Variant C (FoundationPose + SAM + Iterative closed-loop): Based on Variant B, a closed-loop iterative mechanism is introduced; however, only a single “mask matching degree (IoU)” metric is used for confidence assessment, without adopting the multi-dimensional evaluation system proposed in this paper. This variant is used to verify the optimization capability of the closed-loop iterative mechanism and to compare the limitations imposed by single-metric evaluation on the iterative process.
- (4)
- Variant D (FoundationPose + SAM + Iterative closed-loop + Multi-dimensional confidence assessment): The complete framework proposed in this paper, incorporating SAM segmentation, rendering-guided closed-loop iterative feedback, and multi-dimensional confidence assessment. This variant is used to verify the overall performance under the synergistic cooperation of all modules and to demonstrate the gain effect of multi-dimensional assessment on the closed-loop iteration.
3.5.2. Ablation Experiment Results and Analysis
3.5.3. Ablation Experiment Conclusion
3.6. Verification of Robotic Arm Visual Grasping in Real-World Scenarios
3.6.1. Grasping Experiment Procedure
3.6.2. Grasping Result Analysis
4. Discussion
4.1. Innovations of This Work
4.2. Limitations of This Work
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Lee, J.; Kim, H.; Kwon, J.W.; Yun, S.J.; Lee, N.H.; Choi, Y.H.; Chung, G.; Suh, J. Model-Free Transformer Framework for 6-DoF Pose Estimation of Textureless Tableware Objects. Sensors 2025, 25, 6167. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sampath, S.K.; Wang, N.; Yang, C.; Wu, H.; Liu, C.; Pearson, M. A Vision-Guided Deep Learning Framework for Dexterous Robotic Grasping Using Gaussian Processes and Transformers. Appl. Sci. 2025, 15, 2615. [Google Scholar] [CrossRef] [Scilit]
- Sun, H.; Zhang, Y.; Sun, H.; Hashimoto, K. Refined Prior Guided Category-Level 6D Pose Estimation and Its Application on Robotic Grasping. Appl. Sci. 2024, 14, 8009. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Wu, T.; Zou, Q. 6DoF Pose Estimation of Transparent Objects: Dataset and Method. Sensors 2026, 26, 898. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lou, Y.; Zhao, L.; Sui, N.; Gao, X.; Chen, Z.; Zhang, Y. 6D pose estimation method based on hybrid attention mechanism and vector-based local consistency enhancement. Eng. Res. Express 2026, 8, 095407. [Google Scholar] [CrossRef] [Scilit]
- Zheng, D.; Chen, Y. Enhancing Robotic Grasping Detection Using Visual–Tactile Fusion Perception. Sensors 2026, 26, 724. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Chen, Y.; Lai, H.; Zhang, H. Weakly supervised 3D human pose estimation based on PnP projection model. Pattern Recognit. 2025, 163, 111464. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Li, H.; Luo, C. Object Pose Estimation Based on Multi-precision Vectors and Seg-Driven PnP. Int. J. Comput. Vis. 2024, 133, 2620–2634. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Sun, W.; Yang, H.; Zeng, Z.; Liu, C.; Zheng, J.; Liu, X.; Rahmani, H.; Sebe, N.; Mian, A. Deep Learning-Based Object Pose Estimation: A Comprehensive Survey. Int. J. Comput. Vis. 2026, 134, 81. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Xu, D.; Zhu, Y.; Martín-Martín, R.; Lu, C.; Fei-Fei, L.; Savarese, S. DenseFusion: 6D Object Pose Estimation by Iterative Dense Fusion. arXiv 2019, arXiv:1901.04780. [Google Scholar]
- Sijin, L.; Yu, L.; Zhehao, L.; Guoyuan, L.; Can, W.; Xinyu, W. Vision-Guided Object Recognition and 6D Pose Estimation System Based on Deep Neural Network for Unmanned Aerial Vehicles towards Intelligent Logistics. Appl. Sci. 2022, 13, 115. [Google Scholar] [CrossRef] [Scilit]
- Cheng, X.; Wu, L.; Wang, Z.; Hou, J.; Wen, J.; Xu, Y. PVNet: Point-Voxel Interaction LiDAR Scene Upsampling Via Diffusion Models. IEEE Trans. Image Process. 2025, 34, 6895–6910. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wanquan, F.; Juyong, Z.; Yuanfeng, Z.; Shiqing, X. GDR-Net: A Geometric Detail Recovering Network for 3D Scanned Objects. IEEE Trans. Vis. Comput. Graph. 2021, 28, 3959–3973. [Google Scholar] [CrossRef] [Scilit]
- Nguyen, V.T.; Do, C.D.; Dang, T.V.; Bui, T.L.; Tan, P.X. A comprehensive RGB-D dataset for 6D pose estimation for industrial robots pick and place: Creation and real-world validation. Results Eng. 2024, 24, 103459. [Google Scholar] [CrossRef] [Scilit]
- Tian, Z.; Yang, B.; Xu, C. HAND-S3T: Real-time 3D hand pose estimation via hierarchical mesh refinement. Digit. Signal Process. 2026, 183, 106347. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Fang, M.; Wang, B.; Wang, X.; Yang, Y.; Wang, H. ResFuNet: A robust vision-based detection framework for robotic grasp pose estimation. Displays 2026, 94, 103514. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Wang, M.; Cao, J.; Wang, C.; Wu, Z.; Gao, H. A Novel Fish Pose Estimation Method Based on Semi-Supervised Temporal Context Network. Biomimetics 2025, 10, 566. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, W.; Di, N. RSCS6D: Keypoint Extraction-Based 6D Pose Estimation. Appl. Sci. 2025, 15, 6729. [Google Scholar] [CrossRef] [Scilit]
- Li, P.; Zhang, W. Reading recognition for pointer meters based on SAM and MLLM. Neural Comput. Appl. 2026, 38, 331. [Google Scholar] [CrossRef] [Scilit]
- Lang, W.; Xi, L.; Kai, Z.; Zhongwei, L.; Congjun, W.; Yusheng, S. HCCG: Efficient high compatibility correspondence grouping for 3D object recognition and 6D pose estimation in cluttered scenes. Measurement 2022, 197, 111296. [Google Scholar] [CrossRef] [Scilit]
- Rawat, U.; Rai, C.S. Towards geometry-aware attention: Key shift adjustment in vision transformers for image feature extraction. Signal Image Video Process. 2026, 20, 165. [Google Scholar] [CrossRef] [Scilit]
- Jrondi, Z.; Moussaid, A.; Hadi, M.Y. Exploring End-to-End object detection with transformers versus YOLOv8 for enhanced citrus fruit detection within trees. Syst. Soft Comput. 2024, 6, 200103. [Google Scholar] [CrossRef] [Scilit]
- Wen, B.; Yang, W.; Kautz, J.; Birchfield, S. Foundationpose: Unified 6d pose estimation and tracking of novel objects. Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. 2024, 2024, 17868–17879. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; He, L.; He, R.; Kadkhodamohammadi, A.; Stoyanov, D.; Davidson, B.R.; Mazomenos, E.B.; Clarkson, M.J. FoundationPose-Initialized 3D-2D Liver Registration for Surgical Augmented Reality. arXiv 2026, arXiv:2602.17517. [Google Scholar]
- Lee, P.K.; Jang, S.; Kim, C.J.; Kim, G.; Yun, H. 6D Pose Estimation of Reflective and Textureless Object with Improved Accuracy through Multi-view Scanning Using Mobile and Stationary Cameras. Int. J. Precis. Eng. Manuf. 2026; prepublish. [CrossRef] [Scilit]
- Li, Y.; Fang, Y.; Deng, H.; Xu, Y.; Yang, J. High-Fidelity Object Detection and 6D Pose Estimation for Vision-Guided 6-DoF Grasping of Chemical Vials. Signal Image Video Process. 2025, 19, 1447. [Google Scholar] [CrossRef] [Scilit]
- Wang, J.; Liu, G.; Ding, W.; Li, Y.; Song, W. From visual understanding to 6D pose reconstruction: A cutting-edge review of deep learning-based object pose estimation. Displays 2025, 89, 103069. [Google Scholar] [CrossRef] [Scilit]
- Hwang, H.J.; Cho, J.H.; Kim, Y.T. Deep Learning-Based Real-Time 6D Pose Estimation and Multi-Mode Tracking Algorithms for Citrus-Harvesting Robots. Machines 2024, 12, 642. [Google Scholar] [CrossRef] [Scilit]
- Govi, E.; Sapienza, D.; Toscani, S.; Cotti, I.; Franchini, G.; Bertogna, M. Addressing challenges in industrial pick and place: A deep learning-based 6 Degrees-of-Freedom pose estimation solution. Comput. Ind. 2024, 161, 104130. [Google Scholar] [CrossRef] [Scilit]
- Song, Z.; Tang, W.; Deng, W.; Wang, H.; Huang, G.; Wu, H.; Guo, Y.; Liu, J.; Jin, K.; Ma, Z. An FPGA-Based YOLOv5n Accelerator for Online Multi-Track Particle Localization. Electronics 2026, 15, 810. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Tang, F.; Huang, F.; Li, S.; Xu, X.; Xu, Y.; Zhu, L.; Dong, W. Boosting cross-domain semi-supervised medical image segmentation with internal and external regularizations. Pattern Recognit. 2026, 179, 113515. [Google Scholar] [CrossRef] [Scilit]
- Feng, S.; Pan, X.; Zhang, W.; Pan, M.; Han, C.; Lan, R. QuPaS: SAM-based Semi-supervised Histopathological Image Segmentation with Quantum Force Field Finetuning and Adversarial Estimation. IEEE Trans. Med. Imaging 2026, 45, 3150–3162. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, S.; Gong, P.; Zhang, H.; Li, J.; Bi, S.; Li, A.; Luo, Q.; Feng, Z.; Xiao, C. Brain-SAM: A general automatic SAM-based segmentation model for brain science images. Biomed. Opt. Express 2026, 17, 614–632. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Guoyuan, L.; Fan, C.; Yu, L.; Yachun, F.; Can, W.; Xinyu, W. A Manufacturing-Oriented Intelligent Vision System Based on Deep Neural Network for Object Recognition and 6D Pose Estimation. Front. Neurorobot. 2021, 14, 616775. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nasim, H.; Gabriel, L.B.; Harsh, S.; Irene, C. Marker-Less 3d Object Recognition and 6d Pose Estimation for Homogeneous Textureless Objects: An RGB-D Approach. Sensors 2020, 20, 5098. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ren, J.; Li, L.; Li, S.; Liu, M.; Fang, M.; Zhang, S.; Liu, W.; Liu, Y.; Yu, H. Confidence relative off-targets distance-based multi-dimensional transparency evaluation of distribution station area. Front. Energy Res. 2024, 11, 1283775. [Google Scholar] [CrossRef] [Scilit]
- Turco, E.; Bo, V.; Castellani, C.; Salvietti, G.; Malvezzi, M.; Prattichizzo, D.; Pozzi, M. Leveraging Embodied Mechanical Intelligence for Learning Decluttering Tasks: Gripper Design Boosts Learning. IEEE Robot. Autom. Mag. 2026, 33, 39–51. [Google Scholar] [CrossRef] [Scilit]











| Input: RGB-D image I, 3D model Mcad, camera intrinsics K |
| Output: Optimal 6D pose T* |
| Parameters: τ ← 85 (confidence threshold), kmax ← 10 (maximum iterations) |
| 1: k ← 0 |
| 2://Initialization phase |
| 3: bbox ← YOLOv5n(I) ▷ Heuristic bounding box prompt |
| 4: M0 ← SAM(I, bbox) ▷ Initial foreground mask |
| 5: T0 ← FoundationPose(I, M0, M3d) ▷ Initial 6D pose estimation |
| 6: S0 ← ComputeScore(T0, M0, I, M3d, K) ▷ Multi-dimensional confidence score |
| 7://Iterative refinement phase |
| 8: while Sk < τ and k < kmax do |
| 9: k ← k + 1 |
| 10: Mguide ← RenderMask(Tk − 1, M3d, K) ▷ Render geometric prior mask |
| 11: Mk ← SAM(I, Mguide) ▷ Mask refinement with spatial prior |
| 12: Tk ← FoundationPose(I, Mk, M3d) ▷ Pose re-estimation with refined mask |
| 13: Sk ← ComputeScore(Tk, Mk, I, M3d, K) ▷ Update confidence score |
| 14: end while |
| 15: return T* ← Tk |
| Target Object (Category) | Object Type Characteristics | ADD(-S) Recall Rate (%) | Average Translation Error (mm) | Average Rotation Error (°) |
|---|---|---|---|---|
| Headphones | Heterogeneous shape/background texture | 98.2 | 3.2 0.6 | 1.9 0.4 |
| Tape measure | Same-color background confusion | 97.5 | 4.1 0.9 | 2.3 0.5 |
| Detergent bottle | Cylindrical symmetry/high-frequency reflections | 96.8 | 3.8 0.7 | 2.1 0.4 |
| Adhesive tape | Low-contrast edges | 94.1 | 5.2 1.2 | 3.4 0.8 |
| Corn bottle | Complex internal texture | 98.0 | 2.9 0.5 | 1.7 0.3 |
| Ping-pong ball | Textureless solid-color sphere | 99.2 | 1.8 0.3 | 1.2 0.2 |
| Computer mouse | Weak texture/dark color | 98.5 | 2.6 0.4 | 1.8 0.4 |
| Small tire | Strong geometric jagged texture | 97.4 | 3.5 0.6 | 2.0 0.4 |
| Pliers | Slender irregular structure | 96.1 | 4.4 1.0 | 2.6 0.6 |
| Earphone case | Multi-source lighting/specular reflections | 95.3 | 4.8 1.1 | 3.1 0.7 |
| Mean | - | 97.1 | 3.63 0.73 | 2.21 0.47 |
| Method | Input Data Type | ADD-S Recall Rate (%) | Average Translation Error(mm) | Average Rotation Error (°) | Average Inference Time(ms) |
|---|---|---|---|---|---|
| PoseCNN | RGB | 45.2 | 35.6 8.4 | 22.4 5.1 | 45 |
| MegaPose | RGB-D | 71.4 | 19.5 4.2 | 11.2 2.8 | 385 |
| FoundationPose | RGB-D | 79.6 | 14.2 3.1 | 8.7 1.9 | 215 |
| Ours (SAM-FoundationPose) | RGB-D | 91.7 | 3.5 0.7 | 2.1 0.4 | 390 |
| Experimental Variant | SAM Segmentation | Closed-Loop Iteration | Multi-Dimensional Assessment | ADD-S (%) | Average Translation Error (mm) |
|---|---|---|---|---|---|
| A | × | × | × | 72.4 | 16.8 |
| B | √ | × | × | 81.5 | 11.2 |
| C | √ | √ | × | 86.3 | 6.5 |
| D (Ours) | √ | √ | √ | 91.7 | 3.5 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gao, Z.; Li, M.; Bai, H.; Li, J.; Li, S.; Han, J.; Wang, Z. A Robust Visual Grasping Method for Robots in Cluttered and Stacked Scenes. Sensors 2026, 26, 4524. https://doi.org/10.3390/s26144524
Gao Z, Li M, Bai H, Li J, Li S, Han J, Wang Z. A Robust Visual Grasping Method for Robots in Cluttered and Stacked Scenes. Sensors. 2026; 26(14):4524. https://doi.org/10.3390/s26144524
Chicago/Turabian StyleGao, Zhiqiang, Mengqi Li, Huihui Bai, Jinze Li, Sifan Li, Jing Han, and Zhengkai Wang. 2026. "A Robust Visual Grasping Method for Robots in Cluttered and Stacked Scenes" Sensors 26, no. 14: 4524. https://doi.org/10.3390/s26144524
APA StyleGao, Z., Li, M., Bai, H., Li, J., Li, S., Han, J., & Wang, Z. (2026). A Robust Visual Grasping Method for Robots in Cluttered and Stacked Scenes. Sensors, 26(14), 4524. https://doi.org/10.3390/s26144524

