Previous Issue
Volume 15, September
 
 

Robotics, Volume 15, Issue 10 (October 2026) – 1 article

  • Issues are regarded as officially published after their release is announced to the table of contents alert mailing list.
  • You may sign up for e-mail alerts to receive table of contents of newly released issues.
  • PDF is the official format for papers published in both, html and pdf forms. To view the papers in pdf format, click on the "PDF Full-text" link, and use the free Adobe Reader to open them.
Order results
Result details
Section
Select all
Export citation of selected articles as:
26 pages, 7180 KB  
Article
Multi-Step Obstruction Reasoning for Target-Oriented Grasp Sequence Generation in Cluttered Scenes
by Huixuan Yang, Shiqiang Zhu, Yuhua Zheng and Wei Song
Robotics 2026, 15(10), 181; https://doi.org/10.3390/robotics15100181 - 22 Sep 2026
Abstract
Retrieving target objects in severe clutter requires multi-step reasoning to establish valid obstacle removal sequences. While recent Vision–Language Models (VLMs) have advanced instruction-driven clutter grasping, existing paradigms lack explicit construction of graph-constrained grasp sequences encompassing canonical trajectories and valid topological permutations to guide [...] Read more.
Retrieving target objects in severe clutter requires multi-step reasoning to establish valid obstacle removal sequences. While recent Vision–Language Models (VLMs) have advanced instruction-driven clutter grasping, existing paradigms lack explicit construction of graph-constrained grasp sequences encompassing canonical trajectories and valid topological permutations to guide model fine-tuning and evaluation. In addition, comprehensive evaluation requires accounting for the full space of topologically valid clearing sequences while systematically disentangling high-level topological planning errors from low-level physical execution failures. To fulfill these requirements, we introduce a novel DAG-based obstruction reasoning framework coupled with an integrated diagnostic evaluation protocol. Specifically, we model scene-level physical dependencies as Directed Acyclic Graphs (DAGs), explicitly converting graph constraints into topologically feasible sequence permutations to drive VLM fine-tuning. For diagnostic evaluation, we establish a three-part offline protocol comprising Strict Exact Match (EM), Graph-Feasible Accuracy (GFA), and Multi-Reference Normalized Sequence Edit Distance (MR-NSED), paired with online simulation testing. Extensive experiments demonstrate that our topological fine-tuning significantly improves multi-path reasoning performance, outperforming strong foundation model baselines including GPT-4o, GPT-4o-mini, and Llama-3.2-90B-Vision-Instruct, while our diagnostic protocol provides a faithful mechanism to systematically isolate reasoning logic from manipulation mechanics in complex physical clutter. Full article
(This article belongs to the Section AI in Robotics)
►▼ Show Figures

Figure 1

Previous Issue
Back to TopTop