Preprints
https://doi.org/10.5194/ms-2026-165
https://doi.org/10.5194/ms-2026-165
07 Sep 2026
 | 07 Sep 2026
Status: this preprint is currently under review for the journal MS.

A visual grasping method for collaborative robots in unstructured scenes based on object detection and grasp pose estimation

Junxiao Liu, Jun Qian, Yunkai Tan, and Rong Zhou

Abstract. In unstructured multi-object scenes, unclear target regions and interference from backgrounds and adjacent objects can affect grasp point selection. To address these problems, this paper proposes a visual grasping method for collaborative robots based on improved object detection and grasp pose estimation networks. In the object detection stage, the corresponding convolution and upsampling modules in YOLOv8n are replaced with RFAConv and DySample, respectively, to improve the detection of target boundaries and local features. In the grasp pose estimation stage, a residual attention structure and a region-guidance mechanism are incorporated into GR-ConvNet to enhance the stability of grasp prediction in complex backgrounds. To further reduce interference from non-target regions, the detected bounding box is used as a bounding box prompt for MobileSAM to generate a target mask. The target mask is then used to constrain the grasp quality map output by the improved GR-ConvNet and restrict grasp point search to the target region. The proposed method is validated on a collaborative robot visual grasping experimental platform using RGB-D images as input. Experimental results show a target detection success rate of 98 % and an overall grasp success rate of 95 %. The proposed method reduces interference from non-target regions during grasp point selection and improves the stability of target grasping in unstructured multi-object scenes.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Junxiao Liu, Jun Qian, Yunkai Tan, and Rong Zhou

Status: open (until 14 Oct 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Junxiao Liu, Jun Qian, Yunkai Tan, and Rong Zhou
Junxiao Liu, Jun Qian, Yunkai Tan, and Rong Zhou
Metrics will be available soon.
Latest update: 07 Sep 2026
Download
Short summary
Robotic grasping in complex environments is challenging due to variations in object positions, orientations, and backgrounds. This study proposes a vision-based method for collaborative robots that integrates object detection, target mask generation, and grasp position estimation to improve grasping reliability. Experiments on a robot platform demonstrate effective target-specific grasping in unstructured multi-object scenes, showing its potential for intelligent manufacturing applications.
Share