<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" specific-use="SMUR" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">MSD</journal-id>
<journal-title-group>
<journal-title>Mechanical Sciences Discussions</journal-title>
<abbrev-journal-title abbrev-type="publisher">MSD</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Mech. Sci. Discuss.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">-</issn>
<publisher><publisher-name></publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/ms-2026-165</article-id>
<title-group>
<article-title>A visual grasping method for collaborative robots in unstructured scenes based on object detection and grasp pose estimation</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Liu</surname>
<given-names>Junxiao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Qian</surname>
<given-names>Jun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Tan</surname>
<given-names>Yunkai</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhou</surname>
<given-names>Rong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>School of Mechanical Engineering, Hefei University of Technology, Hefei 230009, China</addr-line>
</aff>
<pub-date pub-type="epub">
<day>07</day>
<month>09</month>
<year>2026</year>
</pub-date>
<volume>2026</volume>
<fpage>1</fpage>
<lpage>20</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Junxiao Liu et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://ms.copernicus.org/preprints/ms-2026-165/">This article is available from https://ms.copernicus.org/preprints/ms-2026-165/</self-uri>
<self-uri xlink:href="https://ms.copernicus.org/preprints/ms-2026-165/ms-2026-165.pdf">The full text article is available as a PDF file from https://ms.copernicus.org/preprints/ms-2026-165/ms-2026-165.pdf</self-uri>
<abstract>
<p>In unstructured multi-object scenes, unclear target regions and interference from backgrounds and adjacent objects can affect grasp point selection. To address these problems, this paper proposes a visual grasping method for collaborative robots based on improved object detection and grasp pose estimation networks. In the object detection stage, the corresponding convolution and upsampling modules in YOLOv8n are replaced with RFAConv and DySample, respectively, to improve the detection of target boundaries and local features. In the grasp pose estimation stage, a residual attention structure and a region-guidance mechanism are incorporated into GR-ConvNet to enhance the stability of grasp prediction in complex backgrounds. To further reduce interference from non-target regions, the detected bounding box is used as a bounding box prompt for MobileSAM to generate a target mask. The target mask is then used to constrain the grasp quality map output by the improved GR-ConvNet and restrict grasp point search to the target region. The proposed method is validated on a collaborative robot visual grasping experimental platform using RGB-D images as input. Experimental results show a target detection success rate of 98 % and an overall grasp success rate of 95 %. The proposed method reduces interference from non-target regions during grasp point selection and improves the stability of target grasping in unstructured multi-object scenes.</p>
</abstract>
<counts><page-count count="20"/></counts>
<funding-group>
<award-group id="gs1">
<funding-source>Natural Science Foundation of Anhui Province</funding-source>
<award-id>2208085ME127</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body/>
<back>
</back>
</article>