<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpublishing3.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article" specific-use="SMUR" dtd-version="3.0" xml:lang="en">
<front>
<journal-meta>
<journal-id journal-id-type="publisher">MSD</journal-id>
<journal-title-group>
<journal-title>Mechanical Sciences Discussions</journal-title>
<abbrev-journal-title abbrev-type="publisher">MSD</abbrev-journal-title>
<abbrev-journal-title abbrev-type="nlm-ta">Mech. Sci. Discuss.</abbrev-journal-title>
</journal-title-group>
<issn pub-type="epub">-</issn>
<publisher><publisher-name></publisher-name>
<publisher-loc>Göttingen, Germany</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.5194/ms-2026-168</article-id>
<title-group>
<article-title>A Lightweight BEV Perception Optimization Framework for Real-Time 3D Object Detection in Occluded Scenes</article-title>
</title-group>
<contrib-group><contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhang</surname>
<given-names>Yi</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Zhou</surname>
<given-names>Anhua</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" xlink:type="simple"><name name-style="western"><surname>Li</surname>
<given-names>Jun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
</contrib-group><aff id="aff1">
<label>1</label>
<addr-line>Chongqing Vocational and Technical University of Mechatronics, Chongqing, 402760, China</addr-line>
</aff>
<pub-date pub-type="epub">
<day>22</day>
<month>09</month>
<year>2026</year>
</pub-date>
<volume>2026</volume>
<fpage>1</fpage>
<lpage>21</lpage>
<permissions>
<copyright-statement>Copyright: &#x000a9; 2026 Yi Zhang et al.</copyright-statement>
<copyright-year>2026</copyright-year>
<license license-type="open-access">
<license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri"  xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p>
</license>
</permissions>
<self-uri xlink:href="https://ms.copernicus.org/preprints/ms-2026-168/">This article is available from https://ms.copernicus.org/preprints/ms-2026-168/</self-uri>
<self-uri xlink:href="https://ms.copernicus.org/preprints/ms-2026-168/ms-2026-168.pdf">The full text article is available as a PDF file from https://ms.copernicus.org/preprints/ms-2026-168/ms-2026-168.pdf</self-uri>
<abstract>
<p>Real-time 3D object detection in occluded scenes is a core challenge for pure-vision autonomous driving perception. Existing dense Bird&apos;s Eye View (BEV) detection methods suffer from redundant backbone parameters, inflexible fixed sampling strategies, and coarse-grained temporal modeling, making it difficult to simultaneously satisfy detection accuracy and real-time latency requirements on embedded vehicle platforms. This paper proposes a lightweight BEV perception optimization framework for occluded scenes. An occlusion-aware adaptive voxel feature sampling module dynamically switches between fast ray projection and deformable attention paths according to scene complexity. A temporal grouping fusion module based on Res2Net principles performs grouped cross-frame fusion of consecutive BEV feature maps without introducing additional learnable parameters. A two-stage LiDAR-to-camera knowledge distillation scheme with a geometric compensation module transfers depth geometry knowledge during training while incurring zero inference overhead. Experiments on the nuScenes dataset demonstrate that the proposed method achieves 38.7% mAP and 51.3% NDS under the lightweight configuration, with an inference latency of 38.2 ms and 26.2 FPS on NVIDIA Tesla T4, ranking highest in NDS among comparable methods.</p>
</abstract>
<counts><page-count count="21"/></counts>
</article-meta>
</front>
<body/>
<back>
</back>
</article>