Preprints
https://doi.org/10.5194/ms-2026-168
https://doi.org/10.5194/ms-2026-168
22 Sep 2026
 | 22 Sep 2026
Status: this preprint is currently under review for the journal MS.

A Lightweight BEV Perception Optimization Framework for Real-Time 3D Object Detection in Occluded Scenes

Yi Zhang, Anhua Zhou, and Jun Li

Abstract. Real-time 3D object detection in occluded scenes is a core challenge for pure-vision autonomous driving perception. Existing dense Bird's Eye View (BEV) detection methods suffer from redundant backbone parameters, inflexible fixed sampling strategies, and coarse-grained temporal modeling, making it difficult to simultaneously satisfy detection accuracy and real-time latency requirements on embedded vehicle platforms. This paper proposes a lightweight BEV perception optimization framework for occluded scenes. An occlusion-aware adaptive voxel feature sampling module dynamically switches between fast ray projection and deformable attention paths according to scene complexity. A temporal grouping fusion module based on Res2Net principles performs grouped cross-frame fusion of consecutive BEV feature maps without introducing additional learnable parameters. A two-stage LiDAR-to-camera knowledge distillation scheme with a geometric compensation module transfers depth geometry knowledge during training while incurring zero inference overhead. Experiments on the nuScenes dataset demonstrate that the proposed method achieves 38.7% mAP and 51.3% NDS under the lightweight configuration, with an inference latency of 38.2 ms and 26.2 FPS on NVIDIA Tesla T4, ranking highest in NDS among comparable methods.

Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. While Copernicus Publications makes every effort to include appropriate place names, the final responsibility lies with the authors. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.
Share
Yi Zhang, Anhua Zhou, and Jun Li

Status: open (until 29 Oct 2026)

Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor | : Report abuse
Yi Zhang, Anhua Zhou, and Jun Li
Yi Zhang, Anhua Zhou, and Jun Li
Metrics will be available soon.
Latest update: 22 Sep 2026
Download
Short summary
Autonomous vehicles must detect nearby road users quickly and reliably, even when objects are partly hidden. We developed a lightweight camera-based method that adapts its processing to scene difficulty, combines information across video frames, and learns from laser-sensor data during training. It reached 38.7% on the benchmark detection measure and processed 26.2 frames per second without laser sensors during use, showing strong potential for fast, lower-cost vehicle perception.
Share