DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection
Abstract
In autonomous driving perception, the fusion of LiDAR andcamera modalities has become the dominant paradigm for 3D object de-tection. However, current multi-modal frameworks heavily rely on mas-sive visual backbones pretrained on 2D semantic tasks. This relianceintroduces substantial parameter redundancy and a structural misalign-ment, as 2D priors are ill-equipped to handle the extreme sparsity of Li-DAR projections required for Bird’s-Eye-View geometry. To address this,we present DeGuNet, an ultra-compact and plug-and-play image back-bone explicitly designed for depth-guided representation learning. By in-corporating sparsity-aware feature extraction mechanisms, DeGuNet ef-fectively aligns multi-view images with unstructured LiDAR depth whilestrictly preventing invalid-region contamination. Extensive experimentson the nuScenes dataset demonstrate DeGuNet’s broad plug-and-playapplicability and superior efficiency. When integrated into establishedbaselines, it fundamentally eliminates architectural redundancy, reduc-ing GPU memory consumption by up to 66.5% and achieving a 1.16×inference speedup. Concurrently, DeGuNet delivers up to a 6.20 absolutemAP gain, establishing a new paradigm for parameter-efficient multi-modal 3D perception.