paper-with-me

Papers

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

2026-07-14 · Haifa Zhang, Yijing Wang, Peixi Peng, Zhiqiang Zuo arxiv

In autonomous driving perception, the fusion of LiDAR and camera modalities has become the dominant paradigm for 3D object detection. However, current multi-modal frameworks heavily rely on massive visual backbones pretrained on 2D semantic tasks. This reliance introduces substantial parameter redundancy and a structural misalignment, as 2D priors are ill-equipped to handle the extreme sparsity of LiDAR projections required for Bird's-Eye-View geometry. To address this, we present DeGuNet, an ultra-compact and plug-and-play image backbone explicitly designed for depth-guided representation learning. By incorporating sparsity-aware feature extraction mechanisms, DeGuNet effectively aligns multi-view images with unstructured LiDAR depth while strictly preventing invalid-region contamination. Extensive experiments on the nuScenes dataset demonstrate DeGuNet's broad plug-and-play applicability and superior efficiency. When integrated into established baselines, it fundamentally eliminates architectural redundancy, reducing GPU memory consumption by up to 66.5% and achieving a 1.16x inference speedup. Concurrently, DeGuNet delivers up to a 6.20 absolute mAP gain, establishing a new paradigm for parameter-efficient multi-modal 3D perception.

📄 PDF Abstract BibTeX arXiv:2607.12419

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning3D Object DetectionAutonomous Driving

Similar Papers 제목 키워드 기반

TransTac: Visuo-Tactile Modality Transition via Ultraviolet-Encoded Transparent Elastomers

2026-06-03 · Lingyue Yang, Bin Fang arxiv

Vision-based tactile sensors (VBTS) recover high-resolution contact geometry but typically rely on opaque elastomer layers that prevent visual transparency, while RGB-D cameras provide global depth perception yet degrade…

CodecSplat: Ultra-Compact Latent Coding for Feed-Forward 3D Gaussian Splatting

2026-05-25 · Pengpeng Yu, Runqing Jiang, Qi Zhang, Dingquan Li 외 arxiv

While feed-forward 3D Gaussian splatting reconstructs renderable Gaussian primitives from sparse context views without per-scene optimization, existing pipelines do not provide a compact scene representation for storage …

DepthArb: Training-Free Depth-Arbitrated Generation for Occlusion-Robust Image Synthesis

2026-03-25 · Hongjin Niu, Jiahao Wang, Xirui Hu, Weizhan Zhang 외 arxiv

Text-to-image diffusion models frequently exhibit deficiencies in synthesizing accurate occlusion relationships of multiple objects, particularly within dense overlapping regions. Existing training-free layout-guided met…

Beyond the Limitation of Monocular 3D Detector via Knowledge Distillation

2023-01-01 · ICCV 2023 1 · Yiran Yang, Dongshuo Yin, Xuee Rong, Xian Sun 외

Knowledge distillation (KD) is a promising approach that facilitates the compact student model to learn dark knowledge from the huge teacher model for better results. Although KD methods are well explored in the 2D d…

Knowledge Distillation

Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework

2026-07-02 · Bo Zhao, Yapeng Li, Juhua Liu, Bo Du arxiv

Ultrasound image classification is essential for computer-aided diagnosis. However, current methods often neglect clinical priors, leading to poor generalization in challenging scenarios and a lack of interpretability th…

Image Classification