paper-with-me

홈 › Papers

Geometric-aware Pretraining for Vision-centric 3D Object Detection

2023-04-06 · Linyan Huang, Huijie Wang, Jia Zeng, Shengchuan Zhang, Liujuan Cao, Junchi Yan, Hongyang Li

Multi-camera 3D object detection for autonomous driving is a challenging problem that has garnered notable attention from both academia and industry. An obstacle encountered in vision-based techniques involves the precise extraction of geometry-conscious features from RGB images. Recent approaches have utilized geometric-aware image backbones pretrained on depth-relevant tasks to acquire spatial information. However, these approaches overlook the critical aspect of view transformation, resulting in inadequate performance due to the misalignment of spatial knowledge between the image backbone and view transformation. To address this issue, we propose a novel geometric-aware pretraining framework called GAPretrain. Our approach incorporates spatial and structural cues to camera networks by employing the geometric-rich modality as guidance during the pretraining phase. The transference of modal-specific attributes across different modalities is non-trivial, but we bridge this gap by using a unified bird's-eye-view (BEV) representation and structural hints derived from LiDAR point clouds to facilitate the pretraining process. GAPretrain serves as a plug-and-play solution that can be flexibly applied to multiple state-of-the-art detectors. Our experiments demonstrate the effectiveness and generalization ability of the proposed method. We achieve 46.2 mAP and 55.5 NDS on the nuScenes val set using the BEVFormer method, with a gain of 2.7 and 2.1 points, respectively. We also conduct experiments on various image backbones and view transformations to validate the efficacy of our approach. Code will be released at https://github.com/OpenDriveLab/BEVPerception-Survey-Recipe.

📄 PDF Abstract BibTeX arXiv:2304.03105

Code (1)

opendrivelab/bevperception-survey-recipe 공식 구현 pytorch

Tasks

3D Object DetectionAutonomous DrivingObjectobject-detectionObject Detection

Similar Papers 제목 키워드 기반

ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

2026-06-15 · Hao Li, Ganlong Zhao, Yufei Liu, Haotian Hou 외 arxiv

Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos…

GLaD: Geometric Latent Distillation for Vision-Language-Action Models

2025-12-10 · Minghao Guo, Meng Cao, Jiachen Tao, Rongtao Xu 외 arxiv

Most existing Vision-Language-Action (VLA) models rely primarily on RGB information, while ignoring geometric cues crucial for spatial reasoning and manipulation. In this work, we introduce GLaD, a geometry-aware VLA fra…

Knowledge DistillationSpatial Reasoning

EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining

2025-03-19 · Boshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng 외

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understandin…

Contrastive LearningDecoderDepth EstimationRepresentation Learning

Self-Supervised Learning from Non-Object Centric Images with a Geometric Transformation Sensitive Architecture

2023-04-17 · TaeHo Kim, Jong-Min Lee

Most invariance-based self-supervised methods rely on single object-centric images (e.g., ImageNet images) for pretraining, learning features that invariant to geometric transformation. However, when images are not objec…

image-classificationImage ClassificationInstance SegmentationSelf-Supervised Learning+1

UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models

2026-02-23 · Manish Kumar Govind, Dominick Reilly, Pu Wang, Srijan Das arxiv

Latent action representations learned from unlabeled videos have recently emerged as a promising paradigm for pretraining vision-language-action (VLA) models without explicit robot action supervision. However, latent act…