paper-with-me

홈 › Papers

BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision

2022-11-18 · CVPR 2023 1 · Chenyu Yang, Yuntao Chen, Hao Tian, Chenxin Tao, Xizhou Zhu, Zhaoxiang Zhang, Gao Huang, Hongyang Li, Yu Qiao, Lewei Lu, Jie zhou, Jifeng Dai

We present a novel bird's-eye-view (BEV) detector with perspective supervision, which converges faster and better suits modern image backbones. Existing state-of-the-art BEV detectors are often tied to certain depth pre-trained backbones like VoVNet, hindering the synergy between booming image backbones and BEV detectors. To address this limitation, we prioritize easing the optimization of BEV detectors by introducing perspective space supervision. To this end, we propose a two-stage BEV detector, where proposals from the perspective head are fed into the bird's-eye-view head for final predictions. To evaluate the effectiveness of our model, we conduct extensive ablation studies focusing on the form of supervision and the generality of the proposed detector. The proposed method is verified with a wide spectrum of traditional and modern image backbones and achieves new SoTA results on the large-scale nuScenes dataset. The code shall be released soon.

📄 PDF Abstract BibTeX arXiv:2211.10439

Code (2)

fundamentalvision/BEVFormer pytorch
opengvlab/internimage pytorch

Tasks

3D Object Detection

Methods 이 논문이 사용한 방법론

Batch Normalization 설명 없음
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
One-Shot Aggregation One-Shot Aggregation is an image model block that is an alternative to Dense Blocks, by aggregating intermediate features. It…
VoVNet 설명 없음

Similar Papers 제목 키워드 기반

CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth Flow

2024-03-13 · CVPR 2024 1 · Chenbin Pan, Burhaneddin Yaman, Senem Velipasalar, Liu Ren

Autonomous driving stands as a pivotal domain in computer vision, shaping the future of transportation. Within this paradigm, the backbone of the system plays a crucial role in interpreting the complex environment. Howev…

3D Object DetectionAutonomous DrivingContrastive Learningobject-detection+1

BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

2022-03-31 · Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie 외

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns …

3D Object DetectionAutonomous DrivingBird's-Eye View Semantic SegmentationRobust Camera Only 3D Object Detection

VoxelFormer: Bird's-Eye-View Feature Generation based on Dual-view Attention for Multi-view 3D Object Detection

2023-04-03 · Zhuoling Li, Chuanrui Zhang, Wei-Chiu Ma, Yipin Zhou 외

In recent years, transformer-based detectors have demonstrated remarkable performance in 2D visual perception tasks. However, their performance in multi-view 3D object detection remains inferior to the state-of-the-art (…

3D Object Detectionobject-detectionObject Detection

Geometric-aware Pretraining for Vision-centric 3D Object Detection

2023-04-06 · Linyan Huang, Huijie Wang, Jia Zeng, Shengchuan Zhang 외

Multi-camera 3D object detection for autonomous driving is a challenging problem that has garnered notable attention from both academia and industry. An obstacle encountered in vision-based techniques involves the precis…

3D Object DetectionAutonomous DrivingObjectobject-detection+1

BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection

2022-11-17 · Zehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang 외

3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. Owing to its low cost and high efficiency, multi-view 3D object detection has demonstrated promising app…

3D Object DetectionDepth EstimationDepth PredictionKnowledge Distillation+3