paper-with-me

Papers

CLIP-BEVFormer: Enhancing Multi-View Image-Based BEV Detector with Ground Truth Flow

2024-03-13 · CVPR 2024 1 · Chenbin Pan, Burhaneddin Yaman, Senem Velipasalar, Liu Ren

Autonomous driving stands as a pivotal domain in computer vision, shaping the future of transportation. Within this paradigm, the backbone of the system plays a crucial role in interpreting the complex environment. However, a notable challenge has been the loss of clear supervision when it comes to Bird's Eye View elements. To address this limitation, we introduce CLIP-BEVFormer, a novel approach that leverages the power of contrastive learning techniques to enhance the multi-view image-derived BEV backbones with ground truth information flow. We conduct extensive experiments on the challenging nuScenes dataset and showcase significant and consistent improvements over the SOTA. Specifically, CLIP-BEVFormer achieves an impressive 8.5\% and 9.2\% enhancement in terms of NDS and mAP, respectively, over the previous best BEV model on the 3D object detection task.

📄 PDF Abstract BibTeX arXiv:2403.08919

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous DrivingContrastive Learningobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers

2022-03-31 · Zhiqi Li, Wenhai Wang, Hongyang Li, Enze Xie 외

3D visual perception tasks, including 3D detection and map segmentation based on multi-camera images, are essential for autonomous driving systems. In this work, we present a new framework termed BEVFormer, which learns …

3D Object DetectionAutonomous DrivingBird's-Eye View Semantic SegmentationRobust Camera Only 3D Object Detection

VoxelFormer: Bird's-Eye-View Feature Generation based on Dual-view Attention for Multi-view 3D Object Detection

2023-04-03 · Zhuoling Li, Chuanrui Zhang, Wei-Chiu Ma, Yipin Zhou 외

In recent years, transformer-based detectors have demonstrated remarkable performance in 2D visual perception tasks. However, their performance in multi-view 3D object detection remains inferior to the state-of-the-art (…

3D Object Detectionobject-detectionObject Detection

Multi-View Attentive Contextualization for Multi-View 3D Object Detection

2024-05-20 · CVPR 2024 1 · Xianpeng Liu, Ce Zheng, Ming Qian, Nan Xue 외

We present Multi-View Attentive Contextualization (MvACon), a simple yet effective method for improving 2D-to-3D feature lifting in query-based multi-view 3D (MV3D) object detection. Despite remarkable progress witnessed…

3D Object DetectionObjectobject-detectionObject Detection

QD-BEV : Quantization-aware View-guided Distillation for Multi-view 3D Object Detection

2023-08-21 · ICCV 2023 1 · Yifan Zhang, Zhen Dong, Huanrui Yang, Ming Lu 외

Multi-view 3D detection based on BEV (bird-eye-view) has recently achieved significant improvements. However, the huge memory consumption of state-of-the-art models makes it hard to deploy them on vehicles, and the non-t…

3D Object DetectionModel Compressionobject-detectionObject Detection+1

CLIP3D-AD: Extending CLIP for 3D Few-Shot Anomaly Detection with Multi-View Images Generation

2024-06-27 · Zuo Zuo, Jiahao Dong, Yao Wu, Yanyun Qu 외

Few-shot anomaly detection methods can effectively address data collecting difficulty in industrial scenarios. Compared to 2D few-shot anomaly detection (2D-FSAD), 3D few-shot anomaly detection (3D-FSAD) is still an unex…

Anomaly ClassificationAnomaly DetectionDecoder