paper-with-me

홈 › Papers

PETRv2: A Unified Framework for 3D Perception from Multi-Camera Images

2022-06-02 · ICCV 2023 1 · Yingfei Liu, Junjie Yan, Fan Jia, Shuailin Li, Aqi Gao, Tiancai Wang, Xiangyu Zhang, Jian Sun

In this paper, we propose PETRv2, a unified framework for 3D perception from multi-view images. Based on PETR, PETRv2 explores the effectiveness of temporal modeling, which utilizes the temporal information of previous frames to boost 3D object detection. More specifically, we extend the 3D position embedding (3D PE) in PETR for temporal modeling. The 3D PE achieves the temporal alignment on object position of different frames. A feature-guided position encoder is further introduced to improve the data adaptability of 3D PE. To support for multi-task learning (e.g., BEV segmentation and 3D lane detection), PETRv2 provides a simple yet effective solution by introducing task-specific queries, which are initialized under different spaces. PETRv2 achieves state-of-the-art performance on 3D object detection, BEV segmentation and 3D lane detection. Detailed robustness analysis is also conducted on PETR framework. We hope PETRv2 can serve as a strong baseline for 3D perception. Code is available at \url{https://github.com/megvii-research/PETR}.

📄 PDF Abstract BibTeX arXiv:2206.01256

Code (1)

megvii-research/petr 공식 구현 pytorch

Tasks

3D Lane Detection3D Object DetectionBEV SegmentationBird's-Eye View Semantic SegmentationLane DetectionMulti-Task LearningObjectobject-detectionObject DetectionPositionSegmentation

Similar Papers 제목 키워드 기반

M-BEV: Masked BEV Perception for Robust Autonomous Driving

2023-12-19 · Siran Chen, Yue Ma, Yu Qiao, Yali Wang

3D perception is a critical problem in autonomous driving. Recently, the Bird-Eye-View (BEV) approach has attracted extensive attention, due to low-cost deployment and desirable vision detection capacity. However, the ex…

Autonomous Driving

Geometry-Grounded Unified 3D Perception for Autonomous Driving

2026-08-13 · Longfei Xu, Xiaohui Wang, Zehao Huang, Han Li 외 arxiv

Camera-based autonomous driving perception requires a shared representation that preserves metric 3D structure across synchronized multi-camera streams. However, existing image-based frameworks often rely on backbones pr…

3D Object DetectionAutonomous DrivingDepth Estimation

UniDrive: Towards Universal Driving Perception Across Camera Configurations

2024-10-17 · Ye Li, Wenzhao Zheng, Xiaonan Huang, Kurt Keutzer

Vision-centric autonomous driving has demonstrated excellent performance with economical sensors. As the fundamental step, 3D perception aims to infer 3D information from 2D images based on 3D-2D projection. This makes d…

Autonomous Driving

CoBEVFusion: Cooperative Perception with LiDAR-Camera Bird's-Eye View Fusion

2023-10-09 · Donghao Qiao, Farhana Zulkernine

Autonomous Vehicles (AVs) use multiple sensors to gather information about their surroundings. By sharing sensor data between Connected Autonomous Vehicles (CAVs), the safety and reliability of these vehicles can be impr…

3D Object DetectionAutonomous Vehiclesobject-detectionObject Detection+1

Achelous: A Fast Unified Water-surface Panoptic Perception Framework based on Fusion of Monocular Camera and 4D mmWave Radar

2023-07-14 · Runwei Guan, Shanliang Yao, Xiaohui Zhu, Ka Lok Man 외

Current perception models for different tasks usually exist in modular forms on Unmanned Surface Vehicles (USVs), which infer extremely slowly in parallel on edge devices, causing the asynchrony between perception result…

2D Semantic SegmentationAutonomous NavigationObject DetectionPoint Cloud Segmentation+1