paper-with-me

Papers

HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via Transformer

2025-03-13 · Zhang Zhang, Chao Sun, Chao Yue, Da Wen, Yujie Chen, Tianze Wang, Jianghao Leng

Roadside vision centric 3D object detection has received increasing attention in recent years. It expands the perception range of autonomous vehicles, enhances the road safety. Previous methods focused on predicting per-pixel height rather than depth, making significant gains in roadside visual perception. While it is limited by the perspective property of near-large and far-small on image features, making it difficult for network to understand real dimension of objects in the 3D world. BEV features and voxel features present the real distribution of objects in 3D world compared to the image features. However, BEV features tend to lose details due to the lack of explicit height information, and voxel features are computationally expensive. Inspired by this insight, an efficient framework learning height prediction in voxel features via transformer is proposed, dubbed HeightFormer. It groups the voxel features into local height sequences, and utilize attention mechanism to obtain height distribution prediction. Subsequently, the local height sequences are reassembled to generate accurate 3D features. The proposed method is applied to two large-scale roadside benchmarks, DAIR-V2X-I and Rope3D. Extensive experiments are performed and the HeightFormer outperforms the state-of-the-art methods in roadside vision centric 3D object detection task.

📄 PDF Abstract BibTeX arXiv:2503.10777

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object DetectionAutonomous Vehiclesobject-detectionObject Detection

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

HeightFormer: A Semantic Alignment Monocular 3D Object Detection Method from Roadside Perspective

2024-10-10 · Pei Liu, Zihao Zhang, Haipeng Liu, Nanfang Zheng 외

The on-board 3D object detection technology has received extensive attention as a critical technology for autonomous driving, while few studies have focused on applying roadside sensors in 3D traffic object detection. Ex…

3D Object DetectionAutonomous DrivingMonocular 3D Object DetectionObject+3

HeightFormer: Explicit Height Modeling without Extra Data for Camera-only 3D Object Detection in Bird's Eye View

2023-07-25 · Yiming Wu, Ruixiang Li, Zequn Qin, Xinhai Zhao 외

Vision-based Bird's Eye View (BEV) representation is an emerging perception formulation for autonomous driving. The core challenge is to construct BEV space with multi-camera features, which is a one-to-many ill-posed pr…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection

BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-based Roadside 3D Object Detection

2024-06-13 · CVPR 2024 1 · Wenjie Wang, Yehao Lu, Guangcong Zheng, Shuigen Zhan 외

Vision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it encompasses inherent advantages in reducing blind spots and expanding perception range. While previous work …

3D Object DetectionAutonomous Drivingobject-detectionObject Detection

HeightFormer: A Multilevel Interaction and Image-adaptive Classification-regression Network for Monocular Height Estimation with Aerial Images

2023-10-12 · Zhan Chen, Yidan Zhang, Xiyu Qi, Yongqiang Mao 외

Height estimation has long been a pivotal topic within measurement and remote sensing disciplines, proving critical for endeavours such as 3D urban modelling, MR and autonomous driving. Traditional methods utilise stereo…

Autonomous DrivingregressionStereo Matching

SliceSemOcc: Vertical Slice Based Multimodal 3D Semantic Occupancy Representation

2025-09-04 · Han Huang, Han Sun, Ningzhong Liu, Huiyu Zhou 외 arxiv

Driven by autonomous driving's demands for precise 3D perception, 3D semantic occupancy prediction has become a pivotal research topic. Unlike bird's-eye-view (BEV) methods, which restrict scene representation to a 2D pl…

Autonomous Driving