paper-with-me

홈 › Papers

InverseMatrixVT3D: An Efficient Projection Matrix-Based Approach for 3D Occupancy Prediction

2024-01-23 · Zhenxing Ming, Julie Stephany Berrio, Mao Shan, Stewart Worrall

This paper introduces InverseMatrixVT3D, an efficient method for transforming multi-view image features into 3D feature volumes for 3D semantic occupancy prediction. Existing methods for constructing 3D volumes often rely on depth estimation, device-specific operators, or transformer queries, which hinders the widespread adoption of 3D occupancy models. In contrast, our approach leverages two projection matrices to store the static mapping relationships and matrix multiplications to efficiently generate global Bird's Eye View (BEV) features and local 3D feature volumes. Specifically, we achieve this by performing matrix multiplications between multi-view image feature maps and two sparse projection matrices. We introduce a sparse matrix handling technique for the projection matrices to optimize GPU memory usage. Moreover, a global-local attention fusion module is proposed to integrate the global BEV features with the local 3D feature volumes to obtain the final 3D volume. We also employ a multi-scale supervision mechanism to enhance performance further. Extensive experiments performed on the nuScenes and SemanticKITTI datasets reveal that our approach not only stands out for its simplicity and effectiveness but also achieves the top performance in detecting vulnerable road users (VRU), crucial for autonomous driving and road safety. The code has been made available at: https://github.com/DanielMing123/InverseMatrixVT3D

📄 PDF Abstract BibTeX arXiv:2401.12422

Code (1)

danielming123/inversematrixvt3d 공식 구현 pytorch

Tasks

3D Semantic Occupancy PredictionAutonomous DrivingDepth EstimationGPU

Methods 이 논문이 사용한 방법론

Global-Local Attention 설명 없음

Similar Papers 제목 키워드 기반

FDR-Occ: Factorized Dense Routing for Full-Spectrum 3D Occupancy Prediction

2026-07-04 · Dubing Chen, Huan Zheng, Tianyi Yan, Yucheng Zhou 외 arxiv

Vision-based 3D occupancy prediction fundamentally relies on the 2D-to-3D view transformation. Current paradigms predominantly utilize explicit physical projection, which artificially restricts the routing matrix to stri…

FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction

2026-04-30 · Zeyu Jiang, Changqing Zhou, Xingxing Zuo, Changhao Chen arxiv

Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training-free framework for open-vocabulary occupancy prediction fr…

VGGT-Occ: Geometry-Grounded and Density-Aware Gated Fusion for 3D Occupancy Prediction

2026-05-16 · Xun Chen, Tianchen Deng, Rui Wang, Fangjinhua Wang 외 arxiv

3D semantic occupancy prediction requires accurate 2D-to-3D feature lifting, yet current methods restrict camera geometry to initial projections. Subsequent operations like offset learning, attention weighting, and cross…

PointOcc: Cylindrical Tri-Perspective View for Point-based 3D Semantic Occupancy Prediction

2023-08-31 · Sicheng Zuo, Wenzhao Zheng, Yuanhui Huang, Jie zhou 외

Semantic segmentation in autonomous driving has been undergoing an evolution from sparse point segmentation to dense voxel segmentation, where the objective is to predict the semantic occupancy of each voxel in the conce…

3D Semantic Occupancy PredictionAutonomous DrivingSegmentationSemantic Segmentation

FB-OCC: 3D Occupancy Prediction based on Forward-Backward View Transformation

2023-07-04 · Zhiqi Li, Zhiding Yu, David Austin, Mingsheng Fang 외

This technical report summarizes the winning solution for the 3D Occupancy Prediction Challenge, which is held in conjunction with the CVPR 2023 Workshop on End-to-End Autonomous Driving and CVPR 23 Workshop on Vision-Ce…

Autonomous DrivingPrediction Of Occupancy Grid Maps