paper-with-me

Papers

AdaOcc: Adaptive Forward View Transformation and Flow Modeling for 3D Occupancy and Flow Prediction

2024-07-01 · Dubing Chen, Wencheng Han, Jin Fang, Jianbing Shen

In this technical report, we present our solution for the Vision-Centric 3D Occupancy and Flow Prediction track in the nuScenes Open-Occ Dataset Challenge at CVPR 2024. Our innovative approach involves a dual-stage framework that enhances 3D occupancy and flow predictions by incorporating adaptive forward view transformation and flow modeling. Initially, we independently train the occupancy model, followed by flow prediction using sequential frame integration. Our method combines regression with classification to address scale variations in different scenes, and leverages predicted flow to warp current voxel features to future frames, guided by future frame ground truth. Experimental results on the nuScenes dataset demonstrate significant improvements in accuracy and robustness, showcasing the effectiveness of our approach in real-world scenarios. Our single model based on Swin-Base ranks second on the public leaderboard, validating the potential of our method in advancing autonomous car perception systems.

📄 PDF Abstract BibTeX arXiv:2407.01436

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AdaOcc: Adaptive-Resolution Occupancy Prediction

2024-08-24 · Chao Chen, Ruoyu Wang, Yuliang Guo, Cheng Zhao 외

Autonomous driving in complex urban scenarios requires 3D perception to be both comprehensive and precise. Traditional 3D perception methods focus on object detection, resulting in sparse representations that lack enviro…

3D Reconstruction3D Semantic Occupancy PredictionAutonomous Drivingobject-detection+2

Feedforward 3D Editing Learns from Semantic-Part Transformation

2026-05-26 · Jiawei Weng, Saining Zhang, Zhenxin Diao, Peishuo Li 외 arxiv

3D editing is a fundamental capability for scalable 3D content creation. While image editing has rapidly evolved toward large-scale feedforward generative paradigms, 3D AI generation remains dominated by training-free ed…

Image Editing

FB-BEV: BEV Representation from Forward-Backward View Transformations

2023-08-04 · ICCV 2023 1 · Zhiqi Li, Zhiding Yu, Wenhai Wang, Anima Anandkumar 외

View Transformation Module (VTM), where transformations happen between multi-view image features and Bird-Eye-View (BEV) representation, is a crucial step in camera-based BEV perception systems. Currently, the two most p…

Learning by Analogy: Reliable Supervision from Transformations for Unsupervised Optical Flow Estimation

2020-03-29 · CVPR 2020 6 · Liang Liu, Jiangning Zhang, Ruifei He, Yong liu 외

Unsupervised learning of optical flow, which leverages the supervision from view synthesis, has emerged as a promising alternative to supervised methods. However, the objective of unsupervised learning is likely to be un…

DecoderOptical Flow EstimationSelf-Supervised Learning

Flow Equivariant Recurrent Neural Networks

2025-07-20 · T. Anderson Keller arxiv

Data arrives at our senses as a continuous stream, smoothly transforming from one instant to the next. These smooth transformations can be viewed as continuous symmetries of the environment that we inhabit, defining equi…