paper-with-me

홈 › Papers

LEAP: Enhancing Vision-Based Occupancy Networks with Lightweight Spatio-Temporal Correlation

2025-02-21 · Fengcheng Yu, Haoran Xu, Canming Xia, Guang Tan

Vision-based occupancy networks provide an end-to-end solution for reconstructing the surrounding environment using semantic occupied voxels derived from multi-view images. This technique relies on effectively learning the correlation between pixel-level visual information and voxels. Despite recent advancements, occupancy results still suffer from limited accuracy due to occlusions and sparse visual cues. To address this, we propose a Lightweight Spatio-Temporal Correlation (LEAP)} method, which significantly enhances the performance of existing occupancy networks with minimal computational overhead. LEAP can be seamlessly integrated into various baseline networks, enabling a plug-and-play application. LEAP operates in three stages: 1) it tokenizes information from recent baseline and motion features into a shared, compact latent space; 2) it establishes full correlation through a tri-stream fusion architecture; 3) it generates occupancy results that strengthen the baseline's output. Extensive experiments demonstrate the efficiency and effectiveness of our method, outperforming the latest baseline models. The source code and several demos are available in the supplementary material.

📄 PDF Abstract BibTeX arXiv:2502.15438

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

UnO: Unsupervised Occupancy Fields for Perception and Forecasting

2024-06-12 · CVPR 2024 1 · Ben Agro, Quinlan Sykora, Sergio Casas, Thomas Gilles 외

Perceiving the world and forecasting its future state is a critical task for self-driving. Supervised approaches leverage annotated object labels to learn a model of the world -- traditionally with object detections and …

LEAP: Learning Articulated Occupancy of People

2021-04-14 · CVPR 2021 1 · Marko Mihajlovic, Yan Zhang, Michael J. Black, Siyu Tang

Substantial progress has been made on modeling rigid 3D objects using deep implicit representations. Yet, extending these methods to learn neural models of human shape is still in its infancy. Human bodies are complex an…

SparseOccVLA: Bridging Occupancy and Vision-Language Models via Sparse Queries for Unified 4D Scene Understanding and Planning

2026-01-10 · Chenxu Dang, Jie Wang, Guang Li, Zhiwen Hou 외 arxiv

In autonomous driving, Vision Language Models (VLMs) excel at high-level reasoning , whereas semantic occupancy provides fine-grained details. Despite significant progress in individual fields, there is still no method t…

Scene UnderstandingTrajectory PlanningAutonomous Driving

Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting

2024-11-21 · CVPR 2025 1 · Jingyi Xu, Xieyuanli Chen, Junyi Ma, Jiawei Huang 외

The task of occupancy forecasting (OCF) involves utilizing past and present perception data to predict future occupancy states of autonomous vehicle surrounding environments, which is critical for downstream tasks such a…

GPU

Unified Spatio-Temporal Tri-Perspective View Representation for 3D Semantic Occupancy Prediction

2024-01-24 · Sathira Silva, Savindu Bhashitha Wannigama, Gihan Jayatilaka, Muhammad Haris Khan 외

Holistic understanding and reasoning in 3D scenes play a vital role in the success of autonomous driving systems. The evolution of 3D semantic occupancy prediction as a pretraining task for autonomous driving and robotic…

3D Semantic Occupancy PredictionAutonomous DrivingComputational Efficiency