paper-with-me

홈 › Papers

VEOcc: Voxel-Centric Online Semantic Occupancy Prediction For Embodied Scene Understanding

2026-05-24 · Ruoyu Wang, Yong Liu, Sheng Tao, Yuhang Lin, Yukai Ma arxiv

Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally constructs dense spatial representations on the fly. However, recent Gaussian-centric methods struggle with structural boundary fidelity and rely heavily on predefined scene-size priors, fundamentally limiting their operational efficiency. In this work, we present VEOcc, a voxel-centric framework formulated as a recursive perception-and-assimilation paradigm. By eliminating the need for initial scale estimation, VEOcc enables highly streamlined, open-ended map expansion. Furthermore, to robustly aggregate noisy temporal observations within the discrete voxel space, we propose a Spatio-Temporal-Aware Online Update Strategy. It integrates Cross-Temporal Logit Aggregation (TLA) for temporal consistency, Reliability-Aware Confidence Modulation (RCM) for spatial uncertainty calibration, and Confidence-Driven Incremental State Update (CSU) for robust global state assimilation. % Extensive experiments on Occ-ScanNet and EmbodiedOcc-ScanNet demonstrate that VEOcc establishes new state-of-the-art performance in both local and embodied settings, providing an accurate and efficient solution for real-world exploration. Extensive experiments on Occ-ScanNet and EmbodiedOcc-ScanNet demonstrate that VEOcc establishes new state-of-the-art performance in both local and embodied settings. Notably, zero-shot evaluations on self-collected video sequences further confirm its robust out-of-distribution generalization capability in completely unseen real-world environments. Ultimately, our framework provides an accurate and highly efficient solution for autonomous exploration. Code and supplementary visualizations are available on our project page: https://wryzju.github.io/VEOcc/.

📄 PDF Abstract BibTeX arXiv:2605.25059

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

VoxDet: Rethinking 3D Semantic Occupancy Prediction as Dense Object Detection

2025-06-05 · Wuyang Li, Zhu Yu, Alexandre Alahi

3D semantic occupancy prediction aims to reconstruct the 3D geometry and semantics of the surrounding environment. With dense voxel labels, prior works typically formulate it as a dense segmentation task, independently c…

3D geometry3D Semantic Occupancy PredictionDense Object Detectionobject-detection+1

KRVF: A Source-Aware Semantic Voxel World Representation for Edge Mobile Manipulation

2026-06-24 · Runfeng Ling arxiv

Mobile manipulators need world models that are current, queryable, semantically meaningful, and usable under edge-compute constraints. This technical report presents KRVF, a source-aware semantic voxel world representati…

GaussianSSC: Triplane-Guided Directional Gaussian Fields for 3D Semantic Completion

2026-03-23 · Ruiqi Xian, Jing Liang, He Yin, Xuewei Qi 외 arxiv

We present \emph{GaussianSSC}, a two-stage, grid-native and triplane-guided approach to semantic scene completion (SSC) that injects the benefits of Gaussians without replacing the voxel grid or maintaining a separate Ga…

GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction

2024-05-27 · Yuanhui Huang, Wenzhao Zheng, Yunpeng Zhang, Jie zhou 외

3D semantic occupancy prediction aims to obtain 3D fine-grained geometry and semantics of the surrounding scene and is an important task for the robustness of vision-centric autonomous driving. Most existing methods empl…

3D Semantic Occupancy PredictionAutonomous DrivingDiversityPosition

PanoSSC: Exploring Monocular Panoptic 3D Scene Reconstruction for Autonomous Driving

2024-06-11 · Yining Shi, Jiusi Li, Kun Jiang, Ke Wang 외

Vision-centric occupancy networks, which represent the surrounding environment with uniform voxels with semantics, have become a new trend for safe driving of camera-only autonomous driving perception systems, as they ar…

3D Instance Segmentation3D Scene Reconstruction3D Semantic SegmentationAutonomous Driving+5