paper-with-me

홈 › Papers

Exploring 2D backbone effects for indoor semantic occupancy prediction

2026-09-15 · Shizhang Fanga, Wanling Yea, Qi Zheng arxiv

Semantic occupancy prediction gives an embodied agent a voxel-level account of where space is free, occupied, and semantically meaningful. In RGB-D pipelines such as EmbodiedScan, the image encoder is often left as a default module, even though its features are the visual evidence later sampled into the 3D grid. We study this design choice directly. A central finding is that changing the 2D backbone improves occupancy accuracy more than several carefully designed occupancy architectures or modules. We keep the main RGB-D projection, depth branch, and occupancy head fixed, and replace only the image backbone. The compared encoders are CLIP-ResNet, CLIP-ViT, BLIP2, and DINOv2. Under the controlled setting, the measured mIoU changes substantially: DINOv2 obtains 30.55\%, BLIP2 obtains 29.49\%, CLIP-ViT obtains 24.33\%, and CLIP-ResNet obtains 17.41\%. The stronger encoders also exceed the original EmbodiedScan ResNet-50 baseline without modifying the downstream 3D fusion pipeline. Class-level results give a more detailed picture: DINOv2 is stronger on many layout and structural categories, whereas BLIP2 remains close on several object-centered classes. CLIP-ViT improves clearly over CLIP-ResNet, showing that the way CLIP features are exposed as dense tokens matters for voxel lifting. These results indicate that the image backbone is not a secondary engineering detail in embodied semantic occupancy, but a major source of variation in the final 3D prediction.

📄 PDF Abstract BibTeX arXiv:2609.17257

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

GEM-Occ: From Visual Geometry Evidence to Embodied Semantic Occupancy Memory

2026-07-06 · Hu Zhu, Bohan Li, Xianda Guo, Hongsi Liu 외 arxiv

Semantic occupancy provides a structured spatial memory for embodied indoor agents by jointly representing occupied regions, observed free space, unknown areas, and object semantics. However, existing indoor occupancy be…

SliceOcc: Indoor 3D Semantic Occupancy Prediction with Vertical Slice Representation

2025-01-28 · Jianing Li, Ming Lu, Hao Wang, Chenyang Gu 외

3D semantic occupancy prediction is a crucial task in visual perception, as it requires the simultaneous comprehension of both scene geometry and semantics. It plays a crucial role in understanding 3D scenes and has grea…

3D Semantic Occupancy PredictionAutonomous DrivingPrediction

FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction

2026-04-30 · Zeyu Jiang, Changqing Zhou, Xingxing Zuo, Changhao Chen arxiv

Existing learning-based occupancy prediction methods rely on large-scale 3D annotations and generalize poorly across environments. We present FreeOcc, a training-free framework for open-vocabulary occupancy prediction fr…

Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes

2026-02-26 · Changqing Zhou, Yueru Luo, Han Zhang, Zeyu Jiang 외 arxiv

Open-vocabulary 3D occupancy is vital for embodied agents, which need to understand complex indoor environments where semantic categories are abundant and evolve beyond fixed taxonomies. While recent work has explored op…

Autonomous Exploration and Semantic Updating of Large-Scale Indoor Environments with Mobile Robots

2024-09-23 · Sai Haneesh Allu, Itay Kadosh, Tyler Summers, Yu Xiang

We introduce a new robotic system that enables a mobile robot to autonomously explore an unknown environment, build a semantic map of the environment, and subsequently update the semantic map to reflect environment chang…