paper-with-me

Papers

MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models

2025-10-13 · Subhransu S. Bhattacharjee, Hao Lu, Dylan Campbell, Rahul Shome arxiv

Autonomous robots often view rooms only partially, through a doorway, where the walls and scene structure hide the geometry and task-relevant semantics needed for safe navigation and goal-directed action. We ask whether off-the-shelf pretrained generative vision models can derive this missing structure as zero-shot offline priors for robot reasoning. Such priors should support spatio-semantic queries over unobserved structure, estimating the target object likelihood in hidden regions and the probability that those regions are occupied. Given an egocentric RGB observation and target query, our pipeline uses VLM-guided outpainting, monocular depth estimation, and semantic segmentation to sample semantically labeled 3D point cloud hypotheses of the hidden room. We introduce MatterDoor, a Matterport3D-derived benchmark of doorway-occluded indoor scenes, and evaluate the resulting priors with generative metrics and simulated Stretch robot object-reaching tasks. Our results suggest that useful spatio-semantic priors for planning can be derived without problem-specific fine-tuning.

📄 PDF Abstract BibTeX arXiv:2510.11014

Code (0)

등록된 구현이 없습니다.

Tasks

Monocular Depth EstimationSemantic Segmentation

Similar Papers 제목 키워드 기반

Context-Aware Zero-Shot Anomaly Detection in Surveillance Using Contrastive and Predictive Spatiotemporal Modeling

2025-08-25 · Md. Rashid Shahriar Khan, Md. Abrar Hasan, Mohammod Tareq Aziz Justice arxiv

Detecting anomalies in surveillance footage is inherently challenging due to their unpredictable and context-dependent nature. This work introduces a novel context-aware zero-shot anomaly detection framework that identif…

Anomaly Detection

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning

2026-05-18 · Pawat Chunhachatrachai, Gueter Josmy Faure, Hung-Ting Su, Winston H. Hsu arxiv

Spatial question answering over egocentric video is a challenging task that requires Vision-Language Models (VLMs) to reason about 3D object positions, scene affordances, and directional relationships, particularly in th…

Question AnsweringSpatial Reasoning

Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes

2025-10-31 · Yehna Kim, Young-Eun Kim, Seong-Whan Lee arxiv

Vision-Language Models (VLMs) have demonstrated impressive capabilities in zero-shot action recognition by learning to associate video embeddings with class embeddings. However, a significant challenge arises when relyin…

Zero-Shot Action Recognition

End-to-End Semantic Video Transformer for Zero-Shot Action Recognition

2022-03-10 · Keval Doshi, Yasin Yilmaz

While video action recognition has been an active area of research for several years, zero-shot action recognition has only recently started gaining traction. In this work, we propose a novel end-to-end trained transform…

Action RecognitionTemporal Action LocalizationZero-Shot Action RecognitionZero-Shot Learning

USS-Nav: Unified Spatio-Semantic Scene Graph for Lightweight UAV Zero-Shot Object Navigation

2026-01-31 · Weiqi Gai, Yuman Gao, Yuan Zhou, Yufan Xie 외 arxiv

Zero-Shot Object Navigation in unknown environments poses significant challenges for Unmanned Aerial Vehicles (UAVs) due to the conflict between high-level semantic reasoning requirements and limited onboard computationa…

Computational EfficiencyGraph GenerationGraph Clustering