paper-with-me

Papers

SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes

2026-03-31 · Léopold Maillard, Francis Engelmann, Tom Durand, Boxiao Pan, Yang You, Or Litany, Leonidas Guibas, Maks Ovsjanikov arxiv

Embodied AI depends on interactive 3D environments that support meaningful activities for diverse users, yet assessing their functional affordances remains a core challenge. We introduce SceneTeract, a framework that verifies 3D scene functionality under agent-specific constraints. Our core contribution is a grounded verification engine that couples high-level semantic reasoning with low-level geometric checks. SceneTeract decomposes complex activities into sequences of atomic actions and validates each step against accessibility requirements (e.g., reachability, clearance, and navigability) conditioned on an embodied agent profile, using explicit physical and geometric simulations. We deploy SceneTeract to perform an in-depth evaluation of (i) synthetic indoor environments, uncovering frequent functional failures that prevent basic interactions, and (ii) the ability of frontier Vision-Language Models (VLMs) to reason about and predict functional affordances, revealing systematic mismatches between semantic confidence and physical feasibility even for the strongest current models. Finally, we leverage SceneTeract as a reward engine for VLM post-training, enabling scalable distillation of geometric constraints into reasoning models. We release the SceneTeract verification suite and data to bridge perception and physical reality in embodied 3D scene understanding.

📄 PDF Abstract BibTeX arXiv:2603.29798

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Understanding

Similar Papers 제목 키워드 기반

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

2026-06-02 · Litao Liu, Yifan Han, Pengfei Yi, Wenbo Yu 외 arxiv

Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dependent and often one-to-many in cluttered scenes: the same object ma…

Grounding by Remembering: Cross-Scene and In-Scene Memory for 3D Functional Affordances

2026-05-12 · Qirui Wang, Jingyi He, Yining Pan, Xulei Yang 외 arxiv

Functional affordance grounding requires more than recognizing an object: an agent must localize the specific region that supports an interaction, such as the handle to pull or the button to press. This is difficult for …

AffordanceSAM: Segment Anything Once More in Affordance Grounding

2025-04-22 · Dengyang Jiang, Mengmeng Wang, Teli Ma, Hengzhuang Li 외

Improving the generalization ability of an affordance grounding model to recognize regions for unseen objects and affordance functions is crucial for real-world application. However, current models are still far away fro…

Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action

2025-09-23 · Sacha Morin, Kumaraditya Gupta, Mahtab Sandhu, Charlie Gauthier 외 arxiv

Executing open-ended natural language queries is a core problem in robotics. While recent advances in imitation learning and vision-language-actions models (VLAs) have enabled promising end-to-end policies, these models …

Natural Language QueriesMotion Planning

Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction

2026-08-19 · Zijian Xiao, Zipeng Ye, Jinkun Hao, Xiong Yang 외 arxiv

Indoor scene synthesis provides essential environments for embodied AI, robotic manipulation, and simulation-based policy learning. Recent code-based scene generation methods produce editable and extensible environments,…

Indoor Scene SynthesisScene Generation