paper-with-me

Papers

Grounding 3D Scene Affordance From Egocentric Interactions

2024-09-29 · Cuiyu Liu, Wei Zhai, Yuhang Yang, Hongchen Luo, Sen Liang, Yang Cao, Zheng-Jun Zha

Grounding 3D scene affordance aims to locate interactive regions in 3D environments, which is crucial for embodied agents to interact intelligently with their surroundings. Most existing approaches achieve this by mapping semantics to 3D instances based on static geometric structure and visual appearance. This passive strategy limits the agent's ability to actively perceive and engage with the environment, making it reliant on predefined semantic instructions. In contrast, humans develop complex interaction skills by observing and imitating how others interact with their surroundings. To empower the model with such abilities, we introduce a novel task: grounding 3D scene affordance from egocentric interactions, where the goal is to identify the corresponding affordance regions in a 3D scene based on an egocentric video of an interaction. This task faces the challenges of spatial complexity and alignment complexity across multiple sources. To address these challenges, we propose the Egocentric Interaction-driven 3D Scene Affordance Grounding (Ego-SAG) framework, which utilizes interaction intent to guide the model in focusing on interaction-relevant sub-regions and aligns affordance features from different sources through a bidirectional query decoder mechanism. Furthermore, we introduce the Egocentric Video-3D Scene Affordance Dataset (VSAD), covering a wide range of common interaction types and diverse 3D environments to support this task. Extensive experiments on VSAD validate both the feasibility of the proposed task and the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2409.19650

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Grounded Affordance from Exocentric View

2022-08-28 · Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 외

Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence. Due to the diversity of interactive affordance, the uniqueness of different individual…

DiversityHuman-Object Interaction DetectionObjectTransfer Learning

Learning Affordance Grounding from Exocentric Images

2022-03-18 · CVPR 2022 1 · Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 외

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. H…

DiversityHuman-Object Interaction DetectionObjectTransfer Learning

INTRA: Interaction Relationship-aware Weakly Supervised Affordance Grounding

2024-09-10 · Ji Ha Jang, Hoigi Seo, Se Young Chun

Affordance denotes the potential interactions inherent in objects. The perception of affordance can enable intelligent agents to navigate and interact with new environments efficiently. Weakly supervised affordance groun…

Contrastive LearningLanguage ModelingLanguage ModellingNavigate+1

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

2025-10-20 · Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang arxiv

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions tha…

Knowledge Distillation

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

2026-06-02 · Litao Liu, Yifan Han, Pengfei Yi, Wenbo Yu 외 arxiv

Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dependent and often one-to-many in cluttered scenes: the same object ma…