paper-with-me

홈 › Papers

Text-driven Affordance Learning from Egocentric Vision

2024-04-03 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori

Visual affordance learning is a key component for robots to understand how to interact with objects. Conventional approaches in this field rely on pre-defined objects and actions, falling short of capturing diverse interactions in realworld scenarios. The key idea of our approach is employing textual instruction, targeting various affordances for a wide range of objects. This approach covers both hand-object and tool-object interactions. We introduce text-driven affordance learning, aiming to learn contact points and manipulation trajectories from an egocentric view following textual instruction. In our task, contact points are represented as heatmaps, and the manipulation trajectory as sequences of coordinates that incorporate both linear and rotational movements for various manipulations. However, when we gather data for this task, manual annotations of these diverse interactions are costly. To this end, we propose a pseudo dataset creation pipeline and build a large pseudo-training dataset: TextAFF80K, consisting of over 80K instances of the contact points, trajectories, images, and text tuples. We extend existing referring expression comprehension models for our task, and experimental results show that our approach robustly handles multiple affordances, serving as a new standard for affordance learning in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2404.02523

Code (0)

등록된 구현이 없습니다.

Tasks

Referring ExpressionReferring Expression Comprehension

Similar Papers 제목 키워드 기반

Grounded Affordance from Exocentric View

2022-08-28 · Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 외

Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence. Due to the diversity of interactive affordance, the uniqueness of different individual…

DiversityHuman-Object Interaction DetectionObjectTransfer Learning

Grounding 3D Scene Affordance From Egocentric Interactions

2024-09-29 · Cuiyu Liu, Wei Zhai, Yuhang Yang, Hongchen Luo 외

Grounding 3D scene affordance aims to locate interactive regions in 3D environments, which is crucial for embodied agents to interact intelligently with their surroundings. Most existing approaches achieve this by mappin…

Learning Affordance Grounding from Exocentric Images

2022-03-18 · CVPR 2022 1 · Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 외

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. H…

DiversityHuman-Object Interaction DetectionObjectTransfer Learning

AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant

2022-03-08 · Benita Wong, Joya Chen, You Wu, Stan Weixian Lei 외

A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave for 1 minute?". However, there is still n…

Visual Question Answering (VQA)

MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos

2024-09-04 · Junyi Ma, Xieyuanli Chen, Wentao Bao, Jingyi Xu 외

Understanding human intentions and actions through egocentric videos is important on the path to embodied artificial intelligence. As a branch of egocentric vision techniques, hand trajectory prediction plays a vital rol…

DenoisingMambaRobot ManipulationTrajectory Prediction