paper-with-me

홈 › Papers

Beyond the Contact: Discovering Comprehensive Affordance for 3D Objects from Pre-trained 2D Diffusion Models

2024-01-23 · Hyeonwoo Kim, Sookwan Han, Patrick Kwon, Hanbyul Joo

Understanding the inherent human knowledge in interacting with a given environment (e.g., affordance) is essential for improving AI to better assist humans. While existing approaches primarily focus on human-object contacts during interactions, such affordance representation cannot fully address other important aspects of human-object interactions (HOIs), i.e., patterns of relative positions and orientations. In this paper, we introduce a novel affordance representation, named Comprehensive Affordance (ComA). Given a 3D object mesh, ComA models the distribution of relative orientation and proximity of vertices in interacting human meshes, capturing plausible patterns of contact, relative orientations, and spatial relationships. To construct the distribution, we present a novel pipeline that synthesizes diverse and realistic 3D HOI samples given any 3D object mesh. The pipeline leverages a pre-trained 2D inpainting diffusion model to generate HOI images from object renderings and lifts them into 3D. To avoid the generation of false affordances, we propose a new inpainting framework, Adaptive Mask Inpainting. Since ComA is built on synthetic samples, it can extend to any object in an unbounded manner. Through extensive experiments, we demonstrate that ComA outperforms competitors that rely on human annotations in modeling contact-based affordance. Importantly, we also showcase the potential of ComA to reconstruct human-object interactions in 3D through an optimization framework, highlighting its advantage in incorporating both contact and non-contact properties.

📄 PDF Abstract BibTeX arXiv:2401.12978

Code (1)

snuvclab/coma 공식 구현 pytorch

Tasks

Human-Object Interaction DetectionObjectZero-Shot Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음
Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation

2024-01-15 · Yuanchen Ju, Kaizhe Hu, Guowei Zhang, Gu Zhang 외

Enabling robotic manipulation that generalizes to out-of-distribution scenes is a crucial step toward open-world embodied intelligence. For human beings, this ability is rooted in the understanding of semantic correspond…

Robot ManipulationSemantic correspondence

Text-driven Affordance Learning from Egocentric Vision

2024-04-03 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori

Visual affordance learning is a key component for robots to understand how to interact with objects. Conventional approaches in this field rely on pre-defined objects and actions, falling short of capturing diverse inter…

Referring ExpressionReferring Expression Comprehension

H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows

2025-10-17 · Harry Zhang, Luca Carlone arxiv

Understanding how humans interact with the surrounding environment, and specifically reasoning about object interactions and affordances, is a critical challenge in computer vision, robotics, and AI. Current approaches o…

Point Clouds

End-to-End Affordance Learning for Robotic Manipulation

2022-09-26 · Yiran Geng, Boshi An, Haoran Geng, Yuanpei Chen 외

Learning to manipulate 3D objects in an interactive environment has been a challenging problem in Reinforcement Learning (RL). In particular, it is hard to train a policy that can generalize over objects with different s…

Reinforcement Learning (RL)

AffordanceLLM: Grounding Affordance from Vision Language Models

2024-01-12 · Shengyi Qian, Weifeng Chen, Min Bai, Xiong Zhou 외

Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires the comprehensive understanding of a scene…

Human-Object Interaction DetectionObject