paper-with-me

홈 › Papers

AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis

2026-03-09 · Xiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma, Yujiao Shi, Xuming He arxiv

Generating human grasping poses that accurately reflect both object geometry and user-specified interaction semantics is essential for natural hand-object interactions in AR/VR and embodied AI. However, existing semantic grasping approaches struggle with the large modality gap between 3D object representations and textual instructions, and often lack explicit spatial or semantic constraints, leading to physically invalid or semantically inconsistent grasps. In this work, we present AffordGrasp, a diffusion-based framework that produces physically stable and semantically faithful human grasps with high precision. We first introduce a scalable annotation pipeline that automatically enriches hand-object interaction datasets with fine-grained structured language labels capturing interaction intent. Building upon these annotations, AffordGrasp integrates an affordance-aware latent representation of hand poses with a dual-conditioning diffusion process, enabling the model to jointly reason over object geometry, spatial affordances, and instruction semantics. A distribution adjustment module further enforces physical contact consistency and semantic alignment. We evaluate AffordGrasp across four instruction-augmented benchmarks derived from HO-3D, OakInk, GRAB, and AffordPose, and observe substantial improvements over state-of-the-art methods in grasp quality, semantic accuracy, and diversity.

📄 PDF Abstract BibTeX arXiv:2603.08021

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion

2024-12-19 · Jixuan He, Wanhua Li, Ye Liu, Junsik Kim 외

As a common image editing operation, image composition involves integrating foreground objects into background scenes. In this paper, we expand the application of the concept of Affordance from human-centered image compo…

Object

ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation

2025-11-25 · Yuhan Wu, Tiantian Wei, Shuo Wang, ZhiChao Wang 외 arxiv

Interactive articulated manipulation requires long-horizon, multi-step interactions with appliances while maintaining physical consistency. Existing vision-language and diffusion-based policies struggle to generalize acr…

Affordance-Guided Diffusion Prior for 3D Hand Reconstruction

2025-10-01 · Naru Suzuki, Takehiko Ohkawa, Tatsuro Banno, Jihyun Lee 외 arxiv

How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambiguities by leveraging contextual knowledge -- such as affordances, where an…

Hand Pose Estimation

Object Affordance Recognition and Grounding via Multi-scale Cross-modal Representation Learning

2025-08-02 · Xinhang Wan, Dongqiang Gou, Xinwang Liu, En Zhu 외 arxiv

A core problem of Embodied AI is to learn object manipulation from observation, as humans do. To achieve this, it is important to localize 3D object affordance areas through observation such as images (3D affordance grou…

Representation LearningAffordance Recognition

MAAL: Multimodality-Aware Autoencoder-Based Affordance Learning for 3D Articulated Objects

2023-01-01 · ICCV 2023 1 · Yuanzhi Liang, Xiaohan Wang, Linchao Zhu, Yi Yang

Inferring affordance for 3D articulated objects is a challenging and practical problem. It is a primary problem for applying robots to real-world scenarios. The exploration can be summarized as figuring out where to …

MMEObject