paper-with-me

홈 › Papers

RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation

2026-03-08 · Zhanqi Xiao, Ruiping Wang, Xilin Chen arxiv

Understanding spatial affordances -- comprising the contact regions of object interaction and the corresponding contact poses -- is essential for robots to effectively manipulate objects and accomplish diverse tasks. However, existing spatial affordance prediction methods mainly focus on locating the contact regions while delegating the pose to independent pose estimation approaches, which can lead to task failures due to inconsistencies between predicted contact regions and candidate poses. In this work, we propose RoboPCA, a pose-centered affordance prediction framework that jointly predicts task-appropriate contact regions and poses conditioned on instructions. To enable scalable data collection for pose-centered affordance learning, we devise Human2Afford, a data curation pipeline that automatically recovers scene-level 3D information and infers pose-centered affordance annotations from human demonstrations. With Human2Afford, scene depth and the interaction object's mask are extracted to provide 3D context and object localization, while pose-centered affordance annotations are obtained by tracking object points within the contact region and analyzing hand-object interaction patterns to establish a mapping from the 3D hand mesh to the robot end-effector orientation. By integrating geometry-appearance cues through an RGB-D encoder and incorporating mask-enhanced features to emphasize task-relevant object regions into the diffusion-based framework, RoboPCA outperforms baseline methods on image datasets, simulation, and real robots, and exhibits strong generalization across tasks and categories.

📄 PDF Abstract BibTeX arXiv:2603.07691

Code (0)

등록된 구현이 없습니다.

Tasks

Object LocalizationRobot ManipulationPose Estimation

Similar Papers 제목 키워드 기반

BridgeACT: Bridging Human Demonstrations to Robot Actions via Unified Tool-Target Affordances

2026-04-25 · Yifan Han, Jianxiang Liu, Haoyu Zhang, Yuqi Gu 외 arxiv

Learning robot manipulation from human videos is appealing due to the scale and diversity of human demonstrations, but transferring such demonstrations to executable robot behavior remains challenging. Prior work either …

Robot Manipulation

AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand Pose

2023-09-16 · ICCV 2023 1 · Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu 외

How human interact with objects depends on the functional roles of the target objects, which introduces the problem of affordance-aware hand-object interaction. It requires a large number of human demonstrations for the …

DiversityObject

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

2025-05-17 · Teli Ma, Jia Zheng, Zifan Wang, Ziyao Gao 외

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of actionable affordances. However, transferr…

Benchmarking

The Wilhelm Tell Dataset of Affordance Demonstrations

2025-07-23 · Rachel Ringe, Mihai Pomarlan, Nikolaos Tsiogkas, Stefano De Giorgis 외 arxiv

Affordances - i.e. possibilities for action that an environment or objects in it provide - are important for robots operating in human environments to perceive. Existing approaches train such capabilities on annotated st…

UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation

2026-03-27 · Hao Li, Long Yin Chung, Jack Goler, Ryan Zhang 외 arxiv

Underwater robotic grasping is difficult due to degraded, highly variable imagery and the expense of collecting diverse underwater demonstrations. We introduce a system that (i) autonomously collects successful underwate…

Robotic Grasping