paper-with-me

홈 › Papers

Probing and Bridging Geometry-Interaction Cues for Affordance Reasoning in Vision Foundation Models

2026-02-24 · Qing Zhang, Xuesong Li, Jing Zhang arxiv

What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction, and interaction perception, which models how an agent's actions engage with those parts. To test this hypothesis, we conduct a systematic probing of Visual Foundation Models (VFMs). We find that models like DINO inherently encode part-level geometric structures, while generative models like Flux contain rich, verb-conditioned spatial attention maps that serve as implicit interaction priors. Crucially, we demonstrate that these two dimensions are not merely correlated but are composable elements of affordance. By simply fusing DINO's geometric prototypes with Flux's interaction maps in a training-free and zero-shot manner, we achieve affordance estimation competitive with weakly-supervised methods. This final fusion experiment confirms that geometric and interaction perception are the fundamental building blocks of affordance understanding in VFMs, providing a mechanistic account of how perception grounds action.

📄 PDF Abstract BibTeX arXiv:2602.20501

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

2026-03-02 · Lei Yao, Yong Chen, Yuejiao Su, Yi Wang 외 arxiv

Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate fo…

PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

2026-01-11 · Yuanzhe Liu, Jingyuan Zhu, Yuchen Mo, Gen Li 외 arxiv

Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanism…

VAGNet: Grounding 3D Affordance from Human-Object Interactions in Videos

2026-02-24 · Aihua Mao, Kaihang Huang, Yong-Jin Liu, Chee Seng Chan 외 arxiv

3D object affordance grounding aims to identify regions on 3D objects that support human-object interaction (HOI), a capability essential to embodied visual reasoning. However, most existing approaches rely on static vis…

Visual Reasoning

RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation

2026-03-08 · Zhanqi Xiao, Ruiping Wang, Xilin Chen arxiv

Understanding spatial affordances -- comprising the contact regions of object interaction and the corresponding contact poses -- is essential for robots to effectively manipulate objects and accomplish diverse tasks. How…

Object LocalizationRobot ManipulationPose Estimation

AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

2026-06-04 · Qize Yu, Jiadi You, Yuran Wang, Jiaqi Liang 외 arxiv

Vision-Language-Action (VLA) models leverage the rich world knowledge of pretrained vision-language models (VLMs) to enable instruction-following robotic manipulation. However, the structural mismatch between VLM semanti…

Data Augmentation