paper-with-me

홈 › Papers

Implicit Affordance Acquisition via Causal Action-Effect Modeling in the Video Domain

2023-12-18 · Hsiu-Yu Yang, Carina Silberer

Affordance knowledge is a fundamental aspect of commonsense knowledge. Recent findings indicate that world knowledge emerges through large-scale self-supervised pretraining, motivating our exploration of acquiring affordance knowledge from the visual domain. To this end, we augment an existing instructional video resource to create the new Causal Action-Effect (CAE) dataset and design two novel pretraining tasks -- Masked Action Modeling (MAM) and Masked Effect Modeling (MEM) -- promoting the acquisition of two affordance properties in models: behavior and entity equivalence, respectively. We empirically demonstrate the effectiveness of our proposed methods in learning affordance properties. Furthermore, we show that a model pretrained on both tasks outperforms a strong image-based visual-linguistic foundation model (FLAVA) as well as pure linguistic models on a zero-shot physical reasoning probing task.

📄 PDF Abstract BibTeX arXiv:2312.11345

Code (1)

mallory24/cae_modeling 공식 구현 pytorch

Tasks

World Knowledge

Similar Papers 제목 키워드 기반

Proposition of Affordance-Driven Environment Recognition Framework Using Symbol Networks in Large Language Models

2025-04-02 · Kazuma Arii, Satoshi Kurihara

In the quest to enable robots to coexist with humans, understanding dynamic situations and selecting appropriate actions based on common sense and affordances are essential. Conventional AI systems face challenges in app…

Common Sense Reasoning

A3R: Agentic Affordance Reasoning via Cross-Dimensional Evidence in 3D Gaussian Scenes

2026-04-02 · Di Li, Jie Feng, Guanbin Li, Ronghua Shang 외 arxiv

Affordance reasoning in 3D Gaussian scenes aims to identify the region that supports the action specified by a given text instruction in complex environments. Existing methods typically cast this problem as one-shot pred…

Decision Making

AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment

2026-05-17 · Weijie Kong, Zhian Su, Wei Yu, Huixu Dong arxiv

Recent advances in Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation. However, the visual representations of most VLA models are often dominated by global object app…

Building an Affordances Map with Interactive Perception

2019-03-11 · Leni K. Le Goff, Oussama Yaakoubi, Alexandre Coninx, Stephane Doncieux

Robots need to understand their environment to perform their task. If it is possible to pre-program a visual scene analysis process in closed environments, robots operating in an open environment would benefit from the a…

General ClassificationScene Understanding

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

2026-02-10 · Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA 외 arxiv

3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordance knowledge from static cues such as lan…

Action UnderstandingPoint Clouds