paper-with-me

Papers

Demo2Vec: Reasoning Object Affordances From Online Videos

2018-06-01 · CVPR 2018 6 · Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese, Joseph J. Lim

Watching expert demonstrations is an important way for humans and robots to reason about affordances of unseen objects. In this paper, we consider the problem of reasoning object affordances through the feature embedding of demonstration videos. We design the Demo2Vec model which learns to extract embedded vectors of demonstration videos and predicts the interaction region and the action label on a target image of the same object. We introduce the Online Product Review dataset for Affordance (OPRA) by collecting and labeling diverse YouTube product review videos. Our Demo2Vec model outperforms various recurrent neural network baselines on the collected dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ObjectVideo-to-image Affordance Grounding

Similar Papers 제목 키워드 기반

What Objects Enable, Not What They Are: Functional Latent Spaces for Affordance Reasoning

2026-06-04 · Rohan Siva, Neel P. Bhatt, Yunhao Yang, Seoyoung Lee 외 arxiv

Existing robot planning systems rely on appearance-based reasoning, where visual observations are encoded into latent spaces organized around object appearances (e.g., recognizing a "cart" based on how it looks). However…

Learning Human Activities and Object Affordances from RGB-D Videos

2012-10-04 · Hema Swetha Koppula, Rudhir Gupta, Ashutosh Saxena

Understanding human activities and object affordances are two very important skills, especially for personal robots which operate in human environments. In this work, we consider the problem of extracting a descriptive l…

DescriptiveObjectSkeleton Based Action Recognition

TEXT2AFFORD: Probing Object Affordance Prediction abilities of Language Models solely from Text

2024-02-20 · Sayantan Adak, Daivik Agrawal, Animesh Mukherjee, Somak Aditya

We investigate the knowledge of object affordances in pre-trained language models (LMs) and pre-trained Vision-Language models (VLMs). A growing body of literature shows that PTLMs fail inconsistently and non-intuitively…

Object

GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation

2025-05-17 · Teli Ma, Jia Zheng, Zifan Wang, Ziyao Gao 외

Learning manipulation skills from human demonstration videos offers a promising path toward generalizable and interpretable robotic intelligence-particularly through the lens of actionable affordances. However, transferr…

Benchmarking

Object-agnostic Affordance Categorization via Unsupervised Learning of Graph Embeddings

2023-03-30 · Alexia Toumpa, Anthony G. Cohn

Acquiring knowledge about object interactions and affordances can facilitate scene understanding and human-robot collaboration tasks. As humans tend to use objects in many different ways depending on the scene and the ob…

ObjectScene Understanding