paper-with-me

Papers

Learning Human Activities and Object Affordances from RGB-D Videos

2012-10-04 · Hema Swetha Koppula, Rudhir Gupta, Ashutosh Saxena

Understanding human activities and object affordances are two very important skills, especially for personal robots which operate in human environments. In this work, we consider the problem of extracting a descriptive labeling of the sequence of sub-activities being performed by a human, and more importantly, of their interactions with the objects in the form of associated affordances. Given a RGB-D video, we jointly model the human activities and object affordances as a Markov random field where the nodes represent objects and sub-activities, and the edges represent the relationships between object affordances, their relations with sub-activities, and their evolution over time. We formulate the learning problem using a structural support vector machine (SSVM) approach, where labelings over various alternate temporal segmentations are considered as latent variables. We tested our method on a challenging dataset comprising 120 activity videos collected from 4 subjects, and obtained an accuracy of 79.4% for affordance, 63.4% for sub-activity and 75.0% for high-level activity labeling. We then demonstrate the use of such descriptive labeling in performing assistive tasks by a PR2 robot.

📄 PDF Abstract BibTeX arXiv:1210.1207

Code (0)

등록된 구현이 없습니다.

Tasks

DescriptiveObjectSkeleton Based Action Recognition

Similar Papers 제목 키워드 기반

Predicting Human Activities Using Stochastic Grammar

2017-08-02 · ICCV 2017 10 · Siyuan Qi, Siyuan Huang, Ping Wei, Song-Chun Zhu

This paper presents a novel method to predict future human activities from partially observed RGB-D videos. Human activity prediction is generally difficult due to its non-Markovian property and the rich context between …

Activity Prediction

Learning Asynchronous and Sparse Human-Object Interaction in Videos

2021-03-03 · CVPR 2021 1 · Romero Morais, Vuong Le, Svetha Venkatesh, Truyen Tran

Human activities can be learned from video. With effective modeling it is possible to discover not only the action labels but also the temporal structures of the activities such as the progression of the sub-activities. …

Human-Object Interaction DetectionObject

Learning Spatio-Temporal Structure from RGB-D Videos for Human Activity Detection and Anticipation

2013-02-01 · Proceedings of Machine Learning Research volume 28 2013 2 · Hema S. Koppula, Ashutosh Saxena

We consider the problem of detecting past activities as well as anticipating which activity will happen in the future and how. We start by modeling the rich spatio-temporal relations between human poses and objects (call…

Action DetectionActivity DetectionSkeleton Based Action Recognition

Demo2Vec: Reasoning Object Affordances From Online Videos

2018-06-01 · CVPR 2018 6 · Kuan Fang, Te-Lin Wu, Daniel Yang, Silvio Savarese 외

Watching expert demonstrations is an important way for humans and robots to reason about affordances of unseen objects. In this paper, we consider the problem of reasoning object affordances through the feature embedding…

ObjectVideo-to-image Affordance Grounding

Mining Semantic Affordances of Visual Object Categories

2015-06-01 · CVPR 2015 6 · Yu-Wei Chao, Zhan Wang, Rada Mihalcea, Jia Deng

Affordances are fundamental attributes of objects. Affordances reveal the functionalities of objects and the possible actions that can be performed on them. Understanding affordances is crucial for recognizing human acti…

Collaborative FilteringObject