paper-with-me

홈 › Papers

Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression

2017-03-03 · Jang-Won Lee, Michael S. Ryoo

We design a new approach that allows robot learning of new activities from unlabeled human example videos. Given videos of humans executing the same activity from a human's viewpoint (i.e., first-person videos), our objective is to make the robot learn the temporal structure of the activity as its future regression network, and learn to transfer such model for its own motor execution. We present a new deep learning model: We extend the state-of-the-art convolutional object detection network for the representation/estimation of human hands in training videos, and newly introduce the concept of using a fully convolutional network to regress (i.e., predict) the intermediate scene representation corresponding to the future frame (e.g., 1-2 seconds later). Combining these allows direct prediction of future locations of human hands and objects, which enables the robot to infer the motor control plan using our manipulation network. We experimentally confirm that our approach makes learning of robot activities from unlabeled human interaction videos possible, and demonstrate that our robot is able to execute the learned collaborative activities in real-time directly based on its camera input.

📄 PDF Abstract BibTeX arXiv:1703.01040

Code (0)

등록된 구현이 없습니다.

Tasks

object-detectionObject Detectionregression

Similar Papers 제목 키워드 기반

First-Person Activity Recognition: What Are They Doing to Me?

2013-06-01 · CVPR 2013 6 · Michael S. Ryoo, Larry Matthies

This paper discusses the problem of recognizing interaction-level human activities from a first-person viewpoint. The goal is to enable an observer (e.g., a robot or a wearable camera) to understand 'what activity others…

Activity RecognitionGeneral Classification

Early Recognition of Human Activities from First-Person Videos Using Onset Representations

2014-06-20 · M. S. Ryoo, Thomas J. Fuchs, Lu Xia, J. K. Aggarwal 외

In this paper, we propose a methodology for early recognition of human activities from videos taken with a first-person viewpoint. Early recognition, which is also known as activity prediction, is an ability to infer an …

Activity PredictionPerson RecognitionTime SeriesTime Series Analysis

Learning Human Activities and Object Affordances from RGB-D Videos

2012-10-04 · Hema Swetha Koppula, Rudhir Gupta, Ashutosh Saxena

Understanding human activities and object affordances are two very important skills, especially for personal robots which operate in human environments. In this work, we consider the problem of extracting a descriptive l…

DescriptiveObjectSkeleton Based Action Recognition

A Mobile Robot Generating Video Summaries of Seniors' Indoor Activities

2019-01-30 · Chih-Yuan Yang, Heeseung Yun, Srenavis Varadaraj, Jane Yung-jen Hsu

We develop a system which generates summaries from seniors' indoor-activity videos captured by a social robot to help remote family members know their seniors' daily activities at home. Unlike the traditional video summa…

Action RecognitionHuman DetectionPerson IdentificationPose Estimation+1

Is Tracking really more challenging in First Person Egocentric Vision?

2025-07-21 · Matteo Dunnhofer, Zaira Manigrasso, Christian Micheloni arxiv

Visual object tracking and segmentation are becoming fundamental tasks for understanding human activities in egocentric vision. Recent research has benchmarked state-of-the-art methods and concluded that first person ego…

Visual Object Tracking