Recognizing Actions in Videos from Unseen Viewpoints
Standard methods for video recognition use large CNNs designed to capture spatio-temporal data. However, training these models requires a large amount of labeled training data, containing a wide variety of actions, scenes, settings and camera viewpoints. In this paper, we show that current convolutional neural network models are unable to recognize actions from camera viewpoints not present in their training data (i.e., unseen view action recognition). To address this, we develop approaches based on 3D representations and introduce a new geometric convolutional layer that can learn viewpoint invariant representations. Further, we introduce a new, challenging dataset for unseen view recognition and show the approaches ability to learn viewpoint invariant representations.
Code (0)
등록된 구현이 없습니다.
Tasks
Action ClassificationAction RecognitionVideo RecognitionSimilar Papers 제목 키워드 기반
Learning a Deep Model for Human Action Recognition from Novel Viewpoints
Recognizing human actions from unknown and unseen (novel) views is a challenging problem. We propose a Robust Non-Linear Knowledge Transfer Model (R-NKTM) for human action recognition from novel views. The proposed R-NKT…
Action RecognitionTemporal Action LocalizationTransfer LearningVideo Action Recognition Using spatio-temporal optical flow video frames
Recognizing human actions based on videos has became one of the most popular areas of research in computer vision in recent years. This area has many applications such as surveillance, robotics, health care, video search…
Action RecognitionOptical Flow EstimationTemporal Action LocalizationChop & Learn: Recognizing and Generating Object-State Compositions
Recognizing and generating object-state compositions has been a challenging task, especially when generalizing to unseen compositions. In this paper, we study the task of cutting objects in different styles and the resul…
Action RecognitionImage GenerationObjectView-invariant action recognition
Human action recognition is an important problem in computer vision. It has a wide range of applications in surveillance, human-computer interaction, augmented reality, video indexing, and retrieval. The varying pattern …
Action RecognitionRetrievalTemporal Action LocalizationInferring Unseen Views of People
We pose unseen view synthesis as a probabilistic tensor completion problem. Given images of people organized by their rough viewpoint, we form a 3D appearance tensor indexed by images (pose examples), viewpoints, and ima…