Normalized Human Pose Features for Human Action Video Alignment
We present a novel approach for extracting human pose features from human action videos. The goal is to let the pose features capture only the poses of the action while being invariant to other factors, including video backgrounds, the video subject's anthropometric characteristics and viewpoints. Such human pose features facilitate the comparison of pose similarity and can be used for down-stream tasks, such as human action video alignment and pose retrieval. The key to our approach is to first normalize the poses in the video frames by retargeting the poses onto a pre-defined 3D skeleton to not only disentangle subject physical features, such as bone lengths and ratios, but also to unify global orientations of the poses. Then the normalized poses are mapped to a pose embedding space of high-level features, learned via unsupervised metric learning. We evaluate the effectiveness of our normalized features both qualitatively by visualizations, and quantitatively by a video alignment task on the Human3.6M dataset and an action recognition task on the Penn Action dataset.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionMetric LearningPose RetrievalRetrievalVideo AlignmentSimilar Papers 제목 키워드 기반
Evaluating Low-Level Speech Features Against Human Perceptual Data
We introduce a method for measuring the correspondence between low-level speech features and human perception, using a cognitive model of speech perception implemented directly on speech recordings. We evaluate two speak…
Automatic Speech Recognition (ASR)Representation LearningSpeech RecognitionvalidHuman Emotion Recognition Based On Galvanic Skin Response signal Feature Selection and SVM
A novel human emotion recognition method based on automatically selected Galvanic Skin Response (GSR) signal features and SVM is proposed in this paper. GSR signals were acquired by e-Health Sensor Platform V2.0. Then, t…
Emotion Recognitionfeature selectionReal-Time Facial Expression Recognition using Facial Landmarks and Neural Networks
This paper presents a lightweight algorithm for feature extraction, classification of seven different emotions, and facial expression recognition in a real-time manner based on static images of the human face. In this re…
Facial Expression RecognitionFacial Expression Recognition (FER)Facial Landmark DetectionSimultaneous Joint and Object Trajectory Templates for Human Activity Recognition from 3-D Data
The availability of low-cost range sensors and the development of relatively robust algorithms for the extraction of skeleton joint locations have inspired many researchers to develop human activity recognition methods u…
Activity RecognitionHuman Activity RecognitionHuman-Object Interaction DetectionBird Species Categorization Using Pose Normalized Deep Convolutional Nets
We propose an architecture for fine-grained visual categorization that approaches expert human performance in the classification of bird species. Our architecture first computes an estimate of the object's pose; this is …
ClassificationClusteringFine-Grained Visual CategorizationGeneral Classification