Coding Kendall's Shape Trajectories for 3D Action Recognition
Suitable shape representations as well as their temporal evolution, termed trajectories, often lie to non-linear manifolds. This puts an additional constraint (i.e., non-linearity) in using conventional machine learning techniques for the purpose of classification, event detection, prediction, etc. This paper accommodates the well-known Sparse Coding and Dictionary Learning to the Kendall's shape space and illustrates effective coding of 3D skeletal sequences for action recognition. Grounding on the Riemannian geometry of the shape space, an intrinsic sparse coding and dictionary learning formulation is proposed for static skeletal shapes to overcome the inherent non-linearity of the manifold. As a main result, initial trajectories give rise to sparse code functions with suitable computational properties, including sparsity and vector space representation. To achieve action recognition, two different classification schemes were adopted. A bi-directional LSTM is directly performed on sparse code functions, while a linear SVM is applied after representing sparse code functions using Fourier temporal pyramid. Experiments conducted on three publicly available datasets show the superiority of the proposed approach compared to existing Riemannian representations and its competitiveness with respect to other recently-proposed approaches. When the benefits of invariance are maintained from the Kendall's shape representation, our approach not only overcomes the problem of non-linearity but also yields to discriminative sparse code functions.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Action RecognitionAction RecognitionDictionary LearningEvent DetectionGeneral ClassificationTemporal Action LocalizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Sparse Coding of Shape Trajectories for Facial Expression and Action Recognition
The detection and tracking of human landmarks in video streams has gained in reliability partly due to the availability of affordable RGB-D sensors. The analysis of such time-varying geometric data is playing an importan…
Action RecognitionDictionary LearningMicro Expression RecognitionMicro-Expression Recognition+2KShapeNet: Riemannian network on Kendall shape space for Skeleton based Action Recognition
Deep Learning architectures, albeit successful in most computer vision tasks, were designed for data with an underlying Euclidean structure, which is not usually fulfilled since pre-processed data may lie on a non-linear…
Action RecognitionDeep LearningSkeleton Based Action RecognitionMoving Object Segmentation in Jittery Videos by Stabilizing Trajectories Modeled in Kendall's Shape Space
Moving Object Segmentation is a challenging task for jittery/wobbly videos. For jittery videos, the non-smooth camera motion makes discrimination between foreground objects and background layers hard to solve. While most…
ClusteringObjectSegmentationSemantic Segmentation+2Geometric Deep Neural Network Using Rigid and Non-Rigid Transformations for Human Action Recognition
Deep Learning architectures, albeit successful in mostcomputer vision tasks, were designed for data with an un-derlying Euclidean structure, which is not usually fulfilledsince pre-processed data may lie on a non-lin…
Action RecognitionDeep LearningSkeleton Based Action RecognitionTemporal Action LocalizationAn Elastic Shape Variational Autoencoder for Skeleton Pose Trajectories
Deep generative models provide flexible frameworks for modeling complex, structured data such as images, videos, 3D objects, and texts. However, when applied to sequences of human skeletons, standard variational autoenco…
Action Recognition