Skeletonnet: Mining deep part features for 3-d action recognition
This letter presents SkeletonNet, a deep learning framework for skeleton-based 3-D action recognition. Given a skeleton sequence, the spatial structure of the skeleton joints in each frame and the temporal information between multiple frames are two important factors for action recognition. We first extract body-part-based features from each frame of the skeleton sequence. Compared to the original coordinates of the skeleton joints, the proposed features are translation, rotation, and scale invariant. To learn robust temporal information, instead of treating the features of all frames as a time series, we transform the features into images and feed them to the proposed deep learning network, which contains two parts: one to extract general features from the input images, while the other to generate a discriminative and compact representation for action recognition. The proposed method is tested on the SBU kinect interaction dataset, the CMU dataset, and the large-scale NTU RGB+D dataset and achieves state-of-the-art performance.
Code (0)
등록된 구현이 없습니다.
Tasks
Action RecognitionDeep LearningSkeleton Based Action RecognitionTime SeriesTime Series AnalysisTranslationSimilar Papers 제목 키워드 기반
SkeletonNet: A Topology-Preserving Solution for Learning Mesh Reconstruction of Object Surfaces from RGB Images
This paper focuses on the challenging task of learning 3D object surface reconstructions from RGB images. Existingmethods achieve varying degrees of success by using different surface representations. However, they all h…
Surface ReconstructionSkeleton Transformer Networks: 3D Human Pose and Skinned Mesh from Single RGB Image
In this paper, we present Skeleton Transformer Networks (SkeletonNet), an end-to-end framework that can predict not only 3D joint positions but also 3D angular pose (bone rotations) of a human skeleton from a single colo…
3D Human Pose EstimationMining Mid-level Features for Action Recognition Based on Effective Skeleton Representation
Recently, mid-level features have shown promising performance in computer vision. Mid-level features learned by incorporating class-level information are potentially more discriminative than traditional low-level local f…
3D Action RecognitionAction RecognitionTemporal Action LocalizationInteraction Part Mining: A Mid-Level Approach for Fine-Grained Action Recognition
Modeling human-object interactions and manipulating motions lies in the heart of fine-grained action recognition. Previous methods heavily rely on explicit detection of the object being interacted, which requires intensi…
Action RecognitionFine-grained Action RecognitionHuman-Object Interaction DetectionObject+1Motion Part Regularization: Improving Action Recognition via Trajectory Selection
Dense local motion features such as dense trajectories have been widely used in action recognition. For most actions, only a few local features (e.g., critical movements of the hand, arm, leg etc.) are responsible to the…
Action RecognitionSentenceTemporal Action Localizationtext-classification+1