Fusing Higher-order Features in Graph Neural Networks for Skeleton-based Action Recognition
Skeleton sequences are lightweight and compact, and thus are ideal candidates for action recognition on edge devices. Recent skeleton-based action recognition methods extract features from 3D joint coordinates as spatial-temporal cues, using these representations in a graph neural network for feature fusion to boost recognition performance. The use of first- and second-order features, i.e., joint and bone representations, has led to high accuracy. Nonetheless, many models are still confused by actions that have similar motion trajectories. To address these issues, we propose fusing higher-order features in the form of angular encoding into modern architectures to robustly capture the relationships between joints and body parts. This simple fusion with popular spatial-temporal graph neural networks achieves new state-of-the-art accuracy in two large benchmarks, including NTU60 and NTU120, while employing fewer parameters and reduced run time. Our source code is publicly available at: https://github.com/ZhenyueQin/Angular-Skeleton-Encoding.
Code (1)
Tasks
Action RecognitionGraph Neural NetworkSkeleton Based Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Towards To-a-T Spatio-Temporal Focus for Skeleton-Based Action Recognition
Graph Convolutional Networks (GCNs) have been widely used to model the high-order dynamic dependencies for skeleton-based action recognition. Most existing approaches do not explicitly embed the high-order spatio-tempora…
Action RecognitionSkeleton Based Action RecognitionActional-Structural Graph Convolutional Networks for Skeleton-based Action Recognition
Action recognition with skeleton data has recently attracted much attention in computer vision. Previous studies are mostly based on fixed skeleton graphs, only capturing local physical dependencies among joints, which m…
Action RecognitionDecoderPose PredictionSkeleton Based Action Recognition+1Skeleton-Parted Graph Scattering Networks for 3D Human Motion Prediction
Graph convolutional network based methods that model the body-joints' relations, have recently shown great promise in 3D skeleton-based human motion prediction. However, these methods have two critical issues: first, dee…
Human motion predictionmotion predictionLearning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion
This paper presents a self-supervised temporal video alignment framework which is useful for several fine-grained human activity understanding applications. In contrast with the state-of-the-art method of CASA, where seq…
RetrievalSelf-Supervised LearningVideo AlignmentImproving Skeleton-based Action Recognition with Interactive Object Information
Human skeleton information is important in skeleton-based action recognition, which provides a simple and efficient way to describe human pose. However, existing skeleton-based methods focus more on the skeleton, ignorin…
Action RecognitionData Augmentationgraph constructionObject+1