Modeling Temporal Dynamics and Spatial Configurations of Actions Using Two-Stream Recurrent Neural Networks
Recently, skeleton based action recognition gains more popularity due to cost-effective depth sensors coupled with real-time skeleton estimation algorithms. Traditional approaches based on handcrafted features are limited to represent the complexity of motion patterns. Recent methods that use Recurrent Neural Networks (RNN) to handle raw skeletons only focus on the contextual dependency in the temporal domain and neglect the spatial configurations of articulated skeletons. In this paper, we propose a novel two-stream RNN architecture to model both temporal dynamics and spatial configurations for skeleton based action recognition. We explore two different structures for the temporal stream: stacked RNN and hierarchical RNN. Hierarchical RNN is designed according to human body kinematics. We also propose two effective methods to model the spatial structure by converting the spatial graph into a sequence of joints. To improve generalization of our model, we further exploit 3D transformation based data augmentation techniques including rotation and scaling transformation to transform the 3D coordinates of skeletons during training. Experiments on 3D action recognition benchmark datasets show that our method brings a considerable improvement for a variety of actions, i.e., generic actions, interaction activities and gestures.
Code (0)
등록된 구현이 없습니다.
Tasks
3D Action RecognitionAction RecognitionData AugmentationSkeleton Based Action RecognitionTemporal Action LocalizationSimilar Papers 제목 키워드 기반
Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition
Micro-actions are subtle, localized movements lasting 1-3 seconds such as scratching one's head or tapping fingers. Such subtle actions are essential for social communication, ubiquitously used in natural interactions, a…
Micro-Action RecognitionBeyond Appearance: Transformer-based Person Identification from Conversational Dynamics
This paper investigates the performance of transformer-based architectures for person identification in natural, face-to-face conversation scenario. We implement and evaluate a two-stream framework that separately models…
Person IdentificationTransfer LearningLearning Locally Interacting Discrete Dynamical Systems: Towards Data-Efficient and Scalable Prediction
Locally interacting dynamical systems, such as epidemic spread, rumor propagation through crowd, and forest fire, exhibit complex global dynamics originated from local, relatively simple, and often stochastic interaction…
StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation
Vision-language-action (VLA) models integrate visual observations and language instructions to predict robot actions, demonstrating promising generalization in manipulation tasks. However, most existing approaches primar…
Robot ManipulationSpatial ReasoningZigzag persistence for coral reef resilience using a stochastic spatial model
A complex interplay between species governs the evolution of spatial patterns in ecology. An open problem in the biological sciences is characterising spatio-temporal data and understanding how changes at the local scale…
Topological Data Analysis