Joint-Partition Group Attention for skeleton-based action recognition
Skeleton-based action recognition aims to recognize human actions from the coordinates of human joints. By encoding coordinates as joint tokens, previous methods have successfully utilized the self-attention (SA) mechanism to capture the relationship of each pair of joints. However, the attention map generated from the joint Query and joint Key in SA only captures joint-to-joint correlations at a single granularity, which is obviously insufficient for human actions that express semantics in terms of body parts. In this paper, we argue that SA should have a more comprehensive mechanism to capture correlations in joint-to-joint and joint-to-partition patterns for a higher semantic representation of skeleton-based actions. Therefore, we propose Joint-Partition Group Attention (JPGA) to simultaneously capture correlations between joints and body parts of different granularity sizes. Specifically, JPGA integrates the joint tokens according to the joint’s human body partition attributes and produces different body parts tokens (partition-tokens) with different granularities. Then the attention map of JPGA is computed from joint-token and partition-token of different granularity sizes to represent the relationship between joints and body parts. To adaptively partition human body parts at different granularities, we apply the reparameterization trick to adaptively learn the multi-granularity partitioning matrix. Based on JPGA, we construct our Joint-Partition Former (JPFormer), and conduct extensive experiments on NTU-RGB+D, NTU-RRB+D 120, and Northwestern UCLA datasets and achieve state-of-the-art results, which highlights the effectiveness of our design and practice.
Code (1)
Tasks
Action RecognitionSkeleton Based Action RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition
Skeleton-based action recognition, which classifies human actions based on the coordinates of joints and their connectivity within skeleton data, is widely utilized in various scenarios. While Graph Convolutional Network…
Action RecognitionHuman Interaction RecognitionSkeleton Based Action RecognitionObject Activity Scene Description, Construction and Recognition
Action recognition is a critical task for social robots to meaningfully engage with their environment. 3D human skeleton-based action recognition is an attractive research area in recent years. Although, the existing app…
Action RecognitionGeneral ClassificationObjectSkeleton Based Action Recognition+3Skeleton-based Relational Reasoning for Group Activity Analysis
Research on group activity recognition mostly leans on the standard two-stream approach (RGB and Optical Flow) as their input features. Few have explored explicit pose information, with none using it directly to reason a…
Activity RecognitionGroup Activity RecognitionOptical Flow EstimationRelational ReasoningFineTec: Fine-Grained Action Recognition Under Temporal Corruption via Skeleton Decomposition and Sequence Completion
Recognizing fine-grained actions from temporally corrupted skeleton sequences remains a significant challenge, particularly in real-world scenarios where online pose estimation often yields substantial missing data. Exis…
Action RecognitionPose EstimationSpatiotemporal graph routing for skeleton-based action recognition
With the representation effectiveness, skeleton-based human action recognition has received considerable research attention, and has a wide range of real applications. In this area, many existing methods typically rely o…
Action RecognitionClusteringSkeleton Based Action RecognitionTemporal Action Localization