paper-with-me

홈 › Papers

Skeleton Sequence and RGB Frame Based Multi-Modality Feature Fusion Network for Action Recognition

2022-02-23 · Xiaoguang Zhu, Ye Zhu, Haoyu Wang, Honglin Wen, Yan Yan, Peilin Liu

Action recognition has been a heated topic in computer vision for its wide application in vision systems. Previous approaches achieve improvement by fusing the modalities of the skeleton sequence and RGB video. However, such methods have a dilemma between the accuracy and efficiency for the high complexity of the RGB video network. To solve the problem, we propose a multi-modality feature fusion network to combine the modalities of the skeleton sequence and RGB frame instead of the RGB video, as the key information contained by the combination of skeleton sequence and RGB frame is close to that of the skeleton sequence and RGB video. In this way, the complementary information is retained while the complexity is reduced by a large margin. To better explore the correspondence of the two modalities, a two-stage fusion framework is introduced in the network. In the early fusion stage, we introduce a skeleton attention module that projects the skeleton sequence on the single RGB frame to help the RGB frame focus on the limb movement regions. In the late fusion stage, we propose a cross-attention module to fuse the skeleton feature and the RGB feature by exploiting the correlation. Experiments on two benchmarks NTU RGB+D and SYSU show that the proposed model achieves competitive performance compared with the state-of-the-art methods while reduces the complexity of the network.

📄 PDF Abstract BibTeX arXiv:2202.11374

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…

Similar Papers 제목 키워드 기반

Skeleton-Guided Spatial-Temporal Feature Learning for Video-Based Visible-Infrared Person Re-Identification

2024-11-17 · Wenjia Jiang, Xiaoke Zhu, Jiakang Gao, Di Liao

Video-based visible-infrared person re-identification (VVI-ReID) is challenging due to significant modality feature discrepancies. Spatial-temporal information in videos is crucial, but the accuracy of spatial-temporal i…

Person Re-Identification

Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion

2023-05-31 · Quoc-Huy Tran, Muhammad Ahmed, Murad Popattia, M. Hassan Ahmed 외

This paper presents a self-supervised temporal video alignment framework which is useful for several fine-grained human activity understanding applications. In contrast with the state-of-the-art method of CASA, where seq…

RetrievalSelf-Supervised LearningVideo Alignment

Explore Human Parsing Modality for Action Recognition

2024-01-04 · CAAI Transactions on Intelligence Technology 2023 7 · Jinfu Liu, Runwei Ding, Yuhang Wen, Nan Dai 외

Multimodal-based action recognition methods have achieved high success using pose and RGB modality. However, skeletons sequences lack appearance depiction and RGB images suffer irrelevant noise due to modality limitation…

Action RecognitionHuman Parsing

Action Segmentation Using 2D Skeleton Heatmaps and Multi-Modality Fusion

2023-09-12 · Syed Waleed Hyder, Muhammad Usama, Anas Zafar, Muhammad Naufil 외

This paper presents a 2D skeleton-based action segmentation method with applications in fine-grained human activity recognition. In contrast with state-of-the-art methods which directly take sequences of 3D skeleton coor…

Action SegmentationActivity RecognitionHuman Activity RecognitionSegmentation+1

Skeletons Speak Louder than Text: A Motion-Aware Pretraining Paradigm for Video-Based Person Re-Identification

2025-11-17 · Rifen Lin, Alex Jinpeng Wang, Jiawei Mo, Min Li arxiv

Multimodal pretraining has revolutionized visual understanding, but its impact on video-based person re-identification (ReID) remains underexplored. Existing approaches often rely on video-text pairs, yet suffer from two…

Person Re-IdentificationRepresentation LearningContrastive Learning