paper-with-me

홈 › Papers

Joint Network based Attention for Action Recognition

2016-11-16 · Yemin Shi, Yonghong Tian, Yao-Wei Wang, Tiejun Huang

By extracting spatial and temporal characteristics in one network, the two-stream ConvNets can achieve the state-of-the-art performance in action recognition. However, such a framework typically suffers from the separately processing of spatial and temporal information between the two standalone streams and is hard to capture long-term temporal dependence of an action. More importantly, it is incapable of finding the salient portions of an action, say, the frames that are the most discriminative to identify the action. To address these problems, a \textbf{j}oint \textbf{n}etwork based \textbf{a}ttention (JNA) is proposed in this study. We find that the fully-connected fusion, branch selection and spatial attention mechanism are totally infeasible for action recognition. Thus in our joint network, the spatial and temporal branches share some information during the training stage. We also introduce an attention mechanism on the temporal domain to capture the long-term dependence meanwhile finding the salient portions. Extensive experiments are conducted on two benchmark datasets, UCF101 and HMDB51. Experimental results show that our method can improve the action recognition performance significantly and achieves the state-of-the-art results on both datasets.

📄 PDF Abstract BibTeX arXiv:1611.05215

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Global Context-Aware Attention LSTM Networks for 3D Action Recognition

2017-07-01 · CVPR 2017 7 · Jun Liu, Gang Wang, Ping Hu, Ling-Yu Duan 외

Long Short-Term Memory (LSTM) networks have shown superior performance in 3D human action recognition due to their power in modeling the dynamics and dependencies in sequential data. Since not all joints are informative …

3D Action RecognitionAction AnalysisAction RecognitionOne-Shot 3D Action Recognition+2

Cross-Modal Learning with 3D Deformable Attention for Action Recognition

2022-12-12 · ICCV 2023 1 · Sangwon Kim, Dasom Ahn, Byoung Chul Ko

An important challenge in vision-based action recognition is the embedding of spatiotemporal features with two or more heterogeneous modalities into a single feature. In this study, we propose a new 3D deformable transfo…

Action Recognition

Joint Image-Instance Spatial-Temporal Attention for Few-shot Action Recognition

2025-03-11 · Zefeng Qian, Chongyang Zhang, Yifei HUANG, Gang Wang 외

Few-shot Action Recognition (FSAR) constitutes a crucial challenge in computer vision, entailing the recognition of actions from a limited set of examples. Recent approaches mainly focus on employing image-level features…

Action RecognitionFew-Shot action recognitionFew Shot Action Recognition

Skeleton-Based Human Action Recognition with Global Context-Aware Attention LSTM Networks

2017-07-18 · Jun Liu, Gang Wang, Ling-Yu Duan, Kamila Abdiyeva 외

Human action recognition in 3D skeleton sequences has attracted a lot of research attention. Recently, Long Short-Term Memory (LSTM) networks have shown promising performance in this task due to their strengths in modeli…

Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Context-Aware Cross-Attention for Skeleton-Based Human Action Recognition

2020-01-20 · IEEE Access ( Volume: 8 ) 2020 1 · Yanbo Fan, Shuchen Weng, Yong Zhang, Boxin Shi 외

Skeleton-based human action recognition is becoming popular due to its computational efficiency and robustness. Since not all skeleton joints are informative for action recognition, attention mechanisms are adopted to ex…

Action RecognitionComputational EfficiencySkeleton Based Action RecognitionTemporal Action Localization