Two-stream Flow-guided Convolutional Attention Networks for Action Recognition
This paper proposes a two-stream flow-guided convolutional attention networks for action recognition in videos. The central idea is that optical flows, when properly compensated for the camera motion, can be used to guide attention to the human foreground. We thus develop cross-link layers from the temporal network (trained on flows) to the spatial network (trained on RGB frames). These cross-link layers guide the spatial-stream to pay more attention to the human foreground areas and be less affected by background clutter. We obtain promising performances with our approach on the UCF101, HMDB51 and Hollywood2 datasets.
Code (1)
Tasks
Action RecognitionAction Recognition In VideosTemporal Action LocalizationVocal Bursts Valence PredictionSimilar Papers 제목 키워드 기반
3D Convolutional with Attention for Action Recognition
Human action recognition is one of the challenging tasks in computer vision. The current action recognition methods use computationally expensive models for learning spatio-temporal dependencies of the action. Models uti…
Action RecognitionOptical Flow EstimationTemporal Action LocalizationA Heterogeneous Two-Stream Framework for Video Action Recognition with Comparative Fusion Analysis
Most two-stream action recognition networks apply the same convolutional backbone to both RGB and optical flow streams, ignoring the fact that the two modalities have fundamentally different structural properties. Optica…
Action RecognitionPose-Guided Graph Convolutional Networks for Skeleton-Based Action Recognition
Graph convolutional networks (GCNs), which can model the human body skeletons as spatial and temporal graphs, have shown remarkable potential in skeleton-based action recognition. However, in the existing GCN-based metho…
Action RecognitionSkeleton Based Action RecognitionTemporal Action LocalizationCIR-Net: Cross-modality Interaction and Refinement for RGB-D Salient Object Detection
Focusing on the issue of how to effectively capture and utilize cross-modality information in RGB-D salient object detection (SOD) task, we present a convolutional neural network (CNN) model, named CIR-Net, based on the …
Decoderobject-detectionObject DetectionRGB-D Salient Object Detection+1Attention-guided super-resolution of 4D flow MRI in carotid arteries
Four-dimensional (4D) flow magnetic resonance imaging (MRI) is a powerful non-invasive technique for visualizing and quantifying complex blood flow patterns in vivo. Despite its clinical promise, broader adoption is limi…