paper-with-me

Papers

TSI: Temporal Saliency Integration for Video Action Recognition

2021-06-02 · Haisheng Su, Jinyuan Feng, Dongliang Wang, Weihao Gan, Wei Wu, Yu Qiao

Efficient spatiotemporal modeling is an important yet challenging problem for video action recognition. Existing state-of-the-art methods exploit motion clues to assist in short-term temporal modeling through temporal difference over consecutive frames. However, insignificant noises will be inevitably introduced due to the camera movement. Besides, movements of different actions can vary greatly. In this paper, we propose a Temporal Saliency Integration (TSI) block, which mainly contains a Salient Motion Excitation (SME) module and a Cross-scale Temporal Integration (CTI) module. Specifically, SME aims to highlight the motion-sensitive area through local-global motion modeling, where the saliency alignment and pyramidal feature difference are conducted successively between neighboring frames to capture motion dynamics with less noises caused by misaligned background. CTI is designed to perform multi-scale temporal modeling through a group of separate 1D convolutions respectively. Meanwhile, temporal interactions across different scales are integrated with attention mechanism. Through these two modules, long short-term temporal relationships can be encoded efficiently by introducing limited additional parameters. Extensive experiments are conducted on several popular benchmarks (i.e., Something-Something V1 & V2, Kinetics-400, UCF-101, and HMDB-51), which demonstrate the effectiveness and superiority of our proposed method.

📄 PDF Abstract BibTeX arXiv:2106.01088

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Similar Papers 제목 키워드 기반

Temporal-Spatial Feature Pyramid for Video Saliency Detection

2021-05-10 · Qinyao Chang, Shiping Zhu

Multi-level features are important for saliency detection. Better combination and use of multi-level features with time information can greatly improve the accuracy of the video saliency model. In order to fully combine …

DecoderSaliency DetectionVideo Saliency Detection

Dynamically Encoded Actions Based on Spacetime Saliency

2015-06-01 · CVPR 2015 6 · Christoph Feichtenhofer, Axel Pinz, Richard P. Wildes

Human actions typically occur over a well localized extent in both space and time. Similarly, as typically captured in video, human actions have small spatiotemporal support in image space. This paper capitalizes on thes…

Action RecognitionTemporal Action Localization

Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition

2022-07-20 · Huabin Liu, Weixian Lv, John See, Weiyao Lin

A primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algorithms at the feature level while little a…

Action RecognitionFew-Shot action recognitionFew Shot Action Recognition

Temporal Saliency Query Network for Efficient Video Recognition

2022-07-21 · Boyang xia, Zhihao Wang, Wenhao Wu, Haoran Wang 외

Efficient video recognition is a hot-spot research topic with the explosive growth of multimedia data on the Internet and mobile devices. Most existing methods select the salient frames without awareness of the class-spe…

Action RecognitionVideo Recognition

Convolutions Need Registers Too: HVS-Inspired Dynamic Attention for Video Quality Assessment

2026-01-16 · Mayesha Maliha R. Mithila, Mylene C. Q. Farias arxiv

No-reference video quality assessment (NR-VQA) estimates perceptual quality without a reference video, which is often challenging. While recent techniques leverage saliency or transformer attention, they merely address g…

Video Quality AssessmentSaliency Prediction