paper-with-me

홈 › Papers

NSNet: Non-saliency Suppression Sampler for Efficient Video Recognition

2022-07-21 · Boyang xia, Wenhao Wu, Haoran Wang, Rui Su, Dongliang He, Haosen Yang, Xiaoran Fan, Wanli Ouyang

It is challenging for artificial intelligence systems to achieve accurate video recognition under the scenario of low computation costs. Adaptive inference based efficient video recognition methods typically preview videos and focus on salient parts to reduce computation costs. Most existing works focus on complex networks learning with video classification based objectives. Taking all frames as positive samples, few of them pay attention to the discrimination between positive samples (salient frames) and negative samples (non-salient frames) in supervisions. To fill this gap, in this paper, we propose a novel Non-saliency Suppression Network (NSNet), which effectively suppresses the responses of non-salient frames. Specifically, on the frame level, effective pseudo labels that can distinguish between salient and non-salient frames are generated to guide the frame saliency learning. On the video level, a temporal attention module is learned under dual video-level supervisions on both the salient and the non-salient representations. Saliency measurements from both two levels are combined for exploitation of multi-granularity complementary information. Extensive experiments conducted on four well-known benchmarks verify our NSNet not only achieves the state-of-the-art accuracy-efficiency trade-off but also present a significantly faster (2.4~4.3x) practical inference speed than state-of-the-art methods. Our project page is at https://lawrencexia2008.github.io/projects/nsnet .

📄 PDF Abstract BibTeX arXiv:2207.10388

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionVideo ClassificationVideo Recognition

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Task-adaptive Spatial-Temporal Video Sampler for Few-shot Action Recognition

2022-07-20 · Huabin Liu, Weixian Lv, John See, Weiyao Lin

A primary challenge faced in few-shot action recognition is inadequate video data for training. To address this issue, current methods in this field mainly focus on devising algorithms at the feature level while little a…

Action RecognitionFew-Shot action recognitionFew Shot Action Recognition

Dynamic nsNet2: Efficient Deep Noise Suppression with Early Exiting

2023-08-31 · Riccardo Miccini, Alaa Zniber, Clément Laroche, Tobias Piechowiak 외

Although deep learning has made strides in the field of deep noise suppression, leveraging deep architectures on resource-constrained devices still proved challenging. Therefore, we present an early-exiting model based o…

TransNet: A Transfer Learning-Based Network for Human Action Recognition

2023-09-13 · K. Alomar, X. Cai

Human action recognition (HAR) is a high-level and significant research area in computer vision due to its ubiquitous applications. The main limitations of the current HAR models are their complex structures and lengthy …

Action RecognitionTemporal Action LocalizationTransfer Learning

Poisoning Prompt-Guided Sampling in Video Large Language Models

2025-09-25 · Yuxin Cao, Wei Song, Jingling Xue, Jin Song Dong arxiv

Video Large Language Models (VideoLLMs) are increasingly deployed as automated moderators on user-generated video platforms, where a few unwatched seconds of harmful footage are enough to suppress a safety alert. Because…

Question Answering

MGSampler: An Explainable Sampling Strategy for Video Action Recognition

2021-04-20 · ICCV 2021 10 · Yuan Zhi, Zhan Tong, LiMin Wang, Gangshan Wu

Frame sampling is a fundamental problem in video action recognition due to the essential redundancy in time and limited computation resources. The existing sampling strategy often employs a fixed frame selection and lack…

Action RecognitionTemporal Action Localization