paper-with-me

홈 › Papers

SCSampler: Sampling Salient Clips from Video for Efficient Action Recognition

2019-04-08 · ICCV 2019 10 · Bruno Korbar, Du Tran, Lorenzo Torresani

While many action recognition datasets consist of collections of brief, trimmed videos each containing a relevant action, videos in the real-world (e.g., on YouTube) exhibit very different properties: they are often several minutes long, where brief relevant clips are often interleaved with segments of extended duration containing little change. Applying densely an action recognition system to every temporal clip within such videos is prohibitively expensive. Furthermore, as we show in our experiments, this results in suboptimal recognition accuracy as informative predictions from relevant clips are outnumbered by meaningless classification outputs over long uninformative sections of the video. In this paper we introduce a lightweight "clip-sampling" model that can efficiently identify the most salient temporal clips within a long video. We demonstrate that the computational cost of action recognition on untrimmed videos can be dramatically reduced by invoking recognition only on these most salient clips. Furthermore, we show that this yields significant gains in recognition accuracy compared to analysis of all clips or randomly/uniformly selected clips. On Sports1M, our clip sampling scheme elevates the accuracy of an already state-of-the-art action classifier by 7% and reduces by more than 15 times its computational cost.

📄 PDF Abstract BibTeX arXiv:1904.04289

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionTemporal Action Localization

Methods 이 논문이 사용한 방법론

Average Pooling 설명 없음
Residual Connection 설명 없음
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Batch Normalization 설명 없음
Bottleneck Residual Block A Bottleneck Residual Block is a variant of the residual block that utilises 1x1 convolutions to create a bottleneck. The…
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…

Similar Papers 제목 키워드 기반

TimeGate: Conditional Gating of Segments in Long-range Activities

2020-04-03 · Noureldien Hussein, Mihir Jain, Babak Ehteshami Bejnordi

When recognizing a long-range activity, exploring the entire video is exhaustive and computationally expensive, as it can span up to a few minutes. Thus, it is of great importance to sample only the salient parts of the …

An Image is Worth 16x16 Words, What is a Video Worth?

2021-03-25 · Gilad Sharir, Asaf Noy, Lihi Zelnik-Manor

Leading methods in the domain of action recognition try to distill information from both the spatial and temporal dimensions of an input video. Methods that reach State of the Art (SotA) accuracy, usually make use of 3D …

Action ClassificationAction Recognition

Dynamic Sampling Networks for Efficient Action Recognition in Videos

2020-06-28 · Yin-Dong Zheng, Zhao-Yang Liu, Tong Lu, Li-Min Wang

The existing action recognition methods are mainly based on clip-level classifiers such as two-stream CNNs or 3D CNNs, which are trained from the randomly selected clips and applied to densely sampled clips during testin…

Action RecognitionAction Recognition In Videos

Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering

2021-12-01 · NeurIPS 2021 12 · Weijiang Yu, Haoteng Zheng, Mengfei Li, Lei Ji 외

Recent advances in the video question answering (i.e., VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via s…

Multimodal ReasoningQuestion AnsweringVideo Question Answering

Revisiting Kernel Temporal Segmentation as an Adaptive Tokenizer for Long-form Video Understanding

2023-09-20 · Mohamed Afham, Satya Narayan Shukla, Omid Poursaeed, Pengchuan Zhang 외

While most modern video understanding models operate on short-range clips, real-world videos are often several minutes long with semantically consistent segments of variable length. A common approach to process long vide…

Action LocalizationFormTemporal Action LocalizationVideo Classification+1