Co-Regularized Deep Representations for Video Summarization
Compact keyframe-based video summaries are a popular way of generating viewership on video sharing platforms. Yet, creating relevant and compelling summaries for arbitrarily long videos with a small number of keyframes is a challenging task. We propose a comprehensive keyframe-based summarization framework combining deep convolutional neural networks and restricted Boltzmann machines. An original co-regularization scheme is used to discover meaningful subject-scene associations. The resulting multimodal representations are then used to select highly-relevant keyframes. A comprehensive user study is conducted comparing our proposed method to a variety of schemes, including the summarization currently in use by one of the most popular video sharing websites. The results show that our method consistently outperforms the baseline schemes for any given amount of keyframes both in terms of attractiveness and informativeness. The lead is even more significant for smaller summaries.
Code (0)
등록된 구현이 없습니다.
Tasks
InformativenessVideo SummarizationSimilar Papers 제목 키워드 기반
Unsupervised Video Summarization With Adversarial LSTM Networks
This paper addresses the problem of unsupervised video summarization, formulated as selecting a sparse subset of video frames that optimally represent the input video. Our key idea is to learn a deep summarizer network t…
Unsupervised Video SummarizationVideo SummarizationGPT2MVS: Generative Pre-trained Transformer-2 for Multi-modal Video Summarization
Traditional video summarization methods generate fixed video representations regardless of user interest. Therefore such methods limit users' expectations in content search and exploration scenarios. Multi-modal video su…
Video SummarizationUse of Affective Visual Information for Summarization of Human-Centric Videos
Increasing volume of user-generated human-centric video content and their applications, such as video retrieval and browsing, require compact representations that are addressed by the video summarization literature. Curr…
Emotion RecognitionRetrievalSupervised Video SummarizationVideo Retrieval+1Relational Reasoning Over Spatial-Temporal Graphs for Video Summarization
In this paper, we propose a dynamic graph modeling approach to learn spatial-temporal representations for video summarization. Most existing video summarization methods extract image-level features with ImageNet pre-trai…
Graph ClassificationRelationRelational ReasoningSupervised Video Summarization+1Progressive Video Summarization via Multimodal Self-supervised Learning
Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-…
Self-Supervised LearningSupervised Video SummarizationVideo ClassificationVideo Summarization