Papers Supervised Video Summarization
“Supervised Video Summarization” 태그가 달린 논문 28편 · 필터 해제
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
The increasing ubiquity of video content and the corresponding demand for efficient access to meaningful information have elevated video summarization and video highlights as a vital research area. However, many state-of…
Self-Supervised LearningSupervised Video SummarizationVideo SummarizationFullTransNet: Full Transformer with Local-Global Attention for Video Summarization
Video summarization mainly aims to produce a compact, short, informative, and representative synopsis of raw videos, which is of great importance for browsing, analyzing, and understanding video content. Dominant video s…
DecoderSupervised Video SummarizationVideo SummarizationCSTA: CNN-based Spatiotemporal Attention for Video Summarization
Video summarization aims to generate a concise representation of a video, capturing its essential content and key moments while reducing its overall length. Although several methods employ attention mechanisms to handle …
Supervised Video SummarizationVideo SummarizationLanguage-Guided Self-Supervised Video Summarization Using Text Semantic Matching Considering the Diversity of the Video
Current video summarization methods rely heavily on supervised computer vision techniques, which demands time-consuming and subjective manual annotations. To overcome these limitations, we investigated self-supervised vi…
DiversitySupervised Video SummarizationVideo SummarizationAlign and Attend: Multimodal Summarization with Dual Contrastive Losses
The goal of multimodal summarization is to extract the most important information from different modalities to form output summaries. Unlike the unimodal summarization, the multimodal summarization task explicitly levera…
Extractive Text SummarizationSupervised Video SummarizationVideo SummarizationRelational Reasoning Over Spatial-Temporal Graphs for Video Summarization
In this paper, we propose a dynamic graph modeling approach to learn spatial-temporal representations for video summarization. Most existing video summarization methods extract image-level features with ImageNet pre-trai…
Graph ClassificationRelationRelational ReasoningSupervised Video Summarization+1Progressive Video Summarization via Multimodal Self-supervised Learning
Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-…
Self-Supervised LearningSupervised Video SummarizationVideo ClassificationVideo SummarizationJoint Video Summarization and Moment Localization by Cross-Task Sample Transfer
Video summarization has recently engaged increasing attention in computer vision communities. However, the scarcity of annotated data has been a key obstacle in this task. To address it, this work explores a new solu…
Supervised Video SummarizationVideo SummarizationVideo Joint Modelling Based on Hierarchical Transformer for Co-summarization
Video summarization aims to automatically generate a summary (storyboard or video skim) of a video, which can facilitate large-scale video retrieval and browsing. Most of the existing methods perform video summarization …
RetrievalSupervised Video SummarizationVideo RetrievalVideo Summarization+1Combining Global and Local Attention with Positional Encoding for Video Summarization
This paper presents a new method for supervised video summarization. To overcome drawbacks of existing RNN-based summarization architectures, that relate to the modeling of long-range frames' dependencies and the ability…
Supervised Video SummarizationVideo SummarizationA Stacking Ensemble Approach for Supervised Video Summarization
Video summarization methods are usually classified into shot-level or frame-level methods, which are individually used in a general way. This paper investigates the underlying complementarity between the frame-level and …
Supervised Video SummarizationVideo SummarizationHierarchical Multimodal Transformer to Summarize Videos
Although video summarization has achieved tremendous success benefiting from Recurrent Neural Networks (RNN), RNN-based methods neglect the global dependencies and multi-hop relationships among video frames, which limits…
Machine TranslationSupervised Video SummarizationTranslationVideo Captioning+1Use of Affective Visual Information for Summarization of Human-Centric Videos
Increasing volume of user-generated human-centric video content and their applications, such as video retrieval and browsing, require compact representations that are addressed by the video summarization literature. Curr…
Emotion RecognitionRetrievalSupervised Video SummarizationVideo Retrieval+1CLIP-It! Language-Guided Video Summarization
A generic video summary is an abridged version of a video that conveys the whole story and features the most important scenes. Yet the importance of scenes in a video is often subjective, and users should have the option…
Query-focused SummarizationQuery focused video summarizationSupervised Video SummarizationVideo SummarizationSelf-Attention Recurrent Summarization Network with Reinforcement Learning for Video Summarization Task
With the exponential growth of video data, video summarization techniques are urgently needed for reducing people’s efforts in the videos' content exploration by generating succinct but informative summaries from origina…
reinforcement-learningReinforcement LearningSupervised Video SummarizationUnsupervised Video Summarization+1Supervised Video Summarization via Multiple Feature Sets with Parallel Attention
The assignment of importance scores to particular frames or (short) segments in a video is crucial for summarization, but also a difficult task. Previous work utilizes only one source of visual features. In this paper, w…
Automated Feature Engineeringimage-classificationMultimodal Deep LearningSupervised Video Summarization+1How Good is a Video Summary? A New Benchmarking Dataset and Evaluation Framework Towards Realistic Video Summarization
Automatic video summarization is still an unsolved problem due to several challenges. The currently available datasets either have very short videos or have few long videos of only a particular type. We introduce a new b…
BenchmarkingSupervised Video SummarizationVideo SummarizationDSNet: A Flexible Detect-to-Summarize Network for Video Summarization
In this paper, we propose a Detect-to-Summarize network (DSNet) framework for supervised video summarization. Our DSNet contains anchor-based and anchor-free counterparts. The anchor-based method generates temporal inter…
regressionSupervised Video SummarizationVideo SummarizationQuery Twice: Dual Mixture Attention Meta Learning for Video Summarization
Video summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax function. However, softmax function suffers…
Meta-LearningSupervised Video SummarizationVideo SummarizationWeakly Supervised Video Summarization by Hierarchical Reinforcement Learning
Conventional video summarization approaches based on reinforcement learning have the problem that the reward can only be received after the whole summary is generated. Such kind of reward is sparse and it makes reinforce…
Hierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+2