Audio Summarization with Audio Features and Probability Distribution Divergence
The automatic summarization of multimedia sources is an important task that facilitates the understanding of an individual by condensing the source while maintaining relevant information. In this paper we focus on audio summarization based on audio features and the probability of distribution divergence. Our method, based on an extractive summarization approach, aims to select the most relevant segments until a time threshold is reached. It takes into account the segment's length, position and informativeness value. Informativeness of each segment is obtained by mapping a set of audio features issued from its Mel-frequency Cepstral Coefficients and their corresponding Jensen-Shannon divergence score. Results over a multi-evaluator scheme shows that our approach provides understandable and informative summaries.
Code (0)
등록된 구현이 없습니다.
Tasks
Extractive SummarizationInformativenessPositionSimilar Papers 제목 키워드 기반
Role of Audio in Audio-Visual Video Summarization
Video summarization attracts attention for efficient video representation, retrieval, and browsing to ease volume and traffic surge problems. Although video summarization mostly uses the visual channel for compaction, th…
RetrievalVideo SummarizationAudiovisual Highlight Detection in Videos
In this paper, we test the hypothesis that interesting events in unstructured videos are inherently audiovisual. We combine deep image representations for object recognition and scene understanding with representations f…
Highlight DetectionObject RecognitionScene UnderstandingVideo SummarizationCFSum: A Transformer-Based Multi-Modal Video Summarization Framework With Coarse-Fine Fusion
Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today's video proliferation era. Multi…
Video SummarizationAudioVisual Video Summarization
Audio and vision are two main modalities in video data. Multimodal learning, especially for audiovisual learning, has drawn considerable attention recently, which can boost the performance of various computer vision task…
Video SummarizationImproving Post-Processing of Audio Event Detectors Using Reinforcement Learning
We apply post-processing to the class probability distribution outputs of audio event classification models and employ reinforcement learning to jointly discover the optimal parameters for various stages of a post-proces…
Classificationreinforcement-learningReinforcement LearningReinforcement Learning (RL)