paper-with-me

Papers

Unsupervised Transcript-assisted Video Summarization and Highlight Detection

2025-05-29 · Spyros Barbakos, Charalampos Antoniadis, Gerasimos Potamianos, Gianluca Setti

Video consumption is a key part of daily life, but watching entire videos can be tedious. To address this, researchers have explored video summarization and highlight detection to identify key video segments. While some works combine video frames and transcripts, and others tackle video summarization and highlight detection using Reinforcement Learning (RL), no existing work, to the best of our knowledge, integrates both modalities within an RL framework. In this paper, we propose a multimodal pipeline that leverages video frames and their corresponding transcripts to generate a more condensed version of the video and detect highlights using a modality fusion mechanism. The pipeline is trained within an RL framework, which rewards the model for generating diverse and representative summaries while ensuring the inclusion of video segments with meaningful transcript content. The unsupervised nature of the training allows for learning from large-scale unannotated datasets, overcoming the challenge posed by the limited size of existing annotated datasets. Our experiments show that using the transcript in video summarization and highlight detection achieves superior results compared to relying solely on the visual content of the video.

📄 PDF Abstract BibTeX arXiv:2505.23268

Code (0)

등록된 구현이 없습니다.

Tasks

Highlight DetectionReinforcement Learning (RL)Video Summarization

Similar Papers 제목 키워드 기반

VT-SSum: A Benchmark Dataset for Video Transcript Segmentation and Summarization

2021-06-10 · Tengchao Lv, Lei Cui, Momcilo Vasilijevic, Furu Wei

Video transcript summarization is a fundamental task for video understanding. Conventional approaches for transcript summarization are usually built upon the summarization data for written language such as news articles,…

ArticlesSegmentationText SummarizationVideo Understanding

Learning Summary-Worthy Visual Representation for Abstractive Summarization in Video

2023-05-08 · Zenan Xu, Xiaojun Meng, Yasheng Wang, Qinliang Su 외

Multimodal abstractive summarization for videos (MAS) requires generating a concise textual summary to describe the highlights of a video according to multimodal resources, in our case, the video content and its transcri…

Abstractive Text SummarizationLanguage ModelingLanguage Modelling

Unsupervised Broadcast News Summarization; a comparative study on Maximal Marginal Relevance (MMR) and Latent Semantic Analysis (LSA)

2023-01-05 · Majid Ramezani, Mohammad-Salar Shahryari, Amir-Reza Feizi-Derakhshi, Mohammad-Reza Feizi-Derakhshi

The methods of automatic speech summarization are classified into two groups: supervised and unsupervised methods. Supervised methods are based on a set of features, while unsupervised methods perform summarization based…

News Summarization

SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets

2025-10-07 · Manolis Mylonas, Charalampia Zerva, Evlampios Apostolidis, Vasileios Mezaris arxiv

In this work, we present a method and two large-scale datasets for Script-Driven Multimodal Video Summarization. The proposed method, SD-MVSum, builds on our earlier SD-VSum method for script-driven video summarization, …

Semantic SimilarityVideo Summarization

Unsupervised Video Summarization via Multi-source Features

2021-05-26 · Hussain Kanafani, Junaid Ahmed Ghauri, Sherzod Hakimov, Ralph Ewerth

Video summarization aims at generating a compact yet representative visual summary that conveys the essence of the original video. The advantage of unsupervised approaches is that they do not require human annotations to…

Unsupervised Video SummarizationVideo Summarization