paper-with-me

Papers

Visual Subtitle Feature Enhanced Video Outline Generation

2022-08-24 · Qi Lv, Ziqiang Cao, Wenrui Xie, Derui Wang, Jingwen Wang, Zhiwei Hu, Tangkun Zhang, Ba Yuan, Yuanhang Li, Min Cao, Wenjie Li, Sujian Li, Guohong Fu

With the tremendously increasing number of videos, there is a great demand for techniques that help people quickly navigate to the video segments they are interested in. However, current works on video understanding mainly focus on video content summarization, while little effort has been made to explore the structure of a video. Inspired by textual outline generation, we introduce a novel video understanding task, namely video outline generation (VOG). This task is defined to contain two sub-tasks: (1) first segmenting the video according to the content structure and then (2) generating a heading for each segment. To learn and evaluate VOG, we annotate a 10k+ dataset, called DuVOG. Specifically, we use OCR tools to recognize subtitles of videos. Then annotators are asked to divide subtitles into chapters and title each chapter. In videos, highlighted text tends to be the headline since it is more likely to attract attention. Therefore we propose a Visual Subtitle feature Enhanced video outline generation model (VSENet) which takes as input the textual subtitles together with their visual font sizes and positions. We consider the VOG task as a sequence tagging problem that extracts spans where the headings are located and then rewrites them to form the final outlines. Furthermore, based on the similarity between video outlines and textual outlines, we use a large number of articles with chapter headings to pretrain our model. Experiments on DuVOG show that our model largely outperforms other baseline methods, achieving 77.1 of F1-score for the video segmentation level and 85.0 of ROUGE-L_F0.5 for the headline generation level.

📄 PDF Abstract BibTeX arXiv:2208.11307

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesHeadline GenerationNavigateOptical Character Recognition (OCR)Video SegmentationVideo Semantic SegmentationVideo Understanding

Similar Papers 제목 키워드 기반

Towards Visually-Guided Movie Subtitle Translation for Indic Languages

2026-05-12 · Tarun Chintada, Kshetrimayum Boynao Singh, Asif Ekbal arxiv

Movie subtitle translation is inherently multimodal, yet text-only systems often miss visual cues needed to convey emotion, action, and social nuance, especially for low-resource Indic languages (English to Hindi, Bengal…

Visual Grounding

VISA: An Ambiguous Subtitles Dataset for Visual Scene-Aware Machine Translation

2022-01-20 · LREC 2022 6 · Yihang Li, Shuichiro Shimizu, Weiqi Gu, Chenhui Chu 외

Existing multimodal machine translation (MMT) datasets consist of images and video captions or general subtitles, which rarely contain linguistic ambiguity, making visual information not so effective to generate appropri…

Machine TranslationMultimodal Machine TranslationSentenceTranslation

Towards Visual-Prompt Temporal Answering Grounding in Medical Instructional Video

2022-03-13 · Bin Li, Yixuan Weng, Bin Sun, Shutao Li

The temporal answering grounding in the video (TAGV) is a new task naturally derived from temporal sentence grounding in the video (TSGV). Given an untrimmed video and a text question, this task aims at locating the matc…

Language ModellingQuestion AnsweringSentenceTemporal Sentence Grounding

Video-Helpful Multimodal Machine Translation

2023-10-31 · Yihang Li, Shuichiro Shimizu, Chenhui Chu, Sadao Kurohashi 외

Existing multimodal machine translation (MMT) datasets consist of images and video captions or instructional video subtitles, which rarely contain linguistic ambiguity, making visual information ineffective in generating…

Machine TranslationMultimodal Machine TranslationTranslation

V-SAT: Video Subtitle Annotation Tool

2025-10-28 · Arpita Kundu, Joyita Chakraborty, Anindita Desarkar, Aritra Sen 외 arxiv

The surge of audiovisual content on streaming platforms and social media has heightened the demand for accurate and accessible subtitles. However, existing subtitle generation methods primarily speech-based transcription…

Speech Recognition