paper-with-me

Papers

Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

2023-12-16 · Mingfei Han, Linjie Yang, Xiaojun Chang, Heng Wang

A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark Shot2Story20K with detailed shot-level captions and comprehensive video summaries. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video and narration captioning, multi-shot video summarization, and video retrieval with shot descriptions. Preliminary experiments show some challenges to generate a long and comprehensive video summary. Nevertheless, the generated imperfect summaries can already significantly boost the performance of existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries.

📄 PDF Abstract BibTeX arXiv:2312.10300

Code (1)

bytedance/Shot2Story 공식 구현 pytorch

Tasks

Video Captioningvideo narration captioningVideo Question AnsweringVideo RetrievalVideo SummarizationVideo UnderstandingZero-Shot Video Question Answer

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

STORYWARS: A Dataset and Instruction Tuning Baselines for Collaborative Story Understanding and Generation

2023-05-14 · Yulun Du, Lydia Chilton

Collaborative stories, which are texts created through the collaborative efforts of multiple authors with different writing styles and intentions, pose unique challenges for NLP models. Understanding and generating such …

Multimodal High-order Relation Transformer for Scene Boundary Detection

2023-01-01 · ICCV 2023 1 · Xi Wei, Zhangxiang Shi, Tianzhu Zhang, Xiaoyuan Yu 외

Scene boundary detection breaks down long videos into meaningful story-telling units and plays a crucial role in high-level video understanding. Despite significant advancements in this area, this task remains a chal…

Boundary DetectionDecoderRelationVideo Understanding

Let's Play Across Cultures: A Large Multilingual, Multicultural Benchmark for Assessing Language Models' Understanding of Sports

2025-09-24 · Punit Kumar Singh, Nishant Kumar, Akash Ghosh, Kunal Pasad 외 arxiv

Language Models (LMs) are primarily evaluated on globally popular sports, often overlooking regional and indigenous sporting traditions. To address this gap, we introduce \textbf{\textit{CultSportQA}}, a benchmark design…

Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding

2022-03-11 · Yidan Sun, Qin Chao, Yangfeng Ji, Boyang Li

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), contai…

RetrievalText RetrievalVideo-Text Retrieval

A Video Is Worth 4096 Tokens: Verbalize Videos To Understand Them In Zero Shot

2023-05-16 · Aanisha Bhattacharya, Yaman K Singla, Balaji Krishnamurthy, Rajiv Ratn Shah 외

Multimedia content, such as advertisements and story videos, exhibit a rich blend of creativity and multiple modalities. They incorporate elements like text, visuals, audio, and storytelling techniques, employing devices…

Emotion ClassificationQuestion AnsweringTopic ClassificationVideo Understanding