paper-with-me

홈 › Papers

Recognizing and Presenting the Storytelling Video Structure with Deep Multimodal Networks

2016-10-05 · Lorenzo Baraldi, Costantino Grana, Rita Cucchiara

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable sequences, which can in turn be used to retrieve the most significant parts of it given a textual query and to provide an effective summarization. Previous video decomposition methods mainly employed perceptual cues, tackling the problem either as a story change detection, or as a similarity grouping task, and the lack of semantics limited their ability to identify story boundaries. Our proposal connects together perceptual, audio and semantic cues in a specialized deep network architecture designed with a combination of CNNs which generate an appropriate embedding, and clusters shots into connected sequences of semantic scenes, i.e. stories. A retrieval presentation strategy is also proposed, by selecting the semantically and aesthetically "most valuable" thumbnails to present, considering the query in order to improve the storytelling presentation. Finally, the subjective nature of the task is considered, by conducting experiments with different annotators and by proposing an algorithm to maximize the agreement between automatic results and human annotators.

📄 PDF Abstract BibTeX arXiv:1610.01376

Code (0)

등록된 구현이 없습니다.

Tasks

Change DetectionRetrievalSemantic Segmentation

Similar Papers 제목 키워드 기반

Synopses of Movie Narratives: a Video-Language Dataset for Story Understanding

2022-03-11 · Yidan Sun, Qin Chao, Yangfeng Ji, Boyang Li

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), contai…

RetrievalText RetrievalVideo-Text Retrieval

Video Storytelling: Textual Summaries for Events

2018-07-25 · Junnan Li, Yongkang Wong, Qi Zhao, Mohan S. Kankanhalli

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied …

DiversityReinforcement LearningSentence

Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation

2023-07-13 · Yingqing He, Menghan Xia, Haoxin Chen, Xiaodong Cun 외

Generating videos for visual storytelling can be a tedious and complex process that typically requires either live-action filming or graphics animation rendering. To bypass these challenges, our key idea is to utilize th…

RetrievalVideo GenerationVideo RetrievalVisual Storytelling

The Art of Storytelling: Multi-Agent Generative AI for Dynamic Multimodal Narratives

2024-09-17 · Samee Arif, Taimoor Arif, Muhammad Saad Haroon, Aamina Jamal Khan 외

This paper introduces the concept of an education tool that utilizes Generative Artificial Intelligence (GenAI) to enhance storytelling for children. The system combines GenAI-driven narrative co-creation, text-to-speech…

text-to-speechText to SpeechText-to-Video GenerationVideo Generation

Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video

2025-11-05 · Alexander Htet Kyaw, Lenin Ravindranath Sivalingam arxiv

We present a node-based storytelling system for multimodal content generation. The system represents stories as graphs of nodes that can be expanded, edited, and iteratively refined through direct user edits and natural-…

multimodal generationStory Generation