paper-with-me

홈 › Papers

Narration Generation for Cartoon Videos

2021-01-17 · Nikos Papasarantopoulos, Shay B. Cohen

Research on text generation from multimodal inputs has largely focused on static images, and less on video data. In this paper, we propose a new task, narration generation, that is complementing videos with narration texts that are to be interjected in several places. The narrations are part of the video and contribute to the storyline unfolding in it. Moreover, they are context-informed, since they include information appropriate for the timeframe of video they cover, and also, do not need to include every detail shown in input scenes, as a caption would. We collect a new dataset from the animated television series Peppa Pig. Furthermore, we formalize the task of narration generation as including two separate tasks, timing and content generation, and present a set of models on the new task.

📄 PDF Abstract BibTeX arXiv:2101.06803

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Learning to Generate Long-term Future Narrations Describing Activities of Daily Living

2025-03-03 · Ramanathan Rajendiran, Debaditya Roy, Basura Fernando

Anticipating future events is crucial for various application domains such as healthcare, smart home technology, and surveillance. Narrative event descriptions provide context-rich information, enhancing a system's futur…

Action AnticipationDecision MakingLanguage ModelingLanguage Modelling+1

What You Say Is What You Show: Visual Narration Detection in Instructional Videos

2023-01-05 · Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, Kristen Grauman

Narrated ''how-to'' videos have emerged as a promising data source for a wide range of learning problems, from learning visual representations to training robot policies. However, this data is extremely noisy, as the nar…

Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures

2026-06-10 · Ishani Mondal, Javad Baghirov, Jordan Boyd-Graber arxiv

Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation…

Video Generation

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation

2025-07-27 · Shuolin Xu, Bingyuan Wang, Zeyu Cai, Fangteng Fu 외 arxiv

Generating high-quality cartoon animations multimodal control is challenging due to the complexity of non-human characters, stylistically diverse motions and fine-grained emotions. There is a huge domain gap between real…

Video Generation

Sakuga-42M Dataset: Scaling Up Cartoon Research

2024-05-13 · Zhenglin Pan, Yu Zhu, Yuxuan Mu

Hand-drawn cartoon animation employs sketches and flat-color segments to create the illusion of motion. While recent advancements like CLIP, SVD, and Sora show impressive results in understanding and generating natural v…

MambaText to Video RetrievalVideo to Text Retrieval