paper-with-me

홈 › Papers

DreamRunner: Fine-Grained Storytelling Video Generation with Retrieval-Augmented Motion Adaptation

2024-11-25 · Zun Wang, Jialu Li, Han Lin, Jaehong Yoon, Mohit Bansal

Storytelling video generation (SVG) has recently emerged as a task to create long, multi-motion, multi-scene videos that consistently represent the story described in the input text script. SVG holds great potential for diverse content creation in media and entertainment; however, it also presents significant challenges: (1) objects must exhibit a range of fine-grained, complex motions, (2) multiple objects need to appear consistently across scenes, and (3) subjects may require multiple motions with seamless transitions within a single scene. To address these challenges, we propose DreamRunner, a novel story-to-video generation method: First, we structure the input script using a large language model (LLM) to facilitate both coarse-grained scene planning as well as fine-grained object-level layout and motion planning. Next, DreamRunner presents retrieval-augmented test-time adaptation to capture target motion priors for objects in each scene, supporting diverse motion customization based on retrieved videos, thus facilitating the generation of new videos with complex, scripted motions. Lastly, we propose a novel spatial-temporal region-based 3D attention and prior injection module SR3AI for fine-grained object-motion binding and frame-by-frame semantic control. We compare DreamRunner with various SVG baselines, demonstrating state-of-the-art performance in character consistency, text alignment, and smooth transitions. Additionally, DreamRunner exhibits strong fine-grained condition-following ability in compositional text-to-video generation, significantly outperforming baselines on T2V-ComBench. Finally, we validate DreamRunner's robust ability to generate multi-object interactions with qualitative examples.

📄 PDF Abstract BibTeX arXiv:2411.16657

Code (0)

등록된 구현이 없습니다.

Tasks

Large Language ModelMotion PlanningRetrievalTest-time AdaptationText-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation

2025-12-31 · Bingxuan Li, Yiming Cui, Yicheng He, Yiwei Wang 외 arxiv

Sound effects build an essential layer of multimodal storytelling, shaping the emotional atmosphere and the narrative semantics of videos. Despite recent advancement in video-text-to-audio (VT2A), the current formulation…

StoryMem: Multi-shot Long Video Storytelling with Memory

2025-12-22 · Kaiwen Zhang, Liming Jiang, Angtian Wang, Jacob Zhiyuan Fang 외 arxiv

Visual storytelling requires generating multi-shot videos with cinematic quality and long-range consistency. Inspired by human memory, we propose StoryMem, a paradigm that reformulates long-form video storytelling as ite…

Visual StorytellingStory Generation

The Lost Melody: Empirical Observations on Text-to-Video Generation From A Storytelling Perspective

2024-05-13 · Andrew Shin, Yusuke Mori, Kunitake Kaneko

Text-to-video generation task has witnessed a notable progress, with the generated outcomes reflecting the text prompts with high fidelity and impressive visual qualities. However, current text-to-video generation models…

Text-to-Video GenerationVideo Generation

Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

2025-08-12 · Ao Ma, Jiasong Feng, Ke Cao, Jing Wang 외 arxiv

Storytelling tasks involving generating consistent subjects have gained significant attention recently. However, existing methods, whether training-free or training-based, continue to face challenges in maintaining subje…

Story Generation

CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling

2026-07-29 · Yuyang Huang, Yabo Chen, Wenrui Dai, Ziyang Zheng 외 arxiv

Cinematic video generation is challenging for text-to-video diffusion models due to concurrent requirements on multi-shot generation, fine-grained controllability over characters and scenes, and long-form generation acro…

Video Generation