paper-with-me

홈 › Papers

MovieDreamer: Hierarchical Generation for Coherent Long Visual Sequence

2024-07-23 · Canyu Zhao, MingYu Liu, Wen Wang, Weihua Chen, Fan Wang, Hao Chen, Bo Zhang, Chunhua Shen

Recent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended periods, which is essential for long-form video production like movies. We propose MovieDreamer, a novel hierarchical framework that integrates the strengths of autoregressive models with diffusion-based rendering to pioneer long-duration video generation with intricate plot progressions and high visual fidelity. Our approach utilizes autoregressive models for global narrative coherence, predicting sequences of visual tokens that are subsequently transformed into high-quality video frames through diffusion rendering. This method is akin to traditional movie production processes, where complex stories are factorized down into manageable scene capturing. Further, we employ a multimodal script that enriches scene descriptions with detailed character information and visual style, enhancing continuity and character identity across scenes. We present extensive experiments across various movie genres, demonstrating that our approach not only achieves superior visual and narrative quality but also effectively extends the duration of generated content significantly beyond current capabilities. Homepage: https://aim-uofa.github.io/MovieDreamer/.

📄 PDF Abstract BibTeX arXiv:2407.16655

Code (0)

등록된 구현이 없습니다.

Tasks

Video Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Matching Visual Features to Hierarchical Semantic Topics for Image Paragraph Captioning

2021-05-10 · Dandan Guo, Ruiying Lu, Bo Chen, Zequn Zeng 외

Observing a set of images and their corresponding paragraph-captions, a challenging task is to learn how to produce a semantically coherent paragraph to describe the visual content of an image. Inspired by recent success…

Image Paragraph CaptioningLanguage ModelingLanguage ModellingVariational Inference

AlignTransformer: Hierarchical Alignment of Visual Regions and Disease Tags for Medical Report Generation

2022-03-18 · Di You, Fenglin Liu, Shen Ge, Xiaoxia Xie 외

Recently, medical report generation, which aims to automatically generate a long and coherent descriptive paragraph of a given medical image, has received growing research interests. Different from the general image capt…

DescriptiveImage CaptioningMedical Report Generation

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

2026-01-20 · Xinlei Yin, Xiulian Peng, Xiao Li, Zhiwei Xiong 외 arxiv

Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typi…

Multimodal Reasoning

Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-Enhancement

2024-12-18 · Qianyue Wang, Jinwu Hu, ZhengPing Li, Yufeng Wang 외

Long-form story generation task aims to produce coherent and sufficiently lengthy text, essential for applications such as novel writingand interactive storytelling. However, existing methods, including LLMs, rely on rig…

FormKnowledge GraphsStory Generation

Hierarchically Structured Reinforcement Learning for Topically Coherent Visual Story Generation

2018-05-21 · Qiuyuan Huang, Zhe Gan, Asli Celikyilmaz, Dapeng Wu 외

We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task…

DecoderDeep Reinforcement Learningreinforcement-learningReinforcement Learning+4