paper-with-me

홈 › Papers

CineDance: Towards Next-Generation Multi-Shot Long-Form Cinematic Audio-Video Generation

2026-06-08 · Yuheng Chen, Teng Hu, Yuji Wang, Qingdong He, Zhucun Xue, Qianyu Zhou, Jason Li, Lizhuang Ma, Jiangning Zhang, Dacheng Tao arxiv

The fidelity and structural diversity of training datasets fundamentally determine the capabilities of video generation models. While commercial systems showremarkableabilitytogeneratecinematicnarratives, the progress of open-source models remains limited by the scarcity of high-quality training data. To bridge this gap, we introduce CineDance-1M, a large-scale, open research Text-to-Audio-Video (T2AV) dataset designed specifically for multi-shot, long-form joint audio-video generation. Averaging 92.8 seconds and 24.2 continuous shots per video, it provides configurable, structured annotations for both audio and video modalities. This exceptional quality is achieved through a rigorous three-stage curation pipeline: i) diverse sourcing and comprehensive cleansing, ii) film-theory-inspired narrative parsing, and iii) hierarchical dual-modal captioning. For a comprehensive assessment, we propose CineBench, featuring a diverse prompt suite and a six-dimensional, human-aligned metric system tailored for complex narrative audio-video evaluation. Furthermore, we adapt LTX-2.3 into CineDance, which demonstrates exceptional single-modality quality alongside precise audio-video alignment and robust subject and environment consistency, effectively validating our curation strategy and the high quality of CineDance-1M. We anticipate that this work will serve as a solid foundation for accelerating future research in multi-shot, long-form joint audio-video generation. Our project page is available at https://aliothchen.github.io/projects/CineDance/.

📄 PDF Abstract BibTeX arXiv:2606.09639

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationVideo Alignment

Similar Papers 제목 키워드 기반

OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory

2025-12-08 · Zhaochong An, Menglin Jia, Haonan Qiu, Zijian Zhou 외 arxiv

Storytelling in real-world videos often unfolds through multiple shots -- discontinuous yet semantically connected clips that together convey a coherent narrative. However, existing multi-shot video generation (MSV) meth…

Video Generation

ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling

2026-03-26 · Yawen Luo, Xiaoyu Shi, Junhao Zhuang, Yutian Chen 외 arxiv

Multi-shot video generation is crucial for long narrative storytelling, yet current bidirectional architectures suffer from limited interactivity and high latency. We propose ShotStream, a novel causal multi-shot archite…

Video Generation

Memento: Reconstruct to Remember for Consistent Long Video Generation

2026-06-12 · Xuan Wei, Longbin Ji, Guan Wang, Xiangrui Liu 외 arxiv

Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalability by generating vide…

Video Generation

Cut2Next: Generating Next Shot via In-Context Tuning

2025-08-11 · Jingwen He, Hongbo Liu, Jiajun Li, Ziqi Huang 외 arxiv

Effective multi-shot generation demands purposeful, film-like transitions and strict cinematic continuity. Current methods, however, often prioritize basic visual consistency, neglecting crucial editing patterns (e.g., s…

Next-Scale Autoregressive Models for Text-to-Motion Generation

2026-04-04 · Zhiwei Zheng, Shibo Jin, Lingjie Liu, Mingmin Zhao arxiv

Autoregressive (AR) models offer stable and efficient training, but standard next-token prediction is not well aligned with the temporal structure required for text-conditioned motion generation. We introduce MoScale, a …