paper-with-me

Papers

CI-VID: A Coherent Interleaved Text-Video Dataset

2025-07-02 · Yiming Ju, Jijin Hu, Zhengxiong Luo, Haoge Deng, hanyu Zhao, Li Du, Chengwei Wu, Donglin Hao, Xinlong Wang, Tengfei Pan arxiv

Text-to-video (T2V) generation has recently attracted considerable attention, resulting in the development of numerous high-quality datasets that have propelled progress in this area. However, existing public datasets are primarily composed of isolated text-video (T-V) pairs and thus fail to support the modeling of coherent multi-clip video sequences. To address this limitation, we introduce CI-VID, a dataset that moves beyond isolated text-to-video (T2V) generation toward text-and-video-to-video (TV2V) generation, enabling models to produce coherent, multi-scene video sequences. CI-VID contains over 340,000 samples, each featuring a coherent sequence of video clips with text captions that capture both the individual content of each clip and the transitions between them, enabling visually and textually grounded generation. To further validate the effectiveness of CI-VID, we design a comprehensive, multi-dimensional benchmark incorporating human evaluation, VLM-based assessment, and similarity-based metrics. Experimental results demonstrate that models trained on CI-VID exhibit significant improvements in both accuracy and content consistency when generating video sequences. This facilitates the creation of story-driven content with smooth visual transitions and strong temporal coherence, underscoring the quality and practical utility of the CI-VID dataset We release the CI-VID dataset and the accompanying code for data construction and evaluation at: https://github.com/ymju-BAAI/CI-VID

📄 PDF Abstract BibTeX arXiv:2507.01938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

2025-01-01 · Wenqi Zhang, Hang Zhang, Xin Li, Jiashuo Sun 외

Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges l…

Optical Character Recognition (OCR)

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

2024-06-15 · CVPR 2025 1 · Wei Chen, Lin Li, Yongqi Yang, Bin Wen 외

Interleaved image-text generation has emerged as a crucial multimodal task, aiming at creating sequences of interleaved visual and textual content given a query. Despite notable advancements in recent multimodal large la…

In-Context LearningText GenerationVisual Storytelling

Captain Cinema: Towards Short Movie Generation

2025-07-24 · Junfei Xiao, Ceyuan Yang, Lvmin Zhang, Shengqu Cai 외 arxiv

We present Captain Cinema, a generation framework for short movie generation. Given a detailed textual description of a movie storyline, our approach firstly generates a sequence of keyframes that outline the entire narr…

VINO: A Unified Visual Generator with Interleaved OmniModal Context

2026-01-05 · Junyi Chen, Tong He, Zhoujie Fu, Pengfei Wan 외 arxiv

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a…

Instruction FollowingVideo Generation

Bridging Your Imagination with Audio-Video Generation via a Unified Director

2025-12-29 · Jiaxu Zhang, Tianshu Hu, Yuan Zhang, Zenan Li 외 arxiv

Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter depends on image generation models. We arg…

Logical ReasoningVideo GenerationImage Generation