paper-with-me

홈 › Papers

VIMI: Grounding Video Generation through Multi-modal Instruction

2024-07-08 · Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen, Kuan-Chien Wang, Ivan Skorokhodov, Graham Neubig, Sergey Tulyakov

Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual grounding and restricting their versatility and application in multimodal integration. To address this, we construct a large-scale multimodal prompt dataset by employing retrieval methods to pair in-context examples with the given text prompts and then utilize a two-stage training strategy to enable diverse video generation tasks within the same model. In the first stage, we propose a multimodal conditional video generation framework for pretraining on these augmented datasets, establishing a foundational model for grounded video generation. Secondly, we finetune the model from the first stage on three video generation tasks, incorporating multi-modal instructions. This process further refines the model's ability to handle diverse inputs and tasks, ensuring seamless integration of multi-modal information. After this two-stage train-ing process, VIMI demonstrates multimodal understanding capabilities, producing contextually rich and personalized videos grounded in the provided inputs, as shown in Figure 1. Compared to previous visual grounded video generation methods, VIMI can synthesize consistent and temporally coherent videos with large motion while retaining the semantic control. Lastly, VIMI also achieves state-of-the-art text-to-video generation results on UCF101 benchmark.

📄 PDF Abstract BibTeX arXiv:2407.06304

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo GenerationVisual Grounding

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access

2025-11-23 · Timing Yang, Sucheng Ren, Alan Yuille, Feng Wang arxiv

Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video-text corpus. Existing public datasets typically re…

Video Question AnsweringText-to-Video Generation

WaterVideoQA: ASV-Centric Perception and Rule-Compliant Reasoning via Multi-Modal Agents

2026-02-26 · Runwei Guan, Shaofeng Liang, Ningwei Ouyang, Weichen Fei 외 arxiv

While autonomous navigation has achieved remarkable success in passive perception (e.g., object detection and segmentation), it remains fundamentally constrained by a void in knowledge-driven, interactive environmental c…

Video Question AnsweringObject Detection

GVDIFF: Grounded Text-to-Video Generation with Diffusion Models

2024-07-02 · Huanzhang Dou, Ruixiang Li, Wei Su, Xi Li

In text-to-video (T2V) generation, significant attention has been directed toward its development, yet unifying discrete and continuous grounding conditions in T2V generation remains under-explored. This paper proposes a…

Text-to-Video GenerationVideo Generation

VIMI: Vehicle-Infrastructure Multi-view Intermediate Fusion for Camera-based 3D Object Detection

2023-03-20 · Zhe Wang, Siqi Fan, Xiaoliang Huo, Tongda Xu 외

In autonomous driving, Vehicle-Infrastructure Cooperative 3D Object Detection (VIC3D) makes use of multi-view cameras from both vehicles and traffic infrastructure, providing a global vantage point with rich semantic con…

3D Object DetectionAutonomous DrivingFeature Compressionobject-detection+1

Multi-sentence Video Grounding for Long Video Generation

2024-07-18 · Wei Feng, Xin Wang, Hong Chen, Zeyang Zhang 외

Video generation has witnessed great success recently, but their application in generating long videos still remains challenging due to the difficulty in maintaining the temporal consistency of generated videos and the h…

Moment RetrievalRetrievalSentenceVideo Editing+2