paper-with-me

홈 › Papers

GVDIFF: Grounded Text-to-Video Generation with Diffusion Models

2024-07-02 · Huanzhang Dou, Ruixiang Li, Wei Su, Xi Li

In text-to-video (T2V) generation, significant attention has been directed toward its development, yet unifying discrete and continuous grounding conditions in T2V generation remains under-explored. This paper proposes a Grounded text-to-Video generation framework, termed GVDIFF. First, we inject the grounding condition into the self-attention through an uncertainty-based representation to explicitly guide the focus of the network. Second, we introduce a spatial-temporal grounding layer that connects the grounding condition with target objects and enables the model with the grounded generation capacity in the spatial-temporal domain. Third, our dynamic gate network adaptively skips the redundant grounding process to selectively extract grounding information and semantics while improving efficiency. We extensively evaluate the grounded generation capacity of GVDIFF and demonstrate its versatility in applications, including long-range video generation, sequential prompts, and object-specific editing.

📄 PDF Abstract BibTeX arXiv:2407.01921

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Video GenerationVideo Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

LLM-grounded Video Diffusion Models

2023-09-29 · Long Lian, Baifeng Shi, Adam Yala, Trevor Darrell 외

Text-conditioned diffusion models have emerged as a promising tool for neural video generation. However, current models still struggle with intricate spatiotemporal prompts and often generate restricted or incorrect moti…

Language ModelingLanguage ModellingLarge Language ModelVideo Generation

VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting

2024-12-16 · Muhammet Furkan Ilaslan, Ali Koksal, Kevin Qinhong Lin, Burak Satar 외

Large Language Model (LLM)-based agents have shown promise in procedural tasks, but the potential of multimodal instructions augmented by texts and videos to assist users remains under-explored. To address this gap, we p…

InformativenessLarge Language ModelText GenerationText-to-Video Generation+2

VIMI: Grounding Video Generation through Multi-modal Instruction

2024-07-08 · Yuwei Fang, Willi Menapace, Aliaksandr Siarohin, Tsai-Shien Chen 외

Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual groun…

Text-to-Video GenerationVideo GenerationVisual Grounding

PhysCtrl: Generative Physics for Controllable and Physics-Grounded Video Generation

2025-09-24 · Chen Wang, Chuhao Chen, Yiming Huang, Zhiyang Dou 외 arxiv

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel f…

Video Generation

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

2025-03-18 · Shoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong 외

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we int…

Reasoning SegmentationVideo Editing