paper-with-me

Papers

InstructVid2Vid: Controllable Video Editing with Natural Language Instructions

2023-05-21 · Bosheng Qin, Juncheng Li, Siliang Tang, Tat-Seng Chua, Yueting Zhuang

We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating the need for per-example fine-tuning or inversion. The proposed InstructVid2Vid model modifies a pretrained image generation model, Stable Diffusion, to generate a time-dependent sequence of video frames. By harnessing the collective intelligence of disparate models, we engineer a training dataset rich in video-instruction triplets, which is a more cost-efficient alternative to collecting data in real-world scenarios. To enhance the coherence between successive frames within the generated videos, we propose the Inter-Frames Consistency Loss and incorporate it during the training process. With multimodal classifier-free guidance during the inference stage, the generated videos is able to resonate with both the input video and the accompanying instructions. Experimental results demonstrate that InstructVid2Vid is capable of generating high-quality, temporally coherent videos and performing diverse edits, including attribute editing, background changes, and style transfer. These results underscore the versatility and effectiveness of our proposed method.

📄 PDF Abstract BibTeX arXiv:2305.12328

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage GenerationStyle TransferVideo Editing

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
BLIP Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based…
U-Net 설명 없음

Similar Papers 제목 키워드 기반

InstructVideo: Instructing Video Diffusion Models with Human Feedback

2023-12-19 · CVPR 2024 1 · Hangjie Yuan, Shiwei Zhang, Xiang Wang, Yujie Wei 외

Diffusion models have emerged as the de facto paradigm for video generation. However, their reliance on web-scale data of varied quality often yields results that are visually unappealing and misaligned with the textual …

Video Generation

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

2025-12-19 · Yun He, Francesco Pittaluga, Ziyu Jiang, Matthias Zwicker 외 arxiv

LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represents each video as an explicit 3D scene graph, decomposing the scene into…

Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance

2026-03-02 · Yiqi Lin, Guoqiang Liang, Ziyun Zeng, Zechen Bai 외 arxiv

Instruction-based video editing has witnessed rapid progress, yet current methods often struggle with precise visual control, as natural language is inherently limited in describing complex visual nuances. Although refer…

Instruction Following

FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing

2025-06-05 · Guangzhao Li, Yanming Yang, Chenxi Song, Chi Zhang

Text-driven video editing aims to modify video content according to natural language instructions. While recent training-free approaches have made progress by leveraging pre-trained diffusion models, they typically rely …

Text-to-Video EditingVideo Editing

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

2026-03-30 · Milton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen 외 arxiv

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high pr…