paper-with-me

홈 › Papers

Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance

2026-05-15 · Song Wu, Xinyu Chen, Qian Wang, Liang Li, Zili Yi, Junlan Feng arxiv

Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the rich information embedded within noisy latent, leading to unsatisfactory results. To address this, we propose a \textit{tuning-free, instruction-based} video editing framework. We approach video editing from the perspective of noisy latent: we design a Structural Noise Initialization Strategy (SNIS) to secure a superior editing starting point by assigning higher noise levels to edited regions (to facilitate content change) and lower noise levels to unedited regions (to maintain content consistency). We introduce a Noise Guidance Mechanism (NGM), which leverages the video prior in the generative model and effectively integrates rich information within the noisy latent to guide the denoising process, thereby preserving unedited content and overall visual coherence. Experiments show that our proposed method achieves better visual quality and state-of-the-art performance.

📄 PDF Abstract BibTeX arXiv:2605.15533

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FlowDirector: Training-Free Flow Steering for Precise Text-to-Video Editing

2025-06-05 · Guangzhao Li, Yanming Yang, Chenxi Song, Chi Zhang

Text-driven video editing aims to modify video content according to natural language instructions. While recent training-free approaches have made progress by leveraging pre-trained diffusion models, they typically rely …

Text-to-Video EditingVideo Editing

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

2026-06-30 · Sen Liang, Cong Wang, Zhentao Yu, Fengbin Guan 외 hf

Existing instruction-based video editing datasets commonly focus on single-task appearance editing, failing to meet the complex creative demands of real-world scenarios. To bridge this gap, we present Goku, a large-scale…

Instruction Following

SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing

2026-03-19 · Xinyao Zhang, Wenkai Dong, Yuxin Song, Bo Fang 외 arxiv

Current instruction-guided video editing models struggle to simultaneously balance precise semantic modifications with faithful motion preservation. While existing approaches rely on injecting explicit external priors (e…

Video Restoration

LangDriveCTRL: Natural Language Controllable Driving Scene Editing with Multi-modal Agents

2025-12-19 · Yun He, Francesco Pittaluga, Ziyu Jiang, Matthias Zwicker 외 arxiv

LangDriveCTRL is a natural-language-controllable framework for editing real-world driving videos to synthesize diverse traffic scenarios. It represents each video as an explicit 3D scene graph, decomposing the scene into…

InstructVid2Vid: Controllable Video Editing with Natural Language Instructions

2023-05-21 · Bosheng Qin, Juncheng Li, Siliang Tang, Tat-Seng Chua 외

We introduce InstructVid2Vid, an end-to-end diffusion-based methodology for video editing guided by human language instructions. Our approach empowers video manipulation guided by natural language directives, eliminating…

AttributeImage GenerationStyle TransferVideo Editing