paper-with-me

홈 › Papers

VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization

2025-12-18 · Xiaoyan Cong, Haotian Yang, Angtian Wang, Yizhi Wang, Yiding Yang, Canyu Zhang, Chongyang Ma arxiv

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired data of simple editing operations, which fundamentally limits their ability to generalize to diverse and complex, real-world instructions. To address this generalization gap, we propose VIVA, a scalable framework for instruction-based video editing that leverages VLM-guided encoding and reward optimization. First, we introduce a VLM-based instructor that encodes the textual instruction, the first frame of the source video, and an optional reference image into visually-grounded instruction representations, providing fine-grained spatial and semantic context for the diffusion transformer backbone. Second, we propose a post-training stage, Edit-GRPO, which adapts Group Relative Policy Optimization to the domain of video editing, directly optimizing the model for instruction-faithful, content-preserving, and aesthetically pleasing edits using relative rewards. Furthermore, we propose a data construction pipeline designed to synthetically generate diverse, high-fidelity paired video-instruction data of basic editing operations. Extensive experiments show that VIVA achieves superior instruction following, generalization, and editing quality over state-of-the-art methods. Website: https://viva-paper.github.io

📄 PDF Abstract BibTeX arXiv:2512.16906

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects

2026-04-17 · Xiangbo Gao, Sicong Jiang, Bangya Liu, Xinghao Chen 외 arxiv

As AI-assisted video creation becomes increasingly practical, instruction-guided video editing has become essential for refining generated or captured footage to meet professional requirements. Yet the field still lacks …

Instruction Following

IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment

2025-10-13 · Yinan Chen, Jiangning Zhang, Teng Hu, Yuxiang Zeng 외 arxiv

Instruction-guided video editing has emerged as a rapidly advancing research direction, offering new opportunities for intuitive content transformation while also posing significant challenges for systematic evaluation. …

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

2026-08-18 · Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong 외 arxiv

The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided …

Instruction Following

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

2026-06-02 · Yinan Chen, Chuming Lin, Zhennan Chen, Yuxiang Zeng 외 arxiv

While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge this gap, we present JAVEdit-100k, the fi…

Region-Constrained Group Relative Policy Optimization for Flow-Based Image Editing

2026-04-10 · Zhuohan Ouyang, Zhe Qian, Wenhuo Cui, Chaoqun Wang arxiv

Instruction-guided image editing requires balancing target modification with non-target preservation. Recently, flow-based models have emerged as a strong and increasingly adopted backbone for instruction-guided image ed…

Instruction FollowingImage Editing