paper-with-me

홈 › Papers

Kiwi-Edit: Versatile Video Editing via Instruction and Reference Guidance

2026-03-02 · Yiqi Lin, Guoqiang Liang, Ziyun Zeng, Zechen Bai, Yanzhe Chen, Mike Zheng Shou arxiv

Instruction-based video editing has witnessed rapid progress, yet current methods often struggle with precise visual control, as natural language is inherently limited in describing complex visual nuances. Although reference-guided editing offers a robust solution, its potential is currently bottlenecked by the scarcity of high-quality paired training data. To bridge this gap, we introduce a scalable data generation pipeline that transforms existing video editing pairs into high-fidelity training quadruplets, leveraging image generative models to create synthesized reference scaffolds. Using this pipeline, we construct RefVIE, a large-scale dataset tailored for instruction-reference-following tasks, and establish RefVIE-Bench for comprehensive evaluation. Furthermore, we propose a unified editing architecture, Kiwi-Edit, that synergizes learnable queries and latent visual features for reference semantic guidance. Our model achieves significant gains in instruction following and reference fidelity via a progressive multi-stage training curriculum. Extensive experiments demonstrate that our data and architecture establish a new state-of-the-art in controllable video editing. All datasets, models, and code is released at https://github.com/showlab/Kiwi-Edit.

📄 PDF Abstract BibTeX arXiv:2603.02175

Code (0)

등록된 구현이 없습니다.

Tasks

Instruction Following

Similar Papers 제목 키워드 기반

VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation

2025-03-18 · Shoubin Yu, Difan Liu, Ziqiao Ma, Yicong Hong 외

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we int…

Reasoning SegmentationVideo Editing

RACCooN: A Versatile Instructional Video Editing Framework with Auto-Generated Narratives

2024-05-28 · Jaehong Yoon, Shoubin Yu, Mohit Bansal

Recent video generative models primarily rely on carefully written text prompts for specific tasks, like inpainting or style editing. They require labor-intensive textual descriptions for input videos, hindering their fl…

AttributeVideo Editing

UniVideo: Unified Understanding, Generation, and Editing for Videos

2025-10-09 · Cong Wei, Quande Liu, Zixuan Ye, Qiulin Wang 외 arxiv

Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVideo, a versatile framework that extends un…

Video GenerationText GenerationStyle TransferImage Editing

V2Edit: Versatile Video Diffusion Editor for Videos and 3D Scenes

2025-03-13 · YanMing Zhang, Jun-Kun Chen, Jipeng Lyu, Yu-Xiong Wang

This paper introduces V$^2$Edit, a novel training-free framework for instruction-guided video and 3D scene editing. Addressing the critical challenge of balancing original content preservation with editing task fulfillme…

3D scene EditingDenoisingVideo Editing

In-Context Learning with Unpaired Clips for Instruction-based Video Editing

2025-10-16 · Xinyao Liao, Xianfang Zeng, Ziye Song, Zhoujie Fu 외 arxiv

Despite the rapid progress of instruction-based image editing, its extension to video remains underexplored, primarily due to the prohibitive cost and complexity of constructing large-scale paired video editing datasets.…

Instruction FollowingVideo GenerationImage Editing