paper-with-me

Papers

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation

2026-04-09 · Zhefan Rao, Bin Zou, Haoxuan Che, Xuanhua He, Chong Hou Choi, Yanheng Li, Rui Liu, Qifeng Chen arxiv

Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same time, high-quality video editing data remains scarce. In this paper, we show that a video generation backbone can become a strong video editor without large scale video editing data. We present InsEdit, an instruction-based editing model built on HunyuanVideo-1.5. InsEdit combines a visual editing architecture with a video data pipeline based on Mutual Context Attention (MCA), which creates aligned video pairs where edits can begin in the middle of a clip rather than only from the first frame. With only O(100)K video editing data, InsEdit achieves state-of-the-art results among open-source methods on our video instruction editing benchmarks. In addition, because our training recipe also includes image editing data, the final model supports image editing without any modification.

📄 PDF Abstract BibTeX arXiv:2604.08646

Code (0)

등록된 구현이 없습니다.

Tasks

Video GenerationImage Editing

Similar Papers 제목 키워드 기반

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

2026-07-17 · Yuqing Wen, Yukai Huang, Qianqian Xie, Jiangtao Wu 외 hf

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmark…

Instruction Following

VicEdit: Learning to Edit Videos from Visual In-Context Examples

2026-08-17 · Yuji Wang, Teng Hu, Yuheng Chen, Ran Yi 외 arxiv

Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Edi…

JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation

2026-06-02 · Yinan Chen, Chuming Lin, Zhennan Chen, Yuxiang Zeng 외 arxiv

While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge this gap, we present JAVEdit-100k, the fi…

CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing

2026-08-18 · Fuchen Long, Cong Wang, Zitao Gao, Wenhao Zhong 외 arxiv

The quality and diversity of instruction-based video editing datasets are steadily improving, yet existing datasets mainly focus on single editing operations and fall short in supporting compositional instruction-guided …

Instruction Following

RefVideo-6M: A Reliable Reference-Based Dataset for Instructional Video Editing

2026-08-26 · Bojia Zi, Xiaoyan Yang, Yu Zhou, Ruijie Sun 외 arxiv

Recent advances in video editing have been largely driven by large-scale instruction-based datasets. However, existing datasets still suffer from two critical limitations. First, target videos are commonly produced by au…

Image Editing