paper-with-me

Papers

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks

2026-04-09 · Jiayang Xu, Fan Zhuo, Majun Zhang, Changhao Pan, Zehan Wang, Siyu Chen, Xiaoda Yang, Tao Jin, Zhou Zhao arxiv

Current video editing models often rely on expensive paired video data, which limits their practical scalability. In essence, most video editing tasks can be formulated as a decoupled spatiotemporal process, where the temporal dynamics of the pretrained model are preserved while spatial content is selectively and precisely modified. Based on this insight, we propose ImVideoEdit, an efficient framework that learns video editing capabilities entirely from image pairs. By freezing the pre-trained 3D attention modules and treating images as single-frame videos, we decouple the 2D spatial learning process to help preserve the original temporal dynamics. The core of our approach is a Predict-Update Spatial Difference Attention module that progressively extracts and injects spatial differences. Rather than relying on rigid external masks, we incorporate a Text-Guided Dynamic Semantic Gating mechanism for adaptive and implicit text-driven modifications. Despite training on only 13K image pairs for 5 epochs with exceptionally low computational overhead, ImVideoEdit achieves editing fidelity and temporal consistency comparable to larger models trained on extensive video datasets.

📄 PDF Abstract BibTeX arXiv:2604.07958

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Edit Temporal-Consistent Videos with Image Diffusion Model

2023-08-17 · Yuanzhi Wang, Yong Li, Xiaoya Zhang, Xin Liu 외

Large-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregular…

modelVideo EditingVideo Temporal Consistency

UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance Editing

2024-02-20 · Jianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo 외

Recent advances in text-guided video editing have showcased promising results in appearance editing (e.g., stylization). However, video motion editing in the temporal dimension (e.g., from eating to waving), which distin…

Video Editing

ViDiC: Video Difference Captioning

2025-12-03 · Jiangtao Wu, Shihao Li, Zhaozhou Bian, Jialu Chen 외 arxiv

Understanding visual differences between dynamic scenes requires the comparative perception of compositional, spatial, and temporal changes--a capability that remains underexplored in existing vision-language systems. Wh…

LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning

2025-06-11 · Chenjian Gao, Lihe Ding, Xin Cai, Zhanpeng Huang 외

Video editing using diffusion models has achieved remarkable results in generating high-quality edits for videos. However, current methods often rely on large-scale pretraining, limiting flexibility for specific edits. F…

Video Editing

VideoDirector: Precise Video Editing via Text-to-Video Models

2024-11-26 · CVPR 2025 1 · Yukun Wang, Longguang Wang, Zhiyuan Ma, Qibin Hu 외

Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color fl…

AttributeVideo Editing