paper-with-me

Papers

Taming Flow-based I2V Models for Creative Video Editing

2025-09-26 · Xianghao Kong, Hansheng Chen, Yuwei Guo, Lvmin Zhang, Gordon Wetzstein, Maneesh Agrawala, Anyi Rao arxiv

Although image editing techniques have advanced significantly, video editing, which aims to manipulate videos according to user intent, remains an emerging challenge. Most existing image-conditioned video editing methods either require inversion with model-specific design or need extensive optimization, limiting their capability of leveraging up-to-date image-to-video (I2V) models to transfer the editing capability of image editing models to the video domain. To this end, we propose IF-V2V, an Inversion-Free method that can adapt off-the-shelf flow-matching-based I2V models for video editing without significant computational overhead. To circumvent inversion, we devise Vector Field Rectification with Sample Deviation to incorporate information from the source video into the denoising process by introducing a deviation term into the denoising vector field. To further ensure consistency with the source video in a model-agnostic way, we introduce Structure-and-Motion-Preserving Initialization to generate motion-aware temporally correlated noise with structural information embedded. We also present a Deviation Caching mechanism to minimize the additional computational cost for denoising vector rectification without significantly impacting editing quality. Evaluations demonstrate that our method achieves superior editing quality and consistency over existing approaches, offering a lightweight plug-and-play solution to realize visual creativity.

📄 PDF Abstract BibTeX arXiv:2509.21917

Code (0)

등록된 구현이 없습니다.

Tasks

Image Editing

Similar Papers 제목 키워드 기반

Taming Rectified Flow for Inversion and Editing

2024-11-07 · Jiangshan Wang, Junfu Pu, Zhongang Qi, Jiayi Guo 외

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often st…

Image GenerationText-to-Image GenerationVideo EditingVideo Generation

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

2026-05-05 · Andong Deng, Dawei Du, Zhenfang Chen, Wen Zhong 외 arxiv

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs)…

Multimodal Reasoning

FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis

2023-12-29 · CVPR 2024 1 · Feng Liang, Bichen Wu, Jialiang Wang, Licheng Yu 외

Diffusion models have transformed the image-to-image (I2I) synthesis and are now permeating into videos. However, the advancement of video-to-video (V2V) synthesis has been hampered by the challenge of maintaining tempor…

Optical Flow EstimationVideo-to-Video Synthesis

VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space

2022-11-22 · David Chuan-En Lin, Fabian Caba Heilbron, Joon-Young Lee, Oliver Wang 외

Video editing is a creative and complex endeavor and we believe that there is potential for reimagining a new video editing interface to better support the creative and exploratory nature of video editing. We take inspir…

Asset ManagementManagementVideo Editing

Taming I2V models for Image HOI Editing: A Cognitive Benchmark and Agentic Self-Correcting Framework

2026-06-17 · Jiayi Gao, Qingchao Chen, Yuxin Peng, Yang Liu arxiv

Current image editing methods excel at static attributes but fail at complex Human-Object Interactions (HOI), a critical challenge unaddressed by existing benchmarks that conflate HOI with static attributes, relying on g…

Image Editing