Consistent and Editable: A Balanced Framework for Text-Guided Video Editing
Recently, diffusion models have achieved considerable success in the text-guided video editing domain. However, existing works often struggle to balance the trade-off between temporal consistency and editability in video editing, with consistency and editability typically being inversely related. To address this, we propose a high-quality video editing framework enhanced for consistency and editability, named EquiEdit, which improves coordinatively the temporal consistency and editability of the edited videos while achieving a balance between the two. In terms of temporal consistency, the proposed temporal Mamba module with a tailored temporal-aware scanning scans fused video sequences following four designed directions, effectively enhancing the inter-frame consistency of edited videos. For editability, we design a noise injection strategy based on the spectral transformation to increase editing flexibility, where the Fourier transform is used to preserve the hidden structure in the initial latent noise used for editing, ensuring inter-frame consistency of the edited video and fidelity to the input video. Extensive qualitative and quantitative experiments demonstrate the effectiveness of our method in terms of temporal consistency and editability, as well as its great fidelity to the input video itself.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
OMGTex: One-stage Multi-style Facial Texture Reconstruction without Geometry Guidance
We propose OMGTex, an end-to-end diffusion-based framework for reconstructing high-quality and editable facial UV textures from multi-style facial images. Existing texture reconstruction methods face two major limitation…
Style TransferMorphoSim: An Interactive, Controllable, and Editable Language-guided 4D World Simulator
World models that support controllable and editable spatiotemporal environments are valuable for robotics, enabling scalable training data, repro ducible evaluation, and flexible task design. While recent text-to-video m…
CompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene Layout
Text-to-3D form plays a crucial role in creating editable 3D scenes for AR/VR. Recent advances have shown promise in merging neural radiance fields (NeRFs) with pre-trained diffusion models for text-to-3D object generati…
NeRFObjectScene GenerationText to 3DInstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow
We propose a fast text-guided image editing method called InstantEdit based on the RectifiedFlow framework, which is structured as a few-step editing process that preserves critical content while following closely to tex…
Image EditingAre Watermarked Images Editable? SafeMark for Watermark-Preserving Text-Guided Image Editing
This paper investigates a fundamental yet underexplored question: can watermarked images remain editable without compromising watermark integrity? We propose SafeMark, a framework for watermark-preserving text-guided ima…
Image ManipulationImage Editing