paper-with-me

홈 › Papers

Unveil Inversion and Invariance in Flow Transformer for Versatile Image Editing

2024-11-24 · CVPR 2025 1 · Pengcheng Xu, Boyuan Jiang, Xiaobin Hu, Donghao Luo, Qingdong He, Jiangning Zhang, Chengjie Wang, Yunsheng Wu, Charles Ling, Boyu Wang

Leveraging the large generative prior of the flow transformer for tuning-free image editing requires authentic inversion to project the image into the model's domain and a flexible invariance control mechanism to preserve non-target contents. However, the prevailing diffusion inversion performs deficiently in flow-based models, and the invariance control cannot reconcile diverse rigid and non-rigid editing tasks. To address these, we systematically analyze the \textbf{inversion and invariance} control based on the flow transformer. Specifically, we unveil that the Euler inversion shares a similar structure to DDIM yet is more susceptible to the approximation error. Thus, we propose a two-stage inversion to first refine the velocity estimation and then compensate for the leftover error, which pivots closely to the model prior and benefits editing. Meanwhile, we propose the invariance control that manipulates the text features within the adaptive layer normalization, connecting the changes in the text prompt to image semantics. This mechanism can simultaneously preserve the non-target contents while allowing rigid and non-rigid manipulation, enabling a wide range of editing types such as visual text, quantity, facial expression, etc. Experiments on versatile scenarios validate that our framework achieves flexible and accurate editing, unlocking the potential of the flow transformer for versatile image editing.

📄 PDF Abstract BibTeX arXiv:2411.15843

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Audio synthesizer inversion in symmetric parameter spaces with approximately equivariant flow matching

2025-06-08 · Ben Hayes, Charalampos Saitis, György Fazekas

Many audio synthesizers can produce the same signal given different parameter configurations, meaning the inversion from sound to parameters is an inherently ill-posed problem. We show that this is largely due to intrins…

Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing

2025-09-16 · Weiming Chen, Zhihan Zhu, Yijia Wang, Zhihai He arxiv

Rectified flow (RF) models have recently demonstrated superior generative performance compared to DDIM-based diffusion models. However, in real-world applications, they suffer from two major challenges: (1) low inversion…

Image Reconstruction

Text-to-Image Rectified Flow as Plug-and-Play Priors

2024-06-05 · Xiaofeng Yang, Cheng Chen, Xulei Yang, Fayao Liu 외

Large-scale diffusion models have achieved remarkable performance in generative tasks. Beyond their initial training applications, these models have proven their ability to function as versatile plug-and-play priors. For…

3D GenerationText to 3D

Understanding Invariance via Feedforward Inversion of Discriminatively Trained Classifiers

2021-03-15 · Piotr Teterwak, Chiyuan Zhang, Dilip Krishnan, Michael C. Mozer

A discriminatively trained neural net classifier can fit the training data perfectly if all information about its input other than class membership has been discarded prior to the output layer. Surprisingly, past researc…

High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

2024-07-04 · Gael Le Lan, Bowen Shi, Zhaoheng Ni, Sidd Srinivasan 외

We introduce MelodyFlow, an efficient text-controllable high-fidelity music generation and editing model. It operates on continuous latent representations from a low frame rate 48 kHz stereo variational auto encoder code…

DenoisingMusic Generation