paper-with-me

Papers

Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models

2025-07-14 · Zejian Li, Yize Li, Chenye Meng, Zhongni Liu, Yang Ling, Shengyuan Zhang, Guang Yang, Changyuan Yang, Zhiyuan Yang, Lingyun Sun arxiv

Recent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO

📄 PDF Abstract BibTeX arXiv:2507.11554

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Training-Free Image Editing with Visual Context Integration and Concept Alignment

2026-04-06 · Rui Song, Guo-Hua Wang, Qing-Guo Chen, Weihua Luo 외 arxiv

In image editing, it is essential to incorporate a context image to convey the user's precise requirements, such as subject appearance or image style. Existing training-based visual context-aware editing methods incur da…

Image Editing

ODE-DPS: ODE-based Diffusion Posterior Sampling for Inverse Problems in Partial Differential Equation

2024-04-21 · Enze Jiang, Jishen Peng, Zheng Ma, Xiong-bin Yan

In recent years we have witnessed a growth in mathematics for deep learning, which has been used to solve inverse problems of partial differential equations (PDEs). However, most deep learning-based inversion methods eit…

OSI: One-step Inversion Excels in Extracting Diffusion Watermarks

2026-02-10 · Yuwei Chen, Zhenliang He, Jia Tang, Meina Kan 외 arxiv

Watermarking is an important mechanism for provenance and copyright protection of diffusion-generated images. Training-free methods, exemplified by Gaussian Shading, embed watermarks into the initial noise of diffusion m…

ResetEdit: Precise Text-guided Editing of Generated Image via Resettable Starting Latent

2026-04-28 · Hanyi Wang, Han Fang, Zheng Wang, Shilin Wang 외 arxiv

Recent advances in diffusion models have enabled high-quality image generation, leading to increasing demand for post-generation editing that modifies local regions while preserving global structure. Achieving such flexi…

Image Generation

Training-free image inversion for one-step diffusion models

2026-05-31 · Tao Wu, Senmao Li, Yaxing Wang, Shiqi Yang 외 arxiv

In this work, we introduce a novel training-free inversion (TFinv) framework for one-step diffusion models,addressing key challenges in real image inversion and editing. We first identify two critical factors hamperingre…

Image Editing