paper-with-me

홈 › Papers

BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment

2025-11-24 · Dewei Zhou, Mingwei Li, Zongxin Yang, Yu Lu, Yunqiu Xu, Zhizhong Wang, Zeyi Huang, Yi Yang arxiv

Conditional image generation enhances text-to-image synthesis with structural, spatial, or stylistic priors, but current methods face challenges in handling conflicts between sources. These include 1) input-level conflicts, where the conditioning image contradicts the text prompt, and 2) model-bias conflicts, where generative biases disrupt alignment even when conditions match the text. Addressing these conflicts requires nuanced solutions, which standard supervised fine-tuning struggles to provide. Preference-based optimization techniques like Direct Preference Optimization (DPO) show promise but are limited by gradient entanglement between text and condition signals and lack disentangled training data for multi-constraint tasks. To overcome this, we propose a bidirectionally decoupled DPO framework (BideDPO). Our method creates two disentangled preference pairs-one for the condition and one for the text-to reduce gradient entanglement. The influence of pairs is managed using an Adaptive Loss Balancing strategy for balanced optimization. We introduce an automated data pipeline to sample model outputs and generate conflict-aware data. This process is embedded in an iterative optimization strategy that refines both the model and the data. We construct a DualAlign benchmark to evaluate conflict resolution between text and condition. Experiments show BideDPO significantly improves text success rates (e.g., +35%) and condition adherence. We also validate our approach using the COCO dataset. Project Pages: https://limuloo.github.io/BideDPO/.

📄 PDF Abstract BibTeX arXiv:2511.19268

Code (0)

등록된 구현이 없습니다.

Tasks

Conditional Image Generation

Similar Papers 제목 키워드 기반

PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling

2025-04-19 · Alara Dirik, Tuanfeng Wang, Duygu Ceylan, Stefanos Zafeiriou 외

We present PRISM, a unified framework that enables multiple image generation and editing tasks in a single foundational model. Starting from a pre-trained text-to-image diffusion model, PRISM proposes an effective fine-t…

Conditional Image GenerationImage GenerationIntrinsic Image DecompositionText to Image Generation+1

Entropy Rectifying Guidance for Diffusion and Flow Models

2025-04-18 · Tariq Berrada Ifriqi, Adriana Romero-Soriano, Michal Drozdzal, Jakob Verbeek 외

Guidance techniques are commonly used in diffusion and flow models to improve image quality and consistency for conditional generative tasks such as class-conditional and text-to-image generation. In particular, classifi…

DiversityImage GenerationText to Image GenerationText-to-Image Generation+1

Self-control: A Better Conditional Mechanism for Masked Autoregressive Model

2024-12-18 · Qiaoying Qu, Shiyu Shen

Autoregressive conditional image generation algorithms are capable of generating photorealistic images that are consistent with given textual or image conditions, and have great potential for a wide range of applications…

Conditional Image GenerationImage GenerationQuantization

Guiding a Diffusion Model with a Bad Version of Itself

2024-06-04 · Tero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen 외

The primary axes of interest in image-generating diffusion models are image quality, the amount of variation in the results, and how well the results align with a given condition, e.g., a class label or a text prompt. Th…

Image Generation

V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

2024-06-04 · Cong Wang, Kuan Tian, Jun Zhang, Yonghang Guan 외

In the field of portrait video generation, the use of single images to generate portrait videos has become increasingly prevalent. A common approach involves leveraging generative models to enhance adapters for controlle…

Video Generation