paper-with-me

홈 › Papers

Improving Controllable Generation: Faster Training and Better Performance via $x_0$-Supervision

2026-04-07 · Amadou S. Sangare, Adrien Maglo, Mohamed Chaouch, Bertrand Luvison arxiv

Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisely control image layouts, something that natural language alone cannot reliably express. Controllable generation methods augment the initial T2I model with additional conditions that more easily describe the scene. Prior works straightforwardly train the augmented network with the same loss as the initial network. Although natural at first glance, this can lead to very long training times in some cases before convergence. In this work, we revisit the training objective of controllable diffusion models through a detailed analysis of their denoising dynamics. We show that direct supervision on the clean target image, dubbed $x_0$-supervision, or an equivalent re-weighting of the diffusion loss, yields faster convergence. Experiments on multiple control settings demonstrate that our formulation accelerates convergence by up to 2$\times$ according to our novel metric (mean Area Under the Convergence Curve - mAUCC), while also improving both visual quality and conditioning accuracy. Our code is available at https://github.com/CEA-LIST/x0-supervision

📄 PDF Abstract BibTeX arXiv:2604.05761

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Start Small: Training Controllable Game Level Generators without Training Data by Learning at Multiple Sizes

2022-09-29 · Yahia Zakaria, Magda Fayek, Mayada Hadhoud

A level generator is a tool that generates game levels from noise. Training a generator without a dataset suffers from feedback sparsity, since it is unlikely to generate a playable level via random exploration. A common…

DiversitySokoban

ConZIC: Controllable Zero-shot Image Captioning by Sampling-Based Polishing

2023-03-04 · CVPR 2023 1 · Zequn Zeng, Hao Zhang, Zhengjue Wang, Ruiying Lu 외

Zero-shot capability has been considered as a new revolution of deep learning, letting machines work on tasks without curated training data. As a good start and the only existing outcome of zero-shot image captioning (IC…

DiversityImage CaptioningLanguage ModelingLanguage Modelling

Follow-Your-Emoji-Faster: Towards Efficient, Fine-Controllable, and Expressive Freestyle Portrait Animation

2025-09-20 · Yue Ma, Zexuan Yan, Hongyu Liu, Hongfa Wang 외 arxiv

We present Follow-Your-Emoji-Faster, an efficient diffusion-based framework for freestyle portrait animation driven by facial landmarks. The main challenges in this task are preserving the identity of the reference portr…

CoCoFormer: A controllable feature-rich polyphonic music generation method

2023-10-15 · Jiuyang Zhou, Tengfei Niu, Hong Zhu, Xingping Wang

This paper explores the modeling method of polyphonic music sequence. Due to the great potential of Transformer models in music generation, controllable music generation is receiving more attention. In the task of polyph…

DiversityMusic GenerationRhythm

Topic-Controllable Summarization: Topic-Aware Evaluation and Transformer Methods

2022-06-09 · Tatiana Passali, Grigorios Tsoumakas

Topic-controllable summarization is an emerging research area with a wide range of potential applications. However, existing approaches suffer from significant limitations. For example, the majority of existing methods b…