paper-with-me

홈 › Papers

Learning to Sample Effective and Diverse Prompts for Text-to-Image Generation

2025-02-17 · CVPR 2025 1 · Taeyoung Yun, Dinghuai Zhang, Jinkyoo Park, Ling Pan

Recent advances in text-to-image diffusion models have achieved impressive image generation capabilities. However, it remains challenging to control the generation process with desired properties (e.g., aesthetic quality, user intention), which can be expressed as black-box reward functions. In this paper, we focus on prompt adaptation, which refines the original prompt into model-preferred prompts to generate desired images. While prior work uses reinforcement learning (RL) to optimize prompts, we observe that applying RL often results in generating similar postfixes and deterministic behaviors. To this end, we introduce \textbf{P}rompt \textbf{A}daptation with \textbf{G}FlowNets (\textbf{PAG}), a novel approach that frames prompt adaptation as a probabilistic inference problem. Our key insight is that leveraging Generative Flow Networks (GFlowNets) allows us to shift from reward maximization to sampling from an unnormalized density function, enabling both high-quality and diverse prompt generation. However, we identify that a naive application of GFlowNets suffers from mode collapse and uncovers a previously overlooked phenomenon: the progressive loss of neural plasticity in the model, which is compounded by inefficient credit assignment in sequential prompt generation. To address this critical challenge, we develop a systematic approach in PAG with flow reactivation, reward-prioritized sampling, and reward decomposition for prompt adaptation. Extensive experiments validate that PAG successfully learns to sample effective and diverse prompts for text-to-image generation. We also show that PAG exhibits strong robustness across various reward functions and transferability to different text-to-image models.

📄 PDF Abstract BibTeX arXiv:2502.11477

Code (1)

dbsxodud-11/PAG 공식 구현 pytorch

Tasks

Image GenerationReinforcement Learning (RL)Text to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
PAG 설명 없음
Focus 설명 없음

Similar Papers 제목 키워드 기반

VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models

2024-09-23 · Jingtao Cao, Zheng Zhang, Hongru Wang, Kam-Fai Wong

Progress in Text-to-Image (T2I) models has significantly improved the generation of images from textual descriptions. However, existing evaluation metrics do not adequately assess the models' ability to handle a diverse …

Image Generation

Toward Generalist Anomaly Detection via In-context Residual Learning with Few-shot Sample Prompts

2024-03-11 · CVPR 2024 1 · Jiawen Zhu, Guansong Pang

This paper explores the problem of Generalist Anomaly Detection (GAD), aiming to train one single detection model that can generalize to detect anomalies in diverse datasets from different application domains without any…

Anomaly Detection

Learning Probabilistic Prompt for Continual Learning

2026-07-06 · Hyekang Park, Sanghoon Lee, Geon Lee, Jongyoun Noh 외 arxiv

Continual learning aims to progressively learn from a sequence of tasks, each containing a disjoint subset of classes, while preserving previously learned knowledge. Prompt-based continual learning methods propose to lea…

Continual Learning

INT: Instance-Specific Negative Mining for Task-Generic Promptable Segmentation

2025-01-30 · Jian Hu, Zixu Cheng, Shaogang Gong

Task-generic promptable image segmentation aims to achieve segmentation of diverse samples under a single task description by utilizing only one task-generic prompt. Current methods leverage the generalization capabiliti…

Image SegmentationInstance SegmentationSegmentationSemantic Segmentation

MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation

2024-10-26 · Jialin Luo, Yuanzhi Wang, Ziqi Gu, Yide Qiu 외

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating…

Image GenerationText to Image GenerationText-to-Image Generation