paper-with-me

홈 › Papers

Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?

2026-06-03 · Joong Ho Kim, Keith G. Mills arxiv

Diffusion Models (DM) have revolutionized text-driven generation by enabling the synthesis of high-quality, photorealistic visual content from user prompts. Whereas prior advances in visual generation such as VAEs and GANs were primarily evaluated on perceptual or visual similarity metrics such as FID PSNR, DM advances have fostered the development of more advanced Human Preference Metrics (HPM) that model and quantify human judgment as scalar values. However, DMs synthesize content using an inherently stochastic process where random noise seeds generation. The initial random noise directly affects the quality of generated outputs, both qualitatively and quantitatively. This influence is pronounced in smaller models for local deployment scenarios. Given this phenomenon, we first investigate to what extent we can predict scalar HPM scores prior to committing compute resources for generation. Further, we then investigate to what extent we can leverage such prediction to improve the quality of generated images, and also study which HPMs are best suited for this task. Our investigation reveals that not only is this possible, but that it is feasible to achieve negligible hardware overhead.

📄 PDF Abstract BibTeX arXiv:2606.05478

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EmoCtrl: Controllable Emotional Image Content Generation

2025-12-27 · Jingyuan Yang, Weibin Luo, Hui Huang arxiv

An image conveys meaning through both its visual content and emotional tone, jointly shaping human perception. We introduce Controllable Emotional Image Content Generation (C-EICG), which aims to generate images that rem…

UniAR: A Unified model for predicting human Attention and Responses on visual content

2023-12-15 · Peizhao Li, Junfeng He, Gang Li, Rachit Bhargava 외

Progress in human behavior modeling involves understanding both implicit, early-stage perceptual behavior, such as human attention, and explicit, later-stage behavior, such as subjective preferences or likes. Yet most pr…

Towards NSFW-Free Text-to-Image Generation via Safety-Constraint Direct Preference Optimization

2025-04-19 · Shouwei Ruan, Zhenyu Wu, Yao Huang, Ruochen Zhang 외

Ensuring the safety of generated content remains a fundamental challenge for Text-to-Image (T2I) generation. Existing studies either fail to guarantee complete safety under potentially harmful concepts or struggle to bal…

Contrastive LearningImage GenerationSafety AlignmentText to Image Generation+1

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

2025-02-05 · Zhenglin Zhou, Xiaobo Xia, Fan Ma, Hehe Fan 외

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human p…

3D GenerationText to 3D

Learning User Preferences for Image Generation Model

2025-08-11 · Wenyi Mo, Ying Ba, Tianyu Zhang, Yalong Bai 외 arxiv

User preference prediction requires a comprehensive and accurate understanding of individual tastes. This includes both surface-level attributes, such as color and style, and deeper content-related aspects, such as theme…

Image Generation