paper-with-me

홈 › Papers

HPSv3: Towards Wide-Spectrum Human Preference Score

2025-08-05 · Yuhang Ma, Yunhao Shui, Xiaoshi Wu, Keqiang Sun, Hongsheng Li arxiv

Evaluating text-to-image generation models requires alignment with human perception, yet existing human-centric metrics are constrained by limited data coverage, suboptimal feature extraction, and inefficient loss functions. To address these challenges, we introduce Human Preference Score v3 (HPSv3). (1) We release HPDv3, the first wide-spectrum human preference dataset integrating 1.08M text-image pairs and 1.17M annotated pairwise comparisons from state-of-the-art generative models and low to high-quality real-world images. (2) We introduce a VLM-based preference model trained using an uncertainty-aware ranking loss for fine-grained ranking. Besides, we propose Chain-of-Human-Preference (CoHP), an iterative image refinement method that enhances quality without extra data, using HPSv3 to select the best image at each step. Extensive experiments demonstrate that HPSv3 serves as a robust metric for wide-spectrum image evaluation, and CoHP offers an efficient and human-aligned approach to improve image generation quality. The code and dataset are available at the HPSv3 Homepage.

📄 PDF Abstract BibTeX arXiv:2508.03789

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

HPSv3++: Scaling Reward Models Across the Full Spectrum of Diffusion Model Capabilities

2026-06-12 · Yijun Liu, Jie Huang, Zeyue Xue, Yuming Li 외 arxiv

Reward models guide text-to-image (T2I) systems toward outputs aligned with human preferences. However, typical reward models such as HPSv3 are trained on pre-annotated data from earlier T2I models, without accounting fo…

Reinforcement Learning

Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards

2025-03-25 · Alexander Gambashidze, Konstantin Sobolev, Andrey Kuznetsov, Ivan Oseledets

Can Visual Language Models (VLMs) effectively capture human visual preferences? This work addresses this question by training VLMs to think about preferences at test time, employing reinforcement learning methods inspire…

World Knowledge

Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

2023-06-15 · Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen 외

Recent text-to-image generative models can generate high-fidelity images from text inputs, but the quality of these generated images cannot be accurately evaluated by existing evaluation metrics. To address this issue, w…

Image GenerationPreference Mapping

MIRO: MultI-Reward cOnditioned pretraining improves T2I quality and efficiency

2025-10-29 · Nicolas Dufour, Lucas Degeorge, Arijit Ghosh, Vicky Kalogeiton 외 arxiv

The default paradigm of post-training text-to-image generators includes post-hoc selection of generated images, and subsequent training with one reward model to align the generator to the reward, typically user preferenc…

Diff-Instruct++: Training One-step Text-to-image Generator Model to Align with Human Preferences

2024-10-24 · Weijian Luo

One-step text-to-image generator models offer advantages such as swift inference efficiency, flexible architectures, and state-of-the-art generation performance. In this paper, we study the problem of aligning one-step g…