paper-with-me

홈 › Papers

Learning Multi-dimensional Human Preference for Text-to-Image Generation

2024-05-23 · CVPR 2024 1 · Sixian Zhang, Bohan Wang, Junqiang Wu, Yan Li, Tingting Gao, Di Zhang, Zhongyuan Wang

Current metrics for text-to-image models typically rely on statistical metrics which inadequately represent the real preference of humans. Although recent work attempts to learn these preferences via human annotated images, they reduce the rich tapestry of human preference to a single overall score. However, the preference results vary when humans evaluate images with different aspects. Therefore, to learn the multi-dimensional human preferences, we propose the Multi-dimensional Preference Score (MPS), the first multi-dimensional preference scoring model for the evaluation of text-to-image models. The MPS introduces the preference condition module upon CLIP model to learn these diverse preferences. It is trained based on our Multi-dimensional Human Preference (MHP) Dataset, which comprises 918,315 human preference choices across four dimensions (i.e., aesthetics, semantic alignment, detail quality and overall assessment) on 607,541 images. The images are generated by a wide range of latest text-to-image models. The MPS outperforms existing scoring methods across 3 datasets in 4 dimensions, enabling it a promising metric for evaluating and improving text-to-image generation.

📄 PDF Abstract BibTeX arXiv:2405.14705

Code (1)

kwai-kolors/kolors pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

2024-12-30 · Jiazheng Xu, Yu Huang, Jiale Cheng, Yuanming Yang 외

We present a general strategy to aligning visual generation models -- both image and video generation -- with human preference. To start with, we build VisionReward -- a fine-grained and multi-dimensional reward model. W…

Video GenerationVideo Quality Assessment

Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

2026-01-31 · Zicheng Kong, Dehua Ma, Zhenbo Xu, Alven Yang 외 arxiv

Multimodal large language models (MLLMs) struggle with alignment due to the limitations of existing reward models (RMs), which are predominantly vision-centric, dependent on costly human labels, and provide opaque scalar…

VRM: Teaching Reward Models to Understand Authentic Human Preferences

2026-03-05 · Biao Liu, Ning Xu, Junming Yang, Hao Xu 외 arxiv

Large Language Models (LLMs) have achieved remarkable success across diverse natural language tasks, yet the reward models employed for aligning LLMs often encounter challenges of reward hacking, where the approaches pre…

McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning

2025-11-28 · Qiushi Yang, Yingjie Chen, Yuan Yao, Yifang Men 외 arxiv

Text-to-video (T2V) generation has achieved remarkable progress in producing high-quality videos aligned with textual prompts. However, aligning synthesized videos with nuanced human preference remains challenging due to…

Reinforcement LearningVideo Generation

MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment

2025-09-26 · Tao Wu, Yibo Jiang, Yehao Lu, Zhizhong Wang 외 arxiv

Multi-subject image generation aims to synthesize user-provided subjects in a single image while preserving subject fidelity, ensuring prompt consistency, and aligning with human aesthetic preferences. Existing In-Contex…

Reinforcement LearningImage Generation