paper-with-me

홈 › Papers

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

2024-12-30 · Jiazheng Xu, Yu Huang, Jiale Cheng, Yuanming Yang, Jiajun Xu, YuAn Wang, Wenbo Duan, Shen Yang, Qunlin Jin, Shurun Li, Jiayan Teng, Zhuoyi Yang, Wendi Zheng, Xiao Liu, Ming Ding, Xiaohan Zhang, Xiaotao Gu, Shiyu Huang, Minlie Huang, Jie Tang, Yuxiao Dong

We present a general strategy to aligning visual generation models -- both image and video generation -- with human preference. To start with, we build VisionReward -- a fine-grained and multi-dimensional reward model. We decompose human preferences in images and videos into multiple dimensions, each represented by a series of judgment questions, linearly weighted and summed to an interpretable and accurate score. To address the challenges of video quality assessment, we systematically analyze various dynamic features of videos, which helps VisionReward surpass VideoScore by 17.2% and achieve top performance for video preference prediction. Based on VisionReward, we develop a multi-objective preference learning algorithm that effectively addresses the issue of confounding factors within preference data. Our approach significantly outperforms existing image and video scoring methods on both machine metrics and human evaluation. All code and datasets are provided at https://github.com/THUDM/VisionReward.

📄 PDF Abstract BibTeX arXiv:2412.21059

Code (1)

thudm/visionreward 공식 구현 pytorch

Tasks

Video GenerationVideo Quality Assessment

Similar Papers 제목 키워드 기반

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

2025-08-15 · MengChao Wang, Qiang Wang, Fan Jiang, Mu Xu arxiv

Recent advances in audio-driven portrait animation have demonstrated impressive capabilities. However, existing methods struggle to align with fine-grained human preferences across multiple dimensions, such as motion nat…

UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs

2024-09-30 · Yuho Lee, Taewon Yun, Jason Cai, Hang Su 외

Existing benchmarks for summarization quality evaluation often lack diverse input scenarios, focus on narrowly defined dimensions (e.g., faithfulness), and struggle with subjective and coarse-grained annotation schemes. …

TTS-PRISM: A Perceptual Reasoning and Interpretable Speech Model for Fine-Grained Diagnosis

2026-04-24 · Xi Wang, Jie Wang, Xingchen Song, Baijun Song 외 arxiv

While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To address this, we propose TTS-PRISM, a mult…

Fine-Grained Human Pose Editing Assessment via Layer-Selective MLLMs

2026-01-15 · Ningyu Sun, Zhaolin Cai, Zitong Xu, Peihang Chen 외 arxiv

Text-guided human pose editing has gained significant traction in AIGC applications. However,it remains plagued by structural anomalies and generative artifacts. Existing evaluation metrics often isolate authenticity det…

MHPR: Multidimensional Human Perception and Reasoning Benchmark for Large Vision-Languate Models

2026-05-05 · Kangkang Wang, Qinting Jiang, Wanping Zhang, Bowen Ren 외 arxiv

Multidimensional human understanding is essential for real-world applications such as film analysis and virtual digital humans, yet current LVLM benchmarks largely focus on single-task settings and lack fine-grained, hum…

Reinforcement LearningInstruction Following