PrefScore: Pairwise Preference Learning for Reference-free Single-document Summarization Quality Assessment
Evaluating machine-generated summaries without a human-written reference summary has been a need for a long time. Inspired by preference labeling in existing works of summarization evaluation, we propose to judge summary quality by learning the preference rank of summaries using the Bradley-Terry power ranking model from generated inferior summaries of a base summary. Despite the simplicity of our method, extensive experiments on several datasets show that our weakly supervised scheme can produce scores highly correlate with human ratings.
Code (0)
등록된 구현이 없습니다.
Tasks
Document SummarizationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
PrefScore: Pairwise Preference Learning for Reference-free Summarization Quality Assessment
Evaluating machine-generated summaries without a human-written reference summary has been a need for a long time. Inspired by preference labeling in existing work of summarization evaluation, we propose to judge summary …
Preferences Order, Ratings Anchor: From Fused Expert Aesthetic Ground Truth to Self-Distillation
Pairwise preferences and pointwise ratings are the two dominant annotation protocols in image aesthetic assessment (IAA), yet existing benchmarks adopt only one, leaving their complementarity unmeasured under controlled …
Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling
Reliable reward and preference signals are critical for evaluating and optimizing large language models on open-ended tasks. Rubric-based judges offer a transparent way to decompose such judgments into explicit evaluatio…
Reinforcement LearningClustering and Inference From Pairwise Comparisons
Given a set of pairwise comparisons, the classical ranking problem computes a single ranking that best represents the preferences of all users. In this paper, we study the problem of inferring individual preferences, ari…
ClusteringFrom Reward-Free Representations to Preferences: Rethinking Offline Preference-Based Reinforcement Learning
Preference-based reinforcement learning (PbRL) avoids explicit reward engineering by learning from pairwise human preference feedback. Existing offline PbRL methods typically follow a two-stage pipeline, first learning a…
Representation LearningReinforcement LearningOffline RL