paper-with-me

홈 › Papers

Compare without Despair: Reliable Preference Evaluation with Generation Separability

2024-07-02 · Sayan Ghosh, Tejas Srinivasan, Swabha Swayamdipta

Human evaluation of generated language through pairwise preference judgments is pervasive. However, under common scenarios, such as when generations from a model pair are very similar, or when stochastic decoding results in large variations in generations, it results in inconsistent preference ratings. We address these challenges by introducing a meta-evaluation measure, separability, which estimates how suitable a test instance is for pairwise preference evaluation. For a candidate test instance, separability samples multiple generations from a pair of models, and measures how distinguishable the two sets of generations are. Our experiments show that instances with high separability values yield more consistent preference ratings from both human- and auto-raters. Further, the distribution of separability allows insights into which test benchmarks are more valuable for comparing models. Finally, we incorporate separability into ELO ratings, accounting for how suitable each test instance might be for reliably ranking LLMs. Overall, separability has implications for consistent, efficient and robust preference evaluation of LLMs with both human- and auto-raters.

📄 PDF Abstract BibTeX arXiv:2407.01878

Code (1)

dill-lab/separability 공식 구현

Similar Papers 제목 키워드 기반

Learning and Evaluating Human Preferences for Conversational Head Generation

2023-07-20 · Mohan Zhou, Yalong Bai, Wei zhang, Ting Yao 외

A reliable and comprehensive evaluation metric that aligns with manual preference assessments is crucial for conversational head video synthesis methods development. Existing quantitative evaluations often fail to captur…

Sensitive and Scalable Online Evaluation with Theoretical Guarantees

2017-11-26 · Oosterhuis Harrie, de Rijke Maarten

Multileaved comparison methods generalize interleaved comparison methods to provide a scalable approach for comparing ranking systems based on regular user interactions. Such methods enable the increasingly rapid researc…

From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signals

2026-05-28 · Yeyong Yu, Wenya Hu, Xing Wu, Quan Qian arxiv

As candidate generation and high-throughput experimentation advance, the primary bottleneck in materials discovery is shifting from property prediction to making reliable evaluations among massive candidate sets. We prop…

GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation

2024-01-08 · CVPR 2024 1 · Tong Wu, Guandao Yang, Zhibing Li, Kai Zhang 외

Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the …

3D GenerationText to 3D

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

2026-06-17 · Zilong Zhang, Yi-Ting Hung, Lei Ding, Chi-Kuang Yeh arxiv

Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic quality, most notably verbosity bias. Me…