paper-with-me

홈 › Papers

Favi-Score: A Measure for Favoritism in Automated Preference Ratings for Generative AI Evaluation

2024-06-03 · Pius von Däniken, Jan Deriu, Don Tuggener, Mark Cieliebak

Generative AI systems have become ubiquitous for all kinds of modalities, which makes the issue of the evaluation of such models more pressing. One popular approach is preference ratings, where the generated outputs of different systems are shown to evaluators who choose their preferences. In recent years the field shifted towards the development of automated (trained) metrics to assess generated outputs, which can be used to create preference ratings automatically. In this work, we investigate the evaluation of the metrics themselves, which currently rely on measuring the correlation to human judgments or computing sign accuracy scores. These measures only assess how well the metric agrees with the human ratings. However, our research shows that this does not tell the whole story. Most metrics exhibit a disagreement with human system assessments which is often skewed in favor of particular text generation systems, exposing a degree of favoritism in automated metrics. This paper introduces a formal definition of favoritism in preference metrics, and derives the Favi-Score, which measures this phenomenon. In particular we show that favoritism is strongly related to errors in final system rankings. Thus, we propose that preference-based metrics ought to be evaluated on both sign accuracy scores and favoritism.

📄 PDF Abstract BibTeX arXiv:2406.01131

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Revealing and Utilizing In-group Favoritism for Graph-based Collaborative Filtering

2024-04-23 · Hoin Jung, Hyunsoo Cho, Myungje Choi, Joowon Lee 외

When it comes to a personalized item recommendation system, It is essential to extract users' preferences and purchasing patterns. Assuming that users in the real world form a cluster and there is common favoritism in ea…

ClusteringCollaborative Filtering

Favia: Forensic Agent for Vulnerability-fix Identification and Analysis

2026-02-13 · André Storhaug, Jiamou Sun, Jingyue Li arxiv

Identifying vulnerability-fixing commits corresponding to disclosed CVEs is essential for secure software maintenance but remains challenging at scale, as large repositories contain millions of commits of which only a sm…

Truth or Tribe: How In-group Favoritism Prioritize Facts in Persona Agents

2026-05-02 · Shijun Lei, Hongyu Wang, Yunji Liang, Haowen Zheng 외 arxiv

In-group favoritism refers to the phenomena of favoring members of one's in-group over out-group members and is widely observed in numerous social cooperative behaviors. Recently, in-group favoritism biases have also bee…

Scoring and Favoritism in Optimal Procurement Design

2024-11-19 · Pasha Andreyanov, Ilia Krasikov, Alex Suzdaltsev

We study buyer-optimal procurement mechanisms when quality is contractible. When some costs are borne by every participant of a procurement auction regardless of winning, the classic analysis should be amended. We show t…

Factorization Vision Transformer: Modeling Long Range Dependency with Local Window Cost

2023-12-14 · Haolin Qin, Daquan Zhou, Tingfa Xu, Ziyang Bian 외

Transformers have astounding representational power but typically consume considerable computation which is quadratic with image resolution. The prevailing Swin transformer reduces computational costs through a local win…