paper-with-me

홈 › Papers

ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation

2025-04-02 · Xiao Wang, Daniil Larionov, Siwei Wu, Yiqi Liu, Steffen Eger, Nafise Sadat Moosavi, Chenghua Lin

Evaluating the quality of generated text automatically remains a significant challenge. Conventional reference-based metrics have been shown to exhibit relatively weak correlation with human evaluations. Recent research advocates the use of large language models (LLMs) as source-based metrics for natural language generation (NLG) assessment. While promising, LLM-based metrics, particularly those using smaller models, still fall short in aligning with human judgments. In this work, we introduce ContrastScore, a contrastive evaluation metric designed to enable higher-quality, less biased, and more efficient assessment of generated text. We evaluate ContrastScore on two NLG tasks: machine translation and summarization. Experimental results show that ContrastScore consistently achieves stronger correlation with human judgments than both single-model and ensemble-based baselines. Notably, ContrastScore based on Qwen 3B and 0.5B even outperforms Qwen 7B, despite having only half as many parameters, demonstrating its efficiency. Furthermore, it effectively mitigates common evaluation biases such as length and likelihood preferences, resulting in more robust automatic evaluation.

📄 PDF Abstract BibTeX arXiv:2504.02106

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationText Generation

Similar Papers 제목 키워드 기반

Less Learn Shortcut: Analyzing and Mitigating Learning of Spurious Feature-Label Correlation

2022-05-25 · Yanrui Du, Jing Yan, Yan Chen, Jing Liu 외

Recent research has revealed that deep neural networks often take dataset biases as a shortcut to make decisions rather than understand tasks, leading to failures in real-world applications. In this study, we focus on th…

Natural Language InferenceSentiment Analysis

Train Neural Network by Embedding Space Probabilistic Constraint

2019-03-24 · ICLR Workshop LLD 2019 · Kaiyuan Chen, Zhanyuan Yin

Using higher order knowledge to reduce training data has become a popular research topic. However, the ability for available methods to draw effective decision boundaries is still limited: when training set is small, neu…

Unbiased Pairwise Learning to Rank in Recommender Systems

2021-11-25 · Yi Ren, Hongyan Tang, Siwen Zhu

Nowadays, recommender systems already impact almost every facet of peoples lives. To provide personalized high quality recommendation results, conventional systems usually train pointwise rankers to predict the absolute …

AttributeLearning-To-RankPositionRecommendation Systems

Determination of hysteresis in finite-state random walks using Bayesian cross validation

2017-02-21 · Joshua C. Chang

Consider the problem of modeling hysteresis for finite-state random walks using higher-order Markov chains. This Letter introduces a Bayesian framework to determine, from data, the number of prior states of recent histor…

VarDiU: A Variational Diffusive Upper Bound for One-Step Diffusion Distillation

2025-08-28 · Leyang Wang, Mingtian Zhang, Zijing Ou, David Barber arxiv

Recently, diffusion distillation methods have compressed thousand-step teacher diffusion models into one-step student generators while preserving sample quality. Most existing approaches train the student model using a d…