paper-with-me

홈 › Papers

Unbiased Evaluation of Large Language Models from a Causal Perspective

2025-02-10 · Meilin Chen, Jian Tian, Liang Ma, Di Xie, WeiJie Chen, Jiang Zhu

Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation of questions. Despite their success, the biases in Agents-as-an-Evaluator methods remain largely unexplored. In this paper, we present a theoretical formulation of evaluation bias, providing valuable insights into designing unbiased evaluation protocols. Furthermore, we identify two type of bias in Agents-as-an-Evaluator through carefully designed probing tasks on a minimal Agents-as-an-Evaluator setup. To address these issues, we propose the Unbiased Evaluator, an evaluation protocol that delivers a more comprehensive, unbiased, and interpretable assessment of LLMs.Extensive experiments reveal significant room for improvement in current LLMs. Additionally, we demonstrate that the Unbiased Evaluator not only offers strong evidence of benchmark contamination but also provides interpretable evaluation results.

📄 PDF Abstract BibTeX arXiv:2502.06655

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Online Evaluation Methods for the Causal Effect of Recommendations

2021-07-14 · Masahiro Sato

Evaluating the causal effect of recommendations is an important objective because the causal effect on user interactions can directly leads to an increase in sales and user engagement. To select an optimal recommendation…

Causal Disentanglement for Semantics-Aware Intent Learning in Recommendation

2022-02-05 · Xiangmeng Wang, Qian Li, Dianer Yu, Peng Cui 외

Traditional recommendation models trained on observational interaction data have generated large impacts in a wide range of applications, it faces bias problems that cover users' true intent and thus deteriorate the reco…

Disentanglement

MemeArena: Automating Context-Aware Unbiased Evaluation of Harmfulness Understanding for Multimodal Large Language Models

2025-10-31 · Zixin Chen, Hongzhan Lin, Kaixin Li, Ziyang Luo 외 arxiv

The proliferation of memes on social media necessitates the capabilities of multimodal Large Language Models (mLLMs) to effectively understand multimodal harmfulness. Existing evaluation approaches predominantly focus on…

Binary Classification

The Importance of Causality in Decision Making: A Perspective on Recommender Systems

2024-09-16 · Emanuele Cavenaghi, Alessio Zanga, Fabio Stella, Markus Zanker

Causality is receiving increasing attention in the Recommendation Systems (RSs) community, which has realised that RSs could greatly benefit from causality to transform accurate predictions into effective and explainable…

Decision MakingRecommendation Systems

A Causal Explainable Guardrails for Large Language Models

2024-05-07 · Zhixuan Chu, Yan Wang, Longfei Li, Zhibo Wang 외

Large Language Models (LLMs) have shown impressive performance in natural language tasks, but their outputs can exhibit undesirable attributes or biases. Existing methods for steering LLMs toward desired attributes often…