paper-with-me

홈 › Papers

Bayesian Calibration of Win Rate Estimation with LLM Evaluators

2024-11-07 · Yicheng Gao, Gonghan Xu, Zhe Wang, Arman Cohan

Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evaluators naively to compare or judge between different systems can lead to unreliable results due to the intrinsic win rate estimation bias of LLM evaluators. In order to mitigate this problem, we propose two calibration methods, Bayesian Win Rate Sampling (BWRS) and Bayesian Dawid-Skene, both of which leverage Bayesian inference to more accurately infer the true win rate of generative language models. We empirically validate our methods on six datasets covering story generation, summarization, and instruction following tasks. We show that both our methods are effective in improving the accuracy of win rate estimation using LLMs as evaluators, offering a promising direction for reliable automatic text quality evaluation.

📄 PDF Abstract BibTeX arXiv:2411.04424

Code (1)

yale-nlp/bay-calibration-llm-evaluators 공식 구현

Tasks

Bayesian InferenceInstruction FollowingStory Generation

Similar Papers 제목 키워드 기반

Aligning Model Evaluations with Human Preferences: Mitigating Token Count Bias in Language Model Assessments

2024-07-05 · Roland Daynauth, Jason Mars

The SLAM paper demonstrated that on-device Small Language Models (SLMs) are a viable and cost-effective alternative to API-based Large Language Models (LLMs), such as OpenAI's GPT-4, offering comparable performance and s…

Language ModelingLanguage Modellingmodel

SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models

2025-05-21 · Roland Daynauth, Christopher Clarke, Krisztian Flautner, Lingjia Tang 외

The LLM-as-a-Judge paradigm offers a scalable, reference-free approach for evaluating language models. Although several calibration techniques have been proposed to better align these evaluators with human judgment, prio…

Neural posterior estimation for scalable and accurate inverse parameter inference in Li-ion batteries

2026-04-02 · Malik Hassanaly, Corey R. Randall, Peter J. Weddle, Paul J. Gasper 외 arxiv

Diagnosing the internal state of Li-ion batteries is critical for battery research, operation of real-world systems, and prognostic evaluation of remaining lifetime. By using physics-based models to perform probabilistic…

Aligning with Human Judgement: The Role of Pairwise Preference in Large Language Model Evaluators

2024-03-25 · Yinhong Liu, Han Zhou, Zhijiang Guo, Ehsan Shareghi 외

Large Language Models (LLMs) have demonstrated promising capabilities as automatic evaluators in assessing the quality of generated natural language. However, LLMs still exhibit biases in evaluation and often struggle to…

Language ModelingLanguage ModellingLarge Language Model

Reliable Uncertainties for Bayesian Neural Networks using Alpha-divergences

2020-08-15 · Hector J. Hortua, Luigi Malago, Riccardo Volpi

Bayesian Neural Networks (BNNs) often result uncalibrated after training, usually tending towards overconfidence. Devising effective calibration methods with low impact in terms of computational complexity is thus of cen…

parameter estimationregression