paper-with-me

Papers

SLMEval: Entropy-Based Calibration for Human-Aligned Evaluation of Large Language Models

2025-05-21 · Roland Daynauth, Christopher Clarke, Krisztian Flautner, Lingjia Tang, Jason Mars

The LLM-as-a-Judge paradigm offers a scalable, reference-free approach for evaluating language models. Although several calibration techniques have been proposed to better align these evaluators with human judgment, prior studies focus primarily on narrow, well-structured benchmarks. As a result, it remains unclear whether such calibrations generalize to real-world, open-ended tasks. In this work, we show that SOTA calibrated evaluators often fail in these settings, exhibiting weak or even negative correlation with human judgments. To address this, we propose SLMEval, a novel and efficient calibration method based on entropy maximization over a small amount of human preference data. By estimating a latent distribution over model quality and reweighting evaluator scores accordingly, SLMEval achieves strong correlation with human evaluations across two real-world production use cases and the public benchmark. For example, on one such task, SLMEval achieves a Spearman correlation of 0.57 with human judgments, while G-Eval yields a negative correlation. In addition, SLMEval reduces evaluation costs by 5-30x compared to GPT-4-based calibrated evaluators such as G-eval.

📄 PDF Abstract BibTeX arXiv:2505.16003

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Focus 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

Scoring Rules and Calibration for Imprecise Probabilities

2024-10-30 · Christian Fröhlich, Robert C. Williamson

What does it mean to say that, for example, the probability for rain tomorrow is between 20% and 30%? The theory for the evaluation of precise probabilistic forecasts is well-developed and is grounded in the key concepts…

TruthTensor: Evaluating LLMs through Human Imitation on Prediction Market under Drift and Holistic Reasoning

2026-01-20 · Shirin Shahabi, Spencer Graham, Haruna Isah arxiv

Evaluating language models and AI agents remains fundamentally challenging because static benchmarks fail to capture real-world uncertainty, distribution shift, and the gap between isolated task accuracy and human-aligne…

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

2026-06-29 · Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li 외 arxiv

Calibration evaluates whether a model confidence aligns with its empirical accuracy. Existing studies often compare the calibration of different large language models using global calibration metrics such as Expected Cal…

Fine-Tuning Language Models for Ethical Ambiguity: A Comparative Study of Alignment with Human Responses

2024-10-10 · Pranav Senthilkumar, Visshwa Balasubramanian, Prisha Jain, Aneesa Maity 외

Language models often misinterpret human intentions due to their handling of ambiguity, a limitation well-recognized in NLP research. While morally clear scenarios are more discernible to LLMs, greater difficulty is enco…

Moral Scenarios

Large Language Models are not Fair Evaluators

2023-05-29 · Peiyi Wang, Lei LI, Liang Chen, Zefan Cai 외

In this paper, we uncover a systematic bias in the evaluation paradigm of adopting large language models~(LLMs), e.g., GPT-4, as a referee to score and compare the quality of responses generated by candidate models. We f…

Language ModellingLarge Language ModelPosition