paper-with-me

Papers

Can Many-Shot In-Context Learning Help LLMs as Evaluators? A Preliminary Empirical Study

2024-06-17 · Mingyang Song, Mao Zheng, Xuan Luo, Yue Pan

Utilizing Large Language Models (LLMs) as evaluators to assess the performance of LLMs has garnered attention. However, this kind of evaluation approach is affected by potential biases within LLMs, raising concerns about the accuracy and reliability of the evaluation results of LLMs. To address this problem, we propose and study two many-shot In-Context Learning (ICL) prompt templates to help LLM evaluators mitigate potential biases: Many-Shot with Reference (MSwR) and Many-Shot without Reference (MSoR). Specifically, the former utilizes in-context examples with model-generated evaluation rationales as references, while the latter does not include these references. Using these prompt designs, we investigate the impact of increasing the number of in-context examples on the consistency and quality of the evaluation results. Experimental results show that advanced LLMs, such as GPT-4o, perform better in the many-shot regime than in the zero-shot and few-shot regimes. Furthermore, when using GPT-4o as an evaluator in the many-shot regime, adopting MSwR as the prompt template performs better than MSoR.

📄 PDF Abstract BibTeX arXiv:2406.11629

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context LearningSelection bias

Similar Papers 제목 키워드 기반

Evaluate What You Can't Evaluate: Unassessable Quality for Generated Response

2023-05-24 · Yongkang Liu, Shi Feng, Daling Wang, Yifei Zhang 외

LLMs (large language models) such as ChatGPT have shown remarkable language understanding and generation capabilities. Although reference-free evaluators based on LLMs show better human alignment than traditional referen…

Dialogue Generation

Large Language Models are Inconsistent and Biased Evaluators

2024-05-02 · Rickard Stureborg, Dimitris Alikaniotis, Yoshi Suhara

The zero-shot capability of Large Language Models (LLMs) has enabled highly flexible, reference-free metrics for various tasks, making LLM evaluators common tools in NLP. However, the robustness of these LLM evaluators r…

Attribute

Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments

2024-06-17 · Han Zhou, Xingchen Wan, Yinhong Liu, Nigel Collier 외

Large language models (LLMs) have shown promising abilities as cost-effective and reference-free evaluators for assessing language generation quality. In particular, pairwise LLM evaluators, which compare two generated t…

FairnessLanguage ModelingLanguage ModellingLarge Language Model+2

Aligning Black-box Language Models with Human Judgments

2025-02-07 · Gerrit J. J. van den Burg, Gen Suzuki, Wei Liu, Murat Sensoy

Large language models (LLMs) are increasingly used as automated judges to evaluate recommendation systems, search engines, and other subjective tasks, where relying on human evaluators can be costly, time-consuming, and …

Recommendation Systems

Exploring the Reliability of Large Language Models as Customized Evaluators for Diverse NLP Tasks

2023-10-30 · Qintong Li, Leyang Cui, Lingpeng Kong, Wei Bi

Previous work adopts large language models (LLMs) as evaluators to evaluate natural language process (NLP) tasks. However, certain shortcomings, e.g., fairness, scope, and accuracy, persist for current LLM evaluators. To…

FairnessMathStory GenerationText Generation