paper-with-me

홈 › Papers

Evaluating the Evaluators: Are Current Few-Shot Learning Benchmarks Fit for Purpose?

2023-07-06 · Luísa Shimabucoro, Timothy Hospedales, Henry Gouk

Numerous benchmarks for Few-Shot Learning have been proposed in the last decade. However all of these benchmarks focus on performance averaged over many tasks, and the question of how to reliably evaluate and tune models trained for individual tasks in this regime has not been addressed. This paper presents the first investigation into task-level evaluation -- a fundamental step when deploying a model. We measure the accuracy of performance estimators in the few-shot setting, consider strategies for model selection, and examine the reasons for the failure of evaluators usually thought of as being robust. We conclude that cross-validation with a low number of folds is the best choice for directly estimating the performance of a model, whereas using bootstrapping or cross validation with a large number of folds is better for model selection purposes. Overall, we find that existing benchmarks for few-shot learning are not designed in such a way that one can get a reliable picture of how effectively methods can be used on individual tasks.

📄 PDF Abstract BibTeX arXiv:2307.02732

Code (0)

등록된 구현이 없습니다.

Tasks

Few-Shot LearningModel Selection

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

2026-04-14 · Yuangang Li, Justin Tian Jin Chen, Ethan Yu, David Hong 외 arxiv

Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning evaluators are not designed for coding, and…

Code Generation

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

2026-03-31 · Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski 외 arxiv

Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a pri…

Lost in Translation: Do LVLM Judges Generalize Across Languages?

2026-04-21 · Md Tahmid Rahman Laskar, Mohammed Saidul Islam, Mir Tafseer Nayeem, Amran Bhuiyan 외 arxiv

Automatic evaluators such as reward models play a central role in the alignment and evaluation of large vision-language models (LVLMs). Despite their growing importance, these evaluators are almost exclusively assessed o…

Domain Adaptation

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

2025-08-25 · Khaoula Chehbouni, Mohammed Haddou, Jackie Chi Kit Cheung, Golnoosh Farnadi arxiv

Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aims to be general-purpose. Recentl…

Text Summarization

Becoming Experienced Judges: Selective Test-Time Learning for Evaluators

2025-12-07 · Seungyeon Jwa, Daechul Ahn, Reokyoung Kim, Dongyeop Kang 외 arxiv

Automatic evaluation with large language models, commonly known as LLM-as-a-judge, is now standard across reasoning and alignment tasks. Despite evaluating many samples in deployment, these evaluators typically (i) treat…