paper-with-me

홈 › Papers

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks

2026-06-02 · Alexander Apartsin, Yehudit Aperstein arxiv

Selecting a pretrained language model, or evaluating a fine-tuned one, for a specific application is a high-value decision, yet the public benchmarks used to make it are poorly suited: a generic benchmark need not reflect a particular sub-domain or sub-task, and its scores are suspect when its items have leaked into pretraining and are recalled rather than solved. We present CoEval, an open framework that supplies a trustworthy, task-specific signal through ensemble self-evaluation: from a task or domain description, a pool of models rotates through all three roles, teacher, student, and judge, to generate a fresh, contamination-free benchmark, answer it, and score one another, with no human labels or raters. Because every model also answers as a student, the responses are the data that weight each question by its discriminative power and each judge by its consensus with the panel. Where ground truth exists, CoEval recovers the true ranking and tracks objective correctness at \r{ho}=0.86, and the weighting recovers the gold ranking of thirteen models at Spearman 0.95. Reliability comes from panel composition, not size: this label-free weighting zeroes out broken judges and down-weights saturated questions, so neither distorts the ranking. Generated items show zero verbatim overlap with five public benchmarks, the panel cancels verbosity bias and precludes same-family self-preference, and rankings are domain-specific: three different models top four de-novo domains, so a generic leaderboard misdirects most practitioners. The same pipeline reruns on each model release, giving any team a contamination-free leaderboard for its application.

📄 PDF Abstract BibTeX arXiv:2606.03650

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluation Benchmarks and Learning Criteria for Discourse-Aware Sentence Representations

2019-08-31 · IJCNLP 2019 11 · Mingda Chen, Zewei Chu, Kevin Gimpel

Prior work on pretrained sentence embeddings and benchmarks focus on the capabilities of stand-alone sentences. We propose DiscoEval, a test suite of tasks to evaluate whether sentence representations include broader con…

SentenceSentence Embeddings

LiCoEval: Evaluating LLMs on License Compliance in Code Generation

2024-08-05 · Weiwei Xu, Kai Gao, Hao He, Minghui Zhou

Recent advances in Large Language Models (LLMs) have revolutionized code generation, leading to widespread adoption of AI coding tools by developers. However, LLMs can generate license-protected code without providing th…

Code Generation

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

2026-01-23 · Dongshen Peng, Yi Wang, Austin Schoeffler, Sun-ha Hong 외 arxiv

Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-based guidelines. We developed SycoEval-EM, a multi-agent simulation framework to…

InfiCoEvalChain: A Blockchain-Based Decentralized Framework for Collaborative LLM Evaluation

2026-02-09 · Yifan Yang, Jinjia Li, Kunxi Li, Puhao Zheng 외 arxiv

The rapid advancement of large language models (LLMs) demands increasingly reliable evaluation, yet current centralized evaluation suffers from opacity, overfitting, and hardware-induced variance. Our empirical analysis …

Make Large Language Model a Better Ranker

2024-03-28 · Wen-Shuo Chao, Zhi Zheng, HengShu Zhu, Hao liu

Large Language Models (LLMs) demonstrate robust capabilities across various fields, leading to a paradigm shift in LLM-enhanced Recommender System (RS). Research to date focuses on point-wise and pair-wise recommendation…

Language ModelingLanguage ModellingLarge Language Modelmodel+2