paper-with-me

홈 › Papers

Dynamically Allocating Evaluation Effort for Model Ranking

2026-08-04 · Vilém Zouhar, Julia Kreutzer, Alon Lavie, Tom Kocmi, Matt Post, Ondřej Bojar, Mrinmaya Sachan arxiv

While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing models, typical evaluation protocols waste effort by exhaustively evaluating all models on the entire benchmark, a safe but inefficient approach. In this work, we formalize multi-model human evaluation as a best-arm identification problem in a multi-armed bandit setup with correlated arms, where pulling an arm corresponds to human-evaluating a model. By sampling adaptively based on the intermediate model rankings obtained on the samples so far, we can focus the annotation budget on the most competitive models. We prove the optimality of the proposed algorithms and show that it improves discrimination between top-performing models. This makes evaluations faster, cheaper and more aligned with large-scale competition evaluation goals.

📄 PDF Abstract BibTeX arXiv:2608.03437

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Sharing Credit for Joint Research

2023-07-22 · Nicholas Wu

How can one efficiently share payoffs with collaborators when participating in risky research? First, I show that efficiency can be achieved by allocating payoffs asymmetrically between the researcher who makes a breakth…

Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference

2026-08-19 · Zishan Ahmad, Vishal Vaddina arxiv

Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the…

Consistent Accelerated Inference via Confident Adaptive Transformers

2021-04-18 · EMNLP 2021 11 · Tal Schuster, Adam Fisch, Tommi Jaakkola, Regina Barzilay

We develop a novel approach for confidently accelerating inference in the large and expensive multilayer Transformers that are now ubiquitous in natural language processing (NLP). Amortized or approximate computational m…

Computational EfficiencyConformal PredictionPredictionregression

Effort-Based Criticality Metrics for Evaluating 3D Perception Errors in Autonomous Driving

2026-03-30 · Sharang Kaul, Simon Bultmann, Mario Berk, Abhinav Valada arxiv

Criticality metrics such as time-to-collision (TTC) quantify collision urgency but do not distinguish the operational consequences of false-positive (FP) and false-negative (FN) perception errors. We formulate two error-…

Autonomous Driving

Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms

2025-01-23 · Rajvardhan Oak, Muhammad Haroon, Claire Jo, Magdalena Wojcieszak 외

Social media platforms utilize Machine Learning (ML) and Artificial Intelligence (AI) powered recommendation algorithms to maximize user engagement, which can result in inadvertent exposure to harmful content. Current mo…

Re-Ranking