paper-with-me

Papers

Active Testing of Large Language Models via Approximate Neyman Allocation

2026-05-11 · Zeli Liu, Jiancheng Zhang, Cong Liu, Yinglun Zhu arxiv

Large language models (LLMs) require reliable evaluation from pre-training to test-time scaling, making evaluation a recurring rather than one-off cost. As model scales grow and target tasks increasingly demand expert annotators, both the compute and labeling costs needed for each evaluation rise rapidly. Active testing aims to alleviate this bottleneck by approximating the evaluation result from a small but informative subset of the evaluation pool. However, existing approaches primarily target classification and break down on generative tasks. We introduce a novel active testing algorithm tailored to generative tasks. Our method leverages semantic entropy from surrogate models to stratify the evaluation pool and then conducts approximate Neyman allocation based on signals extracted from these surrogates. Across multiple language and multimodal benchmarks and a range of surrogate-target model pairs, our method significantly improves on baselines and closely tracks Oracle-Neyman, delivering up to 28% MSE reduction over Uniform Sampling and an average of 22.9% budget savings.

📄 PDF Abstract BibTeX arXiv:2605.10075

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Bounding Neyman-Pearson Region with $f$-Divergences

2025-05-13 · Andrew Mullhaupt, Cheng Peng

The Neyman-Pearson region of a simple binary hypothesis testing is the set of points whose coordinates represent the false positive rate and false negative rate of some test. The lower boundary of this region is given by…

LEMMA

Goodness of fit by Neyman-Pearson testing

2023-05-23 · Gaia Grosso, Marco Letizia, Maurizio Pierini, Andrea Wulzer

The Neyman-Pearson strategy for hypothesis testing can be employed for goodness of fit if the alternative hypothesis is selected from data by exploring a rich parametrised family of models, while controlling the impact o…

Stochastic Gradients under Nuisances

2025-08-28 · Facheng Yu, Ronak Mehta, Alex Luedtke, Zaid Harchaoui arxiv

Stochastic gradient optimization is the dominant learning paradigm for a variety of scenarios, from classical supervised learning to modern self-supervised learning. We consider stochastic gradient algorithms for learnin…

Self-Supervised LearningCausal Inference

Optimal Statistical Hypothesis Testing for Social Choice

2020-06-19 · Lirong Xia

We address the following question in this paper: "What are the most robust statistical methods for social choice?'' By leveraging the theory of uniformly least favorable distributions in the Neyman-Pearson framework to f…

Two-sample testing

Variational Inference for Neyman-Scott Processes

2023-03-07 · Chengkuan Hong, Christian R. Shelton

Neyman-Scott processes (NSPs) have been applied across a range of fields to model points or temporal events with a hierarchy of clusters. Markov chain Monte Carlo (MCMC) is typically used for posterior sampling in the mo…

Point ProcessesVariational Inference