paper-with-me

홈 › Papers

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks

2026-04-20 · Taylor Lundy, Narun K. Raman, Kevin Leyton-Brown arxiv

LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effectively unlimited number of question variants. This flexibility is valuable, but it makes evaluation expensive -- especially when the goal is not just determining an average score, but reliably identifying a model's weak spots. This paper introduces a new methodology for identifying hard questions in dynamic benchmarks. It leverages COUP, a recent Bayesian optimization algorithm (Graham, Velez & Leyton-Brown, 2026), after introducing several substantive modifications to make the algorithm suitable for practical LLM pipelines. We also wrap it in a tool that supports flexible choices of datasets and utility functions, enabling users to target the kinds of questions they care about (e.g., low-accuracy questions; questions that are unusually hard relative to their measured complexity). In experiments across a range of benchmarks, we show that our method, dubbed $\texttt{QuickScope}$, discovers truly difficult questions more sample efficiently than standard baselines, while also reducing false positives from noisy outcomes.

📄 PDF Abstract BibTeX arXiv:2604.17842

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Improving Certified Robustness via Statistical Learning with Logical Reasoning

2020-02-28 · Zhuolin Yang, Zhikuan Zhao, Boxin Wang, Jiawei Zhang 외

Intensive algorithmic efforts have been made to enable the rapid improvements of certificated robustness for complex ML models recently. However, current robustness certification methods are only able to certify under a …

BIG-bench Machine LearningLogical Reasoning

Towards Assessment of Randomized Smoothing Mechanisms for Certifying Adversarial Robustness

2020-05-15 · Tianhang Zheng, Di Wang, Baochun Li, Jinhui Xu

As a certified defensive technique, randomized smoothing has received considerable attention due to its scalability to large datasets and neural networks. However, several important questions remain unanswered, such as (…

Adversarial Robustness

A Unified framework for randomized smoothing based certified defenses

2019-09-25 · Tianhang Zheng, Di Wang, Baochun Li, Jinhui Xu

Randomized smoothing, which was recently proved to be a certified defensive technique, has received considerable attention due to its scalability to large datasets and neural networks. However, several important question…

NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

2023-12-22 · Lizhou Fan, Wenyue Hua, Lingyao Li, Haoyang Ling 외

Complex reasoning ability is one of the most important features of current LLMs, which has also been leveraged to play an integral role in complex decision-making tasks. Therefore, the investigation into the reasoning ca…

LMI-Net: Linear Matrix Inequality--Constrained Neural Networks via Differentiable Projection Layers

2026-04-07 · Sunbochen Tang, Andrea Goertzen, Navid Azizan arxiv

Linear matrix inequalities (LMIs) have played a central role in certifying stability, robustness, and forward invariance of dynamical systems. Despite rapid development in learning-based methods for control design and ce…