paper-with-me

홈 › Papers

The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains

2026-05-11 · Jung Min Kang arxiv

Benchmark evaluation across AI and safety-critical domains overwhelmingly relies on simple averaging. We demonstrate that this practice produces substantially misleading rankings when two conditions co-occur: (1) the evaluation matrix is sparse and (2) items vary substantially in difficulty. Through controlled simulation experiments across four domains -- NLP (GLUE), clinical drug trials, autonomous vehicle safety, and cybersecurity -- we show that Spearman rank correlation $ρ$ between simple-average rankings and ground-truth rankings degrades from $ρ= 1.000$ at 100% coverage to $ρ= 0.809$ at 67% coverage with high difficulty heterogeneity (mean over 20 seeds). A standard two-parameter logistic (2PL) Item Response Theory (IRT) model maintains $ρ\geq 0.996$ across all conditions. A 150-condition grid sweep over sparsity $S \in [0, 0.70]$ and difficulty gap $D \in [0.5, 5.0]$ confirms that ranking error forms a failure surface with a strong $S \times D$ interaction ($γ_3 = +0.20$, $t = 13.05$), while IRT maintains $ρ\geq 0.993$ throughout. We discuss implications for Physical AI benchmarking, where evaluation matrices are often incomplete and difficulty gaps are extreme.

📄 PDF Abstract BibTeX arXiv:2605.11205

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

I Can't Believe It's Not Scene Flow!

2024-03-07 · Ishan Khatri, Kyle Vedder, Neehar Peri, Deva Ramanan 외

Current scene flow methods broadly fail to describe motion on small objects, and current scene flow evaluation protocols hide this failure by averaging over many points, with most drawn larger objects. To fix this evalua…

Scaling Limits of Long-Context Transformers

2026-05-08 · Giuseppe Bruno, Shi Chen, Zhengjiang Lin, Yury Polyanskiy 외 arxiv

We study the long-context limit of softmax self-attention with a fixed query and a random context of $n$ i.i.d. keys on the sphere, viewing the inverse temperature $β_n$ as the scaling parameter that decides whether atte…

Depth induces scale-averaging in overparameterized linear Bayesian neural networks

2021-11-23 · Jacob A. Zavatone-Veth, Cengiz Pehlevan

Inference in deep Bayesian neural networks is only fully understood in the infinite-width limit, where the posterior flexibility afforded by increased depth washes out and the posterior predictive collapses to a shallow …

Representation Learning

Scaling Federated Learning for Fine-tuning of Large Language Models

2021-02-01 · Agrin Hilmkil, Sebastian Callh, Matteo Barbieri, Leon René Sütfeld 외

Federated learning (FL) is a promising approach to distributed compute, as well as distributed data, and provides a level of privacy and compliance to legal frameworks. This makes FL attractive for both consumer and heal…

Federated LearningSentiment Analysistext-classificationText Classification

seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs

2025-09-21 · Mohammad Ramezanali, Mo Vazifeh, Paolo Santi arxiv

We introduce seqBench, a parametrized benchmark for probing sequential reasoning limits in Large Language Models (LLMs) through precise, multi-dimensional control over several key complexity dimensions. seqBench allows s…