paper-with-me

홈 › Papers

Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

2026-03-26 · Michael Hardy, Joshua Gilbert, Benjamin Domingue arxiv

The validity of assessments, from large-scale AI benchmarks to human classrooms, depends on the quality of individual items, yet modern evaluation instruments often contain thousands of items with minimal psychometric vetting. We introduce a new family of nonparametric scalability coefficients based on interitem isotonic regression for efficiently detecting globally bad items (e.g., miskeyed, ambiguously worded, or construct-misaligned). The central contribution is the signed isotonic $R^2$, which measures the maximal proportion of variance in one item explainable by a monotone function of another while preserving the direction of association via Kendall's $τ$. Aggregating these pairwise coefficients yields item-level scores that sharply separate problematic items from acceptable ones without assuming linearity or committing to a parametric item response model. We show that the signed isotonic $R^2$ is extremal among monotone predictors (it extracts the strongest possible monotone signal between any two items) and show that this optimality property translates directly into practical screening power. Across three AI benchmark datasets (HS Math, GSM8K, MMLU) and two human assessment datasets, the signed isotonic $R^2$ consistently achieves top-tier AUC for ranking bad items above good ones, outperforming or matching a comprehensive battery of classical test theory, item response theory, and dimensionality-based diagnostics. Crucially, the method remains robust under the small-n/large-p conditions typical of AI evaluation, requires only bivariate monotone fits computable in seconds, and handles mixed item types (binary, ordinal, continuous) without modification. It is a lightweight, model-agnostic filter that can materially reduce the reviewer effort needed to find flawed items in modern large-scale evaluation regimes.

📄 PDF Abstract BibTeX arXiv:2603.24999

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

S-Walk: Accurate and Scalable Session-based Recommendationwith Random Walks

2022-01-04 · Minjin Choi, jinhong Kim, Joonsek Lee, Hyunjung Shim 외

Session-based recommendation (SR) predicts the next items from a sequence of previous items consumed by an anonymous user. Most existing SR models focus only on modeling intra-session characteristics but pay less attenti…

Computational EfficiencyRecommendation SystemsSession-Based Recommendations

Mitigating Homophily Disparity in Graph Anomaly Detection: A Scalable and Adaptive Approach

2026-03-09 · Yunhui Liu, Qizhuo Xie, Yinfeng Chen, Xudong Jin 외 arxiv

Graph anomaly detection (GAD) aims to identify nodes that deviate from normal patterns in structure or features. While recent GNN-based approaches have advanced this task, they struggle with two major challenges: 1) homo…

Graph Anomaly Detection

Sari Sandbox: A Virtual Retail Store Environment for Embodied AI Agents

2025-08-01 · Janika Deborah Gajo, Gerarld Paul Merales, Jerome Escarcha, Brenden Ashley Molina 외 arxiv

We present Sari Sandbox, a high-fidelity, photorealistic 3D retail store simulation for benchmarking embodied agents against human performance in shopping tasks. Addressing a gap in retail-specific sim environments for e…

Unified Conversational Recommendation Policy Learning via Graph-based Reinforcement Learning

2021-05-20 · Yang Deng, Yaliang Li, Fei Sun, Bolin Ding 외

Conversational recommender systems (CRS) enable the traditional recommender systems to explicitly acquire user preferences towards items and attributes through interactive conversations. Reinforcement learning (RL) is wi…

AttributeConversational RecommendationDecision MakingRecommendation Systems+3

Dynamics of a birth-death process based on combinatorial innovation

2019-08-19

A feature of human creativity is the ability to take a subset of existing items (e.g. objects, ideas, or techniques) and combine them in various ways to give rise to new items, which, in turn, fuel further growth. Occasi…