paper-with-me

홈 › Papers

Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization

2026-08-13 · Jinhyung Bae arxiv

Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convention, identical for every instance. Whether a non-uniform allocation of a fixed total budget would buy anything has not been measured. We measure it, and we audit the measurement itself. First, on in-distribution workloads the allocation headroom is not detectable. Across three pretrained solvers (POMO, AM, SymNCO) on uniform TSP-100, an oracle allocation computed and evaluated on the same stored samples reports a 2.2-2.6% gain with intervals excluding zero; measured out of sample the same gain is indistinguishable from zero (0.457, 0.015, -0.512 percent). Following the customary in-sample procedure, all three solvers would have supported a published 2%-level gain that does not exist. We calibrate this bias against an instance-wise null in which the true gain is zero by construction; over the ranges we test it does not shrink with more samples or more instances. Second, the same correction that removes the phantom gains preserves a real one. Under distribution shift (a workload mixing uniform and clustered instances), a pre-registered confirmatory experiment finds that allocation guided by held-out sample statistics improves best-of-k by 11.5% (AM, primary endpoint; 95% CI [7.4, 19.7]) and 12.0% (SymNCO, replication) at equal evaluation budget, with the signal-acquisition cost not charged; a pre-registered negative control (POMO, an order of magnitude more robust to shift) shows -0.3% [-0.7, 0.24]. The gain exceeds a frozen distribution-label baseline by 4.2 points [1.9, 7.7]. An exploratory policy charging a 20-sample probe against the same budget retains 3.4% (AM) and 4.6% (SymNCO). We give a correction procedure and a reporting checklist, and release all data, code, and the pre-registration record.

📄 PDF Abstract BibTeX arXiv:2608.13087

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Auditing Marketing Budget Allocation with Hindsight Regret

2026-04-28 · Nilavra Pathak, Olivier Jeunen, Eric Lambert arxiv

Organizations routinely make strategic budget allocations under operational constraints, but often lack a principled way to assess whether realized allocations were close to the best feasible choices in hindsight. We pre…

Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs

2026-01-06 · David Hartmann, Lena Pohlmann, Lelia Hanslik, Noah Gießing 외 arxiv

Large Language Models (LLMs) exhibit systematic biases across demographic groups. Auditing is proposed as an accountability tool for black-box LLM applications, but suffers from resource-intensive query access. We concep…

Sweeping through the Topic Space: Bad luck? Roll again!

2012-04-01 · WS 2012 4 · Martin Riedl, Chris Biemann
Topic Models

Overcoming the Incentive Collapse Paradox

2026-03-27 · Qichuan Yin, Ziwei Su, Shuangning Li arxiv

AI-assisted task delegation is increasingly common, yet human effort in such systems is costly and typically unobserved. Recent work by Bastani and Cachon (2025); Sambasivan et al. (2021) shows that accuracy-based paymen…

Active Learning

Retrying vs Resampling in AI Control

2026-05-25 · James Lucassen, Adam Kaufman arxiv

AI coding scaffolds like Claude Code and Codex use retrying: blocking actions flagged as risky and continuing the trajectory. We study retrying from an AI control perspective, which treats the model as potentially advers…