paper-with-me

홈 › Papers

Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging

2026-04-29 · Jon-Paul Cacioli arxiv

A predecessor pilot (Cacioli, 2026) found that Llama-3-8B implements prompted sandbagging as positional collapse rather than answer avoidance. However, fixed option ordering in MMLU-Pro left open whether this reflected a model-level position-dominant policy or dataset-level distractor structure. This pre-registered follow-up (3 models, 2,000 MMLU-Pro items, 4 conditions, 24,000 primary trials) added cyclic option-order randomisation as the critical control. The pre-registered item-level same-letter diagnostic did not confirm deterministic position-tracking (same-letter rate 37.3%, below the 50% threshold). However, pre-specified supporting analyses revealed that the response-position distribution under sandbagging was highly stable under complete content rotation (Pearson r = 0.9994; Jensen-Shannon divergence = 0.027, compared to 0.386 between honest and sandbagging conditions). Accuracy spiked to 72.1% when the correct answer coincidentally occupied the preferred position E, and fell to 4.3% at position A. The data provide strong evidence for a soft distributional attractor: under sandbagging instruction, the model enters a low-entropy response-position basin centred on E/F/G that is highly stable and largely content-invariant at the aggregate level. Qwen-2.5-7B served as a negative control (non-compliant, no distributional shift). These results provide evidence, at the 7-9 billion parameter scale, that response-position entropy is a promising black-box behavioural signature of this sandbagging mode.

📄 PDF Abstract BibTeX arXiv:2604.26206

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Control randomisation approach for policy gradient and application to reinforcement learning in optimal switching

2024-04-27 · Robert Denkert, Huyên Pham, Xavier Warin

We propose a comprehensive framework for policy gradient methods tailored to continuous time reinforcement learning. This is based on the connection between stochastic control problems and randomised problems, enabling a…

Policy Gradient Methods

On the importance of block randomisation when designing proteomics experiments

2020-07-13

Randomisation is used in experimental design to reduce the prevalence of unanticipated confounders. Complete randomisation can however create unbalanced designs, for example, grouping all samples of the same condition in…

BlockingExperimental Design

Deep combinatorial optimisation for optimal stopping time problems : application to swing options pricing

2020-01-30 · Thomas Deschatre, Joseph Mikael

A new method for stochastic control based on neural networks and using randomisation of discrete random variables is proposed and applied to optimal stopping time problems. The method models directly the policy and does …

Analysing Deep Reinforcement Learning Agents Trained with Domain Randomisation

2019-12-18 · Tianhong Dai, Kai Arulkumaran, Tamara Gerbert, Samyakh Tukra 외

Deep reinforcement learning has the potential to train robots to perform complex tasks in the real world without requiring accurate models of the robot or its environment. A practical approach is to train agents in simul…

Deep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)

Game Theory with Simulation in the Presence of Unpredictable Randomisation

2024-10-18 · Vojtech Kovarik, Nathaniel Sauerberg, Lewis Hammond, Vincent Conitzer

AI agents will be predictable in certain ways that traditional agents are not. Where and how can we leverage this predictability in order to improve social welfare? We study this question in a game-theoretic setting wher…