paper-with-me

홈 › Papers

Retrying vs Resampling in AI Control

2026-05-25 · James Lucassen, Adam Kaufman arxiv

AI coding scaffolds like Claude Code and Codex use retrying: blocking actions flagged as risky and continuing the trajectory. We study retrying from an AI control perspective, which treats the model as potentially adversarial. We find that while retrying reduces honest suspicion scores, the untrusted model can exploit monitor rationale to construct sneakier attacks, negating safety gains. We also study resampling: drawing multiple samples from the same context, which does not leak exploitable information. We disentangle design choices that previous work on resampling had bundled together. In BashArena, with Claude Opus 4.6 as the untrusted model and MiMo-V2-Flash as the trusted monitor, drawing five samples per step and auditing on the maximum suspicion score raises safety from 61% to 71% at a 0.3% audit budget, at no cost to usefulness. Selectively resampling only the steps that look suspicious on the first draw recovers 6.2 percentage points of the gain while drawing only 10% as many extra samples. Two of our findings in this setting contradict earlier work on resampling. The first is that auditing based on the maximum across resampled suspicion scores outperforms using the minimum, which is the opposite of what Ctrl-Z found. The second is that executing the least suspicious sample, which is the central mechanism in earlier defer-to-resample protocols, gives only a small empirical safety gain in our setting (+3.9 pp, with the confidence interval overlapping zero).

📄 PDF Abstract BibTeX arXiv:2605.26047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating Control Protocols for Untrusted AI Agents

2025-11-04 · Jon Kutasov, Chloe Loughridge, Yuqi Sun, Henry Sleight 외 arxiv

As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions …

Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying

2026-05-29 · Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno 외 arxiv

In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greed…

Reinforcement Learning

Weighted Poisson-disk Resampling on Large-Scale Point Clouds

2024-12-12 · Xianhe Jiao, Chenlei Lv, Junli Zhao, Ran Yi 외

For large-scale point cloud processing, resampling takes the important role of controlling the point number and density while keeping the geometric consistency. % in related tasks. However, current methods cannot balance…

Advancing Diffusion Models: Alias-Free Resampling and Enhanced Rotational Equivariance

2024-11-14 · Md Fahim Anjum

Recent advances in image generation, particularly via diffusion models, have led to impressive improvements in image synthesis quality. Despite this, diffusion models are still challenged by model-induced artifacts and l…

Computational EfficiencyImage Generation

Resampling Forgery Detection Using Deep Learning and A-Contrario Analysis

2018-03-01 · Arjuna Flenner, Lawrence Peterson, Jason Bunk, Tajuddin Manhar Mohammed 외

The amount of digital imagery recorded has recently grown exponentially, and with the advancement of software, such as Photoshop or Gimp, it has become easier to manipulate images. However, most images on the internet ha…

Deep LearningTwo-sample testing