paper-with-me

홈 › Papers

AI Control: Improving Safety Despite Intentional Subversion

2023-12-12 · Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien Roger

As large language models (LLMs) become more powerful and are deployed more autonomously, it will be increasingly important to prevent them from causing harmful outcomes. Researchers have investigated a variety of safety techniques for this purpose, e.g. using models to review the outputs of other models, or red-teaming techniques to surface subtle failure modes. However, researchers have not evaluated whether such techniques still ensure safety if the model is itself intentionally trying to subvert them. In this paper, we develop and evaluate pipelines of safety techniques ("protocols") that are robust to intentional subversion. We investigate a scenario in which we want to solve a sequence of programming problems, using access to a powerful but untrusted model (in our case, GPT-4), access to a less powerful trusted model (in our case, GPT-3.5), and limited access to high-quality trusted labor. We investigate protocols that aim to never submit solutions containing backdoors, which we operationalize here as logical errors that are not caught by test cases. We investigate a range of protocols and test each against strategies that the untrusted model could use to subvert them. One protocol is what we call trusted editing. This protocol first asks GPT-4 to write code, and then asks GPT-3.5 to rate the suspiciousness of that code. If the code is below some suspiciousness threshold, it is submitted. Otherwise, GPT-3.5 edits the solution to remove parts that seem suspicious and then submits the edited code. Another protocol is untrusted monitoring. This protocol asks GPT-4 to write code, and then asks another instance of GPT-4 whether the code is backdoored, using various techniques to prevent the GPT-4 instances from colluding. These protocols improve substantially on simple baselines.

📄 PDF Abstract BibTeX arXiv:2312.06942

Code (1)

rgreenblatt/control-evaluations 공식 구현

Tasks

Red Teaming

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Position-Wise Feed-Forward Layer 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…

Similar Papers 제목 키워드 기반

Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion

2026-04-11 · Vishal Pramanik, Maisha Maliha, Susmit Jha, Sumit Kumar Jha arxiv

Large language models remain vulnerable to jailbreak attacks -- inputs designed to bypass safety mechanisms and elicit harmful responses -- despite advances in alignment and instruction tuning. We propose Head-Masked Nul…

Subversion Strategy Eval: Evaluating AI's stateless strategic capabilities against control protocols

2024-12-17 · Alex Mallen, Charlie Griffin, Alessandro Abate, Buck Shlegeris

AI control protocols are plans for usefully deploying AI systems in a way that is safe, even if the AI intends to subvert the protocol. Previous work evaluated protocols by subverting them with a human-AI red team, where…

Diffuse AI Control on Fuzzy Tasks

2026-06-08 · Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton arxiv

AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned with mitigating risks from AI sabotage dist…

Password-Activated Shutdown Protocols for Misaligned Frontier Agents

2025-11-29 · Kai Williams, Rohan Subramani, Francis Rhys Ward arxiv

Frontier AI developers may fail to align or control highly-capable AI agents. In many cases, it could be useful to have emergency shutdown mechanisms which effectively prevent misaligned agents from carrying out harmful …

Governable AI: Provable Safety Under Extreme Threat Models

2025-08-28 · Donglin Wang, Weiyun Liang, Chunyuan Chen, Jing Xu 외 arxiv

As AI rapidly advances, the security risks posed by AI are becoming increasingly severe, especially in critical scenarios, including those posing existential risks. If AI becomes uncontrollable, manipulated, or actively …