paper-with-me

Papers

Training on Documents About Monitoring Leads to CoT Obfuscation

2026-05-14 · Reilly Haskins, Bilal Chughtai, Joshua Engels arxiv

Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study whether monitor-aware models are capable of obfuscating their reasoning to evade detection. We use synthetic document finetuning to expose eight models to realistic pre-training-style documents describing a CoT monitor and find that monitor-aware models consistently achieve higher rates of undetected misbehavior compared to unaware controls. This effect is weaker but still present on a harder agentic task. We also show that CoT controllability, a model's ability to reshape its own reasoning trace under an imposed constraint, is closely correlated with obfuscation success across the eight models studied ($r=0.800$, $p=0.017$). Monitor-aware models placed under equal reinforcement learning optimization pressure also learn to reward-hack without triggering a CoT monitor substantially faster than unaware controls. Together, these results suggest that knowledge of monitoring combined with high CoT controllability poses a risk to CoT-based monitoring.

📄 PDF Abstract BibTeX arXiv:2605.15257

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

A Systematic Exploration of Text Decomposition and Budget Distribution in Differentially Private Text Obfuscation

2026-05-01 · Stephen Meisenbacher, Angelo Kleinert, Florian Matthes arxiv

The goal of differentially private text obfuscation is to obfuscate, or "perturb", input texts with Differential Privacy (DP) guarantees, such that the private output texts are quantifiably indistinguishable from the ori…

Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability

2025-10-21 · Artur Zolkowski, Wen Xing, David Lindner, Florian Tramèr 외 arxiv

Recent findings suggest that misaligned models may exhibit deceptive behavior, raising concerns about output trustworthiness. Chain-of-thought (CoT) is a promising tool for alignment monitoring: when models articulate th…

Adversarial Authorship Attribution for Deobfuscation

2022-05-01 · ACL 2022 5 · Wanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan

Recent advances in natural language processing have enabled powerful privacy-invasive authorship attribution. To counter authorship attribution, researchers have proposed a variety of rule-based and learning-based text o…

Authorship Attribution

A Girl Has A Name, And It's ... Adversarial Authorship Attribution for Deobfuscation

2022-03-22 · Wanyue Zhai, Jonathan Rusert, Zubair Shafiq, Padmini Srinivasan

Recent advances in natural language processing have enabled powerful privacy-invasive authorship attribution. To counter authorship attribution, researchers have proposed a variety of rule-based and learning-based text o…

Authorship Attribution

Prompt Obfuscation for Large Language Models

2024-09-17 · David Pape, Sina Mavali, Thorsten Eisenhofer, Lea Schönherr

System prompts that include detailed instructions to describe the task performed by the underlying LLM can easily transform foundation models into tools and services with minimal overhead. Because of their crucial impact…

Large Language ModelSemantic SimilaritySemantic Textual Similarity