paper-with-me

Papers

Steerability of Instrumental-Convergence Tendencies in LLMs

2026-01-04 · Jakub Hoscilowicz arxiv

We examine two properties of AI systems: capability (what a system can do) and steerability (how reliably one can shift behavior toward intended outcomes). A central question is whether capability growth reduces steerability and risks control collapse. We also distinguish between authorized steerability (builders reliably reaching intended behaviors) and unauthorized steerability (attackers eliciting disallowed behaviors). This distinction highlights a fundamental safety--security dilemma of AI models: safety requires high steerability to enforce control (e.g., stop/refuse), while security requires low steerability for malicious actors to elicit harmful behaviors. This tension presents a significant challenge for open-weight models, which currently exhibit high steerability via common techniques like fine-tuning or adversarial attacks. Using Qwen3 and InstrumentalEval, we find that a short anti-instrumental prompt suffix sharply reduces the measured convergence rate (e.g., shutdown avoidance, self-replication). For Qwen3-30B Instruct, the convergence rate drops from 81.69% under a pro-instrumental suffix to 2.82% under an anti-instrumental suffix. Under anti-instrumental prompting, larger aligned models show lower convergence rates than smaller ones (Instruct: 2.82% vs. 4.23%; Thinking: 4.23% vs. 9.86%). Code is available at github.com/j-hoscilowicz/instrumental_steering.

📄 PDF Abstract BibTeX arXiv:2601.01584

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?

2025-02-16 · Yufei He, Yuexin Li, Jiaying Wu, Yuan Sui 외

As large language models (LLMs) continue to evolve, ensuring their alignment with human goals and values remains a pressing challenge. A key concern is \textit{instrumental convergence}, where an AI system, in optimizing…

reinforcement-learningReinforcement LearningReinforcement Learning (RL)

A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs

2025-05-27 · Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens

Despite advances in large language models (LLMs) on reasoning and instruction-following benchmarks, it remains unclear whether they can reliably produce outputs aligned with a broad variety of user goals, a concept we re…

Instruction FollowingPrompt Engineering

Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

2026-05-07 · Jonas Wiedermann-Möller, Leonard Dung, Maksym Andriushchenko arxiv

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful fo…

An Aristotelian ontology of instrumental goals: Structural features to be managed and not failures to be eliminated

2025-10-29 · Willem Fourie arxiv

Instrumental goals such as resource acquisition, power-seeking, and self-preservation are key to contemporary AI alignment research, yet the phenomenon's ontology remains under-theorised. This article develops an ontolog…

ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies

2026-03-18 · Zhenyang Chen, Alan Tian, Liquan Wang, Benjamin Joffe 외 arxiv

Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction ``put the bowl in the sink" when moving towards the oven, execu…