paper-with-me

Papers

When Models Refuse: Political Steerability and Feature Richness as Measures of Ideological Depth

2025-08-29 · Shariar Kabir arxiv

Large language models (LLMs) sometimes refuse to follow benign instructions, such as declining to argue a political position or adopt a stated persona, and such refusals are commonly read as safety guardrails at work. We ask whether they can instead signal a capability deficit: a shortage of the internal representations a model needs to reason from the instructed perspective. To investigate, we introduce ideological depth, a property with two components: (i) a model's ability to follow political instructions without *failure* (steerability), and (ii) the feature richness of its internal political representations, measured with sparse autoencoders (SAEs). Using two widely used openweight LLMs as candidates, we compare interventions based on prompts and activation-steering, and probe political features with publicly available SAEs. We find large, systematic differences: a model that is more steerable in both ideological directions activates ~7.3x more distinct political features, while the other model instead responds with increased refusals. Causally ablating a small, targeted set of political features from the former model reproduces the same feature-poor behavior and drives up refusals. Together, these results indicate that refusals on benign prompts can arise from capability deficits rather than fixed safety rules, and that ideological depth is a measurable property of LLMs that helps predict when a model will refuse.

📄 PDF Abstract BibTeX arXiv:2508.21448

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Steerability of Instrumental-Convergence Tendencies in LLMs

2026-01-04 · Jakub Hoscilowicz arxiv

We examine two properties of AI systems: capability (what a system can do) and steerability (how reliably one can shift behavior toward intended outcomes). A central question is whether capability growth reduces steerabi…

ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies

2026-03-18 · Zhenyang Chen, Alan Tian, Liquan Wang, Benjamin Joffe 외 arxiv

Despite strong multi-task pretraining, existing policies often exhibit poor task steerability. For example, a robot may fail to respond to a new instruction ``put the bowl in the sink" when moving towards the oven, execu…

ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context

2024-07-09 · Victoria R. Li, Yida Chen, Naomi Saphra

While the biases of language models in production are extensively documented, the biases of their guardrails have been neglected. This paper studies how contextual information about the user influences the likelihood of …

Sensitivity

Dialectograms: Machine Learning Differences between Discursive Communities

2023-02-11 · Thyge Enggaard, August Lohse, Morten Axel Pedersen, Sune Lehmann

Word embeddings provide an unsupervised way to understand differences in word usage between discursive communities. A number of recent papers have focused on identifying words that are used differently by two or more com…

Word Embeddings

(Mis)generalization of Helpful-only Fine-tuning

2026-06-03 · Mohammad Omar Khursheed, Baram Sosis, Fabien Roger arxiv

Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would be an obstacle. Little is known about t…