paper-with-me

Papers

Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering

2026-02-03 · Eitan Sprejer, Oscar Agustín Stanchi, María Victoria Carro, Denise Alejandra Mester, Iván Arcuschin arxiv

Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineering. However, its practical effectiveness in real-world applications remains poorly understood, particularly regarding potential trade-offs with output quality. We show that feature steering methods substantially degrade model performance even when successfully controlling target behaviors, a critical trade-off. Specifically, we evaluate Goodfire's Auto Steer against prompt engineering baselines across 14 steering queries (covering innocuous and safety-relevant behaviors) on 171 Massive Multitask Language Understanding (MMLU) questions using Llama-8B and Llama-70B, measuring accuracy, coherence, and behavioral control. Our findings show that Auto Steer successfully modifies target behaviors (achieving scores of 3.33 vs. 2.98 for prompting on Llama-8B and 3.57 vs. 3.10 on Llama-70B), but causes dramatic performance degradation: accuracy on the MMLU questions drops from 66% to 46% on Llama-8B and 87% to 73% on Llama-70B, with coherence falling from 4.62 to 2.24 and 4.94 to 3.89 respectively. Simple prompting achieves the best overall balance. These findings highlight limitations of current feature steering methods for practical deployment where task performance cannot be sacrificed. More broadly, our work demonstrates that mechanistic control methods face fundamental capability-behavior trade-offs that must be empirically characterized before deployment.

📄 PDF Abstract BibTeX arXiv:2602.04903

Code (0)

등록된 구현이 없습니다.

Tasks

Prompt Engineering

Similar Papers 제목 키워드 기반

Heterogeneity in Multi-Robot Environmental Monitoring for Resolving Time-Conflicting Tasks

2025-12-09 · Connor York, Zachary R Madin, Paul O'Dowd, Edmund R Hunt arxiv

Multi-robot systems performing continuous tasks face a performance trade-off when interrupted by urgent, time-critical sub-tasks. We investigate this trade-off in a scenario where a team must balance area patrolling with…

Inside you are many wolves: Using cognitive models to interpret value trade-offs in LLMs

2025-06-25 · Sonia K. Murthy, Rosie Zhao, Jennifer Hu, Sham Kakade 외

Navigating everyday social situations often requires juggling conflicting goals, such as conveying a harsh truth, maintaining trust, all while still being mindful of another person's feelings. These value trade-offs are …

Mathematical Reasoning

Burn-in, bias, and the rationality of anchoring

2012-12-01 · NeurIPS 2012 12 · Falk Lieder, Tom Griffiths, Noah Goodman

Bayesian inference provides a unifying framework for addressing problems in machine learning, artificial intelligence, and robotics, as well as the problems facing the human mind. Unfortunately, exact Bayesian inference …

Bayesian Inference

Beyond Mimicry: Preference Coherence in LLMs

2025-11-17 · Luhan Mikaelson, Derek Shiller, Hayley Clatterbuck arxiv

We investigate whether large language models exhibit genuine preference structures by testing their responses to AI-specific trade-offs involving GPU reduction, capability restrictions, shutdown, deletion, oversight, and…

Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects

2025-09-05 · Gunmay Handa, Zekun Wu, Adriano Koshiyama, Philip Treleaven arxiv

Personality manipulation in large language models (LLMs) is increasingly applied in customer service and agentic scenarios, yet its mechanisms and trade-offs remain unclear. We present a systematic study of personality c…

parameter-efficient fine-tuning