paper-with-me

홈 › Papers

Position: Capability Control Should be a Separate Goal From Alignment

2026-02-05 · Shoaib Ahmed Siddiqui, Eleni Triantafillou, David Krueger, Adrian Weller arxiv

Foundation models are trained on broad data distributions, yielding generalist capabilities that enable many downstream applications but also expand the space of potential misuse and failures. This position paper argues that capability control -- imposing restrictions on permissible model behavior -- should be treated as a distinct goal from alignment. While alignment is often context and preference-driven, capability control aims to impose hard operational limits on permissible behaviors, including under adversarial elicitation. We organize capability control mechanisms across the model lifecycle into three layers: (i) data-based control of the training distribution, (ii) learning-based control via weight- or representation-level interventions, and (iii) system-based control via post-deployment guardrails over inputs, outputs, and actions. Because each layer has characteristic failure modes when used in isolation, we advocate for a defense-in-depth approach that composes complementary controls across the full stack. We further outline key open challenges in achieving such control, including the dual-use nature of knowledge and compositional generalization.

📄 PDF Abstract BibTeX arXiv:2602.05164

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Composition Collapse: Stable Factual Knowledge Does Not Imply Compositional Reasoning

2026-05-26 · Zhe Yu, Wenpeng Xing, Yunzhao Wei, Jie Chen 외 arxiv

Post-training is routinely evaluated through aggregate benchmark scores that treat multi-hop reasoning as a single capability -- as if a model that answers more questions correctly must be better at assembling facts. We …

Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels

2026-07-26 · Haining Zheng, Qian Dong, Rodolfo K. Depena, Jonathan D. Bhatia 외 arxiv

As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing with what they should be permitted to do in practice. This paper introduces a gove…

Diagnosing Capability Gaps in Fine-Tuning Data

2026-04-30 · Saeid Asgari Taghanaki, Rakshanda Agarwal, Bruce Sun, Rohan Jha 외 arxiv

Fine-tuning large language models (LLMs) for domain-specific tasks requires training datasets that comprehensively cover the target capabilities a practitioner needs. Yet identifying which capabilities a dataset fails to…

Code Generation

Modularization of End-to-End Learning: Case Study in Arcade Games

2019-01-27 · Andrew Melnik, Sascha Fleer, Malte Schilling, Helge Ritter

Complex environments and tasks pose a difficult problem for holistic end-to-end learning approaches. Decomposition of an environment into interacting controllable and non-controllable objects allows supervised learning f…

Atari GamesReinforcement Learning

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

2026-07-30 · Ming Wang, Jiaqi Wu Young, Wenfang Wu, Daling Wang 외 arxiv

Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for t…