paper-with-me

Papers

What Is the Alignment Tax?

2026-02-09 · Robin Young arxiv

The alignment tax is widely discussed but has not been formally characterized. We provide a geometric theory of the alignment tax in representation space. Under linear representation assumptions, we define the alignment tax rate as the squared projection of the safety direction onto the capability subspace and derive the Pareto frontier governing safety-capability tradeoffs, parameterized by a single quantity of the principal angle between the safety and capability subspaces. We prove this frontier is tight and show it has a recursive structure. safety-safety tradeoffs under capability constraints are governed by the same equation, with the angle replaced by the partial correlation between safety objectives given capability directions. We derive a scaling law decomposing the alignment tax into an irreducible component determined by data structure and a packing residual that vanishes as $O(m'/d)$ with model dimension $d$, and establish conditions under which capability preservation mediates or resolves conflicts between safety objectives.

📄 PDF Abstract BibTeX arXiv:2603.00047

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interactive AI Alignment: Specification, Process, and Evaluation Alignment

2023-10-23 · Michael Terry, Chinmay Kulkarni, Martin Wattenberg, Lucas Dixon 외

Modern AI enables a high-level, declarative form of interaction: Users describe the intended outcome they wish an AI to produce, but do not actually create the outcome themselves. In contrast, in traditional user interfa…

Concept Space Alignment in Multilingual LLMs

2024-10-01 · Qiwei Peng, Anders Søgaard

Multilingual large language models (LLMs) seem to generalize somewhat across languages. We hypothesize this is a result of implicit vector space alignment. Evaluating such alignment, we see that larger models exhibit ver…

Word Embeddings

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

2026-05-31 · Partha Pratim Saha, Samarth Raina, Mayur Parvatikar, Amit Dhanda 외 arxiv

Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally. Aligned systems still fail under jailbreaks, prompt injection, a…

Concept Alignment as a Prerequisite for Value Alignment

2023-10-30 · Sunayana Rane, Mark Ho, Ilia Sucholutsky, Thomas L. Griffiths

Value alignment is essential for building AI systems that can safely and reliably interact with people. However, what a person values -- and is even capable of valuing -- depends on the concepts that they are currently u…

Concept Alignment

What does Attention in Neural Machine Translation Pay Attention to?

2017-10-09 · IJCNLP 2017 11 · Hamidreza Ghader, Christof Monz

Attention in neural machine translation provides the possibility to encode relevant parts of the source sentence at each translation step. As a result, attention is considered to be an alignment model as well. However, t…

Machine TranslationSentenceTranslation