paper-with-me

홈 › Papers

Trait-space Monitoring for Emergent Misalignment During Supervised Finetuning

2026-05-31 · Huy Nghiem, Sy-Tuyen Ho, Sarah Wiegreffe, Hal Daumé arxiv

Emergent misalignment (EM) occurs when narrow finetuning causes a model to behave dangerously outside the finetuning task. Standard training signals can miss this shift, making reliable detection costly if it depends on repeated behavioral evaluation. We ask whether emergent misalignment can instead be detected from internal representations during finetuning. Using seven alignment-relevant traits encoded as linear directions in activation space, we track representational drift across training checkpoints in four open-source 7-9B LLMs. EM-relevant drift concentrates on a low-dimensional axis that explains 65.5% of the variance, revealing a geometric signature in the studied regime. A low-overhead monitor built on this drift profile detects dangerous checkpoints with 2.2% false negative rate, 2.9% false positive rate, and 0.990 AUROC on held-out perturbation types, outperforming unsupervised PCA and SAE baselines. Stress tests on two 14B models, longer finetuning runs, and misaligned starting points identify key deployment boundaries. These results position trait-space monitoring as a practical complement to behavioral evaluation for EM detection during LoRA-based finetuning, while showing that deployment across substantially different regimes may require recalibration.

📄 PDF Abstract BibTeX arXiv:2606.07631

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Monitoring Emergent Reward Hacking During Generation via Internal Activations

2026-03-04 · Patrick Wilhelm, Thorsten Wittkopp, Odej Kao arxiv

Fine-tuned large language models can exhibit reward-hacking behavior arising from emergent misalignment, which is difficult to detect from final outputs alone. While prior work has studied reward hacking at the level of …

Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models

2026-01-06 · Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik 외 arxiv

Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, given that their training data contain…

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

2026-06-07 · Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng 외 arxiv

Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs). By constructing a steering vector from examples of a target behavior and injecting it int…

Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors

2026-06-29 · Maxime Riché, Daniel Tan, Vili Kohonen, Niels Warncke arxiv

Inoculation prompting is a selective-generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), a family of methods that similarly reduce the optimization pressure to learn undes…

BLOCK-EM: Preventing Emergent Misalignment via Latent Blocking

2026-01-31 · Muhammed Ustaomeroglu, Guannan Qu arxiv

Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior, yet also develops undesirable out-of-domain behaviors. We investigate a…