paper-with-me

홈 › Papers

Overtrained, Not Misaligned

2026-05-12 · Joel Schreiber, Ariel Goldstein arxiv

Emergent misalignment (EM), where fine-tuning on a narrow task (like insecure code) causes broad misalignment across unrelated domains, was first demonstrated by Betley et al. (2025). We conduct the most comprehensive EM study to date, reproducing the original GPT-4o finding and expanding to 12 open-source models across 4 families (Llama, Qwen, DeepSeek, GPT-OSS) ranging from 8B to 671B parameters, evaluating over one million model responses with multiple random seeds. We find that EM replicates in GPT-4o but is far from universal: only 2 of 12 open-source models (17%) exhibit consistent EM across seeds, with a significant correlation between model size and EM susceptibility. Through checkpoint-level analysis during fine-tuning, we demonstrate that EM emerges late in training, distinct from and subsequent to near convergence of the primary task, suggesting EM emerges from continued training past task convergence. This yields practical mitigations: early stopping eliminates EM while retaining an average of 93% of task performance, and careful learning rate selection further minimizes risk. Cross-domain validation on medical fine-tuning confirms these patterns generalize: the size-EM correlation strengthens (r = 0.90), and overgeneralization to untruthfulness remains avoidable via early stopping in 67% of cases, though semantically proximate training domains produce less separable misalignment. As LLMs become increasingly integrated into real-world systems, fine-tuning and reinforcement learning remain the primary methods for adapting model behavior. Our findings demonstrate that with proper training practices, EM can be avoided, reframing it from an unforeseen fine-tuning risk to an avoidable training artifact.

📄 PDF Abstract BibTeX arXiv:2605.12199

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Reusing Overtrained Language Models Saturates Scaling

2025-10-08 · Seng Pei Liew, Takuya Kato arxiv

Reusing pretrained base models for further pretraining, such as continual pretraining or model growth, is promising at reducing the cost of training language models from scratch. However, the effectiveness remains unclea…

Continual Pretraining

Quantifying Overfitting: Evaluating Neural Network Performance through Analysis of Null Space

2023-05-30 · Hossein Rezaei, Mohammad Sabokrou

Machine learning models that are overfitted/overtrained are more vulnerable to knowledge leakage, which poses a risk to privacy. Suppose we download or receive a model from a third-party collaborator without knowing its …

Multiple Descents in Deep Learning as a Sequence of Order-Chaos Transitions

2025-05-26 · Wenbo Wei, Nicholas Chong Jia Le, Choy Heng Lai, Ling Feng

We observe a novel 'multiple-descent' phenomenon during the training process of LSTM, in which the test loss goes through long cycles of up and down trend multiple times after the model is overtrained. By carrying out as…

Will we run out of data? Limits of LLM scaling based on human-generated data

2022-10-26 · Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu 외

We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data. We forecast the growing demand for training data based on current trends and estimate the total stock…

Language ModelingLanguage ModellingSynthetic Data GenerationTransfer Learning

Towards Characterizing Cyber Networks with Large Language Models

2024-11-11 · Alaric Hartsock, Luiz Manella Pereira, Glenn Fink

Threat hunting analyzes large, noisy, high-dimensional data to find sparse adversarial behavior. We believe adversarial activities, however they are disguised, are extremely difficult to completely obscure in high dimens…

Language ModelingLanguage Modelling