paper-with-me

홈 › Papers

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

2026-01-15 · Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien arxiv

Pretraining corpora contain extensive discourse about AI systems, yet the causal influence of this discourse on downstream alignment remains poorly understood. If prevailing descriptions of AI behaviour are predominantly negative, LLMs may internalise corresponding behavioural priors, giving rise to self-fulfilling misalignment. This paper provides the first controlled study of this hypothesis by pretraining 6.9B-parameter LLMs with varying amounts of (mis)alignment discourse. We find that discussion of AI contributes to misalignment. Upsampling synthetic training documents about AI misalignment leads to a notable increase in misaligned behaviour. Conversely, upsampling documents about aligned behaviour reduces misalignment scores from 45% to 9%. We consider this evidence of self-fulfilling alignment. These effects are dampened, but persist through post-training. Our findings establish the study of how pretraining data shapes alignment priors, or alignment pretraining, as a complement to post-training. We recommend practitioners consider pretraining for alignment alongside capabilities. We share our models, data, and evaluations at AlignmentPretraining.ai.

📄 PDF Abstract BibTeX arXiv:2601.10160

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Self-Fulfilling Alignment: Shaping Safety-Oriented Personas via Threat-Related Images

2026-03-09 · Qishun Yang, Shu Yang, Lijie Hu, Di Wang arxiv

Multimodal large language models (MLLMs) face safety misalignment, where visual inputs enable harmful outputs. To address this, existing methods require explicit safety labels or contrastive data; yet, threat-related con…

Identifying Predictions That Influence the Future: Detecting Performative Concept Drift in Data Streams

2024-12-13 · Brandon Gower-Winter, Georg Krempl, Sergey Dragomiretskiy, Tineke Jelsma 외

Concept Drift has been extensively studied within the context of Stream Learning. However, it is often assumed that the deployed model's predictions play no role in the concept drift the system experiences. Closer inspec…

Drift Detection

Code Pretraining Improves Entity Tracking Abilities of Language Models

2024-05-31 · Najoung Kim, Sebastian Schuster, Shubham Toshniwal

Recent work has provided indirect evidence that pretraining language models on code improves the ability of models to track state changes of discourse entities expressed in natural language. In this work, we systematical…

Math

Scaling Laws for Downstream Task Performance of Large Language Models

2024-02-06 · Berivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas 외

Scaling laws provide important insights that can guide the design of large language models (LLMs). Existing work has primarily focused on studying scaling laws for pretraining (upstream) loss. However, in transfer learni…

Machine TranslationTransfer LearningTranslation

Unleashing the Power of Neural Discourse Parsers -- A Context and Structure Aware Approach Using Large Scale Pretraining

2020-11-06 · Grigorii Guz, Patrick Huber, Giuseppe Carenini

RST-based discourse parsing is an important NLP task with numerous downstream applications, such as summarization, machine translation and opinion mining. In this paper, we demonstrate a simple, yet highly accurate disco…

Discourse ParsingMachine TranslationOpinion MiningTranslation