paper-with-me

홈 › Papers

Layer of Truth: Probing Belief Shifts under Continual Pre-Training Poisoning

2025-10-29 · Svetlana Churina, Niranjan Chebrolu, Kokil Jaidka arxiv

We show that continual pretraining on plausible misinformation can overwrite specific factual knowledge in large language models without degrading overall performance. Unlike prior poisoning work under static pretraining, we study repeated exposure to counterfactual claims during continual updates. Using paired fact-counterfact items with graded poisoning ratios, we track how internal preferences between competing facts evolve across checkpoints, layers, and model scales. Even moderate poisoning (50-100%) flips over 55% of responses from correct to counterfactual while leaving ambiguity nearly unchanged. These belief flips emerge abruptly, concentrate in late layers (e.g., Layers 29-36 in 3B models), and are partially reversible via patching (up to 56.8%). The corrupted beliefs generalize beyond poisoned prompts, selectively degrading commonsense reasoning while leaving alignment benchmarks largely intact and transferring imperfectly across languages. These results expose a failure mode of continual pre-training in which targeted misinformation replaces internal factual representations without triggering broad performance collapse, motivating representation-level monitoring of factual integrity during model updates.

📄 PDF Abstract BibTeX arXiv:2510.26829

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Pretraining

Similar Papers 제목 키워드 기반

Representational and Behavioral Stability of Truth in Large Language Models

2025-11-24 · Samantha Dies, Courtney Maynard, Germans Savcisens, Tina Eliassi-Rad arxiv

Large language models (LLMs) are increasingly used as information sources, yet small changes in semantic framing can destabilize their truth judgments. We propose P-StaT (Perturbation Stability of Truth), an evaluation f…

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

2026-03-05 · Siddharth Boppana, Annabel Ma, Max Loeffler, Raphael Sarfati 외 arxiv

We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer, but continues generating tokens without revealing its internal belief. Our ana…

Probing the Geometry of Truth: Consistency and Generalization of Truth Directions in LLMs Across Logical Transformations and Question Answering Tasks

2025-06-01 · Yuntai Bao, Xuhong Zhang, Tianyu Du, Xinkui Zhao 외

Large language models (LLMs) are trained on extensive datasets that encapsulate substantial world knowledge. However, their outputs often include confidently stated inaccuracies. Earlier works suggest that LLMs encode tr…

In-Context LearningNegationQuestion AnsweringWorld Knowledge

BASIL: Bayesian Assessment of Sycophancy in LLMs

2025-08-23 · Katherine Atwell, Pedram Heydari, Anthony Sicilia, Malihe Alikhani arxiv

Sycophancy (overly agreeable or flattering behavior) poses a fundamental challenge for human-AI collaboration, particularly in high-stakes decision-making domains such as health, law, and education. A central difficulty …

Multimodal Belief-Space Covariance Steering with Active Probing and Influence for Interactive Driving

2026-02-16 · Devodita Chakravarty, John Dolan, Yiwei Lyu arxiv

Autonomous driving in complex traffic requires reasoning under uncertainty. Common approaches rely on prediction-based planning or risk-aware control, but these are typically treated in isolation, limiting their ability …

Bayesian InferenceAutonomous Driving