paper-with-me

홈 › Papers

What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs

2026-01-07 · Seyed Mahed Mousavi, Simone Alghisi, Giuseppe Riccardi arxiv

Continual Pre-Training (CPT) is widely used for acquiring and updating factual knowledge in LLMs. This practice treats loss as a proxy for knowledge learning, while offering no grounding into how it changes during training. We study CPT as a knowledge learning process rather than a solely optimization problem. We construct a controlled, distribution-matched benchmark of factual documents and interleave diagnostic probes directly into the CPT loop, enabling epoch-level measurement of knowledge acquisition dynamics and changes in Out-Of-Domain (OOD) general skills (e.g., math). We further analyze how CPT reshapes knowledge circuits during training. Across three instruction-tuned LLMs and multiple CPT strategies, optimization and learning systematically diverge as loss decreases monotonically while factual learning is unstable and non-monotonic. Acquired facts are rarely consolidated, learning is strongly conditioned on prior exposure, and OOD performance degrades from early epochs. Circuit analysis reveals rapid reconfiguration of knowledge pathways across epochs, providing an explanation for narrow acquisition windows and systematic forgetting. These results show that loss optimization is misaligned with learning progress in CPT and motivate evaluation of stopping criteria based on task-level learning dynamics.

📄 PDF Abstract BibTeX arXiv:2601.03858

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

First, Do No Harm: AI Supervisor Scaffolds Novice Growth in Counselor Education

2025-08-12 · Chen Xu, Zhenyu Lyu, Tian Lan, Yi Yang 외 arxiv

The most dangerous mistakes a novice counselor makes are not the obvious ones: they are utterances that sound caring while quietly violating professional ethics and leaving vulnerable clients less protected. We build an …

Optimization of rule-based energy management strategies for hybrid vehicles using dynamic programming

2022-07-08 · Di Zhu, Ewan Pritchard, Sumanth Reddy Dadam, Vivek Kumar 외

Reducing energy consumption is a key focus for hybrid electric vehicle (HEV) development. The popular vehicle dynamic model used in many energy management optimization studies does not capture the vehicle dynamics that t…

energy managementManagement

D-STEER - Preference Alignment Techniques Learn to Behave, not to Believe -- Beneath the Surface, DPO as Steering Vector Perturbation in Activation Space

2025-12-03 · Samarth Raina, Saksham Aggarwal, Aman Chadha, Vinija Jain 외 arxiv

Direct Preference Optimization (DPO) has become a standard recipe for aligning large language models, yet it is still unclear what kind of change it actually induces inside the network. This paper argues that DPO does no…

Machine Teaching for Bayesian Learners in the Exponential Family

2013-06-20 · NeurIPS 2013 · Xiaojin Zhu

What if there is a teacher who knows the learning goal and wants to design good training data for a machine learner? We propose an optimal teaching framework aimed at learners who employ Bayesian models. Our framework is…

Machine Teaching for Bayesian Learners in the Exponential Family

2013-12-01 · NeurIPS 2013 12 · Jerry Zhu

What if there is a teacher who knows the learning goal and wants to design good training data for a machine learner? We propose an optimal teaching framework aimed at learners who employ Bayesian models. Our framework …