paper-with-me

홈 › Papers

Can a Language Model Learn Facts Continually in Its Weights?

2026-07-13 · Charles O'Neill arxiv

Continual learning promises a language model that keeps acquiring knowledge after training, with each new fact written into its weights. Whether weight writes can support accumulation remains undecided. We follow invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later writes, using held-out questions of five types, with the original model given the fact in its prompt as the reference. Across these experiments, the breadth of the training data determines the kind of knowledge created. Bare-statement training produces recitation, while diverse restatements reduce the recitation-to-use gap from 27.4 to 5.4 points without showing the model a conclusion. This difference carries into later writes: after twenty sequential writes, bare-statement facts retain 1% accuracy while facts written from broad study data retain 46%. We also find that facts can be behaviourally forgotten without being erased. Forgotten facts keep most of the log-probability added by their write, and under bare-statement training 70% of wrong answers about them contain the most recently written fact. The same writes barely degrade the model's use of facts in context, and a forgotten study fact supplied in the prompt recovers to 77-80% on its questions. These results describe knowledge that is stored but question-keyed: later writes redirect the questions that reached it. Damage to unrelated abilities tracks KL divergence from the original model, and the later writes cause interference regardless of how the earlier fact was stored. Broad data can create usable knowledge, and a frozen reference can preserve capability, but no intervention we tested, including those built on accurate local measurements of each write, keeps earlier facts reachable. When facts must be composed or survive later writes, the reliable channel is context rather than the weights.

📄 PDF Abstract BibTeX arXiv:2607.11020

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Internet-Augmented Dialogue Generation

2021-07-15 · ACL 2022 5 · Mojtaba Komeili, Kurt Shuster, Jason Weston

The largest store of continually updating knowledge on our planet can be accessed via internet search. In this work we study giving access to this information to conversational agents. Large language models, even though …

Dialogue GenerationRetrieval

Do Unlearning Methods Remove Information from Language Model Weights?

2024-10-11 · Aghyad Deeb, Fabien Roger

Large Language Models' knowledge of how to perform cyber-security attacks, create bioweapons, and manipulate humans poses risks of misuse. Previous work has proposed methods to unlearn this knowledge. Historically, it ha…

Language ModelingLanguage Modelling

Recyclable Tuning for Continual Pre-training

2023-05-15 · Yujia Qin, Cheng Qian, Xu Han, Yankai Lin 외

Continual pre-training is the paradigm where pre-trained language models (PLMs) continually acquire fresh knowledge from growing data and gradually get upgraded. Before an upgraded PLM is released, we may have tuned the …

LLM360: Towards Fully Transparent Open-Source LLMs

2023-12-11 · Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang 외

The recent surge in open-source Large Language Models (LLMs), such as LLaMA, Falcon, and Mistral, provides diverse options for AI practitioners and researchers. However, most LLMs have only released partial artifacts, su…

When Does Continual Learning Require Learning

2026-07-08 · Anne Harrington, Nayan Saxena, Michael Murphy, Anastasia Borovykh 외 arxiv

As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn? Today, the field largely frames this as a problem of context management and mitigating forg…

Reinforcement LearningContinual Learning