paper-with-me

홈 › Papers

Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines

2026-06-17 · Jianzhe Lin, Fei Wang, Xiaolin Li, Rajeshkumar Golani, Jubin Chheda arxiv

Industrial LLM teams often ship behavior updates by repeatedly DPO-training a base model on sequences of related preference-data campaigns. The dominant failure mode in this regime is not always classical catastrophic forgetting: a pipeline may preserve previously learned behaviors while still failing to accumulate reusable methodological knowledge about how to train the next campaign. We call this failure mode scientific amnesia. This paper turns that practitioner intuition into a measurable industrial problem. We contribute: (i) a diagnostic suite for amnesia, (ii) a Program-based pipeline that chains FSDP-sharded DPO checkpoints across Qwen2.5-7B-Instruct runs, (iii) a 30-campaign HumanEval subdomain benchmark, and (iv) a comparative diagnostic study of five strategy proposers: random memory, rule-based scheduling, retrieval-only memory, warm-start Bayesian optimization, and MSCL, a meta-scientific memory and reasoner candidate. Across a single-seed 5-condition * 3-step real-LM chain, 4 of 5 candidates degrade in step-level peak pass@1, including MSCL; only the deliberately conservative rule-based schedule improves. Follow-up pilots qualify rather than overturn this finding: in a heterogeneous chain, MSCL is the only completed candidate that improves, whereas in a small multi-seed homogeneous sweep, retrieval-only has the best mean Delta and no pairwise candidate gap is statistically distinguishable. The contribution is therefore diagnostic, not a claim that MSCL solves the problem: scientific amnesia is observable in a production-like continual-DPO pipeline, and conclusions about interventions depend sharply on chain regime, evaluator design, and seed coverage.

📄 PDF Abstract BibTeX arXiv:2606.21089

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

2026-07-28 · Fanqing Meng, Lingxiao Du, Qiguang Chen, Ziqi Zhao 외 arxiv

Recursive self-improvement requires turning evidence of model failures into better models. Data-centric post-training research entails diagnosing capability gaps, designing and validating training-data strategies, and le…

Question Answering

A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

2025-10-21 · Mengqi Li, Lei Zhao, Anthony Man-Cho So, Ruoyu Sun 외 arxiv

Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolving Post-Training (SePT), a simple post-t…

AI scientists produce results without reasoning scientifically

2026-04-20 · Martiño Ríos-García, Nawaf Alampara, Chandan Gupta, Indrajeet Mandal 외 arxiv

Large language model (LLM)-based systems are increasingly deployed to conduct scientific research autonomously, yet whether their reasoning adheres to the epistemic norms that make scientific inquiry self-correcting is p…

Scientific Paper Classification Based on Graph Neural Network with Hypergraph Self-attention Mechanism

2022-10-07 · Jiashun Liu, Zhe Xue, Ang Li

The number of scientific papers has increased rapidly in recent years. How to make good use of scientific papers for research is very important. Through the high-quality classification of scientific papers, researchers c…

Graph Neural NetworkManagement

Catching the Imposter: Self-Supervised Learning of Physical Coherence with Cross-Entity Feature Permutations

2026-08-14 · Aleksei Rozanov, Arvind Renganathan, Vipin Kumar arxiv

Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. We introduce imposter, a …

Self-Supervised Learning