paper-with-me

홈 › Papers

Continual Learning in Large Language Models: Methods, Challenges, and Opportunities

2026-03-13 · Hongyang Chen, Zhongwu Sun, Hongfei Ye, Kunchi Li, Xuemin Lin arxiv

Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge and sequential tasks while mitigating catastrophic forgetting-a critical limitation of the static pre-training paradigm inherent to modern LLMs. This survey presents a comprehensive overview of CL methodologies tailored for LLMs, structured around three core training stages: continual pre-training, continual fine-tuning, and continual alignment.Beyond the canonical taxonomy of rehearsal-, regularization-, and architecture-based methods, we further subdivide each category by its distinct forgetting mitigation mechanisms and conduct a rigorous comparative analysis of the adaptability and critical improvements of traditional CL methods for LLMs. In doing so, we explicitly highlight core distinctions between LLM CL and traditional machine learning, particularly with respect to scale, parameter efficiency, and emergent capabilities. Our analysis covers essential evaluation metrics, including forgetting rates and knowledge transfer efficiency, along with emerging benchmarks for assessing CL performance. This survey reveals that while current methods demonstrate promising results in specific domains, fundamental challenges persist in achieving seamless knowledge integration across diverse tasks and temporal scales. This systematic review contributes to the growing body of knowledge on LLM adaptation, providing researchers and practitioners with a structured framework for understanding current achievements and future opportunities in lifelong learning for language models.

📄 PDF Abstract BibTeX arXiv:2603.12658

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training

2025-03-04 · Vaibhav Singh, Paul Janson, Paria Mehrbod, Adam Ibrahim 외

The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has emerged as a powerful paradigm for extract…

Self-Supervised Learning

Continual Learning on Graphs: Challenges, Solutions, and Opportunities

2024-02-18 · Xikun Zhang, Dongjin Song, DaCheng Tao

Continual learning on graph data has recently attracted paramount attention for its aim to resolve the catastrophic forgetting problem on existing tasks while adapting the sequentially updated model to newly emerged grap…

Continual LearningGraph Learning

CVPR 2020 Continual Learning in Computer Vision Competition: Approaches, Results, Current Challenges and Future Directions

2020-09-14 · Vincenzo Lomonaco, Lorenzo Pellegrini, Pau Rodriguez, Massimo Caccia 외

In the last few years, we have witnessed a renewed and fast-growing interest in continual learning with deep neural networks with the shared objective of making current AI systems more adaptive, efficient and autonomous.…

BenchmarkingContinual Learning

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

2024-08-14 · Enneng Yang, Li Shen, Guibing Guo, Xingwei Wang 외

Model merging is an efficient empowerment technique in the machine learning community that does not require the collection of raw training data and does not require expensive computation. As model merging becomes increas…

Continual LearningFew-Shot LearningMulti-Task Learning

Online Continual Learning: A Systematic Literature Review of Approaches, Challenges, and Benchmarks

2025-01-09 · Seyed Amir Bidaki, Amir Mohammadkhah, Kiyan Rezaee, Faeze Hassani 외

Online Continual Learning (OCL) is a critical area in machine learning, focusing on enabling models to adapt to evolving data streams in real-time while addressing challenges such as catastrophic forgetting and the stabi…

Continual Learningimage-classificationImage Classificationobject-detection+3