paper-with-me

홈 › Papers

Online Curvature-Aware Replay: Leveraging $\mathbf{2^{nd}}$ Order Information for Online Continual Learning

2025-02-03 · Edoardo Urettini, Antonio Carta

Online Continual Learning (OCL) models continuously adapt to nonstationary data streams, usually without task information. These settings are complex and many traditional CL methods fail, while online methods (mainly replay-based) suffer from instabilities after the task shift. To address this issue, we formalize replay-based OCL as a second-order online joint optimization with explicit KL-divergence constraints on replay data. We propose Online Curvature-Aware Replay (OCAR) to solve the problem: a method that leverages second-order information of the loss using a K-FAC approximation of the Fisher Information Matrix (FIM) to precondition the gradient. The FIM acts as a stabilizer to prevent forgetting while also accelerating the optimization in non-interfering directions. We show how to adapt the estimation of the FIM to a continual setting stabilizing second-order optimization for non-iid data, uncovering the role of the Tikhonov regularization in the stability-plasticity tradeoff. Empirical results show that OCAR outperforms state-of-the-art methods in continual metrics achieving higher average accuracy throughout the training process in three different benchmarks.

📄 PDF Abstract BibTeX arXiv:2502.01866

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Similar Papers 제목 키워드 기반

Design Criteria for SGD Preconditioners: Local Conditioning, Noise Floors, and Basin Stability

2025-11-24 · Mitchell Scott, Tianshi Xu, Ziyuan Tang, Alexandra Pichette-Emmons 외 arxiv

Stochastic Gradient Descent (SGD) often slows in the late stage of training due to anisotropic curvature and gradient noise. We analyze preconditioned SGD in the geometry induced by a symmetric positive definite matrix $…

On the Superlinear Relationship between SGD Noise Covariance and Loss Landscape Curvature

2026-02-05 · Yikuan Zhang, Ning Yang, Yuhai Tu arxiv

Stochastic Gradient Descent (SGD) introduces anisotropic noise that is correlated with the local curvature of the loss landscape, thereby biasing optimization toward flat minima. Prior work often assumes an equivalence b…

FlatVPR: Plug-and-play Geo-linear Residual Adapter for Geometric Rectification of Foundation Model Feature Manifolds

2026-06-01 · Rai Hisada, Kanji Tanaka arxiv

This paper proposes ``FlatVPR,'' a novel geometric rectification paradigm that effectively bridges the trade-off between map lightweightness and localization accuracy in visual place recognition (VPR) by enforcing a feat…

Visual Place Recognition

Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning

2026-07-25 · Tao Zhang, Qixuan Fan, Yiyuan Liang, Yanjie Wang 외 arxiv

Class-incremental learning (CIL) requires models to continuously acquire new knowledge while avoiding catastrophic forgetting. While exemplar replay is effective, it raises concerns regarding privacy and storage. Thus, g…

class-incremental learningClass Incremental Learning

Stateful Reasoning via Insight Replay

2026-05-14 · Bin Lei, Caiwen Ding, Jiachen Yang, Ang Li 외 arxiv

Chain-of-Thought (CoT) reasoning has become a foundation for eliciting multi-step reasoning in large language models, but recent studies show that its benefits do not scale monotonically with chain length: while longer C…