paper-with-me

홈 › Papers

On Local Overfitting and Forgetting in Deep Neural Networks

2024-12-17 · Uri Stern, Tomer Yaacoby, Daphna Weinshall

The infrequent occurrence of overfitting in deep neural networks is perplexing: contrary to theoretical expectations, increasing model size often enhances performance in practice. But what if overfitting does occur, though restricted to specific sub-regions of the data space? In this work, we propose a novel score that captures the forgetting rate of deep models on validation data. We posit that this score quantifies local overfitting: a decline in performance confined to certain regions of the data space. We then show empirically that local overfitting occurs regardless of the presence of traditional overfitting. Using the framework of deep over-parametrized linear models, we offer a certain theoretical characterization of forgotten knowledge, and show that it correlates with knowledge forgotten by real deep models. Finally, we devise a new ensemble method that aims to recover forgotten knowledge, relying solely on the training history of a single network. When combined with self-distillation, this method enhances the performance of any trained model without adding inference costs. Extensive empirical evaluations demonstrate the efficacy of our method across multiple datasets, contemporary neural network architectures, and training protocols.

📄 PDF Abstract BibTeX arXiv:2412.12968

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Forget Me Not: Fighting Local Overfitting with Knowledge Fusion and Distillation

2025-07-11 · Uri Stern, Eli Corn, Daphna Weinshall arxiv

Overfitting in deep neural networks occurs less frequently than expected. This is a puzzling observation, as theory predicts that greater model capacity should eventually lead to overfitting -- yet this is rarely seen in…

Knowledge Distillation

Making Pre-trained Language Models Better Continual Few-Shot Relation Extractors

2024-02-24 · Shengkun Ma, Jiale Han, Yi Liang, Bo Cheng

Continual Few-shot Relation Extraction (CFRE) is a practical problem that requires the model to continuously learn novel relations while avoiding forgetting old ones with few labeled training data. The primary challenges…

Contrastive LearningPrompt LearningRelationRelation Extraction

ProCal: Probability Calibration for Neighborhood-Guided Source-Free Domain Adaptation

2026-03-19 · Ying Zheng, Yiyi Zhang, Yi Wang, Lap-Pui Chau arxiv

Source-Free Domain Adaptation (SFDA) adapts pre-trained models to unlabeled target domains without requiring access to source data. Although state-of-the-art methods leveraging local neighborhood structures show promise …

Source-Free Domain Adaptation

Lifelong Event Detection with Embedding Space Separation and Compaction

2024-04-03 · Chengwei Qin, Ruirui Chen, Ruochen Zhao, Wenhan Xia 외

To mitigate forgetting, existing lifelong event detection methods typically maintain a memory module and replay the stored memory data during the learning of a new task. However, the simple combination of memory data and…

Event DetectionTransfer Learning

Adversarially Diversified Rehearsal Memory (ADRM): Mitigating Memory Overfitting Challenge in Continual Learning

2024-05-20 · Hikmat Khan, Ghulam Rasool, Nidhal Carla Bouaynaya

Continual learning focuses on learning non-stationary data distribution without forgetting previous knowledge. Rehearsal-based approaches are commonly used to combat catastrophic forgetting. However, these approaches suf…

Continual LearningDiversity