paper-with-me

홈 › Papers

Early Period of Training Impacts Adaptation for Out-of-Distribution Generalization: An Empirical Study

2024-03-22 · Chen Cecilia Liu, Iryna Gurevych

Prior research shows that differences in the early period of neural network training significantly impact the performance of in-distribution (ID) data of tasks. Yet, the implications of early learning dynamics on out-of-distribution (OOD) generalization remain poorly understood, primarily due to the complexities and limitations of existing analytical techniques. In this work, we investigate the relationship between learning dynamics, OOD generalization under covariate shift and the early period of neural network training. We utilize the trace of Fisher Information and sharpness, focusing on gradual unfreezing (i.e., progressively unfreezing parameters during training) as our methodology for investigation. Through a series of empirical experiments, we show that 1) changing the number of trainable parameters during the early period of training via gradual unfreezing can significantly improve OOD results; 2) the trace of Fisher Information and sharpness can be used as indicators for the removal of gradual unfreezing during the early period of training for better OOD generalization. Our experiments on both image and text data show that the early period of training is a general phenomenon that can provide Pareto improvements in ID and OOD performance with minimal complexity. Our work represents a first step towards understanding how early learning dynamics affect neural network OOD generalization under covariate shift and suggests a new avenue to improve and study this problem.

📄 PDF Abstract BibTeX arXiv:2403.15210

Code (0)

등록된 구현이 없습니다.

Tasks

Out-of-Distribution Generalization

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Toward more accurate and generalizable brain deformation estimators for traumatic brain injury detection with unsupervised domain adaptation

2023-06-08 · Xianghao Zhan, Jiawei Sun, Yuzhe Liu, Nicholas J. Cecchi 외

Machine learning head models (MLHMs) are developed to estimate brain deformation for early detection of traumatic brain injury (TBI). However, the overfitting to simulated impacts and the lack of generalizability caused …

Domain AdaptationUnsupervised Domain Adaptation

Artificial Intelligence for Climate Adaptation: Reinforcement Learning for Climate Change-Resilient Transport

2026-03-06 · Miguel Costa, Arthur Vandervoort, Carolin Schmidt, João Miranda 외 arxiv

Climate change is expected to intensify rainfall and, consequently, pluvial flooding, leading to increased disruptions in urban transportation systems over the coming decades. Designing effective adaptation strategies is…

Reinforcement Learning

The challenges of temporal alignment on Twitter during crises

2021-04-17 · Aniket Pramanick, Tilman Beck, Kevin Stowe, Iryna Gurevych

Language use changes over time, and this impacts the effectiveness of NLP systems. This phenomenon is even more prevalent in social media data during crisis events where meaning and frequency of word usage may change ove…

Domain Adaptation

RDumb: A simple approach that questions our progress in continual test-time adaptation

2023-06-08 · NeurIPS 2023 11 · Ori Press, Steffen Schneider, Matthias Kümmerer, Matthias Bethge

Test-Time Adaptation (TTA) allows to update pre-trained models to changing data distributions at deployment time. While early work tested these algorithms for individual fixed distribution shifts, recent work proposed an…

Test-time Adaptation

On the Use of Sparse Filtering for Covariate Shift Adaptation

2016-07-22 · Fabio Massimo Zennaro, Ke Chen

In this paper we formally analyse the use of sparse filtering algorithms to perform covariate shift adaptation. We provide a theoretical analysis of sparse filtering by evaluating the conditions required to perform covar…