paper-with-me

홈 › Papers

Stochastic Gradient Descent-Induced Drift of Representation in a Two-Layer Neural Network

2023-02-06 · Farhad Pashakhanloo, Alexei Koulakov

Representational drift refers to over-time changes in neural activation accompanied by a stable task performance. Despite being observed in the brain and in artificial networks, the mechanisms of drift and its implications are not fully understood. Motivated by recent experimental findings of stimulus-dependent drift in the piriform cortex, we use theory and simulations to study this phenomenon in a two-layer linear feedforward network. Specifically, in a continual online learning scenario, we study the drift induced by the noise inherent in the Stochastic Gradient Descent (SGD). By decomposing the learning dynamics into the normal and tangent spaces of the minimum-loss manifold, we show the former corresponds to a finite variance fluctuation, while the latter could be considered as an effective diffusion process on the manifold. We analytically compute the fluctuation and the diffusion coefficients for the stimuli representations in the hidden layer as functions of network parameters and input distribution. Further, consistent with experiments, we show that the drift rate is slower for a more frequently presented stimulus. Overall, our analysis yields a theoretical framework for better understanding of the drift phenomenon in biological and artificial neural networks.

📄 PDF Abstract BibTeX arXiv:2302.02563

Code (0)

등록된 구현이 없습니다.

Tasks

Continual Learning

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

On the Provable Suboptimality of Momentum SGD in Nonstationary Stochastic Optimization

2026-01-18 · Sharan Sahu, Cameron J. Hogan, Martin T. Wells arxiv

In this paper, we provide a comprehensive theoretical analysis of Stochastic Gradient Descent (SGD) and its momentum variants (Polyak Heavy-Ball and Nesterov) for tracking time-varying optima under strong convexity and s…

Stochastic Optimization

Noise-induced degeneration in online learning

2020-08-24 · Yuzuru Sato, Daiji Tsutsui, Akio Fujiwara

In order to elucidate the plateau phenomena caused by vanishing gradient, we herein analyse stability of stochastic gradient descent near degenerated subspaces in a multi-layer perceptron. In stochastic gradient descent …

Contribution of task-irrelevant stimuli to drift of neural representations

2025-10-24 · Farhad Pashakhanloo arxiv

Biological and artificial learners are inherently exposed to a stream of data and experience throughout their lifetimes and must constantly adapt to, learn from, or selectively ignore the ongoing input. Recent findings r…

Mitigating Heterogeneity-Induced Drift in Hierarchical Sign-Based Federated Learning

2026-02-02 · Amirreza Kazemi, Seyed Mohammad Azimi-Abarghouyi, Gabor Fodor, Carlo Fischione arxiv

Hierarchical federated learning (HFL) is well suited for large-scale wireless and Internet of Things systems, where devices communicate with nearby edge servers before reaching the cloud. In these environments, uplink ba…

Federated Learning

Second-Order Path Kernel Interpolation Formulas in Machine Learning

2026-06-05 · Jin Guo, Roy Y. He, Jean-Michel Morel arxiv

Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gr…

Stochastic Optimization