Three Mechanisms of Feature Learning in the Exact Solution of a Latent Variable Model
We identify and exactly solve the learning dynamics of a one-hidden-layer linear model at any finite width whose limits exhibit both the kernel phase and the feature learning phase. We analyze the phase diagram of this model in different limits of common hyperparameters including width, layer-wise learning rates, scale of output, and scale of initialization. Our solution identifies three novel prototype mechanisms of feature learning: (1) learning by alignment, (2) learning by disalignment, and (3) learning by rescaling. In sharp contrast, none of these mechanisms is present in the kernel regime of the model. We empirically demonstrate that these discoveries also appear in deep nonlinear networks in real tasks.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The WTP-WTA Gap for Public Goods: New Insights from Compensating and Equivalent Variation Closed-Form Solutions
This study finds exact closed-form solutions for compensating variation (CV) and equivalent variation (EV) for both marginal and non-marginal changes in public goods given homothetic utility. The parameters for these sol…
FormTransformers Learn Latent Mixture Models In-Context via Mirror Descent
Sequence modelling requires determining which past tokens are causally relevant from the context and their importance: a process inherent to the attention layers in transformers, yet whose underlying learned mechanisms r…
Properties from Mechanisms: An Equivariance Perspective on Identifiable Representation Learning
A key goal of unsupervised representation learning is "inverting" a data generating process to recover its latent properties. Existing work that provably achieves this goal relies on strong assumptions on relationships b…
Representation LearningQuantitative analysis of cell size control mechanisms
Cell size control is crucial for maintaining cellular function and homeostasis. In this study, we develop a first-order partial differential equation model to examine the effects of three key size control mechanisms: the…
SEMIR: Semantic Minor-Induced Representation Learning on Graphs for Visual Segmentation
Segmenting small and sparse structures in large-scale images is fundamentally constrained by voxel-level, lattice-bound computation and extreme class imbalance -- dense, full-resolution inference scales poorly and forces…
Representation LearningGraph Neural NetworkTumor Segmentation