Homogenized Transformers
We study a random model of deep multi-head self-attention in which the weights are resampled independently across layers and heads, as at initialization of training. Viewing depth as a time variable, the residual stream defines a discrete-time interacting particle system on the unit sphere. We prove that, under suitable joint scalings of the depth, the residual step size, and the number of heads, this dynamics admits a nontrivial homogenized limit. Depending on the scaling, the limit is either deterministic or stochastic with common noise; in the mean-field regime, the latter leads to a stochastic nonlinear Fokker--Planck equation for the conditional law of a representative token. In the Gaussian setting, the limiting drift vanishes, making the homogenized dynamics explicit enough to study representation collapse. This yields quantitative trade-offs between dimension, context length, and temperature, and identifies regimes in which clustering can be mitigated.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
HomoFormer: Homogenized Transformer for Image Shadow Removal
The spatial non-uniformity and diverse patterns of shadow degradation conflict with the weight sharing manner of dominant models which may lead to an unsatisfactory compromise. To tackle with this issue we present a …
DiversityImage Shadow RemovalShadow RemovalHomogenization of Ordinary Differential Equations for the Fast Prediction of Diabetes Progression
The impact of physical activity on a person's progression to type 2 diabetes is multifaceted. Systems of ordinary differential equations have been crucial in simulating this progression. However, such models often operat…
Homogenization of SGD in high-dimensions: Exact dynamics and generalization properties
We develop a stochastic differential equation, called homogenized SGD, for analyzing the dynamics of stochastic gradient descent (SGD) on a high-dimensional random least squares problem with $\ell^2$-regularization. We s…
Vocal Bursts Intensity PredictionNeural Network Layers for Prediction of Positive Definite Elastic Stiffness Tensors
Machine learning models can be used to predict physical quantities like homogenized elasticity stiffness tensors, which must always be symmetric positive definite (SPD) based on conservation arguments. Two datasets of ho…
BIG-bench Machine LearningNeural Network Accelerated Process Design of Polycrystalline Microstructures
Computational experiments are exploited in finding a well-designed processing path to optimize material structures for desired properties. This requires understanding the interplay between the processing-(micro)structure…