paper-with-me

홈 › Papers

Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization

2024-10-11 · Yang Chen, Yitao Liang, Zhouchen Lin

Low-Dimension-to-High-Dimension (LDHD) generalization is a special case of Out-of-Distribution (OOD) generalization, where the training data are restricted to a low-dimensional subspace of the high-dimensional testing space. Assuming that each instance is generated from a latent variable and the dimension of the latent variable reflects the problem scale, the inherent scaling challenge in length generalization can be captured by the LDHD generalization in the latent space. We theoretically demonstrate that LDHD generalization is generally unattainable without exploiting prior knowledge to provide appropriate inductive bias. Specifically, we explore LDHD generalization in Boolean functions. We verify that different architectures trained with (S)GD converge to \emph{min-degree interpolators w.r.t. different independent sets}. LDHD generalization is achievable if and only if the target function coincides with this inductive bias. Applying the insights from LDHD generalization to length generalization, we explain the effectiveness of CoT as changing the structure latent space to enable better LDHD generalization. We also propose a principle for position embedding design to handle both the inherent LDHD generalization and the nuisances such as the data format. Following the principle, we propose a novel position embedding called RPE-Square that remedies the RPE for dealing with the data format nuisance.

📄 PDF Abstract BibTeX arXiv:2410.08898

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive BiasPosition

Similar Papers 제목 키워드 기반

Approximate Description Length, Covering Numbers, and VC Dimension

2022-09-26 · Amit Daniely, Gal Katzhendler

Recently, Daniely and Granot [arXiv:1910.05697] introduced a new notion of complexity called Approximate Description Length (ADL). They used it to derive novel generalization bounds for neural networks, that despite subs…

Generalization Bounds

High-dimensional Penalty Selection via Minimum Description Length Principle

2018-04-26 · Kohei Miyaguchi, Kenji Yamanishi

We tackle the problem of penalty selection of regularization on the basis of the minimum description length (MDL) principle. In particular, we consider that the design space of the penalty function is high-dimensional. I…

Vocal Bursts Intensity Prediction

Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling Laws

2025-09-29 · Fabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu 외 arxiv

Trained attention layers exhibit striking and reproducible spectral structure of the weights, including low-rank collapse, bulk deformation, and isolated spectral outliers, yet the origin of these phenomena and their imp…

Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count

2024-10-21 · Hanseul Cho, Jaeyoung Cha, Srinadh Bhojanapalli, Chulhee Yun

Transformers often struggle with length generalization, meaning they fail to generalize to sequences longer than those encountered during training. While arithmetic tasks are commonly used to study length generalization,…

Position

Kernel Methods and Multi-layer Perceptrons Learn Linear Models in High Dimensions

2022-01-20 · Mojtaba Sahraee-Ardakan, Melikasadat Emami, Parthe Pandit, Sundeep Rangan 외

Empirical observation of high dimensional phenomena, such as the double descent behaviour, has attracted a lot of interest in understanding classical techniques such as kernel methods, and their implications to explain g…