paper-with-me

홈 › Papers

On Rank-Dependent Generalisation Error Bounds for Transformers

2024-10-15 · Lan V. Truong

In this paper, we introduce various covering number bounds for linear function classes, each subject to different constraints on input and matrix norms. These bounds are contingent on the rank of each class of matrices. We then apply these bounds to derive generalization errors for single layer transformers. Our results improve upon several existing generalization bounds in the literature and are independent of input sequence length, highlighting the advantages of employing low-rank matrices in transformer design. More specifically, our achieved generalisation error bound decays as $O(1/\sqrt{n})$ where $n$ is the sample length, which improves existing results in research literature of the order $O((\log n)/(\sqrt{n}))$. It also decays as $O(\log r_w)$ where $r_w$ is the rank of the combination of query and and key matrices.

📄 PDF Abstract BibTeX arXiv:2410.11500

Code (0)

등록된 구현이 없습니다.

Tasks

Generalization Bounds

Similar Papers 제목 키워드 기반

Quantum Reservoir Computing and Risk Bounds

2025-01-15 · Naomi Mona Chmielewski, Nina Amini, Joseph Mikael

We propose a way to bound the generalisation errors of several classes of quantum reservoirs using the Rademacher complexity. We give specific, parameter-dependent bounds for two particular quantum reservoir classes. We …

Exact Generalisation Error Exposes Benchmarks Skew Graph Neural Networks Success (or Failure)

2025-09-12 · Nil Ayday, Mahalakshmi Sabanayagam, Debarghya Ghoshdastidar arxiv

Graph Neural Networks (GNNs) have become the standard method for learning from networks across fields ranging from biology to social systems, yet a principled understanding of what enables them to extract meaningful repr…

Sharper Generalization Bounds for Transformer

2026-03-23 · Yawen Li, Tao Hu, Zhouhui Lian, Wan Tian 외 arxiv

This paper studies generalization error bounds for Transformer models. Based on the offset Rademacher complexity, we derive sharper generalization bounds for different Transformer architectures, including single-layer si…

Chained Generalisation Bounds

2022-03-02 · Eugenio Clerico, Amitis Shidani, George Deligiannidis, Arnaud Doucet

This work discusses how to derive upper bounds for the expected generalisation error of supervised learning algorithms by means of the chaining technique. By developing a general theoretical framework, we establish a dua…

Generalization Bounds for Label Noise Stochastic Gradient Descent

2023-11-01 · Jung Eun Huh, Patrick Rebeschini

We develop generalization error bounds for stochastic gradient descent (SGD) with label noise in non-convex settings under uniform dissipativity and smoothness conditions. Under a suitable choice of semimetric, we establ…

Generalization Bounds