paper-with-me

Papers

Attention Enables Zero Approximation Error

2022-02-24 · Zhiying Fang, Yidong Ouyang, Ding-Xuan Zhou, Guang Cheng

Deep learning models have been widely applied in various aspects of daily life. Many variant models based on deep learning structures have achieved even better performances. Attention-based architectures have become almost ubiquitous in deep learning structures. Especially, the transformer model has now defeated the convolutional neural network in image classification tasks to become the most widely used tool. However, the theoretical properties of attention-based models are seldom considered. In this work, we show that with suitable adaptations, the single-head self-attention transformer with a fixed number of transformer encoder blocks and free parameters is able to generate any desired polynomial of the input with no error. The number of transformer encoder blocks is the same as the degree of the target polynomial. Even more exciting, we find that these transformer encoder blocks in this model do not need to be trained. As a direct consequence, we show that the single-head self-attention transformer with increasing numbers of free parameters is universal. These surprising theoretical results clearly explain the outstanding performances of the transformer model and may shed light on future modifications in real applications. We also provide some experiments to verify our theoretical result.

📄 PDF Abstract BibTeX arXiv:2202.12166

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Learningimage-classificationImage Classification

Similar Papers 제목 키워드 기반

What and How does In-Context Learning Learn? Bayesian Model Averaging, Parameterization, and Generalization

2023-05-30 · Yufeng Zhang, Fengzhuo Zhang, Zhuoran Yang, Zhaoran Wang

In this paper, we conduct a comprehensive study of In-Context Learning (ICL) by addressing several open questions: (a) What type of ICL estimator is learned by large language models? (b) What is a proper performance metr…

In-Context Learning

Zeroth-order Riemannian Averaging Stochastic Approximation Algorithms

2023-09-25 · Jiaxiang Li, Krishnakumar Balasubramanian, Shiqian Ma

We present Zeroth-order Riemannian Averaging Stochastic Approximation (\texttt{Zo-RASA}) algorithms for stochastic optimization on Riemannian manifolds. We show that \texttt{Zo-RASA} achieves optimal sample complexities …

Stochastic Optimization

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

2026-08-26 · Gerard Conangla Planes arxiv

Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. …

Learning Sparsity and Randomness for Data-driven Low Rank Approximation

2022-12-15 · Tiejin Chen, Yicheng Tao

Learning-based low rank approximation algorithms can significantly improve the performance of randomized low rank approximation with sketch matrix. With the learned value and fixed non-zero positions for sketch matrices …

The equivalent constant-elasticity-of-variance (CEV) volatility of the stochastic-alpha-beta-rho (SABR) model

2019-11-29 · Jaehyuk Choi, Lixin Wu

This study presents new analytic approximations of the stochastic-alpha-beta-rho (SABR) model. Unlike existing studies that focus on the equivalent Black-Scholes (BS) volatility, we instead derive the equivalent constant…