paper-with-me

홈 › Papers

Standard Transformers Achieve the Minimax Rate in Nonparametric Regression with $C^{s,λ}$ Targets

2026-02-24 · Yanming Lai, Defeng Sun arxiv

The tremendous success of Transformer models in fields such as large language models and computer vision necessitates a rigorous theoretical investigation. To the best of our knowledge, this paper is the first work proving that standard Transformers can approximate Hölder functions $ C^{s,λ}\left([0,1]^{d\times n}\right) $$ (s\in\mathbb{N}_{\geq0},0<λ\leq1) $ under the $L^t$ distance ($t \in [1, \infty]$) with arbitrary precision. Building upon this approximation result, we demonstrate that standard Transformers achieve the minimax optimal rate in nonparametric regression for Hölder target functions. It is worth mentioning that, by introducing two metrics: the size tuple and the dimension vector, we provide a fine-grained characterization of Transformer structures, which facilitates future research on the generalization and optimization errors of Transformers with different structures. As intermediate results, we also derive the upper bounds for the Lipschitz constant of standard Transformers and their memorization capacity, which may be of independent interest. These findings provide theoretical justification for the powerful capabilities of Transformer models.

📄 PDF Abstract BibTeX arXiv:2602.20555

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Efficient and Minimax Optimal In-context Nonparametric Regression with Transformers

2026-01-21 · Michelle Ching, Ioana Popescu, Nico Smith, Tianyi Ma 외 arxiv

We study in-context learning for nonparametric regression with $α$-Hölder smooth regression functions, for some $α>0$. We prove that, with $n$ in-context examples and $d$-dimensional regression covariates, a pretrained t…

Minimax rates of convergence for nonparametric regression under adversarial attacks

2024-10-12 · Jingfu Peng, Yuhong Yang

Recent research shows the susceptibility of machine learning models to adversarial attacks, wherein minor but maliciously chosen perturbations of the input can significantly degrade model performance. In this paper, we t…

regression

Transformers are Minimax Optimal Nonparametric In-Context Learners

2024-08-22 · Juno Kim, Tai Nakamaki, Taiji Suzuki

In-context learning (ICL) of large language models has proven to be a surprisingly effective method of learning a new task from only a few demonstrative examples. In this paper, we study the efficacy of ICL from the view…

DiversityIn-Context LearningLearning TheoryRepresentation Learning

Transfer Learning for Nonparametric Regression: Non-asymptotic Minimax Analysis and Adaptive Procedure

2024-01-22 · T. Tony Cai, Hongming Pu

Transfer learning for nonparametric regression is considered. We first study the non-asymptotic minimax risk for this problem and develop a novel estimator called the confidence thresholding estimator, which is shown to …

regressionTransfer Learning

Adversarial learning for nonparametric regression: Minimax rate and adaptive estimation

2025-06-02 · Jingfu Peng, Yuhong Yang

Despite tremendous advancements of machine learning models and algorithms in various application domains, they are known to be vulnerable to subtle, natural or intentionally crafted perturbations in future input data, kn…

regression