paper-with-me

Papers

How Transformers Get Rich: Approximation and Dynamics Analysis

2024-10-15 · Mingze Wang, Ruoxi Yu, Weinan E, Lei Wu

Transformers have demonstrated exceptional in-context learning capabilities, yet the theoretical understanding of the underlying mechanisms remains limited. A recent work (Elhage et al., 2021) identified a `rich'' in-context mechanism known as induction head, contrasting with `lazy'' $n$-gram models that overlook long-range dependencies. In this work, we provide both approximation and dynamics analyses of how transformers implement induction heads. In the {\em approximation} analysis, we formalize both standard and generalized induction head mechanisms, and examine how transformers can efficiently implement them, with an emphasis on the distinct role of each transformer submodule. For the {\em dynamics} analysis, we study the training dynamics on a synthetic mixed target, composed of a 4-gram and an in-context 2-gram component. This controlled setting allows us to precisely characterize the entire training process and uncover an {\em abrupt transition} from lazy (4-gram) to rich (induction head) mechanisms as training progresses.

📄 PDF Abstract BibTeX arXiv:2410.11474

Code (0)

등록된 구현이 없습니다.

Tasks

In-Context Learning

Similar Papers 제목 키워드 기반

On the Limitations of First-Order Approximation in GAN Dynamics

2017-06-29 · ICML 2018 7 · Jerry Li, Aleksander Madry, John Peebles, Ludwig Schmidt

While Generative Adversarial Networks (GANs) have demonstrated promising performance on multiple vision tasks, their learning dynamics are not yet well understood, both in theory and in practice. To address this issue, w…

Provable optimal transport with transformers: The essence of depth and prompt engineering

2024-10-25 · Hadi Daneshmand

Can we establish provable performance guarantees for transformers? Establishing such theoretical guarantees is a milestone in developing trustworthy generative AI. In this paper, we take a step toward addressing this que…

Prompt Engineering

Mamba Neural Operator: Who Wins? Transformers vs. State-Space Models for PDEs

2024-10-03 · Chun-Wun Cheng, Jiahao Huang, Yi Zhang, Guang Yang 외

Partial differential equations (PDEs) are widely used to model complex physical systems, but solving them efficiently remains a significant challenge. Recently, Transformers have emerged as the preferred architecture for…

MambaState Space Models

Precise Verification of Transformers through ReLU-Catalyzed Abstraction Refinement

2026-05-14 · Hengjie Liu, Zhenya Zhang, Jianjun Zhao arxiv

Formal verification of transformers has become increasingly important due to their widespread deployment in safety-critical applications. Compared to classic neural networks, the inferences of transformers involve highly…

Sentiment Analysis

Dynamics of Transient Structure in In-Context Linear Regression Transformers

2025-01-29 · Liam Carroll, Jesse Hoogland, Matthew Farrugia-Roberts, Daniel Murfet

Modern deep neural networks display striking examples of rich internal computational structure. Uncovering principles governing the development of such structure is a priority for the science of deep learning. In this pa…

DiversityModel Selectionregression