A Unified Geometric Field Theory Framework for Transformers: From Manifold Embeddings to Kernel Modulation
The Transformer architecture has achieved tremendous success in natural language processing, computer vision, and scientific computing through its self-attention mechanism. However, its core components-positional encoding and attention mechanisms-have lacked a unified physical or mathematical interpretation. This paper proposes a structural theoretical framework that integrates positional encoding, kernel integral operators, and attention mechanisms for in-depth theoretical investigation. We map discrete positions (such as text token indices and image pixel coordinates) to spatial functions on continuous manifolds, enabling a field-theoretic interpretation of Transformer layers as kernel-modulated operators acting over embedded manifolds.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
From Self-Attention to Connection Laplacian: A Unified Operator View of Transformers
Self-attention is a ubiquitous primitive in modern sequence models, yet its operator-level geometry is only partially understood. We view a token sequence as a vector field over the token-position graph and identify atte…
Geometric Formulation for Discrete Points and its Applications
We introduce a novel formulation for geometry on discrete points. It is based on a universal differential calculus, which gives a geometric description of a discrete set by the algebra of functions. We expand this mathem…
BIG-bench Machine LearningGraph Signal Processing for Geometric Data and Beyond: Theory and Applications
Geometric data acquired from real-world scenes, e.g., 2D depth images, 3D point clouds, and 4D dynamic point clouds, have found a wide range of applications including immersive telepresence, autonomous driving, surveilla…
Autonomous DrivingA Probabilistic Interpretation of Transformers
We propose a probabilistic interpretation of exponential dot product attention of transformers and contrastive learning based off of exponential families. The attention sublayer of transformers is equivalent to a gradien…
Contrastive LearningA Unified Framework for Interpretable Transformers Using PDEs and Information Theory
This paper presents a novel unified theoretical framework for understanding Transformer architectures by integrating Partial Differential Equations (PDEs), Neural Information Flow Theory, and Information Bottleneck Theor…