paper-with-me

홈 › Papers

Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions

2025-08-26 · Zhihang Xin, Xitong Hu, Rui Wang arxiv

Vision Transformers have demonstrated remarkable success in computer vision tasks, yet their reliance on learnable one-dimensional positional embeddings fundamentally disrupts the inherent two-dimensional spatial structure of images through patch flattening procedures. Traditional positional encoding approaches lack geometric constraints and fail to establish monotonic correspondence between Euclidean spatial distances and sequential index distances, thereby limiting the model's capacity to leverage spatial proximity priors effectively. We propose Weierstrass Elliptic Function Positional Encoding (WEF-PE), a mathematically principled approach that directly addresses two-dimensional coordinates through natural complex domain representation, where the doubly periodic properties of elliptic functions align remarkably with translational invariance patterns commonly observed in visual data. Our method exploits the non-linear geometric nature of elliptic functions to encode spatial distance relationships naturally, while the algebraic addition formula enables direct derivation of relative positional information between arbitrary patch pairs from their absolute encodings. Comprehensive experiments demonstrate that WEF-PE achieves superior performance across diverse scenarios, including 63.78\% accuracy on CIFAR-100 from-scratch training with ViT-Tiny architecture, 93.28\% on CIFAR-100 fine-tuning with ViT-Base, and consistent improvements on VTAB-1k benchmark tasks. Theoretical analysis confirms the distance-decay property through rigorous mathematical proof, while attention visualization reveals enhanced geometric inductive bias and more coherent semantic focus compared to conventional approaches.The source code implementing the methods described in this paper is publicly available on GitHub.

📄 PDF Abstract BibTeX arXiv:2508.19167

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Weierstrass Positional Encoding for Vision Transformers

2026-05-20 · Zhihang Xin, Rui Wang, Xitong Hu, Xiaojun Wu arxiv

Vision Transformers have achieved remarkable success in computer vision, but their common use of learnable one-dimensional positional encodings weakens the inherent two-dimensional spatial structure of images after patch…

Spatial Reasoning is Not a Free Lunch: A Controlled Study on LLaVA

2026-03-13 · Nahid Alam, Leema Krishna Murali, Siddhant Bharadwaj, Patrick Liu 외 arxiv

Vision-language models (VLMs) have advanced rapidly, yet they still struggle with basic spatial reasoning. Despite strong performance on general benchmarks, modern VLMs remain brittle at understanding 2D spatial relation…

Spatial Reasoning

Linearized Relative Positional Encoding

2023-07-18 · Zhen Qin, Weixuan Sun, Kaiyue Lu, Hui Deng 외

Relative positional encoding is widely used in vanilla and linear transformers to represent positional information. However, existing encoding methods of a vanilla transformer are not always directly applicable to a line…

image-classificationImage ClassificationLanguage ModelingLanguage Modelling+2

Conceptors for Semantic Steering

2026-05-06 · Ilias Triantafyllopoulos, Young-Min Cho, Ren Tao, Miranda Muqing Miao 외 arxiv

Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction whose geometry is left largely unexamined. Rather than selecting a single…

ViT-NeBLa: A Hybrid Vision Transformer and Neural Beer-Lambert Framework for Single-View 3D Reconstruction of Oral Anatomy from Panoramic Radiographs

2025-06-16 · Bikram Keshari Parida, Anusree P. Sunilkumar, Abhijit Sen, Wonsang You

Dental diagnosis relies on two primary imaging modalities: panoramic radiographs (PX) providing 2D oral cavity representations, and Cone-Beam Computed Tomography (CBCT) offering detailed 3D anatomical information. While …

3D ReconstructionAnatomyDiagnosticSingle-View 3D Reconstruction