paper-with-me

홈 › Papers

SP-ViT: Learning 2D Spatial Priors for Vision Transformers

2022-06-15 · Yuxuan Zhou, Wangmeng Xiang, Chao Li, Biao Wang, Xihan Wei, Lei Zhang, Margret Keuper, Xiansheng Hua

Recently, transformers have shown great potential in image classification and established state-of-the-art results on the ImageNet benchmark. However, compared to CNNs, transformers converge slowly and are prone to overfitting in low-data regimes due to the lack of spatial inductive biases. Such spatial inductive biases can be especially beneficial since the 2D structure of an input image is not well preserved in transformers. In this work, we present Spatial Prior-enhanced Self-Attention (SP-SA), a novel variant of vanilla Self-Attention (SA) tailored for vision transformers. Spatial Priors (SPs) are our proposed family of inductive biases that highlight certain groups of spatial relations. Unlike convolutional inductive biases, which are forced to focus exclusively on hard-coded local regions, our proposed SPs are learned by the model itself and take a variety of spatial relations into account. Specifically, the attention score is calculated with emphasis on certain kinds of spatial relations at each head, and such learned spatial foci can be complementary to each other. Based on SP-SA we propose the SP-ViT family, which consistently outperforms other ViT models with similar GFlops or parameters. Our largest model SP-ViT-L achieves a record-breaking 86.3% Top-1 accuracy with a reduction in the number of parameters by almost 50% compared to previous state-of-the-art model (150M for SP-ViT-L vs 271M for CaiT-M-36) among all ImageNet-1K models trained on 224x224 and fine-tuned on 384x384 resolution w/o extra data.

📄 PDF Abstract BibTeX arXiv:2206.07662

Code (1)

ZhouYuxuanYX/SP-ViT 공식 구현 pytorch

Tasks

image-classificationImage Classification

Similar Papers 제목 키워드 기반

Learning Spatial Decay for Vision Transformers

2025-08-13 · Yuxin Mao, Zhen Qin, Jinxing Zhou, Bin Fan 외 arxiv

Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approa…

ZACH-ViT: Regime-Dependent Inductive Bias in Compact Vision Transformers for Medical Imaging

2026-02-20 · Athanasios Angelakis arxiv

Vision Transformers rely on positional embeddings and class tokens encoding fixed spatial priors. While effective for natural images, these priors may be suboptimal when spatial layout is weakly informative, a frequent c…

Learning Priors of Human Motion With Vision Transformers

2025-01-30 · Placido Falqueto, Alberto Sanfeliu, Luigi Palopoli, Daniele Fontanelli

A clear understanding of where humans move in a scenario, their usual paths and speeds, and where they stop, is very important for different applications, such as mobility studies in urban areas or robot navigation tasks…

Robot Navigation

Spatial Priors via Space Filling Curves for Small and Limited Data Vision Transformers

2026-06-08 · Leyla Naz Candogan, Arshia Afzal, Pol Puigdemont, Volkan Cevher arxiv

Though Vision Transformers (ViTs) have become the dominant backbone in many computer vision tasks, due to permutation equivariance, their attention mechanism lacks explicit spatial inductive biases. This become particula…

parameter-efficient fine-tuning

Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning

2026-01-29 · Jian Shi, Michael Birsak, Wenqing Cui, Zhenyu Li 외 arxiv

This paper revisits the role of positional embeddings (PEs) within vision transformers (ViTs) from a geometric perspective. We show that PEs are not mere token indices but effectively function as geometric priors that sh…

Spatial Reasoning