paper-with-me

홈 › Papers

Learning Spatial Decay for Vision Transformers

2025-08-13 · Yuxin Mao, Zhen Qin, Jinxing Zhou, Bin Fan, Jing Zhang, Yiran Zhong, Yuchao Dai arxiv

Vision Transformers (ViTs) have revolutionized computer vision, yet their self-attention mechanism lacks explicit spatial inductive biases, leading to suboptimal performance on spatially-structured tasks. Existing approaches introduce data-independent spatial decay based on fixed distance metrics, applying uniform attention weighting regardless of image content and limiting adaptability to diverse visual scenarios. Inspired by recent advances in large language models where content-aware gating mechanisms (e.g., GLA, HGRN2, FOX) significantly outperform static alternatives, we present the first successful adaptation of data-dependent spatial decay to 2D vision transformers. We introduce \textbf{Spatial Decay Transformer (SDT)}, featuring a novel Context-Aware Gating (CAG) mechanism that generates dynamic, data-dependent decay for patch interactions. Our approach learns to modulate spatial attention based on both content relevance and spatial proximity. We address the fundamental challenge of 1D-to-2D adaptation through a unified spatial-content fusion framework that integrates manhattan distance-based spatial priors with learned content representations. Extensive experiments on ImageNet-1K classification and generation tasks demonstrate consistent improvements over strong baselines. Our work establishes data-dependent spatial decay as a new paradigm for enhancing spatial attention in vision transformers.

📄 PDF Abstract BibTeX arXiv:2508.09525

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RMT: Retentive Networks Meet Vision Transformers

2023-09-20 · CVPR 2024 1 · Qihang Fan, Huaibo Huang, Mingrui Chen, Hongmin Liu 외

Vision Transformer (ViT) has gained increasing attention in the computer vision community in recent years. However, the core component of ViT, Self-Attention, lacks explicit spatial priors and bears a quadratic computati…

Instance Segmentationobject-detectionObject DetectionSemantic Segmentation

Spatial Priors via Space Filling Curves for Small and Limited Data Vision Transformers

2026-06-08 · Leyla Naz Candogan, Arshia Afzal, Pol Puigdemont, Volkan Cevher arxiv

Though Vision Transformers (ViTs) have become the dominant backbone in many computer vision tasks, due to permutation equivariance, their attention mechanism lacks explicit spatial inductive biases. This become particula…

parameter-efficient fine-tuning

Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating

2026-04-07 · Dipan Maity, Suman Mondal, Arindam Roy arxiv

We introduce Gated-SwinRMT, a family of hybrid vision transformers that combine the shifted-window attention of the Swin Transformer with the Manhattan-distance spatial decay of Retentive Networks (RMT), augmented by inp…

Colinearity Decay: Training Quantization-Friendly ViTs with Outlier Decay

2026-05-02 · Jin Tong, Guang Liang, Peilin Sun, Jianxin Wu arxiv

Low-bit quantization is a practical route for efficiently deploying vision Transformers, yet activation outliers complicate fully quantized deployment. Existing methods either handle quantization post-training or suppres…

Beyond flattening: a geometrically principled positional encoding for vision transformers with Weierstrass elliptic functions

2025-08-26 · Zhihang Xin, Xitong Hu, Rui Wang arxiv

Vision Transformers have demonstrated remarkable success in computer vision tasks, yet their reliance on learnable one-dimensional positional embeddings fundamentally disrupts the inherent two-dimensional spatial structu…