paper-with-me

홈 › Papers

SPoT: Subpixel Placement of Tokens in Vision Transformers

2025-07-02 · Martine Hjelkrem-Tan, Marius Aasan, Gabriel Y. Arteaga, Adín Ramírez Rivera arxiv

Vision Transformers naturally accommodate sparsity, yet standard tokenization methods confine features to discrete patch grids. This constraint prevents models from fully exploiting sparse regimes, forcing awkward compromises. We propose Subpixel Placement of Tokens (SPoT), a novel tokenization strategy that positions tokens continuously within images, effectively sidestepping grid-based limitations. With our proposed oracle-guided search, we uncover substantial performance gains achievable with ideal subpixel token positioning, drastically reducing the number of tokens necessary for accurate predictions during inference. SPoT provides a new direction for flexible, efficient, and interpretable ViT architectures, redefining sparsity as a strategic advantage rather than an imposed limitation.

📄 PDF Abstract BibTeX arXiv:2507.01654

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Extracting full-field subpixel structural displacements from videos via deep learning

2020-08-31 · Lele Luan, Jingwei Zheng, Yongchao Yang, Ming L. Wang 외

This paper develops a deep learning framework based on convolutional neural networks (CNNs) that enable real-time extraction of full-field subpixel structural displacements from videos. In particular, two new CNN archite…

Deep Learning

SPOT: Sparsification with Attention Dynamics via Token Relevance in Vision Transformers

2025-11-13 · Oded Schlesinger, Amirhossein Farzam, J. Matias Di Martino, Guillermo Sapiro arxiv

While Vision Transformers (ViT) have demonstrated remarkable performance across diverse tasks, their computational demands are substantial, scaling quadratically with the number of processed tokens. Compact attention rep…

Computational Efficiency

CoMatch: Dynamic Covisibility-Aware Transformer for Bilateral Subpixel-Level Semi-Dense Image Matching

2025-03-31 · Zizhuo Li, Yifan Lu, Linfeng Tang, Shihua Zhang 외

This prospective study proposes CoMatch, a novel semi-dense image matcher with dynamic covisibility awareness and bilateral subpixel accuracy. Firstly, observing that modeling context interaction over the entire coarse f…

Computational Efficiency

SpotEdit: Selective Region Editing in Diffusion Transformers

2025-12-26 · Zhibin Qin, Zhenxiong Tan, Zeqing Wang, Songhua Liu 외 arxiv

Diffusion Transformer models have significantly advanced image editing by encoding conditional images and integrating them into transformer layers. However, most edits involve modifying only small regions, while current …

Image Editing

Image Registration Based Flicker Solving in Video Face Replacement and Analysis Based Sub-pixel Image Registration

2018-03-09 · Xiaofang Wang, Guoqiang Xiang, Xinyue Zhang, Wei Wei

In this paper, a framework of video face replacement is proposed and it deals with the flicker of swapped face in video sequence. This framework contains two main innovations: 1) the technique of image registration is ex…

Image Registration