paper-with-me

홈 › Papers

Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers

2023-09-23 · Adam Pardyl, Grzegorz Kurzejamski, Jan Olszewski, Tomasz Trzciński, Bartosz Zieliński

Vision transformers have excelled in various computer vision tasks but mostly rely on rigid input sampling using a fixed-size grid of patches. It limits their applicability in real-world problems, such as active visual exploration, where patches have various scales and positions. Our paper addresses this limitation by formalizing the concept of input elasticity for vision transformers and introducing an evaluation protocol for measuring this elasticity. Moreover, we propose modifications to the transformer architecture and training regime, which increase its elasticity. Through extensive experimentation, we spotlight opportunities and challenges associated with such architecture.

📄 PDF Abstract BibTeX arXiv:2309.13353

Code (1)

apardyl/beyondgrids 공식 구현 pytorch

Similar Papers 제목 키워드 기반

SetONet: A Deep Set-based Operator Network for Solving PDEs with permutation invariant variable input sampling

2025-05-07 · Stepan Tretiakov, Xingjian Li, Krishna Kumar

Neural operators, particularly the Deep Operator Network (DeepONet), have shown promise in learning mappings between function spaces for solving differential equations. However, standard DeepONet requires input functions…

Operator learning

A Complement to Neural Networks for Anisotropic Inelasticity at Finite Strains

2025-10-05 · Hagen Holthusen, Ellen Kuhl arxiv

We propose a complement to constitutive modeling that augments neural networks with material principles to capture anisotropy and inelasticity at finite strains. The key element is a dual potential that governs dissipati…

PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding

2026-05-28 · Selim Kuzucu, Alessio Tonioni, Vasile Lup, Bernt Schiele 외 arxiv

Large Vision-Language Models (LVLMs) map visual inputs into dense token sequences, imposing a quadratic computational bottleneck for inference. Elastic visual-token compression addresses this by training a single model t…

Towards Compressive and Scalable Recurrent Memory

2026-02-11 · Yunchong Song, Jushi Kai, Liming Lu, Kaixi Qiu 외 arxiv

Transformers face a quadratic bottleneck in attention when scaling to long contexts. Recent approaches introduce recurrent memory to extend context beyond the current window, yet these often face a fundamental trade-off …

Domain Elastic Transform: Bayesian Function Registration for High-Dimensional Scientific Data

2026-03-22 · Osamu Hirose, Emanuele Rodola arxiv

Nonrigid registration is conventionally divided into point set registration, which aligns sparse geometries, and image registration, which aligns continuous intensity fields on regular grids. However, this dichotomy crea…

Image Registration