paper-with-me

홈 › Papers

Multi-Dimensional Hyena for Spatial Inductive Bias

2023-09-24 · Itamar Zimerman, Lior Wolf

In recent years, Vision Transformers have attracted increasing interest from computer vision researchers. However, the advantage of these transformers over CNNs is only fully manifested when trained over a large dataset, mainly due to the reduced inductive bias towards spatial locality within the transformer's self-attention mechanism. In this work, we present a data-efficient vision transformer that does not rely on self-attention. Instead, it employs a novel generalization to multiple axes of the very recent Hyena layer. We propose several alternative approaches for obtaining this generalization and delve into their unique distinctions and considerations from both empirical and theoretical perspectives. Our empirical findings indicate that the proposed Hyena N-D layer boosts the performance of various Vision Transformer architectures, such as ViT, Swin, and DeiT across multiple datasets. Furthermore, in the small dataset regime, our Hyena-based ViT is favorable to ViT variants from the recent literature that are specifically designed for solving the same challenge, i.e., working with small datasets or incorporating image-specific inductive bias into the self-attention mechanism. Finally, we show that a hybrid approach that is based on Hyena N-D for the first layers in ViT, followed by layers that incorporate conventional attention, consistently boosts the performance of various vision transformer architectures.

📄 PDF Abstract BibTeX arXiv:2309.13600

Code (0)

등록된 구현이 없습니다.

Tasks

Inductive Bias

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Feedforward Network A Feedforward Network, or a Multilayer Perceptron (MLP), is a neural network with solely densely connected layers. This is the classic neural network architecture of the…
DeiT 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

HyenaPixel: Global Image Context with Convolutions

2024-02-29 · Julian Spravil, Sebastian Houben, Sven Behnke

In computer vision, a larger effective receptive field (ERF) is associated with better performance. While attention natively supports global context, its quadratic complexity limits its applicability to tasks that benefi…

Image ClassificationObject DetectionSemantic Segmentation

2-D SSM: A General Spatial Layer for Visual Transformers

2023-06-11 · Ethan Baron, Itamar Zimerman, Lior Wolf

A central objective in computer vision is to design models with appropriate 2-D inductive bias. Desiderata for 2D inductive bias include two-dimensional position awareness, dynamic spatial locality, and translation and p…

Inductive BiasPosition

SE(3)-Hyena Operator for Scalable Equivariant Learning

2024-07-01 · Artem Moskalev, Mangal Prakash, Rui Liao, Tommaso Mansi

Modeling global geometric context while maintaining equivariance is crucial for accurate predictions in many fields such as biology, chemistry, or vision. Yet, this is challenging due to the computational demands of proc…

PDE-SSM: A Spectral State Space Approach to Spatial Mixing in Diffusion Transformers

2026-03-14 · Eshed Gal, Moshe Eliasof, Siddharth Rout, Eldad Haber arxiv

The success of vision transformers-especially for generative modeling-is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces atte…

Contrastive Learning Relies More on Spatial Inductive Bias Than Supervised Learning: An Empirical Study

2023-01-01 · ICCV 2023 1 · Yuanyi Zhong, Haoran Tang, Jun-Kun Chen, Yu-Xiong Wang

Though self-supervised contrastive learning (CL) has shown its potential to achieve state-of-the-art accuracy without any supervision, its behavior still remains under investigated by academia. Different from most pr…

Contrastive LearningInductive Bias