paper-with-me

홈 › Papers

Hyena Hierarchy: Towards Larger Convolutional Language Models

2023-02-21 · Michael Poli, Stefano Massaroli, Eric Nguyen, Daniel Y. Fu, Tri Dao, Stephen Baccus, Yoshua Bengio, Stefano Ermon, Christopher Ré

Recent advances in deep learning have relied heavily on the use of large Transformers due to their ability to learn at scale. However, the core building block of Transformers, the attention operator, exhibits quadratic cost in sequence length, limiting the amount of context accessible. Existing subquadratic methods based on low-rank and sparse approximations need to be combined with dense attention layers to match Transformers, indicating a gap in capability. In this work, we propose Hyena, a subquadratic drop-in replacement for attention constructed by interleaving implicitly parametrized long convolutions and data-controlled gating. In recall and reasoning tasks on sequences of thousands to hundreds of thousands of tokens, Hyena improves accuracy by more than 50 points over operators relying on state-spaces and other implicit and explicit methods, matching attention-based models. We set a new state-of-the-art for dense-attention-free architectures on language modeling in standard datasets (WikiText103 and The Pile), reaching Transformer quality with a 20% reduction in training compute required at sequence length 2K. Hyena operators are twice as fast as highly optimized attention at sequence length 8K, and 100x faster at sequence length 64K.

📄 PDF Abstract BibTeX arXiv:2302.10866

Code (7)

hazyresearch/safari 공식 구현 pytorch
MindSpore-scientific-2/code-4/tree/main/Hyena-A-Convolutional-Neural-Network-for-Modelling-Sentences mindspore
Suro-One/Hyena-Hierarchy pytorch
expz/annotated-hyena
i404788/s5-pytorch jax
lindermanlab/S5 jax
togethercomputer/stripedhyena pytorch

Tasks

2k8kLanguage ModelingLanguage ModellingQuestion Answering

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Position-Wise Feed-Forward Layer 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

HyenaPixel: Global Image Context with Convolutions

2024-02-29 · Julian Spravil, Sebastian Houben, Sven Behnke

In computer vision, a larger effective receptive field (ERF) is associated with better performance. While attention natively supports global context, its quadratic complexity limits its applicability to tasks that benefi…

Image ClassificationObject DetectionSemantic Segmentation

SE(3)-Hyena Operator for Scalable Equivariant Learning

2024-07-01 · Artem Moskalev, Mangal Prakash, Rui Liao, Tommaso Mansi

Modeling global geometric context while maintaining equivariance is crucial for accurate predictions in many fields such as biology, chemistry, or vision. Yet, this is challenging due to the computational demands of proc…

How do Hyenas deal with Human Speech? Speech Recognition and Translation with ConfHyena

2024-02-20 · Marco Gaido, Sara Papi, Matteo Negri, Luisa Bentivogli

The attention mechanism, a cornerstone of state-of-the-art neural models, faces computational hurdles in processing long sequences due to its quadratic complexity. Consequently, research efforts in the last few years foc…

Automatic Speech Recognitionimage-classificationImage ClassificationLanguage Modeling+3

Hyena Neural Operator for Partial Differential Equations

2023-06-28 · Saurabh Patil, Zijie Li, Amir Barati Farimani

Numerically solving partial differential equations typically requires fine discretization to resolve necessary spatiotemporal scales, which can be computationally expensive. Recent advances in deep learning have provided…

HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution

2023-06-27 · NeurIPS 2023 11 · Eric Nguyen, Michael Poli, Marjan Faizi, Armin Thomas 외

Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation models in genomics to learn generalizab…

4kIn-Context LearningLanguage ModellingLarge Language Model