paper-with-me

홈 › Papers

Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems

2025-09-18 · Saeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito Koishida arxiv

Transformers and their attention mechanism have been revolutionary in the field of Machine Learning. While originally proposed for the language data, they quickly found their way to the image, video, graph, etc. data modalities with various signal geometries. Despite this versatility, generalizing the attention mechanism to scenarios where data is presented at different scales from potentially different modalities is not straightforward. The attempts to incorporate hierarchy and multi-modality within transformers are largely based on ad hoc heuristics, which are not seamlessly generalizable to similar problems with potentially different structures. To address this problem, in this paper, we take a fundamentally different approach: we first propose a mathematical construct to represent multi-modal, multi-scale data. We then mathematically derive the neural attention mechanics for the proposed construct from the first principle of entropy minimization. We show that the derived formulation is optimal in the sense of being the closest to the standard Softmax attention while incorporating the inductive biases originating from the hierarchical/geometric information of the problem. We further propose an efficient algorithm based on dynamic programming to compute our derived attention mechanism. By incorporating it within transformers, we show that the proposed hierarchical attention mechanism not only can be employed to train transformer models in hierarchical/multi-modal settings from scratch, but it can also be used to inject hierarchical information into classical, pre-trained transformer models post training, resulting in more efficient models in zero-shot manner.

📄 PDF Abstract BibTeX arXiv:2509.15448

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cell Attention Networks

2022-09-16 · Lorenzo Giusti, Claudio Battiloro, Lucia Testa, Paolo Di Lorenzo 외

Since their introduction, graph attention networks achieved outstanding results in graph representation learning tasks. However, these networks consider only pairwise relationships among nodes and then they are not able …

Graph AttentionGraph ClassificationGraph Representation LearningRepresentation Learning

Mechanics of Next Token Prediction with Self-Attention

2024-03-12 · Yingcong Li, Yixiao Huang, M. Emrullah Ildiz, Ankit Singh Rawat 외

Transformer-based language models are trained on large datasets to predict the next token given an input sequence. Despite this simple training objective, they have led to revolutionary advances in natural language proce…

PredictionRetrieval

Self-Attention Graph Pooling

2019-04-17 · Junhyun Lee, Inyeop Lee, Jaewoo Kang

Advanced methods of applying deep learning to structured data such as graphs have been proposed in recent years. In particular, studies have focused on generalizing convolutional neural networks to graph data, which incl…

Graph Classification

Theoretical Limitations of Self-Attention in Neural Sequence Models

2019-06-16 · TACL 2020 1 · Michael Hahn

Transformers are emerging as the new workhorse of NLP, showing great success across tasks. Unlike LSTMs, transformers process input sequences entirely through self-attention. Previous work has suggested that the computat…

Hard Attention

Paying U-Attention to Textures: Multi-Stage Hourglass Vision Transformer for Universal Texture Synthesis

2022-02-23 · Shouchang Guo, Valentin Deschaintre, Douglas Noll, Arthur Roullier

We present a novel U-Attention vision Transformer for universal texture synthesis. We exploit the natural long-range dependencies enabled by the attention mechanism to allow our approach to synthesize diverse textures wh…

Texture Synthesis