paper-with-me

홈 › Papers

Hierarchical Global Attention (HGA)

2026-06-29 · Woernle Frank, Fedosov Vladimir, Grinenko Artemiy arxiv

Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preserves the original checkpoint parameters: the pretrained $W_Q$, $W_K$, $W_V$, and $W_O$ projections remain unchanged, no calibration parameters are introduced, and no retraining is required. Applied to Qwen3-30B-A3B-Instruct-2507-FP8 on a single RTX~5090 (32GB), the patched model runs out of the box at a 64K-token context, where token-level K/V storage is not feasible on this hardware. Unlike previous sparse-attention methods, HGA performs hierarchical two-level routing. It first retrieves relevant chunks using compact RoPE-aware summaries and then refines the selection by routing only the most relevant groups before performing exact token-level attention. This hierarchical retrieval significantly reduces the number of fetched tokens while preserving exact attention over the retrieved token set, making RAM- and NVMe-backed storage practical. The full historical token K/V resides in host RAM or NVMe storage, while only a small routed working set is transferred to GPU memory during attention. Consequently, GPU memory consumption depends primarily on model weights and the routed working set rather than on the total context length. Across all tested context lengths (4K - 64K tokens), routed attention remains within approximately $0.01$--$0.02$ nats of dense attention while the sparsity used is just about 3%. These results suggest that the approximation introduced by hierarchical routing is small, and that the remaining quality gap is likely dominated by long-context positional encoding rather than by the routing algorithm itself.

📄 PDF Abstract BibTeX arXiv:2606.30709

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LA-HCN: Label-based Attention for Hierarchical Multi-label TextClassification Neural Network

2020-09-23 · Xinyi Zhang, Jiahao Xu, Charlie Soh, Lihui Chen

Hierarchical multi-label text classification (HMTC) has been gaining popularity in recent years thanks to its applicability to a plethora of real-world applications. The existing HMTC algorithms largely focus on the desi…

Multi Label Text ClassificationMulti-Label Text Classificationtext-classificationText Classification

Global2Local: A Joint-Hierarchical Attention for Video Captioning

2022-03-13 · Chengpeng Dai, Fuhai Chen, Xiaoshuai Sun, Rongrong Ji 외

Recently, automatic video captioning has attracted increasing attention, where the core challenge lies in capturing the key semantic items, like objects and actions as well as their spatial-temporal correlations from the…

Video Captioning

Trajectory-User Linking via Hierarchical Spatio-Temporal Attention Networks

2023-02-11 · Wei Chen, Chao Huang, Yanwei Yu, Yongguo Jiang 외

Trajectory-User Linking (TUL) is crucial for human mobility modeling by linking diferent trajectories to users with the exploration of complex mobility patterns. Existing works mainly rely on the recurrent neural framewo…

HiCI: Hierarchical Construction-Integration for Long-Context Attention

2026-03-21 · Xiangyu Zeng, Qi Xu, Yunke Wang, Chang Xu arxiv

Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information structuring remains largely implicit in existing approaches. Drawing on cognitive the…

Fusion of regional and sparse attention in Vision Transformers

2024-06-13 · Nabil Ibtehaz, Ning Yan, Masood Mortazavi, Daisuke Kihara

Modern vision transformers leverage visually inspired local interaction between pixels through attention computed within window or grid regions, in contrast to the global attention employed in the original ViT. Regional …