paper-with-me

홈 › Papers

Optimizing Native Sparse Attention with Latent Attention and Local Global Alternating Strategies

2025-11-02 · Yuxuan Hu, Jianchao Tan, Jiaqi Zhang, Wen Zan, Pingwei Sun, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai, Jing Zhang arxiv

In this work, we conduct a systematic analysis of Native Sparse Attention (NSA) and propose targeted improvements that enhance long-context modeling. A key insight is that alternating between local (sliding-window) and global (compression, selective) attention across layers, rather than using fixed patterns, enables more effective propagation of long-range dependencies and substantially boosts performance on long-sequence tasks. Meanwhile, we further refine NSA's branches with Latent Attention that the sliding-window branch is enhanced with Multi-head Latent Attention (MLA) while compression and selective branches adopt Group-head Latent Attention (GLA). These changes reduce KV-cache memory by 50\% versus NSA while improving the model's common-sense reasoning and long-text understanding capabilities. Experiments on models from 340M to 1.3B parameters (trained on 15B and 100B tokens) show our method matches or exceeds full attention and native sparse attention in both common-sense reasoning and long-context understanding tasks.

📄 PDF Abstract BibTeX arXiv:2511.00819

Code (0)

등록된 구현이 없습니다.

Tasks

Long-Context Understanding

Similar Papers 제목 키워드 기반

Latent-Condensed Transformer for Efficient Long Context Modeling

2026-04-14 · Zeng You, Yaofo Chen, Qiuwu Chen, Ying Sun 외 arxiv

Large language models (LLMs) face significant challenges in processing long contexts due to the linear growth of the key-value (KV) cache and quadratic complexity of self-attention. Existing approaches address these bott…

Sparse Attention Vectors: Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers

2024-11-28 · Chancharik Mitra, Brandon Huang, Tianning Chai, Zhiqiu Lin 외

Generative Large Multimodal Models (LMMs) like LLaVA and Qwen-VL excel at a wide variety of vision-language (VL) tasks such as image captioning or visual question answering. Despite strong performance, LMMs are not direc…

Image Captioningimage-classificationImage ClassificationMultiple-choice+3

FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel

2025-08-25 · Ran Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai 외 arxiv

Recent advances in sparse attention mechanisms have demonstrated strong potential for reducing the computational cost of long-context training and inference in large language models (LLMs). Native Sparse Attention (NSA),…

Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation

2025-09-01 · Jiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang 외 arxiv

Open-vocabulary semantic segmentation (OVSS) conducts pixel-level classification via text-driven alignment, where the domain discrepancy between base category training and open-vocabulary inference poses challenges in di…

Semantic Segmentation

SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference

2025-02-25 · Jintao Zhang, Chendong Xiang, Haofeng Huang, Jia Wei 외

An efficient attention implementation is essential for large models due to its quadratic time complexity. Fortunately, attention commonly exhibits sparsity, i.e., many values in the attention map are near zero, allowing …

modelVideo Generation