paper-with-me

홈 › Papers

Overcoming Long-Context Limitations of State-Space Models via Context-Dependent Sparse Attention

2025-07-01 · Zhihao Zhan, Jianan Zhao, Zhaocheng Zhu, Jian Tang arxiv

Efficient long-context modeling remains a critical challenge for natural language processing (NLP), as the time complexity of the predominant Transformer architecture scales quadratically with the sequence length. While state-space models (SSMs) offer alternative sub-quadratic solutions, they struggle to capture long-range dependencies effectively. In this work, we focus on analyzing and improving the long-context modeling capabilities of SSMs. We show that the widely used synthetic task, associative recall, which requires a model to recall a value associated with a single key without context, insufficiently represents the complexities of real-world long-context modeling. To address this limitation, we extend the associative recall to a novel synthetic task, \emph{joint recall}, which requires a model to recall the value associated with a key given in a specified context. Theoretically, we prove that SSMs do not have the expressiveness to solve multi-query joint recall in sub-quadratic time complexity. To resolve this issue, we propose a solution based on integrating SSMs with Context-Dependent Sparse Attention (CDSA), which has the expressiveness to solve multi-query joint recall with sub-quadratic computation. To bridge the gap between theoretical analysis and real-world applications, we propose locality-sensitive Hashing Attention with sparse Key Selection (HAX), which instantiates the theoretical solution and is further tailored to natural language domains. Extensive experiments on both synthetic and real-world long-context benchmarks show that HAX consistently outperforms SSM baselines and SSMs integrated with context-independent sparse attention (CISA).

📄 PDF Abstract BibTeX arXiv:2507.00449

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OWL: Overcoming Window Length-Dependence in Speculative Decoding for Long-Context Inputs

2025-10-08 · Jaeseong Lee, seung-won hwang, Aurick Qiao, Gabriele Oliaro 외 arxiv

Speculative decoding promises faster inference for large language models (LLMs), yet existing methods fail to generalize to real-world settings. Benchmarks typically assume short contexts (e.g., 2K tokens), whereas pract…

Guided Code Generation with LLMs: A Multi-Agent Framework for Complex Code Tasks

2025-01-11 · Amr Almorsi, Mohanned Ahmed, Walid Gomaa

Large Language Models (LLMs) have shown remarkable capabilities in code generation tasks, yet they face significant limitations in handling complex, long-context programming challenges and demonstrating complex compositi…

Code GenerationHumanEvalLong-Context Understanding

In Defense of RAG in the Era of Long-Context Language Models

2024-09-03 · Tan Yu, Anbang Xu, Rama Akkiraju

Overcoming the limited context limitations in early-generation LLMs, retrieval-augmented generation (RAG) has been a reliable solution for context-based answer generation in the past. Recently, the emergence of long-cont…

Answer GenerationRAGRetrievalRetrieval-augmented Generation

GNN-Assisted Phase Space Integration with Application to Atomistics

2023-03-20 · Shashank Saxena, Jan-Hendrik Bastek, Miguel Spinola, Prateek Gupta 외

Overcoming the time scale limitations of atomistics can be achieved by switching from the state-space representation of Molecular Dynamics (MD) to a statistical-mechanics-based representation in phase space, where approx…

Computational Efficiency

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

2025-10-21 · Howard Yen, Ashwin Paranjape, Mengzhou Xia, Thejas Venkatesh 외 arxiv

Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling powerful applications like deep research systems. In this work, we show…