paper-with-me

홈 › Papers

Some Attention is All You Need for Retrieval

2025-10-21 · Felix Michalak, Steven Abreu arxiv

We demonstrate complete functional segregation in hybrid SSM-Transformer architectures: retrieval depends exclusively on self-attention layers. Across RecurrentGemma-2B/9B and Jamba-Mini-1.6, attention ablation causes catastrophic retrieval failure (0% accuracy), while SSM layers show no compensatory mechanisms even with improved prompting. Conversely, sparsifying attention to just 15% of heads maintains near-perfect retrieval while preserving 84% MMLU performance, suggesting self-attention specializes primarily for retrieval tasks. We identify precise mechanistic requirements for retrieval: needle tokens must be exposed during generation and sufficient context must be available during prefill or generation. This strict functional specialization challenges assumptions about redundancy in hybrid architectures and suggests these models operate as specialized modules rather than integrated systems, with immediate implications for architecture optimization and interpretability.

📄 PDF Abstract BibTeX arXiv:2510.19861

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Compositional Attention: Disentangling Search and Retrieval

2021-10-18 · ICLR 2022 4 · Sarthak Mittal, Sharath Chandra Raparthy, Irina Rish, Yoshua Bengio 외

Multi-head, key-value attention is the backbone of the widely successful Transformer model and its variants. This attention mechanism uses multiple parallel key-value attention blocks (called heads), each performing two …

Retrieval

Current Limitations of Language Models: What You Need is Retrieval

2020-09-15 · Aran Komatsuzaki

We classify and re-examine some of the current approaches to improve the performance-computes trade-off of language models, including (1) non-causal models (such as masked language models), (2) extension of batch length …

RetrievalText Generation

MOON: Multi-Hash Codes Joint Learning for Cross-Media Retrieval

2021-08-17 · Donglin Zhang, Xiao-Jun Wu, He-Feng Yin, Josef Kittler

In recent years, cross-media hashing technique has attracted increasing attention for its high computation efficiency and low storage cost. However, the existing approaches still have some limitations, which need to be e…

Retrieval

Diverse legal case search

2023-01-29 · Ruizhe Zhang, Qingyao Ai, Yueyue Wu, Yixiao Ma 외

In last decades, legal case search has received more and more attention. Legal practitioners need to work or enhance their efficiency by means of class case search. In the process of searching, legal practitioners often …

DiversityRetrievalSpecificity

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

2024-09-16 · Di Liu, Meng Chen, Baotong Lu, Huiqiang Jiang 외

Transformer-based Large Language Models (LLMs) have become increasingly important. However, due to the quadratic time complexity of attention computation, scaling LLMs to longer contexts incurs extremely slow inference s…

CPUGPURetrieval