paper-with-me

홈 › Papers

Attention Reveals More Than Tokens: Training-Free Long-Context Reasoning with Attention-guided Retrieval

2025-03-12 · Yuwei Zhang, Jayanth Srinivasa, Gaowen Liu, Jingbo Shang

Large Language Models (LLMs) often exhibit substantially shorter effective context lengths than their claimed capacities, especially when handling complex reasoning tasks that require integrating information from multiple parts of a long context and performing multi-step reasoning. Although Chain-of-Thought (CoT) prompting has shown promise in reducing task complexity, our empirical analysis reveals that it does not fully resolve this limitation. Through controlled experiments, we identify poor recall of implicit facts as the primary cause of failure, which significantly hampers reasoning performance. Interestingly, we observe that the internal attention weights from the generated CoT tokens can effectively ground implicit facts, even when these facts are not explicitly recalled. Building on this insight, we propose a novel training-free algorithm, Attrieval, which leverages attention weights to retrieve relevant facts from the long context and incorporates them into the reasoning process. Additionally, we find that selecting context tokens from CoT tokens further improves performance. Our results demonstrate that Attrieval enhances long-context reasoning capability notably on both synthetic and real-world QA datasets with various models.

📄 PDF Abstract BibTeX arXiv:2503.09819

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Devils in Middle Layers of Large Vision-Language Models: Interpreting, Detecting and Mitigating Object Hallucinations via Attention Lens

2024-11-23 · CVPR 2025 1 · Zhangqi Jiang, Junkai Chen, Beier Zhu, Tingjin Luo 외

Hallucinations in Large Vision-Language Models (LVLMs) significantly undermine their reliability, motivating researchers to explore the causes of hallucination. However, most studies primarily focus on the language aspec…

Hallucination

Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning

2026-05-08 · Gengyang Li, Zheng-Fan Wu, Siqi Bao, Yunfang Wu arxiv

Reinforcement-learning-based post-training has become a key approach for improving the reasoning ability of large language models, but its token-level learning signals remain poorly understood. This work studies their he…

Hierarchical and Step-Layer-Wise Tuning of Attention Specialty for Multi-Instance Synthesis in Diffusion Transformers

2025-04-14 · ChunYang Zhang, Zhenhong Sun, Zhicheng Zhang, Junyan Wang 외

Text-to-image (T2I) generation models often struggle with multi-instance synthesis (MIS), where they must accurately depict multiple distinct instances in a single image based on complex prompts detailing individual feat…

AttributeLayout Generation

RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

2024-07-22 · Hanlin Tang, Yang Lin, Jing Lin, Qingsen Han 외

The memory and computational demands of Key-Value (KV) cache present significant challenges for deploying long-context language models. Previous approaches attempt to mitigate this issue by selectively dropping tokens, w…

Retrieval

ToDRE: Visual Token Pruning via Diversity and Task Awareness for Efficient Large Vision-Language Models

2025-05-24 · Duo Li, Zuhao Yang, Shijian Lu

The representation of visual inputs of large vision-language models (LVLMs) usually involves substantially more tokens than that of textual inputs, leading to significant computational overhead. Several recent studies st…

DecoderDiversityLarge Language Model