paper-with-me

홈 › Papers

NeedleChain: Measuring Intact Context Comprehension Capability of Large Language Models

2025-07-30 · Hyeonseok Moon, Heuiseok Lim arxiv

Recent reports suggest that LLMs can handle increasingly long contexts. However, many existing benchmarks for context understanding embed substantial query-irrelevant content, which shifts evaluation toward retrieving relevant snippets rather than fully integrating all provided information. Under this setting, we view that current benchmarks can overestimate true context-understanding ability of LLMs. In particular, we demonstrate that when the context consists entirely of query-relevant text, even advanced models such as GPT-4o fail to reliably integrate inputs as short as 200 tokens. To evaluate this capability more rigorously, we introduce NeedleChain, a benchmark designed to test whether models can faithfully incorporate all given evidence. NeedleChain includes three variants that differ in the required order of comprehension, along with a parallel benchmark based on the needle-in-a-haystack(NIAH) paradigm. By comparing these variants, NeedleChain enables a more comprehensive assessment of context understanding. We further propose a training-free strategy that encourages models to reflect all available information, ROPE contraction, highlighting the importance of full-context integration and pointing to new directions for improving reliable reasoning over context.

📄 PDF Abstract BibTeX arXiv:2507.22411

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ReadBench: Measuring the Dense Text Visual Reading Ability of Vision-Language Models

2025-05-25 · Benjamin Clavié, Florian Brand

Recent advancements in Large Vision-Language Models (VLMs), have greatly enhanced their capability to jointly process text and images. However, despite extensive benchmarks evaluating visual comprehension (e.g., diagrams…

Optical Character Recognition (OCR)Reading Comprehension

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

2026-07-13 · Brenda Lelis, Rodrigo Cabral-Carvalho arxiv

Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instructions, inside the same finite prompt used for the current task. This…

The Thin Line Between Comprehension and Persuasion in LLMs

2025-07-02 · Adrian de Wynter, Tangming Yuan arxiv

Large language models (LLMs) are excellent at maintaining high-level, convincing dialogue, but it remains unclear whether their persuasive success reflects genuine understanding of the discourse. We examine this question…

The grip of grammar on meaning uncertainty: cross-linguistic evidence, neural correlates, and clinical relevance

2026-05-02 · Rui He, Claudio Palominos, Samuele Vallisa, Ni Yang 외 arxiv

Isolated word meanings are inherently uncertain. This uncertainty reduces when they are combined and anchored in context. We propose that grammar compresses meaning uncertainty cross-linguistically, which is reflected in…

Rethinking LoRA Memory Through the Lens of KV Cache Compression

2026-06-04 · Chunsheng Zuo, Liaoyaqi Wang, William Jurayj, William Fleshman 외 arxiv

Parametric retrieval augmentation encodes document information into lightweight, document-specific modules such as LoRA adapters, reducing the need to include all evidence as input context. However, it remains unclear ho…

Question AnsweringAnswer Generation