paper-with-me

홈 › Papers

The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning

2026-05-11 · Muhan Gao, Zih-Ching Chen, Kuan-Hao Huang arxiv

As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understanding how distracting information affects long-context performance becomes critical. Prior work shows that semantically relevant yet misleading documents degrade performance, but the quantitative relationship between the proportion of distractors and performance remains unstudied. In this work, we systematically vary the hard-distractor proportion in fixed-length contexts, revealing a striking nonlinear pattern: as the proportion of hard distractors increases, performance drops sharply within the first small fraction, while the remainder of the range yields only marginal additional decline. We term this ''The First Drop of Ink'' effect, analogous to how a single drop of ink contaminates water. Our theoretical and empirical analyses grounded in attention mechanics show that hard distractors capture disproportionate attention even at small proportions, with diminishing marginal impact as their proportion grows. Controlled experiments further show that filtering gains mainly come from context-length reduction rather than distractor removal; substantial recovery requires reducing the hard-distractor proportion to near zero, highlighting the importance of upstream retrieval precision.

📄 PDF Abstract BibTeX arXiv:2605.10828

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Iterative LLM-Based Generation and Refinement of Distracting Conditions in Math Word Problems

2025-10-08 · Kaiqi Yang, Hang Li, Yucheng Chu, Zitao Liu 외 arxiv

Mathematical reasoning serves as a crucial testbed for the intelligence of large language models (LLMs), and math word problems (MWPs) are a popular type of math problems. Most MWP datasets consist of problems containing…

Mathematical Reasoning

Do RAG Systems Suffer From Positional Bias?

2025-05-21 · Florin Cuconasu, Simone Filice, Guy Horowitz, Yoelle Maarek 외

Retrieval Augmented Generation enhances LLM accuracy by adding passages retrieved from an external corpus to the LLM prompt. This paper investigates how positional bias - the tendency of LLMs to weight information differ…

RAGRetrievalRetrieval-augmented Generation

Sorting through the noise: Testing robustness of information processing in pre-trained language models

2021-09-25 · EMNLP 2021 11 · Lalchand Pandia, Allyson Ettinger

Pre-trained LMs have shown impressive performance on downstream NLP tasks, but we have yet to establish a clear understanding of their sophistication when it comes to processing, retaining, and applying information prese…

Semantic SimilaritySemantic Textual Similarity

Abstract Reasoning with Distracting Features

2019-12-02 · NeurIPS 2019 12 · Kecheng Zheng, Zheng-Jun Zha, Wei Wei

Abstraction reasoning is a long-standing challenge in artificial intelligence. Recent studies suggest that many of the deep architectures that have triumphed over other domains failed to work well in abstract reasoning. …

Reinforcement Learning

The Distracting Effect: Understanding Irrelevant Passages in RAG

2025-05-11 · Chen Amiraz, Florin Cuconasu, Simone Filice, Zohar Karnin

A well-known issue with Retrieval Augmented Generation (RAG) is that retrieved passages that are irrelevant to the query sometimes distract the answer-generating LLM, causing it to provide an incorrect response. In this …

Binary ClassificationRAGRetrieval-augmented Generation