paper-with-me

홈 › Papers

The Distracting Effect: Understanding Irrelevant Passages in RAG

2025-05-11 · Chen Amiraz, Florin Cuconasu, Simone Filice, Zohar Karnin

A well-known issue with Retrieval Augmented Generation (RAG) is that retrieved passages that are irrelevant to the query sometimes distract the answer-generating LLM, causing it to provide an incorrect response. In this paper, we shed light on this core issue and formulate the distracting effect of a passage w.r.t. a query (and an LLM). We provide a quantifiable measure of the distracting effect of a passage and demonstrate its robustness across LLMs. Our research introduces novel methods for identifying and using hard distracting passages to improve RAG systems. By fine-tuning LLMs with these carefully selected distracting passages, we achieve up to a 7.5% increase in answering accuracy compared to counterparts fine-tuned on conventional RAG datasets. Our contribution is two-fold: first, we move beyond the simple binary classification of irrelevant passages as either completely unrelated vs. distracting, and second, we develop and analyze multiple methods for finding hard distracting passages. To our knowledge, no other research has provided such a comprehensive framework for identifying and utilizing hard distracting passages.

📄 PDF Abstract BibTeX arXiv:2505.06914

Code (0)

등록된 구현이 없습니다.

Tasks

Binary ClassificationRAGRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…

Similar Papers 제목 키워드 기반

Do RAG Systems Suffer From Positional Bias?

2025-05-21 · Florin Cuconasu, Simone Filice, Guy Horowitz, Yoelle Maarek 외

Retrieval Augmented Generation enhances LLM accuracy by adding passages retrieved from an external corpus to the LLM prompt. This paper investigates how positional bias - the tendency of LLMs to weight information differ…

RAGRetrievalRetrieval-augmented Generation

QPaug: Question and Passage Augmentation for Open-Domain Question Answering of LLMs

2024-06-20 · Minsang Kim, Cheoneum Park, Seungjun Baek

Retrieval-augmented generation (RAG) has received much attention for Open-domain question-answering (ODQA) tasks as a means to compensate for the parametric knowledge of large language models (LLMs). While previous appro…

Open-Domain Question AnsweringQuestion AnsweringRAGRetrieval+1

Making Retrieval-Augmented Language Models Robust to Irrelevant Context

2023-10-02 · Ori Yoran, Tomer Wolfson, Ori Ram, Jonathan Berant

Retrieval-augmented language models (RALMs) hold promise to produce language understanding systems that are are factual, efficient, and up-to-date. An important desideratum of RALMs, is that retrieved information helps m…

Language ModellingNatural Language InferenceOpen-Domain Question AnsweringQuestion Answering+1

DTP: A Simple yet Effective Distracting Token Pruning Framework for Vision-Language Action Models

2026-01-22 · Chenyang Li, Jieyuan Liu, Bin Li, Bo Gao 외 arxiv

Vision-Language Action (VLA) models have shown remarkable progress in robotic manipulation by leveraging the powerful perception abilities of Vision-Language Models (VLMs) to understand environments and directly output a…

Sorting through the noise: Testing robustness of information processing in pre-trained language models

2021-09-25 · EMNLP 2021 11 · Lalchand Pandia, Allyson Ettinger

Pre-trained LMs have shown impressive performance on downstream NLP tasks, but we have yet to establish a clear understanding of their sophistication when it comes to processing, retaining, and applying information prese…

Semantic SimilaritySemantic Textual Similarity