paper-with-me

홈 › Papers

ACoRN: Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models

2025-04-17 · Singon Kim, Gunho Jung, Seong-Whan Lee

Abstractive compression utilizes smaller langauge models to condense query-relevant context, reducing computational costs in retrieval-augmented generation (RAG). However,retrieved documents often include information that is either irrelevant to answering the query or misleading due to factual incorrect content, despite having high relevance scores. This behavior indicates that abstractive compressors are more likely to omit important information essential for the correct answer, especially in long contexts where attention dispersion occurs. To address this issue, we categorize retrieved documents in a more fine-grained manner and propose Abstractive Compression Robust against Noise (ACoRN), which introduces two novel training steps. First, we use offline data augmentation on the training dataset to enhance compressor robustness against two distinct types of retrieval noise. Second, since the language modelbased compressor cannot fully utilize information from multiple retrieved documents and exhibits positional bias, we perform finetuning to generate summaries centered around key information that directly supports the correct answer. Our experiments demonstrate that T5-large, trained with ACoRN as a compressor, improves EM and F1 scores while preserving the answer string, which could serve as direct evidence. ACoRN excels on datasets with many accuracy-reducing documents, making it highly useful in real-world scenarios.

📄 PDF Abstract BibTeX arXiv:2504.12673

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationRAGRetrievalRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

Noise-Robust Abstractive Compression in Retrieval-Augmented Language Models

2025-11-19 · Singon Kim arxiv

Abstractive compression utilizes smaller langauge models to condense query-relevant context, reducing computational costs in retrieval-augmented generation (RAG). However, retrieved documents often include information th…

Data Augmentation

EXIT: Context-Aware Extractive Compression for Enhancing Retrieval-Augmented Generation

2024-12-17 · Taeho Hwang, Sukmin Cho, Soyeong Jeong, Hoyun Song 외

We introduce EXIT, an extractive context compression framework that enhances both the effectiveness and efficiency of retrieval-augmented generation (RAG) in question answering (QA). Current RAG systems often struggle wh…

Question AnsweringRAGRetrievalRetrieval-augmented Generation+1

An Information Bottleneck Perspective for Effective Noise Filtering on Retrieval-Augmented Generation

2024-06-03 · Kun Zhu, Xiaocheng Feng, Xiyuan Du, Yuxuan Gu 외

Retrieval-augmented generation integrates the capabilities of large language models with relevant information retrieved from an extensive corpus, yet encounters challenges when confronted with real-world noisy data. One …

Answer GenerationQuestion AnsweringRetrievalRetrieval-augmented Generation

Recursive Abstractive Processing for Retrieval in Dynamic Datasets

2024-10-02 · Charbel Chucri, Rami Azouz, Joachim Ott

Recent retrieval-augmented models enhance basic methods by building a hierarchical structure over retrieved text chunks through recursive embedding, clustering, and summarization. The most relevant information is then re…

ClusteringRetrieval

ACORN: Performant and Predicate-Agnostic Search Over Vector Embeddings and Structured Data

2024-03-07 · Liana Patel, Peter Kraft, Carlos Guestrin, Matei Zaharia

Applications increasingly leverage mixed-modality data, and must jointly search over vector data, such as embedded images, text and video, as well as structured data, such as attributes and keywords. Proposed methods for…