paper-with-me

홈 › Papers

Separate the Wheat from the Chaff: Winnowing Down Divergent Views in Retrieval Augmented Generation

2025-11-01 · Song Wang, Zihan Chen, Peng Wang, Zhepei Wei, Zhen Tan, Yu Meng, Cong Shen, Jundong Li arxiv

Retrieval-augmented generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources to address their limitations in accessing up-to-date or specialized information. A natural strategy to increase the likelihood of retrieving relevant information is to expand the number of retrieved documents. However, involving more documents could introduce significant noise, as many documents may be irrelevant or misleading, thereby reducing the overall accuracy of the generated responses. To overcome the challenge associated with handling a larger number of documents, we propose WinnowRAG, a novel RAG framework designed to systematically filter out noisy documents while preserving valuable content -- a process we refer to as winnowing. WinnowRAG operates in two stages: In Stage I, we perform query-aware clustering to group similar documents and form distinct topic clusters. Each cluster is assigned to an LLM agent for generating a unique answer. In Stage II, we perform winnowing, wherein a critic LLM evaluates the outputs of multiple agents and iteratively separates useful documents from noisy ones. To retain useful documents when discarding agents, we propose two strategic merging techniques to ensure that only relevant knowledge is used for generating the final response. Crucially, WinnowRAG is model-agnostic and does not require any model fine-tuning, making it easily adaptable to various tasks. Extensive experiments on various realistic datasets demonstrate the effectiveness of WinnowRAG over state-of-the-art baselines.

📄 PDF Abstract BibTeX arXiv:2511.04700

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automating Document Discovery in the Systematic Review Process: How to Use Chaff to Extract Wheat

2018-05-01 · LREC 2018 5 · Christopher Norman, Mariska Leeflang, Pierre Zweigenbaum, Aur{\'e}lie N{\'e}v{\'e}ol
Decision Making

Separate the Wheat from the Chaff: A Post-Hoc Approach to Safety Re-Alignment for Fine-Tuned Language Models

2024-12-15 · Di wu, Xin Lu, Yanyan Zhao, Bing Qin

Although large language models (LLMs) achieve effective safety alignment at the time of release, they still face various safety challenges. A key issue is that fine-tuning often compromises the safety alignment of LLMs. …

Safety Alignment

SVALex: a CEFR-graded Lexical Resource for Swedish Foreign and Second Language Learners

2016-05-01 · LREC 2016 5 · Thomas Fran{\c{c}}ois, Elena Volodina, Ildik{\'o} Pil{\'a}n, Ana{\"\i}s Tack

The paper introduces SVALex, a lexical resource primarily aimed at learners and teachers of Swedish as a foreign and second language that describes the distribution of 15,681 words and expressions across the Common Europ…

WheaCha: A Method for Explaining the Predictions of Models of Code

2021-02-09 · Yu Wang, Ke Wang, Linzhang Wang

Attribution methods have emerged as a popular approach to interpreting model predictions based on the relevance of input features. Although the feature importance ranking can provide insights of how models arrive at a pr…

BIG-bench Machine LearningCode SummarizationFeature ImportanceImage Classification+2

Separate the Wheat from the Chaff: Model Deficiency Unlearning via Parameter-Efficient Module Operation

2023-08-16 · Xinshuo Hu, Dongfang Li, Baotian Hu, Zihao Zheng 외

Large language models (LLMs) have been widely used in various applications but are known to suffer from issues related to untruthfulness and toxicity. While parameter-efficient modules (PEMs) have demonstrated their effe…

Language ModelingLanguage ModellingMathematical Reasoning