Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions
This position paper argues that Retrieval-Augmented Generation (RAG) systems exhibit a factual bias-optimizing for epistemic uncertainty reduction while ignoring the aleatoric uncertainty inherent in opinion-rich content. This misalignment demands a paradigm shift in RAG system design. A survey of 34 major RAG benchmarks reveals that only one addresses opinion synthesis, confirming that the bias is structural and embedded in datasets, retrieval-generation objectives, and evaluation metrics alike. Beyond technical limitations, this bias poses risks to transparent and accountable AI. Namely, echo chamber effects that amplify dominant viewpoints, which can lead to opinion manipulation and under-representation of minority voices. We formalize the problem through the lens of uncertainty quantification, showing that factual queries should minimize posterior entropy while opinion queries must preserve it. We derive a unified objective over coverage, fidelity, and fairness using the Wasserstein distance. As an existence proof, we present Opinion-Aware RAG (O-RAG), an architecture featuring LLM-based opinion extraction and entity-linked opinion metadata. We evaluate it across two domains -- e-commerce seller forums and public hotel reviews. Experiments demonstrate 18-48% reduction in Wasserstein distance to corpus-level sentiment distributions, +26.8% sentiment diversity, and +42.7% entity match rate. Human evaluators preferred opinion-enriched generation 79.2% of the time. We propose a research agenda and argue that as RAG systems increasingly mediate access to information, their ability to represent diverse perspectives is of the essence.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Retrieval Augmented Spelling Correction for E-Commerce Applications
The rapid introduction of new brand names into everyday language poses a unique challenge for e-commerce spelling correction services, which must distinguish genuine misspellings from novel brand names that use unconvent…
Language ModelingLanguage ModellingLarge Language ModelRAG+3Backdoored Retrievers for Prompt Injection Attacks on Retrieval Augmented Generation of Large Language Models
Large Language Models (LLMs) have demonstrated remarkable capabilities in generating coherent text but remain limited by the static nature of their training data. Retrieval Augmented Generation (RAG) addresses this issue…
Backdoor AttackInformation RetrievalMisinformationRAG+2Beyond Perplexity: Let the Reader Select Retrieval Summaries via Spectrum Projection Score
Large Language Models (LLMs) have shown improved generation performance through retrieval-augmented generation (RAG) following the retriever-reader paradigm, which supplements model inputs with externally retrieved knowl…
GE-Chat: A Graph Enhanced RAG Framework for Evidential Response Generation of LLMs
Large Language Models are now key assistants in human decision-making processes. However, a common note always seems to follow: "LLMs can make mistakes. Be careful with important info." This points to the reality that no…
RAGResponse GenerationRetrievalRetrieval-augmented Generation+1ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
Retrieval-Augmented Generation (RAG) pipelines must address challenges beyond simple single-document retrieval, such as interpreting visual elements (tables, charts, images), synthesizing information across documents, an…
Answer GenerationVisual Grounding