paper-with-me

홈 › Papers

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation

2026-06-09 · Ian van Dort, Maria Heuss arxiv

Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in external documents, often using inline citations for verifiability. However, the faithfulness of these citations -- whether the model genuinely uses a source to generate an answer -- remains a critical, unverified assumption. This paper offers the first mechanistic account of how a large language model decides whether to attach an inline citation while answering a factoid question. Using the Llama-3.1-8B-Instruct model in a controlled experimental environment based on the PopQA dataset, we employ an activation patching approach. We map the underlying mechanism responsible for citation, discovering that it is not a single, localized component but a distributed, multi-stage "attributional ensemble" of attention heads and MLP layers. We show that amplifying or attenuating only those critical heads and MLPs repairs over 90% of missed citations and eliminates 69% of spurious ones on PopQA without harming answer accuracy. Although gains on the multi-document HotpotQA benchmark are modest, the same component set still moves citation rates in the intended direction, indicating that the underlying mechanism is not dataset-specific. The results reveal a potential disconnect between the model's apparent reasoning and its internal computational pathway, suggesting that inline citations can create a false sense of security.

📄 PDF Abstract BibTeX arXiv:2606.28358

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Generation-Time vs. Post-hoc Citation: A Holistic Evaluation of LLM Attribution

2025-09-25 · Yash Saxena, Raviteja Bommireddy, Ankur Padia, Manas Gaur arxiv

Trustworthy Large Language Models (LLMs) must cite human-verifiable sources in high-stakes domains such as healthcare, law, academia, and finance, where even small errors can have severe consequences. Practitioners and r…

Attributing Response to Context: A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation

2025-05-22 · Ruizhe Li, Chen Chen, Yuchen Hu, Yanjun Gao 외

Retrieval-Augmented Generation (RAG) leverages large language models (LLMs) combined with external contexts to enhance the accuracy and reliability of generated responses. However, reliably attributing generated content …

ARCAttributeComputational EfficiencyRAG+1

ContextCite: Attributing Model Generation to Context

2024-09-01 · Benjamin Cohen-Wang, Harshay Shah, Kristian Georgiev, Aleksander Madry

How do language models use information provided as context when generating a response? Can we infer whether a particular generated statement is actually grounded in the context, a misinterpretation, or fabricated? To hel…

Language ModelingLanguage Modellingmodel

Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents

2026-05-07 · Hailey Onweller, Elias Lumer, Austin Huber, Pia Ramchandani 외 arxiv

Large language models (LLMs) power deep research agents that synthesize information from hundreds of web sources into cited reports, yet these citations cannot be reliably verified. Current approaches either trust models…

Towards Fair RAG: On the Impact of Fair Ranking in Retrieval-Augmented Generation

2024-09-17 · To Eun Kim, Fernando Diaz

Modern language models frequently include retrieval components to improve their outputs, giving rise to a growing number of retrieval-augmented generation (RAG) systems. Yet, most existing work in RAG has underemphasized…

FairnessRAGRetrievalRetrieval-augmented Generation