paper-with-me

홈 › Papers

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels

2026-07-27 · Zhuchenyang Liu, Yao Zhang, Yu Xiao hf

Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investigates whether this failure is partially limited by what the model can express through coordinates. On a verified bilingual CiteVQA subset, we compare the coordinate interface with a language interface in which the model outputs only text, quoting its evidence verbatim, and a multimodal retriever returns the location of each quote as a page region proposed by a layout parser (tables and figures are quoted through their captions or notes); the comparison is repeated over six open vision-language models. Compared with the coordinate interface, evidence recall rises from at most 8 points to between 26 and 47 and the hallucination rate roughly halves, with little change in answer quality. Building on this comparison, we use the same quote-and-retrieve pipeline as a training scaffold: because region-level evidence labels are expensive to collect for long documents, we introduce a GRPO recipe whose reward is a judge's reading of the gold answer and crops of the retrieved regions, training the model to quote better evidence without any region labels and raising an 8B backbone's strict attributed accuracy from 22.4 to 33.8. These findings indicate a practical path to improve attribution"without a coordinate interface and without costly region-level supervision.

📄 PDF Abstract BibTeX arXiv:2607.24651

Code (3)

InsomaniacElf/sg-tamil-tts-resources- ★ 1
Noblegasesgoo/report-review ★ 2
Ryenhails/quote-and-retrieve

Similar Papers 제목 키워드 기반

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation

2026-05-02 · Peiyang Liu, Ziqiang Cui, Xi Wang, Di Liang 외 arxiv

Iterative Retrieval-Augmented Generation (iRAG) has emerged as a powerful paradigm for answering complex multi-hop questions by progressively retrieving and reasoning over external documents. However, current systems pre…

Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning

2025-11-15 · Shuochen Liu, Pengfei Luo, Chao Zhang, Yuhao Chen 외 arxiv

Aiming to identify precise evidence sources from visual documents, visual evidence attribution for visual document retrieval-augmented generation (VD-RAG) ensures reliable and verifiable predictions from vision-language …

Reinforcement LearningQuestion Answering

VISA: Retrieval Augmented Generation with Visual Source Attribution

2024-12-19 · Xueguang Ma, Shengyao Zhuang, Bevan Koopman, Guido Zuccon 외

Generation with source attribution is important for enhancing the verifiability of retrieval-augmented generation (RAG) systems. However, existing approaches in RAG primarily link generated content to document-level refe…

Answer GenerationRAGRetrievalRetrieval-augmented Generation

Understanding Retrieval Augmentation for Long-Form Question Answering

2023-10-18 · Hung-Ting Chen, Fangyuan Xu, Shane A. Arora, Eunsol Choi

We present a study of retrieval-augmented language models (LMs) on long-form question answering. We analyze how retrieval augmentation impacts different LMs, by comparing answers generated from models while using the sam…

FormLong Form Question AnsweringQuestion AnsweringRetrieval+1

Attribute or Abstain: Large Language Models as Long Document Assistants

2024-07-10 · Jan Buchmann, Xiao Liu, Iryna Gurevych

LLMs can help humans working with long documents, but are known to hallucinate. Attribution can increase trust in LLM responses: The LLM provides evidence that supports its response, which enhances verifiability. Existin…

AttributeRAGResponse GenerationRetrieval