paper-with-me

홈 › Papers

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation

2026-05-02 · Peiyang Liu, Ziqiang Cui, Xi Wang, Di Liang, Wei Ye arxiv

Iterative Retrieval-Augmented Generation (iRAG) has emerged as a powerful paradigm for answering complex multi-hop questions by progressively retrieving and reasoning over external documents. However, current systems predominantly operate on parsed text, which creates two critical bottlenecks: (1) \textit{Coarse-grained attribution}, where users are burdened with manually locating evidence within lengthy documents based on vague text-level citations; and (2) \textit{Visual semantic loss}, where the conversion of visually rich documents (e.g., slides, PDFs with charts) into text discards spatial logic and layout cues essential for reasoning. To bridge this gap, we present \textbf{Chain of Evidence (CoE)}, a retriever-agnostic visual attribution framework that leverages Vision-Language Models to reason directly over screenshots of retrieved document candidates. CoE eliminates format-specific parsing and outputs precise bounding boxes, visualizing the complete reasoning chain within the retrieved candidate set. We evaluate CoE on two distinct benchmarks: \textbf{Wiki-CoE}, a large-scale dataset of structured web pages derived from 2WikiMultiHopQA, and \textbf{SlideVQA}, a challenging dataset of presentation slides featuring complex diagrams and free-form layouts. Experiments demonstrate that fine-tuned Qwen3-VL-8B-Instruct achieves robust performance, significantly outperforming text-based baselines in scenarios requiring visual layout understanding, while establishing a retriever-agnostic solution for pixel-level interpretable iRAG. Our code is available at https://github.com/PeiYangLiu/CoE.git.

📄 PDF Abstract BibTeX arXiv:2605.01284

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Look as You Think: Unifying Reasoning and Visual Evidence Attribution for Verifiable Document RAG via Reinforcement Learning

2025-11-15 · Shuochen Liu, Pengfei Luo, Chao Zhang, Yuhao Chen 외 arxiv

Aiming to identify precise evidence sources from visual documents, visual evidence attribution for visual document retrieval-augmented generation (VD-RAG) ensures reliable and verifiable predictions from vision-language …

Reinforcement LearningQuestion Answering

Toward Faithful Segmentation Attribution via Benchmarking and Dual-Evidence Fusion

2026-03-23 · Abu Noman Md Sakib, OFM Riaz Rahman Aranya, Kevin Desai, Zijie Zhang arxiv

Attribution maps for semantic segmentation are almost always judged by visual plausibility. Yet looking convincing does not guarantee that the highlighted pixels actually drive the model's prediction, nor that attributio…

Semantic Segmentation

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models

2025-09-26 · Jiawei Liang, Ruoyu Chen, Xianghao Jiao, Siyuan Liang 외 arxiv

Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains difficult to inspect. Recent logit-lens attribution methods project each visual-tok…

Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Models

2026-05-19 · Guangzhi Xiong, Qiao Jin, Sanchit Sinha, Zhiyong Lu 외 arxiv

Large Vision Language Models (LVLMs) show promise in medical applications, but their inability to faithfully ground responses in visual evidence raises serious concerns about clinical trustworthiness. While visual attrib…

Object-Level Explanations for Image Geolocation Models: a GeoGuessr use-case

2026-04-29 · Emilie Durrieu, Christophe Hurter, Philippe Muller, Victor Boutin arxiv

When humans play geolocation games such as GeoGuessr, they rely on concrete visual cues, such as road markings, vegetation, or architectural details, to infer where an image was captured. Whether image geolocation models…