paper-with-me

홈 › Papers

BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation

2026-01-12 · Xuan Li, Yining Wang, Haocai Luo, Shengping Liu, Jerry Liang, Ying Fu, Weihuang, Jun Yu, Junnan Zhu arxiv

Retrieval-Augmented Generation (RAG) has become a pivotal paradigm for Large Language Models (LLMs), yet current approaches struggle with visually rich documents by treating text and images as isolated retrieval targets. Existing methods relying solely on cosine similarity often fail to capture the semantic reinforcement provided by cross-modal alignment and layout-induced coherence. To address these limitations, we propose BayesRAG, a novel multimodal retrieval framework grounded in Bayesian inference and Dempster-Shafer evidence theory. Unlike traditional approaches that rank candidates strictly by similarity, BayesRAG models the intrinsic consistency of retrieved candidates across modalities as probabilistic evidence to refine retrieval confidence. Specifically, our method computes the posterior association probability for combinations of multimodal retrieval results, prioritizing text-image pairs that mutually corroborate each other in terms of both semantics and layout. Extensive experiments demonstrate that BayesRAG significantly outperforms state-of-the-art (SOTA) methods on challenging multimodal benchmarks. This study establishes a new paradigm for multimodal retrieval fusion that effectively resolves the isolation of heterogeneous modalities through an evidence fusion mechanism and enhances the robustness of retrieval outcomes. Our code is available at https://github.com/TioeAre/BayesRAG.

📄 PDF Abstract BibTeX arXiv:2601.07329

Code (0)

등록된 구현이 없습니다.

Tasks

Bayesian Inference

Similar Papers 제목 키워드 기반

M-ArtAgent: Evidence-Based Multimodal Agent for Implicit Art Influence Discovery

2026-04-08 · Hanyi Liu, Zhonghao Jiu, Minghao Wang, Yuhang Xie 외 arxiv

Implicit artistic influence, although visually plausible, is often undocumented and thus poses a historically constrained attribution problem: resemblance is necessary but not sufficient evidence. Most prior systems redu…

Uncertainty-Aware Web-Conditioned Scientific Fact-Checking

2026-04-13 · Ashwin Vinod, Katrin Erk arxiv

Scientific fact-checking is vital for assessing claims in specialized domains such as biomedicine and materials science, yet existing systems often hallucinate or apply inconsistent reasoning, especially when verifying t…

Conditional Evidence Reconstruction and Decomposition for Interpretable Multimodal Diagnosis

2026-04-18 · Shaowen Wan, Yanjun Lv, Lu Zhang, Dajiang Zhu 외 arxiv

Neurobiological and neurodegenerative diseases are inherently multifactorial, arising from coupled influences spanning genetic susceptibility, brain alterations, and environmental and behavioral factors. Multimodal model…

M-MRE: Extending the Mutual Reinforcement Effect to Multimodal Information Extraction

2025-04-24 · Chengguang Gan, Zhixi Cai, Yanbin Wei, Yunhao Liang 외

Mutual Reinforcement Effect (MRE) is an emerging subfield at the intersection of information extraction and model interpretability. MRE aims to leverage the mutual understanding between tasks of different granularities, …

Mutual Gaze and Linguistic Repetition in a Multimodal Corpus

2022-06-01 · LREC 2022 6 · Anais Murat, Maria Koutsombogera, Carl Vogel

This paper investigates the correlation between mutual gaze and linguistic repetition, a form of alignment, which we take as evidence of mutual understanding. We focus on a multimodal corpus made of three-party conversat…

Mutual Gaze