paper-with-me

Papers

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs

2025-07-28 · Xueyao Wan, Hang Yu arxiv

Large Language Models (LLMs) often suffer from hallucinations, which Retrieval-Augmented Generation (RAG) and GraphRAG mitigate by incorporating external knowledge and knowledge graphs (KGs). However, GraphRAG remains text-centric due to the difficulty of constructing fine-grained Multimodal KGs (MMKGs). Existing fusion methods, such as shared embeddings or captioning, require task-specific training and fail to preserve visual structural knowledge or cross-modal reasoning paths. To bridge this gap, we propose MMGraphRAG, which integrates visual scene graphs with text KGs via a novel cross-modal fusion approach. It introduces SpecLink, a method leveraging spectral clustering for accurate cross-modal entity linking and path-based retrieval to guide generation. We also release the CMEL dataset, specifically designed for fine-grained multi-entity alignment in complex multimodal scenarios. Evaluations on CMEL, DocBench, and MMLongBench demonstrate that MMGraphRAG achieves state-of-the-art performance, showing robust domain adaptability and superior multimodal information processing capabilities.

📄 PDF Abstract BibTeX arXiv:2507.20804

Code (0)

등록된 구현이 없습니다.

Tasks

Knowledge GraphsEntity AlignmentEntity Linking

Similar Papers 제목 키워드 기반

Towards Interpreting Visual Information Processing in Vision-Language Models

2024-10-09 · Clement Neo, Luke Ong, Philip Torr, Mor Geva 외

Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA, a prominent VLM. Our approach focuses …

Language ModelingLanguage ModellingObject

VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events

2026-03-18 · Mohammad Qazim Bhat, Yufan Huang, Niket Agarwal, Hao Wang 외 arxiv

The rapid growth of ego-centric dashcam footage presents a major challenge for detecting safety-critical events such as collisions and near-collisions, scenarios that are brief, rare, and difficult for generic vision mod…

Visual Question AnsweringAutonomous DrivingAnomaly Detection

MMRQA: Signal-Enhanced Multimodal Large Language Models for MRI Quality Assessment

2025-09-29 · Fankai Jia, Daisong Gan, Zhe Zhang, Zhaochi Wen 외 arxiv

Magnetic resonance imaging (MRI) quality assessment is crucial for clinical decision-making, yet remains challenging due to data scarcity and protocol variability. Traditional approaches face fundamental trade-offs: sign…

Zero-shot Generalization

AgriChain Visually Grounded Expert Verified Reasoning for Interpretable Agricultural Vision Language Models

2026-04-09 · Hazza Mahmood, Yongqiang Yu, Rao Anwer arxiv

Accurate and interpretable plant disease diagnosis remains a major challenge for vision-language models (VLMs) in real-world agriculture. We introduce AgriChain, a dataset of approximately 11,000 expert-curated leaf imag…

Griffon-G: Bridging Vision-Language and Vision-Centric Tasks via Large Multimodal Models

2024-10-21 · Yufei Zhan, Hongyin Zhao, Yousong Zhu, Fan Yang 외

Large Multimodal Models (LMMs) have achieved significant breakthroughs in various vision-language and vision-centric tasks based on auto-regressive modeling. However, these models typically focus on either vision-centric…

Instruction Followingobject-detectionObject DetectionQuestion Answering+5