paper-with-me

홈 › Papers

Towards Multi-modal Graph Large Language Model

2025-06-11 · Xin Wang, Zeyang Zhang, Linxin Xiao, Haibo Chen, Chendi Ge, Wenwu Zhu

Multi-modal graphs, which integrate diverse multi-modal features and relations, are ubiquitous in real-world applications. However, existing multi-modal graph learning methods are typically trained from scratch for specific graph data and tasks, failing to generalize across various multi-modal graph data and tasks. To bridge this gap, we explore the potential of Multi-modal Graph Large Language Models (MG-LLM) to unify and generalize across diverse multi-modal graph data and tasks. We propose a unified framework of multi-modal graph data, task, and model, discovering the inherent multi-granularity and multi-scale characteristics in multi-modal graphs. Specifically, we present five key desired characteristics for MG-LLM: 1) unified space for multi-modal structures and attributes, 2) capability of handling diverse multi-modal graph tasks, 3) multi-modal graph in-context learning, 4) multi-modal graph interaction with natural language, and 5) multi-modal graph reasoning. We then elaborate on the key challenges, review related works, and highlight promising future research directions towards realizing these ambitious characteristics. Finally, we summarize existing multi-modal graph datasets pertinent for model training. We believe this paper can contribute to the ongoing advancement of the research towards MG-LLM for generalization across multi-modal graph data and tasks.

📄 PDF Abstract BibTeX arXiv:2506.09738

Code (0)

등록된 구현이 없습니다.

Tasks

Graph LearningIn-Context LearningLanguage ModelingLanguage ModellingLarge Language Modelmodel

Similar Papers 제목 키워드 기반

ELMM: Efficient Lightweight Multimodal Large Language Models for Multimodal Knowledge Graph Completion

2025-10-19 · Wei Huang, Peining Li, Meiyu Liang, Xu Hou 외 arxiv

Multimodal Knowledge Graphs (MKGs) extend traditional knowledge graphs by incorporating visual and textual modalities, enabling richer and more expressive entity representations. However, existing MKGs often suffer from …

Knowledge Graph CompletionKnowledge Graphs

GraphextQA: A Benchmark for Evaluating Graph-Enhanced Large Language Models

2023-10-12 · Yuanchun Shen, Ruotong Liao, Zhen Han, Yunpu Ma 외

While multi-modal models have successfully integrated information from image, video, and audio modalities, integrating graph modality into large language models (LLMs) remains unexplored. This discrepancy largely stems f…

Answer GenerationHallucinationLanguage ModelingLanguage Modelling+2

GraphVLM: Benchmarking Vision Language Models for Multimodal Graph Learning

2026-03-09 · Jiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha 외 arxiv

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in aligning and understanding multimodal signals, yet their potential to reason over structured data, where multimodal entities are connected throug…

Recommendation SystemsGraph Learning

Mario: Multimodal Graph Reasoning with Large Language Models

2026-03-05 · Yuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu 외 arxiv

Recent advances in large language models (LLMs) have opened new avenues for multimodal reasoning. Yet, most existing methods still rely on pretrained vision-language models (VLMs) to encode image-text pairs in isolation,…

Multimodal ReasoningContrastive LearningNode ClassificationLink Prediction

VL-KGE: Vision-Language Models Meet Knowledge Graph Embeddings

2026-03-02 · Athanasios Efthymiou, Stevan Rudinac, Monika Kackovic, Nachoem Wijnberg 외 arxiv

Real-world multimodal knowledge graphs (MKGs) are inherently heterogeneous, modeling entities that are associated with diverse modalities. Traditional knowledge graph embedding (KGE) methods excel at learning continuous …

Knowledge Graph EmbeddingKnowledge GraphsLink Prediction