paper-with-me

홈 › Papers

Graph-to-Vision: Multi-graph Understanding and Reasoning using Vision-Language Models

2025-03-27 · Ruizhou Li, Haiyun Jiang

Graph Neural Networks (GNNs), as the dominant paradigm for graph-structured learning, have long faced dual challenges of exponentially escalating computational complexity and inadequate cross-scenario generalization capability. With the rapid advancement of multimodal learning, Vision-Language Models (VLMs) have demonstrated exceptional cross-modal relational reasoning capabilities and generalization capacities, thereby opening up novel pathways for overcoming the inherent limitations of conventional graph learning paradigms. However, current research predominantly concentrates on investigating the single-graph reasoning capabilities of VLMs, which fundamentally fails to address the critical requirement for coordinated reasoning across multiple heterogeneous graph data in real-world application scenarios. To address these limitations, we propose the first multi-graph joint reasoning benchmark for VLMs. Our benchmark encompasses four graph categories: knowledge graphs, flowcharts, mind maps, and route maps,with each graph group accompanied by three progressively challenging instruction-response pairs. Leveraging this benchmark, we conducted comprehensive capability assessments of state-of-the-art VLMs and performed fine-tuning on open-source models. This study not only addresses the underexplored evaluation gap in multi-graph reasoning for VLMs but also empirically validates their generalization superiority in graph-structured learning.

📄 PDF Abstract BibTeX arXiv:2503.21435

Code (0)

등록된 구현이 없습니다.

Tasks

Graph LearningKnowledge GraphsRelational Reasoning

Similar Papers 제목 키워드 기반

HyperGVL: Benchmarking and Improving Large Vision-Language Models in Hypergraph Understanding and Reasoning

2026-04-17 · Yanbin Wei, Chun Kang, Siwei Li, Haoxuan Che 외 arxiv

Large Vision-Language Models (LVLMs) consistently require new arenas to guide their expanding boundaries, yet their capabilities with hypergraphs remain unexplored. In the real world, hypergraphs have significant practic…

Open-World 3D Scene Graph Generation for Retrieval-Augmented Reasoning

2025-11-08 · Fei Yu, Quan Deng, Shengeng Tang, Yuehua Li 외 arxiv

Understanding 3D scenes in open-world settings poses fundamental challenges for vision and robotics, particularly due to the limitations of closed-vocabulary supervision and static annotations. To address this, we propos…

Scene Graph GenerationScene UnderstandingQuestion AnsweringVisual Grounding

VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context

2024-05-08 · Yunxin Li, Baotian Hu, Haoyuan Shi, Wei Wang 외

Large Multimodal Models (LMMs) have achieved impressive success in visual understanding and reasoning, remarkably improving the performance of mathematical reasoning in a visual context. Yet, a challenging type of visual…

MathMathematical Reasoning

Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning

2024-12-18 · Yingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang 외

Large Vision-Language Models (LVLMs) have demonstrated remarkable performance across diverse tasks. Despite great success, recent studies show that LVLMs encounter substantial limitations when engaging with visual graphs…

BenchmarkingGraph LearningSelf-Supervised Learning

Graphhopper: Multi-Hop Scene Graph Reasoning for Visual Question Answering

2021-07-13 · Rajat Koner, Hang Li, Marcel Hildebrandt, Deepan Das 외

Visual Question Answering (VQA) is concerned with answering free-form questions about an image. Since it requires a deep semantic and linguistic understanding of the question and the ability to associate it with various …

NavigateQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)