paper-with-me

홈 › Papers

GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning

2024-02-03 · Yanbin Wei, Shuai Fu, Weisen Jiang, Zejian Zhang, Zhixiong Zeng, Qi Wu, James T. Kwok, Yu Zhang

Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reasoning. The potential benefits and capabilities of representing graph structures as visual images (i.e., $\textit{visual graph}$) are still unexplored. To fill the gap, we innovatively propose an end-to-end framework, called $\textbf{G}$raph to v$\textbf{I}$sual and $\textbf{T}$extual Integr$\textbf{A}$tion (GITA), which firstly incorporates visual graphs into general graph reasoning. Besides, we establish $\textbf{G}$raph-based $\textbf{V}$ision-$\textbf{L}$anguage $\textbf{Q}$uestion $\textbf{A}$nswering (GVLQA) dataset from existing graph data, which is the first vision-language dataset for general graph reasoning purposes. Extensive experiments on the GVLQA dataset and five real-world datasets show that GITA outperforms mainstream LLMs in terms of general graph reasoning capabilities. Moreover, We highlight the effectiveness of the layout augmentation on visual graphs and pretraining on the GVLQA dataset.

📄 PDF Abstract BibTeX arXiv:2402.02130

Code (1)

WEIYanbin1999/GITA 공식 구현 pytorch

Tasks

Link PredictionNode Classification

Similar Papers 제목 키워드 기반

Multimodal Contextualized Semantic Parsing from Speech

2024-06-10 · Jordan Voas, Raymond Mooney, David Harwath

We introduce Semantic Parsing in Contextual Environments (SPICE), a task designed to enhance artificial agents' contextual awareness by integrating multimodal inputs with prior contexts. SPICE goes beyond traditional sem…

Data Integrationgraph constructionSemantic Parsing

Benchmarking Graph Neural Networks for Document Layout Analysis in Public Affairs

2025-05-12 · Miguel Lopez-Duran, Julian Fierrez, Aythami Morales, Ruben Tolosana 외

The automatic analysis of document layouts in digital-born PDF documents remains a challenging problem due to the heterogeneous arrangement of textual and nontextual elements and the imprecision of the textual metadata i…

BenchmarkingDocument Layout AnalysisFeature Engineeringgraph construction+1

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

2026-09-04 · Bahar Uddin Mahmud, Sumit Barua, Guan Yue Hong, Ajay Gupta 외 arxiv

Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning…

Object RecognitionKnowledge Graphs

Toward Intelligent Scene Augmentation for Context-Aware Object Placement and Sponsor-Logo Integration

2025-12-25 · Unnati Saraswat, Tarun Rao, Namah Gupta, Shweta Swami 외 arxiv

Intelligent image editing increasingly relies on advances in computer vision, multimodal reasoning, and generative modeling. While vision-language models (VLMs) and diffusion models enable guided visual manipulation, exi…

Multimodal ReasoningImage Editing

Digital Twin Buildings: 3D Modeling, GIS Integration, and Visual Descriptions Using Gaussian Splatting, ChatGPT/Deepseek, and Google Maps Platform

2025-02-09 · Kyle Gao, Dening Lu, Liangzhi Li, Nan Chen 외

Urban digital twins are virtual replicas of cities that use multi-source data and data analytics to optimize urban planning, infrastructure management, and decision-making. Towards this, we propose a framework focused on…

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model+1