paper-with-me

홈 › Papers

Mario: Multimodal Graph Reasoning with Large Language Models

2026-03-05 · Yuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu, Qiaoyu Tan arxiv

Recent advances in large language models (LLMs) have opened new avenues for multimodal reasoning. Yet, most existing methods still rely on pretrained vision-language models (VLMs) to encode image-text pairs in isolation, ignoring the relational structure that real-world multimodal data naturally form. This motivates reasoning on multimodal graphs (MMGs), where each node has textual and visual attributes and edges provide structural cues. Enabling LLM-based reasoning on such heterogeneous multimodal signals while preserving graph topology introduces two key challenges: resolving weak cross-modal consistency and handling heterogeneous modality preference. To address this, we propose Mario, a unified framework that simultaneously resolves the two above challenges and enables effective LLM-based reasoning over MMGs. Mario consists of two innovative stages. Firstly, a graph-conditioned VLM design that jointly refines textual and visual features through fine-grained cross-modal contrastive learning guided by graph topology. Secondly, a modality-adaptive graph instruction tuning mechanism that organizes aligned multimodal features into graph-aware instruction views and employs a learnable router to surface, for each node and its neighborhood, the most informative modality configuration to the LLM. Extensive experiments across diverse MMG benchmarks demonstrate that Mario consistently outperforms state-of-the-art graph models in both supervised and zero-shot scenarios for node classification and link prediction. The code will be made available at https://github.com/sunyuanfu/Mario.

📄 PDF Abstract BibTeX arXiv:2603.05181

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningContrastive LearningNode ClassificationLink Prediction

Similar Papers 제목 키워드 기반

MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

2024-01-16 · Minpeng Liao, Wei Luo, Chengxi Li, Jing Wu 외

Large language models (LLMs) have seen considerable advancements in natural language understanding tasks, yet there remains a gap to bridge before attaining true artificial general intelligence, especially concerning sho…

GSM8KMathMathematical ReasoningNatural Language Understanding

MARIOH: Multiplicity-Aware Hypergraph Reconstruction

2025-04-01 · Kyuhan Lee, Geon Lee, Kijung Shin

Hypergraphs offer a powerful framework for modeling higher-order interactions that traditional pairwise graphs cannot fully capture. However, practical constraints often lead to their simplification into projected graphs…

Step-level Value Preference Optimization for Mathematical Reasoning

2024-06-16 · Guoxin Chen, Minpeng Liao, Chengxi Li, Kai Fan

Direct Preference Optimization (DPO) using an implicit reward model has proven to be an effective alternative to reinforcement learning from human feedback (RLHF) for fine-tuning preference aligned large language models …

Learning-To-RankMathMathematical Reasoning

MarioGPT: Open-Ended Text2Level Generation through Large Language Models

2023-02-12 · NeurIPS 2023 11 · Shyam Sudhakaran, Miguel González-Duque, Claire Glanois, Matthias Freiberger 외

Procedural Content Generation (PCG) is a technique to generate complex and diverse environments in an automated way. However, while generating content with PCG methods is often straightforward, generating meaningful cont…

MARIO: Modular and Extensible Architecture for Computing Visual Statistics in RoboCup SPL

2022-09-20 · Domenico D. Bloisi, Andrea Pennisi, Cristian Zampino, Flavio Biancospino 외

This technical report describes a modular and extensible architecture for computing visual statistics in RoboCup SPL (MARIO), presented during the SPL Open Research Challenge at RoboCup 2022, held in Bangkok (Thailand). …

Camera CalibrationPose EstimationRobot Pose Estimation