paper-with-me

Papers

Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models

2025-03-03 · Boyu Jia, Junzhe Zhang, Huixuan Zhang, Xiaojun Wan

In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating knowledge across these modalities during multimodal knowledge reasoning, leading to inconsistencies in reasoning outcomes. To systematically explore this issue, we propose four evaluation tasks and construct a new dataset. We conduct a series of experiments on this dataset to analyze and compare the extent of consistency degradation in multimodal knowledge reasoning within MLLMs. Based on the experimental results, we identify factors contributing to the observed degradation in consistency. Our research provides new insights into the challenges of multimodal knowledge reasoning and offers valuable guidance for future efforts aimed at improving MLLMs.

📄 PDF Abstract BibTeX arXiv:2503.04801

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

2025-05-30 · Ai Jian, Weijie Qiu, Xiaokun Wang, Peiyu Wang 외

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remains inadequately assessed. Current multimodal benchmarks predominantly …

DiagnosticImage Comprehensionvalid

When Modalities Remember: Continual Learning for Multimodal Knowledge Graphs

2026-04-03 · Linyu Li, Zhi Jin, Yichi Zhang, Dongming Jin 외 arxiv

Real-world multimodal knowledge graphs (MMKGs) are dynamic, with new entities, relations, and multimodal knowledge emerging over time. Existing continual knowledge graph reasoning (CKGR) methods focus on structural tripl…

Continual LearningKnowledge Graphs

Comparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks

2023-11-14 · Melanie Mitchell, Alessandro B. Palmarini, Arseny Moskvichev

We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concept…

MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning

2026-03-02 · Jiachun Li, Shaoping Huang, Zhuoran Jin, Chenlong Zhang 외 arxiv

Recent progress in the reasoning capabilities of multimodal large language models (MLLMs) has empowered them to address more complex tasks such as scientific analysis and mathematical reasoning. Despite their promise, ML…

Mathematical ReasoningMultimodal Reasoning

Exploring Failure Cases in Multimodal Reasoning About Physical Dynamics

2024-02-24 · Sadaf Ghaffari, Nikhil Krishnaswamy

In this paper, we present an exploration of LLMs' abilities to problem solve with physical reasoning in situated environments. We construct a simple simulated environment and demonstrate examples of where, in a zero-shot…

Language ModelingLanguage ModellingMultimodal ReasoningObject+1