Exploring and Evaluating Multimodal Knowledge Reasoning Consistency of Multimodal Large Language Models
In recent years, multimodal large language models (MLLMs) have achieved significant breakthroughs, enhancing understanding across text and vision. However, current MLLMs still face challenges in effectively integrating knowledge across these modalities during multimodal knowledge reasoning, leading to inconsistencies in reasoning outcomes. To systematically explore this issue, we propose four evaluation tasks and construct a new dataset. We conduct a series of experiments on this dataset to analyze and compare the extent of consistency degradation in multimodal knowledge reasoning within MLLMs. Based on the experimental results, we identify factors contributing to the observed degradation in consistency. Our research provides new insights into the challenges of multimodal knowledge reasoning and offers valuable guidance for future efforts aimed at improving MLLMs.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remains inadequately assessed. Current multimodal benchmarks predominantly …
DiagnosticImage ComprehensionvalidWhen Modalities Remember: Continual Learning for Multimodal Knowledge Graphs
Real-world multimodal knowledge graphs (MMKGs) are dynamic, with new entities, relations, and multimodal knowledge emerging over time. Existing continual knowledge graph reasoning (CKGR) methods focus on structural tripl…
Continual LearningKnowledge GraphsComparing Humans, GPT-4, and GPT-4V On Abstraction and Reasoning Tasks
We explore the abstract reasoning abilities of text-only and multimodal versions of GPT-4, using the ConceptARC benchmark [10], which is designed to evaluate robust understanding and reasoning with core-knowledge concept…
MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
Recent progress in the reasoning capabilities of multimodal large language models (MLLMs) has empowered them to address more complex tasks such as scientific analysis and mathematical reasoning. Despite their promise, ML…
Mathematical ReasoningMultimodal ReasoningExploring Failure Cases in Multimodal Reasoning About Physical Dynamics
In this paper, we present an exploration of LLMs' abilities to problem solve with physical reasoning in situated environments. We construct a simple simulated environment and demonstrate examples of where, in a zero-shot…
Language ModelingLanguage ModellingMultimodal ReasoningObject+1