paper-with-me

Papers

Multimodal Causal Reasoning Benchmark: Challenging Vision Large Language Models to Infer Causal Links Between Siamese Images

2024-08-15 · Zhiyuan Li, Heng Wang, Dongnan Liu, Chaoyi Zhang, Ao Ma, Jieting Long, Weidong Cai

Large Language Models (LLMs) have showcased exceptional ability in causal reasoning from textual information. However, will these causalities remain straightforward for Vision Large Language Models (VLLMs) when only visual hints are provided? Motivated by this, we propose a novel Multimodal Causal Reasoning benchmark, namely MuCR, to challenge VLLMs to infer semantic cause-and-effect relationship when solely relying on visual cues such as action, appearance, clothing, and environment. Specifically, we introduce a prompt-driven image synthesis approach to create siamese images with embedded semantic causality and visual cues, which can effectively evaluate VLLMs' causal reasoning capabilities. Additionally, we develop tailored metrics from multiple perspectives, including image-level match, phrase-level understanding, and sentence-level explanation, to comprehensively assess VLLMs' comprehension abilities. Our extensive experiments reveal that the current state-of-the-art VLLMs are not as skilled at multimodal causal reasoning as we might have hoped. Furthermore, we perform a comprehensive analysis to understand these models' shortcomings from different views and suggest directions for future research. We hope MuCR can serve as a valuable resource and foundational benchmark in multimodal causal reasoning research. The project is available at: https://github.com/Zhiyuan-Li-John/MuCR

📄 PDF Abstract BibTeX arXiv:2408.08105

Code (1)

zhiyuan-li-john/mucr 공식 구현

Tasks

Image GenerationSentence

Similar Papers 제목 키워드 기반

ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

2025-07-30 · Ananya Sadana, Yash Kumar Lal, Jiawei Zhou arxiv

Understanding causal relationships across modalities is a core challenge for multimodal models operating in real-world environments. We introduce ISO-Bench, a benchmark for evaluating whether models can infer causal depe…

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

2025-08-08 · Keummin Ka, Junhyeong Park, Jaehyun Jeon, Youngjae Yu arxiv

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive capabilities in perception and reasoning. However, the ability to perform causal inference -- a core aspect of human cognition -- remains unde…

Causal InferenceVisual Grounding

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

2026-05-28 · Junzhe Zhang, Huixuan Zhang, Guirong Wang, Xingyao Zhang 외 arxiv

With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However, whether MLLMs trained through statistical learning can truly unders…

Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective

2024-03-27 · Meiqi Chen, Yixin Cao, Yan Zhang, Chaochao Lu

Recent advancements in Large Language Models (LLMs) have facilitated the development of Multimodal LLMs (MLLMs). Despite their impressive capabilities, MLLMs often suffer from over-reliance on unimodal biases (e.g., lang…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification

2026-01-30 · Rory Driscoll, Alexandros Christoforos, Chadbourne Davis arxiv

While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insuffi…

Multimodal Reasoning