paper-with-me

Papers

DMC-CF: Dynamic Multimodal CounterFactual QA benchmark for Causal Reasoning

2026-05-28 · Junzhe Zhang, Huixuan Zhang, Guirong Wang, Xingyao Zhang, Pei Liu, Lin Qu, Hu Wei, Xiaojun Wan arxiv

With the rapid advancement of multimodal large language models (MLLMs), models have demonstrated increasingly powerful multimodal capabilities. However, whether MLLMs trained through statistical learning can truly understand the causal relationships underlying the real world remains a key research question. In recent years, numerous multimodal causal reasoning datasets have been proposed. Nevertheless, these datasets are either limited in scale or constructed from synthetic images and videos, cartoon-based content, or other non-realistic multimodal sources. To address these limitations, we collect real-world videos and construct DMC-CF-Static, a large-scale benchmark for multimodal causal counterfactual reasoning. Furthermore, to mitigate issues such as data contamination in traditional static evaluation, we represent causal events using causal graphs and propose the Dynamic Graph Intervention (DGI) framework to build the dynamic evaluation benchmark DMC-CF-Dynamic from DMC-CF-Static. Experimental results on the overall DMC-CF, which includes both static and dynamic evaluation benchmarks, demonstrate that the multimodal causal reasoning capabilities of current multimodal large language models in real-world scenarios still require substantial improvement.

📄 PDF Abstract BibTeX arXiv:2605.29339

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for Counterfactual Question Answering

2025-06-12 · IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) 2024 1 · Adam Ishay, Zhun Yang, Joohyung Lee, Ilgu Kang 외

Causal and temporal reasoning about video dynamics is a challenging problem. While neuro-symbolic models that combine symbolic reasoning with neural-based perception and prediction have shown promise, they exhibit limita…

counterfactualCounterfactual ReasoningQuestion Answering

Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions

2026-06-04 · Wenhao Mu, Facundo Yan, Anik Mumssen, Marisa Eisenberg 외 arxiv

Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic benchmarks with observable counterfactual outcomes. Existing datasets either rely …

Causal Inference

Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency

2026-02-18 · Victoria Lin, Xinnuo Xu, Rachel Lawrence, Risa Ueno 외 arxiv

Despite their strong performance on reasoning benchmarks, large language models (LLMs) have proven brittle when presented with counterfactual questions, suggesting weaknesses in their causal reasoning ability. While rece…

InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?

2025-08-08 · Keummin Ka, Junhyeong Park, Jaehyun Jeon, Youngjae Yu arxiv

Recent advances in Vision-Language Models (VLMs) have demonstrated impressive capabilities in perception and reasoning. However, the ability to perform causal inference -- a core aspect of human cognition -- remains unde…

Causal InferenceVisual Grounding

LogicGaze: Benchmarking Causal Consistency in Visual Narratives via Counterfactual Verification

2026-01-30 · Rory Driscoll, Alexandros Christoforos, Chadbourne Davis arxiv

While sequential reasoning enhances the capability of Vision-Language Models (VLMs) to execute complex multimodal tasks, their reliability in grounding these reasoning chains within actual visual evidence remains insuffi…

Multimodal Reasoning