paper-with-me

홈 › Papers

MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models

2025-12-09 · Jusheng Zhang, Kaitong Cai, Xiaoyang Guo, Sidi Liu, Qinhan Lv, Ruiqi Chen, Jing Yang, Yijia Fan, Xiaofei Sun, Jian Wang, Ziliang Chen, Liang Lin, Keze Wang arxiv

The ability to perform Chain-of-Thought (CoT) reasoning marks a major milestone for multimodal models (MMs), enabling them to solve complex visual reasoning problems. Yet a critical question remains: is such reasoning genuinely grounded in visual evidence and logically coherent? Existing benchmarks emphasize generation but neglect verification, i.e., the capacity to assess whether a reasoning chain is both visually consistent and logically valid. To fill this gap, we introduce MM-CoT, a diagnostic benchmark specifically designed to probe the visual grounding and logical coherence of CoT reasoning in MMs. Instead of generating free-form explanations, models must select the sole event chain that satisfies two orthogonal constraints: (i) visual consistency, ensuring all steps are anchored in observable evidence, and (ii) logical coherence, ensuring causal and commonsense validity. Adversarial distractors are engineered to violate one of these constraints, exposing distinct reasoning failures. We evaluate leading vision-language models on MM-CoT and find that even the most advanced systems struggle, revealing a sharp discrepancy between generative fluency and true reasoning fidelity. MM-CoT shows low correlation with existing benchmarks, confirming that it measures a unique combination of visual grounding and logical reasoning. This benchmark provides a foundation for developing future models that reason not just plausibly, but faithfully and coherently within the visual world.

📄 PDF Abstract BibTeX arXiv:2512.08228

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningVisual GroundingVisual Reasoning

Similar Papers 제목 키워드 기반

AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning

2025-09-30 · Xiping Li, Jianghong Ma arxiv

Interleaved-Modal Chain-of-Thought (I-MCoT) advances vision-language reasoning, such as Visual Question Answering (VQA). This paradigm integrates specially selected visual evidence from the input image into the context o…

Visual Question Answering

Hallucination Detection via Internal States and Structured Reasoning Consistency in Large Language Models

2025-10-13 · Yusheng Song, Lirong Qiu, Xi Zhang, Zhihao Tang arxiv

The detection of sophisticated hallucinations in Large Language Models (LLMs) is hampered by a ``Detection Dilemma'': methods probing internal states (Internal State Probing) excel at identifying factual inconsistencies …

Mathematical ReasoningLogical Fallacies

Let's Think with Images Efficiently! An Interleaved-Modal Chain-of-Thought Reasoning Framework with Dynamic and Precise Visual Thoughts

2026-03-23 · Xu Liu, Yongheng Zhang, Qiguang Chen, Yao Li 외 arxiv

Recently, Interleaved-modal Chain-of-Thought (ICoT) reasoning has achieved remarkable success by leveraging both multimodal inputs and outputs, attracting increasing attention. While achieving promising performance, curr…

The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task

2023-11-15 · Yifan Wu, Pengchuan Zhang, Wenhan Xiong, Barlas Oguz 외

The study explores the effectiveness of the Chain-of-Thought approach, known for its proficiency in language tasks by breaking them down into sub-tasks and intermediate steps, in improving vision-language tasks that dema…

Visual Reasoning

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

2024-12-17 · Zihui Cheng, Qiguang Chen, Jin Zhang, Hao Fei 외

Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks s…

Multimodal Reasoning