paper-with-me

Papers

MM-THEBench: Do Reasoning MLLMs Think Reasonably?

2026-01-30 · Zhidian Huang, Zijun Yao, Ji Qi, Shangqing Tu, Junxian Ma, Jinxin Liu, Weichuan Liu, Xiaoyin Che, Lei Hou, Juanzi Li arxiv

Recent advances in multimodal large language models (MLLMs) mark a shift from non-thinking models to post-trained reasoning models capable of solving complex problems through thinking. However, whether such thinking mitigates hallucinations in multimodal perception and reasoning remains unclear. Self-reflective reasoning enhances robustness but introduces additional hallucinations, and subtle perceptual errors still result in incorrect or coincidentally correct answers. Existing benchmarks primarily focus on models before the emergence of reasoning MLLMs, neglecting the internal thinking process and failing to measure the hallucinations that occur during thinking. To address these challenges, we introduce MM-THEBench, a comprehensive benchmark for assessing hallucinations of intermediate CoTs in reasoning MLLMs. MM-THEBench features a fine-grained taxonomy grounded in cognitive dimensions, diverse data with verified reasoning annotations, and a multi-level automated evaluation framework. Extensive experiments on mainstream reasoning MLLMs reveal insights into how thinking affects hallucination and reasoning capability in various multimodal tasks.

📄 PDF Abstract BibTeX arXiv:2601.22735

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning

2025-10-27 · Shijian Wang, Jiarui Jin, Xingjian Wang, Linxin Song 외 arxiv

Recent advances in image reasoning methods, particularly "Thinking with Images", have demonstrated remarkable success in Multimodal Large Language Models (MLLMs); however, this dynamic reasoning paradigm has not yet been…

Reinforcement Learning

Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

2025-01-03 · Yifan Du, Zikang Liu, YiFan Li, Wayne Xin Zhao 외

Recently, slow-thinking reasoning systems, built upon large language models (LLMs), have garnered widespread attention by scaling the thinking time during inference. There is also growing interest in adapting this capabi…

Language ModelingLanguage ModellingVisual Reasoning

SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning

2025-11-04 · Fangxun Shu, Yongjie Ye, Yue Liao, Zijian Kang 외 arxiv

We introduce SAIL-RL, a reinforcement learning (RL) post-training framework that enhances the reasoning capabilities of multimodal large language models (MLLMs) by teaching them when and how to think. Existing approaches…

Reinforcement Learning

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

2025-11-05 · Jindong Hong, Tianjie Chen, Lingjie Luo, Chuanyang Zheng 외 arxiv

A recent advancement in Multimodal Large Language Models (MLLMs) research is the emergence of "reasoning MLLMs" that offer explicit control over their internal thinking processes (normally referred as the "thinking mode"…

Linguistic Analysis, Description, and Typological Exploration with Categorial Grammar (TheBench Guide)

2024-06-03 · Cem Bozsahin

TheBench is a tool to study monadic structures in natural language. It is for writing monadic grammars to explore analyses, compare diverse languages through their categories, and to train models of grammar from form-mea…