paper-with-me

Papers

Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models

2026-01-27 · Jialong Wu, Xiaoying Zhang, Hongyi Yuan, Xiangcheng Zhang, Tianhao Huang, Changjing He, Chaoyi Deng, Renrui Zhang, Youbin Wu, Mingsheng Long arxiv

Humans construct internal world models and reason by manipulating the concepts within these models. Recent advances in AI, particularly chain-of-thought (CoT) reasoning, approximate such human cognitive abilities, where world models are believed to be embedded within large language models. Expert-level performance in formal and abstract domains such as mathematics and programming has been achieved in current systems by relying predominantly on verbal reasoning. However, they still lag far behind humans in domains like physical and spatial intelligence, which require richer representations and prior knowledge. The emergence of unified multimodal models (UMMs) capable of both verbal and visual generation has therefore sparked interest in more human-like reasoning grounded in complementary multimodal pathways, though their benefits remain unclear. From a world-model perspective, this paper presents the first principled study of when and how visual generation benefits reasoning. Our key position is the visual superiority hypothesis: for certain tasks--particularly those grounded in the physical world--visual generation more naturally serves as world models, whereas purely verbal world models encounter bottlenecks arising from representational limitations or insufficient prior knowledge. Theoretically, we formalize internal world modeling as a core component of CoT reasoning and analyze distinctions among different forms of world models. Empirically, we identify tasks that necessitate interleaved visual-verbal CoT reasoning, constructing a new evaluation suite, VisWorld-Eval. Controlled experiments on a state-of-the-art UMM show that interleaved CoT significantly outperforms purely verbal CoT on tasks that favor visual world modeling, but offers no clear advantage otherwise. Together, this work clarifies the potential of multimodal world modeling for more powerful, human-like multimodal AI.

📄 PDF Abstract BibTeX arXiv:2601.19834

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

2025-06-20 · Zeyuan Yang, Xueyang Yu, Delin Chen, Maohao Shen 외

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train…

Image GenerationMultimodal ReasoningVisual Reasoning

Thinking with Generated Images

2025-05-28 · Ethan Chern, Zhulin Hu, Steffi Chern, Siqi Kou 외

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to natively think across text and vision modaliti…

Visual Reasoning

VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models

2025-11-14 · Xinlei Yu, Chengming Xu, Guibin Zhang, Zhangquan Chen 외 arxiv

Despite the remarkable success of Vision-Language Models (VLMs), their performance on a range of complex visual tasks is often hindered by a "visual processing bottleneck": a propensity to lose grounding in visual eviden…

Free Lunch for Unified Multimodal Models: Enhancing Generation via Reflective Rectification with Inherent Understanding

2026-04-15 · Yibo Jiang, Tao Wu, Rui Jiang, Yehao Lu 외 arxiv

Unified Multimodal Models (UMMs) aim to integrate visual understanding and generation within a single structure. However, these models exhibit a notable capability mismatch, where their understanding capability significa…

Visual Reasoning

SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

2024-01-22 · CVPR 2024 1 · Boyuan Chen, Zhuo Xu, Sean Kirmani, Brian Ichter 외

Understanding and reasoning about spatial relationships is a fundamental capability for Visual Question Answering (VQA) and robotics. While Vision Language Models (VLM) have demonstrated remarkable performance in certain…

Question AnsweringSpatial ReasoningVisual Question AnsweringVisual Question Answering (VQA)