paper-with-me

Papers

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

2025-01-13 · Chengzu Li, Wenshan Wu, Huanyu Zhang, Yan Xia, Shaoguang Mao, Li Dong, Ivan Vulić, Furu Wei

Chain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning.

📄 PDF Abstract BibTeX arXiv:2501.07542

Code (1)

janhq/visual-thinker pytorch

Tasks

Spatial Reasoning

Similar Papers 제목 키워드 기반

FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation

2026-01-20 · Jing Zuo, Lingzhou Mu, Fan Jiang, Chengcheng Ma 외 arxiv

Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to jointly understand multimodal instructions and visual-spatial context while reasoning over long action sequences. Re…

Vision-Language Navigation

Imagination Helps Visual Reasoning, But Not Yet in Latent Space

2026-02-26 · You Li, Chi Chen, Yanghao Li, Fanhu Zeng 외 arxiv

Latent visual reasoning aims to mimic human's imagination process by meditating through hidden states of Multimodal Large Language Models. While recognized as a promising paradigm for visual reasoning, the underlying mec…

Visual Reasoning

Enhancing Visual Reasoning with Autonomous Imagination in Multimodal Large Language Models

2024-11-27 · Jingming Liu, Yumeng Li, Boyuan Xiao, Yichang Jian 외

There have been recent efforts to extend the Chain-of-Thought (CoT) paradigm to Multimodal Large Language Models (MLLMs) by finding visual clues in the input scene, advancing the visual reasoning ability of MLLMs. Howeve…

Visual Reasoning

AbductiveMLLM: Boosting Visual Abductive Reasoning Within MLLMs

2026-01-06 · Boyu Chang, Qi Wang, Xi Guo, Zhixiong Nan 외 arxiv

Visual abductive reasoning (VAR) is a challenging task that requires AI systems to infer the most likely explanation for incomplete visual observations. While recent MLLMs develop strong general-purpose multimodal reason…

Multimodal Reasoning

Imagine2Act: Leveraging Object-Action Motion Consistency from Imagined Goals for Robotic Manipulation

2025-09-21 · Liang Heng, Jiadong Xu, Yiwen Wang, Xiaoqi Li 외 arxiv

Relational object rearrangement (ROR) tasks (e.g., insert flower to vase) require a robot to manipulate objects with precise semantic and geometric reasoning. Existing approaches either rely on pre-collected demonstratio…

Object RearrangementPoint Clouds