paper-with-me

홈 › Papers

From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models

2025-09-29 · Chenyue Zhou, Mingxuan Wang, Yanbiao Ma, Chenxu Wu, Wanyi Chen, Zhe Qian, Xinyu Liu, Yiwei Zhang, Junhao Wang, Hengbo Xu, Fei Luo, Xiaohua Chen, Xiaoshuai Hao, Hehan Li, Andi Zhang, Wenxuan Wang, Kaiyan Zhang, Guoli Jia, Lingling Li, Zhiwu Lu, Yang Lu, Yike Guo arxiv

Multimodal Large Language Models (MLLMs) strive to achieve a profound, human-like understanding of and interaction with the physical world, but often exhibit a shallow and incoherent integration when acquiring information (Perception) and conducting reasoning (Cognition). This disconnect leads to a spectrum of reasoning failures, with hallucination being the most prominent. Collectively, these issues expose a fundamental challenge: the ability to process pixels does not yet confer the ability to construct a coherent, credible internal world model. To systematically dissect and address this challenge, this survey introduces a novel and unified analytical framework: ``From Perception to Cognition." We deconstruct the complex process of vision-language interactive understanding into two interdependent layers: Perception, the foundational ability to accurately extract visual information and achieve fine-grained alignment with textual instructions; and Cognition, the higher-order capability for proactive, multi-step, goal-oriented reasoning built upon this perceptual foundation, the core of which is the formation of a dynamic observe-think-verify reasoning loop. Guided by this framework, this paper systematically analyzes the key bottlenecks of current MLLMs at both layers. It surveys the landscape of cutting-edge methods designed to address these challenges, spanning from techniques that enhance low-level visual representations to those that improve high-level reasoning paradigms. Furthermore, we review critical benchmarks and delineate future research directions. This survey aims to provide the research community with a clear, structured perspective for understanding the intrinsic limitations of current MLLMs and to illuminate the path toward building next-generation models capable of deep reasoning and a genuine understanding of the world.

📄 PDF Abstract BibTeX arXiv:2509.25373

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis

2026-04-03 · Dexiang Li, Zhenning Che, Haijun Zhang, Dongliang Zhou 외 arxiv

Pavement condition assessment is essential for road safety and maintenance. Existing research has made significant progress. However, most studies focus on conventional computer vision tasks such as classification, detec…

Visual Question AnsweringSemantic SegmentationObject Detection

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models

2026-06-24 · Haoxiang Sun, Tao Wang, Li Yuan, Jian Zhao 외 arxiv

Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, especially following the introduction of models such as OpenAI's O-series and DeepS…

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

2025-02-12 · Shixiang Tang, Yizhou Wang, Lu Chen, YuAn Wang 외

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs) inspired by the success of generalist models, such as large language…

Survey

OrigamiBench: An Interactive Environment to Synthesize Flat-Foldable Origamis

2026-03-14 · Naaisha Agarwal, Yihan Wu, Yichang Jian, Yikuan Hu 외 arxiv

Building AI systems that can plan, act, and create in the physical world requires more than pattern recognition. Such systems must understand the causal mechanisms and constraints governing physical processes in order to…

From 2D to 3D Cognition: A Brief Survey of General World Models

2025-06-25 · Ningwei Xie, Zizi Tian, Lei Yang, Xiao-Ping Zhang 외

World models have garnered increasing attention in the development of artificial general intelligence (AGI), serving as computational frameworks for learning representations of the external world and forecasting future s…

Autonomous DrivingScene GenerationSpatial ReasoningWorld Knowledge