paper-with-me

Papers

DeepPerception: Advancing R1-like Cognitive Visual Perception in MLLMs for Knowledge-Intensive Visual Grounding

2025-03-17 · Xinyu Ma, Ziyang Ding, Zhicong Luo, Chi Chen, Zonghao Guo, Derek F. Wong, Xiaoyi Feng, Maosong Sun

Human experts excel at fine-grained visual discrimination by leveraging domain knowledge to refine perceptual features, a capability that remains underdeveloped in current Multimodal Large Language Models (MLLMs). Despite possessing vast expert-level knowledge, MLLMs struggle to integrate reasoning into visual perception, often generating direct responses without deeper analysis. To bridge this gap, we introduce knowledge-intensive visual grounding (KVG), a novel visual grounding task that requires both fine-grained perception and domain-specific knowledge integration. To address the challenges of KVG, we propose DeepPerception, an MLLM enhanced with cognitive visual perception capabilities. Our approach consists of (1) an automated data synthesis pipeline that generates high-quality, knowledge-aligned training samples, and (2) a two-stage training framework combining supervised fine-tuning for cognitive reasoning scaffolding and reinforcement learning to optimize perception-cognition synergy. To benchmark performance, we introduce KVG-Bench a comprehensive dataset spanning 10 domains with 1.3K manually curated test cases. Experimental results demonstrate that DeepPerception significantly outperforms direct fine-tuning, achieving +8.08\% accuracy improvements on KVG-Bench and exhibiting +4.60\% superior cross-domain generalization over baseline approaches. Our findings highlight the importance of integrating cognitive processes into MLLMs for human-like visual perception and open new directions for multimodal reasoning research. The data, codes, and models are released at https://github.com/thunlp/DeepPerception.

📄 PDF Abstract BibTeX arXiv:2503.12797

Code (1)

thunlp/deepperception 공식 구현 pytorch

Tasks

Domain GeneralizationMultimodal ReasoningVisual Grounding

Similar Papers 제목 키워드 기반

Multimodal LLM Augmented Reasoning for Interpretable Visual Perception Analysis

2025-04-16 · Shravan Chaudhari, Trilokya Akula, Yoon Kim, Tom Blake

In this paper, we advance the study of AI-augmented reasoning in the context of Human-Computer Interaction (HCI), psychology and cognitive science, focusing on the critical task of visual perception. Specifically, we inv…

Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language

2025-05-28 · Guangfu Hao, Haojie Wen, Liangxuna Guo, Yang Chen 외

Flexible tool selection reflects a complex cognitive ability that distinguishes humans from other species, yet computational models that capture this ability remain underdeveloped. We developed a framework using low-dime…

Attribute

SPHINX: A Synthetic Environment for Visual Perception and Reasoning

2025-11-25 · Md Tanvirul Alam, Saksham Aggarwal, Justin Yang Chae, Nidhi Rastogi arxiv

We present Sphinx, a synthetic environment for visual perception and reasoning that targets core cognitive primitives. Sphinx procedurally generates puzzles using motifs, tiles, charts, icons, and geometric primitives, e…

Reinforcement LearningMultimodal ReasoningSymmetry DetectionSpatial Reasoning

A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs

2025-01-23 · Mohit Vaishnav, Tanel Tammet

A fundamental challenge in artificial intelligence involves understanding the cognitive mechanisms underlying visual reasoning in sophisticated models like Vision-Language Models (VLMs). How do these models integrate vis…

DescriptiveDiagnosticFew-Shot Image ClassificationVisual Reasoning

From 2D to 3D Cognition: A Brief Survey of General World Models

2025-06-25 · Ningwei Xie, Zizi Tian, Lei Yang, Xiao-Ping Zhang 외

World models have garnered increasing attention in the development of artificial general intelligence (AGI), serving as computational frameworks for learning representations of the external world and forecasting future s…

Autonomous DrivingScene GenerationSpatial ReasoningWorld Knowledge