paper-with-me

Papers

Toward Cognitive Supersensing in Multimodal Large Language Model

2026-02-02 · Boyi Li, Yifan Shen, Yuanzhe Liu, Yifan Xu, Jiateng Liu, Xinzhuo Li, Zhengyuan Li, Jingyuan Zhu, Yunhan Zhong, Fangzhou Lan, Jianguo Cao, James M. Rehg, Heng Ji, Ismini Lourentzou, Xu Cao arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abstract and require visual memory. Current approaches primarily scale Chain-of-Thought (CoT) reasoning in the text space, even when language alone is insufficient for clear and structured reasoning, and largely neglect visual reasoning mechanisms analogous to the human visuospatial sketchpad and visual imagery. To mitigate this deficiency, we introduce Cognitive Supersensing, a novel training paradigm that endows MLLMs with human-like visual imagery capabilities by integrating a Latent Visual Imagery Prediction (LVIP) head that jointly learns sequences of visual cognitive latent embeddings and aligns them with the answer, thereby forming vision-based internal reasoning chains. We further introduce a reinforcement learning stage that optimizes text reasoning paths based on this grounded visual latent. To evaluate the cognitive capabilities of MLLMs, we present CogSense-Bench, a comprehensive visual question answering (VQA) benchmark assessing five cognitive dimensions. Extensive experiments demonstrate that MLLMs trained with Cognitive Supersensing significantly outperform state-of-the-art baselines on CogSense-Bench and exhibit superior generalization on out-of-domain mathematics and science VQA benchmarks, suggesting that internal visual imagery is potentially key to bridging the gap between perceptual recognition and cognitive understanding. We will open-source the CogSense-Bench and our model weights.

📄 PDF Abstract BibTeX arXiv:2602.01541

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question AnsweringReinforcement LearningVisual Reasoning

Similar Papers 제목 키워드 기반

Cambrian-S: Towards Spatial Supersensing in Video

2025-11-06 · Shusheng Yang, Jihan Yang, Pinzhi Huang, Ellis Brown 외 arxiv

We argue that progress in true multimodal intelligence calls for a shift from reactive, task-driven systems and brute-force long context towards a broader paradigm of supersensing. We frame spatial supersensing as four s…

Event Segmentation

Solving Spatial Supersensing Without Spatial Supersensing

2025-11-20 · Vishaal Udandarao, Shyamgopal Karthik, Surabhi S. Nath, Andreas Hochlehnert 외 arxiv

Cambrian-S aims to take the first steps towards improving video world models with spatial supersensing by introducing (i) two benchmarks, VSI-Super-Recall (VSR) and VSI-Super-Counting (VSC), and (ii) bespoke predictive s…

PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World

2026-05-13 · Changpeng Wang, Xin Lin, Junhan Liu, Yuheng Liu 외 arxiv

Multimodal large laboratory models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human-like perception. For navigation, roboti…

Scene UnderstandingSpatial Reasoning

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding

2026-03-19 · Yinghui Li, Jiayi Kuang, Peng Xing, Daixian Liu 외 arxiv

Multimodal large language models (MLLMs) perform strongly on natural images, yet their ability to understand discrete visual symbols remains unclear. We present a multi-domain benchmark spanning language, culture, mathem…

Visual Grounding

Cognitive Insights Across Languages: Enhancing Multimodal Interview Analysis

2024-06-11 · David Ortiz-Perez, Jose Garcia-Rodriguez, David Tomás

Cognitive decline is a natural process that occurs as individuals age. Early diagnosis of anomalous decline is crucial for initiating professional treatment that can enhance the quality of life of those affected. To addr…