paper-with-me

홈 › Papers

Are MLMs Trapped in the Visual Room?

2025-05-29 · Yazhou Zhang, Chunwang Zou, Qimeng Liu, Lu Rong, Ben Yao, Zheng Lian, Qiuchi Li, Peng Zhang, Jing Qin

Can multi-modal large models (MLMs) that can `see'' an image be said to `understand'' it? Drawing inspiration from Searle's Chinese Room, we propose the \textbf{Visual Room} argument: a system may process and describe every detail of visual inputs by following algorithmic rules, without genuinely comprehending the underlying intention. This dilemma challenges the prevailing assumption that perceptual mastery implies genuine understanding. In implementation, we introduce a two-tier evaluation framework spanning perception and cognition. The perception component evaluates whether MLMs can accurately capture the surface-level details of visual contents, where the cognitive component examines their ability to infer sarcasm polarity. To support this framework, We further introduce a high-quality multi-modal sarcasm dataset comprising both 924 static images and 100 dynamic videos. All sarcasm labels are annotated by the original authors and verified by independent reviewers to ensure clarity and consistency. We evaluate eight state-of-the-art (SoTA) MLMs. Our results highlight three key findings: (1) MLMs perform well on perception tasks; (2) even with correct perception, models exhibit an average error rate of ~16.1\% in sarcasm understanding, revealing a significant gap between seeing and understanding; (3) error analysis attributes this gap to deficiencies in emotional reasoning, commonsense inference, and context alignment. This work provides empirical grounding for the proposed Visual Room argument and offers a new evaluation paradigm for MLMs.

📄 PDF Abstract BibTeX arXiv:2505.23272

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Simple Aerial Detection Baseline of Multimodal Language Models

2025-01-16 · Qingyun Li, Yushi Chen, Xinya Shu, Dong Chen 외

The multimodal language models (MLMs) based on generative pre-trained Transformer are considered powerful candidates for unifying various domains and tasks. MLMs developed for remote sensing (RS) have demonstrated outsta…

object-detectionObject DetectionQuestion AnsweringVisual Grounding+1

VisLingInstruct: Elevating Zero-Shot Learning in Multi-Modal Language Models with Autonomous Instruction Optimization

2024-02-12 · Dongsheng Zhu, Xunzhu Tang, Weidong Han, Jinghui Lu 외

This paper presents VisLingInstruct, a novel approach to advancing Multi-Modal Language Models (MMLMs) in zero-shot learning. Current MMLMs show impressive zero-shot abilities in multi-modal tasks, but their performance …

In-Context LearningTextVQAZero-Shot Learning

Template Matters: Understanding the Role of Instruction Templates in Multimodal Language Model Evaluation and Training

2024-12-11 · Shijian Wang, Linxin Song, Jieyu Zhang, Ryotaro Shimizu 외

Current multimodal language models (MLMs) evaluation and training approaches overlook the influence of instruction format, presenting an elephant-in-the-room problem. Previous research deals with this problem by manually…

Language Model EvaluationLanguage ModelingLanguage Modelling

Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

2024-12-04 · CVPR 2025 1 · Mahtab Bigverdi, Zelun Luo, Cheng-Yu Hsieh, Ethan Shen 외

Multimodal language models (MLMs) still face challenges in fundamental visual perception tasks where specialized models excel. Tasks requiring reasoning about 3D structures benefit from depth estimation, and reasoning ab…

Depth Estimationobject-detectionObject DetectionVisual Reasoning

A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering

2023-11-13 · Yunxin Li, Longyue Wang, Baotian Hu, Xinyu Chen 외

The emergence of multimodal large models (MLMs) has significantly advanced the field of visual understanding, offering remarkable capabilities in the realm of visual question answering (VQA). Yet, the true challenge lies…

Decision MakingExplanation GenerationGeneral KnowledgeQuestion Answering+4