paper-with-me

Papers

VisFactor: Benchmarking Fundamental Visual Cognition in Multimodal Large Language Models

2025-02-23 · Jen-tse Huang, Dasen Dai, Jen-Yuan Huang, Youliang Yuan, Xiaoyuan Liu, Wenxuan Wang, Wenxiang Jiao, Pinjia He, Zhaopeng Tu

Multimodal Large Language Models (MLLMs) have demonstrated remarkable advancements in multimodal understanding; however, their fundamental visual cognitive abilities remain largely underexplored. To bridge this gap, we introduce VisFactor, a novel benchmark derived from the Factor-Referenced Cognitive Test (FRCT), a well-established psychometric assessment of human cognition. VisFactor digitalizes vision-related FRCT subtests to systematically evaluate MLLMs across essential visual cognitive tasks including spatial reasoning, perceptual speed, and pattern recognition. We present a comprehensive evaluation of state-of-the-art MLLMs, such as GPT-4o, Gemini-Pro, and Qwen-VL, using VisFactor under diverse prompting strategies like Chain-of-Thought and Multi-Agent Debate. Our findings reveal a concerning deficiency in current MLLMs' fundamental visual cognition, with performance frequently approaching random guessing and showing only marginal improvements even with advanced prompting techniques. These results underscore the critical need for focused research to enhance the core visual reasoning capabilities of MLLMs. To foster further investigation in this area, we release our VisFactor benchmark at https://github.com/CUHK-ARISE/VisFactor.

📄 PDF Abstract BibTeX arXiv:2502.16435

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingSpatial ReasoningVisual Reasoning

Similar Papers 제목 키워드 기반

Benchmarking Multimodal Sentiment Analysis

2017-07-29 · Erik Cambria, Devamanyu Hazarika, Soujanya Poria, Amir Hussain 외

We propose a framework for multimodal sentiment analysis and emotion recognition using convolutional neural network-based feature extraction from text and visual modalities. We obtain a performance improvement of 10% ove…

BenchmarkingEmotion RecognitionMultimodal Sentiment AnalysisSentiment Analysis

Benchmarking Large Vision-Language Models on CFMME: A Comprehensive Chinese Financial Multimodal Evaluation Dataset

2026-05-28 · Qian Chen, Xianyin Zhang, Yanzhi Liu, Lifan Guo 외 arxiv

The emergence of Large Vision-Language Models (LVLMs) has substantially expanded model capabilities beyond text-only understanding, enabling unified inference across both visual and textual modalities and supporting a br…

Information ExtractionQuestion Answering

Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

2025-04-24 · Zhikai Wang, Jiashuo Sun, Wenqi Zhang, Zhiqiang Hu 외

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, cap…

BenchmarkingMathMathematical ReasoningObject Recognition+2

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning

2026-04-04 · Zeyu Wang, Jingye Xu, Xiaogang Li, Peiyao Xiao 외 arxiv

Current multimodal benchmarks for scientific reasoning primarily evaluate local information extraction -- models recognize symbols and values and then perform textual inference. They do not assess whether models can reas…

Information ExtractionMultimodal Reasoning

Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline

2024-08-06 · Tianqi Wei, Zhi Chen, Zi Huang, Xin Yu

Existing plant disease classification models have achieved remarkable performance in recognizing in-laboratory diseased images. However, their performance often significantly degrades in classifying in-the-wild images. F…

Benchmarking