paper-with-me

Papers

NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

2024-07-15 · Pranshu Pandya, Vatsal Gupta, Agney S Talwarr, Tushar Kataria, Dan Roth, Vivek Gupta

Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. Due to extensive training on vast amounts of human-curated data, LLMs and VLMs excel in common-sense reasoning tasks, however still struggle with more complex reasoning that demands deeper cognitive understanding. We introduce NTSEBench, a new dataset designed to evaluate cognitive multi-modal reasoning and problem-solving skills of large models. The dataset contains 2728 multiple-choice questions, accompanied by a total of 4,642 images, categorized into 26 different types. These questions are drawn from the nationwide NTSE examination in India and feature a mix of visual and textual general aptitude challenges, designed to assess intelligence and critical thinking skills beyond mere rote learning. We establish baselines on the dataset using state-of-the-art LLMs and VLMs. To facilitate a comparison between open source and propriety models, we propose four distinct modeling strategies to handle different modalities -- text and images -- in the dataset instances.

📄 PDF Abstract BibTeX arXiv:2407.10380

Code (0)

등록된 구현이 없습니다.

Tasks

Common Sense ReasoningMultiple-choiceVisual Reasoning

Similar Papers 제목 키워드 기반

A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models

2024-02-28 · Xiujie Song, Mengyue Wu, Kenny Q. Zhu, Chunhao Zhang 외

Large Vision-Language Models (LVLMs), despite their recent success, are hardly comprehensively tested for their cognitive abilities. Inspired by the prevalent use of the "Cookie Theft" task in human cognition test, we pr…

Image DescriptionQuestion AnsweringVisual Question Answering

Toward Cognitive Supersensing in Multimodal Large Language Model

2026-02-02 · Boyi Li, Yifan Shen, Yuanzhe Liu, Yifan Xu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abs…

Visual Question AnsweringReinforcement LearningVisual Reasoning

MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity

2025-11-05 · Kaiyuan Zhang, Chenghao Yang, Zhoufutu Wen, Sihang Yuan 외 arxiv

As reasoning models scale rapidly, the essential role of multimodality in human cognition has come into sharp relief, driving a growing need to probe vision-centric cognitive behaviors. Yet, existing multimodal benchmark…

OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models

2025-06-03 · Mengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 외

Spatial reasoning is a key aspect of cognitive psychology and remains a major bottleneck for current vision-language models (VLMs). While extensive research has aimed to evaluate or improve VLMs' understanding of basic s…

Object CountingSpatial Reasoning

Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods

2025-11-14 · Weichen Liu, Qiyao Xue, Haoming Wang, Xiangyu Yin 외 arxiv

Spatial reasoning, which requires ability to perceive and manipulate spatial relationships in the 3D world, is a fundamental aspect of human intelligence, yet remains a persistent challenge for Multimodal large language …

Spatial Reasoning