paper-with-me

Papers

BloomVQA: Assessing Hierarchical Multi-modal Comprehension

2023-12-20 · Yunye Gong, Robik Shrestha, Jared Claypoole, Michael Cogswell, Arijit Ray, Christopher Kanan, Ajay Divakaran

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus on fact-based memorization and simple reasoning tasks without theoretical grounding, we collect multiple-choice samples based on picture stories that reflect different levels of comprehension, as laid out in Bloom's Taxonomy, a classic framework for learning assessment widely adopted in education research. Our data maps to a novel hierarchical graph representation which enables automatic data augmentation and novel measures characterizing model consistency. We perform graded evaluation and reliability analysis on recent multi-modal models. In comparison to low-level tasks, we observe decreased performance on tasks requiring advanced comprehension and cognitive skills with up to 38.0\% drop in VQA accuracy. In comparison to earlier models, GPT-4V demonstrates improved accuracy over all comprehension levels and shows a tendency of bypassing visual inputs especially for higher-level tasks. Current models also show consistency patterns misaligned with human comprehension in various scenarios, demonstrating the need for improvement based on theoretically-grounded criteria.

📄 PDF Abstract BibTeX arXiv:2312.12716

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationMemorizationMultiple-choiceVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark

2024-08-14 · Minxuan Zhou, Hao Liang, Tianpeng Li, Zhiyu Wu 외

With the development of Multimodal Large Language Models (MLLMs), the evaluation of multimodal models in the context of mathematical problems has become a valuable research field. Multimodal visual-textual mathematical r…

MathMathematical Reasoning

H2VU-Benchmark: A Comprehensive Benchmark for Hierarchical Holistic Video Understanding

2025-03-31 · Qi Wu, Quanlong Zheng, Yanhao Zhang, Junlin Xie 외

With the rapid development of multimodal models, the demand for assessing video understanding capabilities has been steadily increasing. However, existing benchmarks for evaluating video understanding exhibit significant…

Video Understanding

SEED-Bench: Benchmarking Multimodal Large Language Models

2024-01-01 · CVPR 2024 1 · Bohao Li, Yuying Ge, Yixiao Ge, Guangzhi Wang 외

Multimodal large language models (MLLMs) building upon the foundation of powerful large language models (LLMs) have recently demonstrated exceptional capabilities in generating not only texts but also images given in…

BenchmarkingImage GenerationMultiple-choice

MM-BigBench: Evaluating Multimodal Models on Multimodal Content Comprehension Tasks

2023-10-13 · Xiaocui Yang, Wenfang Wu, Shi Feng, Ming Wang 외

The popularity of multimodal large language models (MLLMs) has triggered a recent surge in research efforts dedicated to evaluating these models. Nevertheless, existing evaluation studies of MLLMs primarily focus on the …

multimodal interactionMultimodal Reasoning

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

2025-01-02 · Yaxian Wang, Henghui Ding, Shuting He, Xudong Jiang 외

In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC ext…

Generalized Referring Expression ComprehensionGeneralized Referring Expression SegmentationObject CountingPhrase Grounding+3