paper-with-me

Papers

Benchmarking Multi-Image Understanding in Vision and Language Models: Perception, Knowledge, Reasoning, and Multi-Hop Reasoning

2024-06-18 · Bingchen Zhao, Yongshuo Zong, Letian Zhang, Timothy Hospedales

The advancement of large language models (LLMs) has significantly broadened the scope of applications in natural language processing, with multi-modal LLMs extending these capabilities to integrate and interpret visual data. However, existing benchmarks for visual language models (VLMs) predominantly focus on single-image inputs, neglecting the crucial aspect of multi-image understanding. In this paper, we introduce a Multi-Image Relational Benchmark MIRB, designed to evaluate VLMs' ability to compare, analyze, and reason across multiple images. Our benchmark encompasses four categories: perception, visual world knowledge, reasoning, and multi-hop reasoning. Through a comprehensive evaluation of a wide range of open-source and closed-source models, we demonstrate that while open-source VLMs were shown to approach the performance of GPT-4V in single-image tasks, a significant performance gap remains in multi-image reasoning tasks. Our findings also reveal that even the state-of-the-art GPT-4V model struggles with our benchmark, underscoring the need for further research and development in this area. We believe our contribution of MIRB could serve as a testbed for developing the next-generation multi-modal models.

📄 PDF Abstract BibTeX arXiv:2406.12742

Code (1)

dtennant/mirb_eval 공식 구현 pytorch

Tasks

BenchmarkingWorld Knowledge

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

UVLM: A Universal Vision-Language Model Loader for Reproducible Multimodal Benchmarking

2026-03-14 · Joan Perez, Giovanni Fusco arxiv

Vision-Language Models (VLMs) have emerged as powerful tools for image understanding tasks, yet their practical deployment remains hindered by significant architectural heterogeneity across model families. This paper int…

Prompt Engineering

Benchmarking Vision Language Models for Cultural Understanding

2024-07-15 · Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 외

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed…

BenchmarkingQuestion AnsweringScene UnderstandingVisual Question Answering

MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding

2024-10-25 · Fengbin Zhu, Ziyang Liu, Xiang Yao Ng, Haohui Wu 외

Large Vision-Language Models (LVLMs) have achieved remarkable performance in many vision-language tasks, yet their capabilities in fine-grained visual understanding remain insufficiently evaluated. Existing benchmarks ei…

Benchmarkingdocument understandingOptical Character Recognition (OCR)

TDBench: Benchmarking Vision-Language Models in Understanding Top-Down Images

2025-04-01 · Kaiyuan Hou, Minghui Zhao, Lilin Xu, Yuang Fan 외

The rapid emergence of Vision-Language Models (VLMs) has significantly advanced multimodal understanding, enabling applications in scene comprehension and visual reasoning. While these models have been primarily evaluate…

Autonomous NavigationBenchmarkingVisual Reasoning

YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models

2024-09-20 · Abhilash Nandy, Yash Agarwal, Ashish Patwa, Millon Madhur Das 외

Understanding satire and humor is a challenging task for even current Vision-Language models. In this paper, we propose the challenging tasks of Satirical Image Detection (detecting whether an image is satirical), Unders…

BenchmarkingImage Captioning