paper-with-me

Papers

MMSciBench: Benchmarking Language Models on Multimodal Scientific Problems

2025-02-27 · Xinwu Ye, Chengfan Li, Siming Chen, Xiangru Tang, Wei Wei

Recent advances in large language models (LLMs) and vision-language models (LVLMs) have shown promise across many tasks, yet their scientific reasoning capabilities remain untested, particularly in multimodal settings. We present MMSciBench, a benchmark for evaluating mathematical and physical reasoning through text-only and text-image formats, with human-annotated difficulty levels, solutions with detailed explanations, and taxonomic mappings. Evaluation of state-of-the-art models reveals significant limitations, with even the best model achieving only \textbf{63.77\%} accuracy and particularly struggling with visual reasoning tasks. Our analysis exposes critical gaps in complex reasoning and visual-textual integration, establishing MMSciBench as a rigorous standard for measuring progress in multimodal scientific understanding. The code for MMSciBench is open-sourced at GitHub, and the dataset is available at Hugging Face.

📄 PDF Abstract BibTeX arXiv:2503.01891

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingVisual Reasoning

Similar Papers 제목 키워드 기반

SAIBench: Benchmarking AI for Science

2022-06-11 · Yatao Li, Jianfeng Zhan

Scientific research communities are embracing AI-based solutions to target tractable scientific tasks and improve research workflows. However, the development and evaluation of such solutions are scattered across multipl…

BenchmarkingFriction

MicroVQA: A Multimodal Reasoning Benchmark for Microscopy-Based Scientific Research

2025-03-17 · CVPR 2025 1 · James Burgess, Jeffrey J Nirschl, Laura Bravo-Sánchez, Alejandro Lozano 외

Scientific research demands sophisticated reasoning over multimodal data, a challenge especially prevalent in biology. Despite recent advances in multimodal large language models (MLLMs) for AI-assisted research, existin…

ArticlesBenchmarkingMultimodal ReasoningMultiple-choice+4

Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

2024-03-01 · Lei LI, Yuqi Wang, Runxin Xu, Peiyi Wang 외

Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains l…

BenchmarkingMathematical ReasoningQuestion Answering

Does Table Source Matter? Benchmarking and Improving Multimodal Scientific Table Understanding and Reasoning

2025-01-22 · Bohao Yang, Yingji Zhang, Dong Liu, André Freitas 외

Recent large language models (LLMs) have advanced table understanding capabilities but rely on converting tables into text sequences. While multimodal large language models (MLLMs) enable direct visual processing, they f…

Benchmarking

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

2026-07-16 · Patrick Phuoc Do, Chau M. Ta, Chaoli Wang arxiv

Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-centric and provide limited evidence of understanding of scientific visualization (…