paper-with-me

홈 › Papers

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models

2025-08-19 · Jiacheng Ruan, Dan Jiang, Xian Gao, Ting Liu, Yuzhuo Fu, Yangyang Kang arxiv

Recently, multimodal large language models (MLLMs) have achieved significant advancements across various domains, and corresponding evaluation benchmarks have been continuously refined and improved. In this process, benchmarks in the scientific domain have played an important role in assessing the reasoning capabilities of MLLMs. However, existing benchmarks still face three key challenges: 1) Insufficient evaluation of models' reasoning abilities in multilingual scenarios; 2) Inadequate assessment of MLLMs' comprehensive modality coverage; 3) Lack of fine-grained annotation of scientific knowledge points. To address these gaps, we propose MME-SCI, a comprehensive and challenging benchmark. We carefully collected 1,019 high-quality question-answer pairs, which involve 3 distinct evaluation modes. These pairs cover four subjects, namely mathematics, physics, chemistry, and biology, and support five languages: Chinese, English, French, Spanish, and Japanese. We conducted extensive experiments on 16 open-source models and 4 closed-source models, and the results demonstrate that MME-SCI is widely challenging for existing MLLMs. For instance, under the Image-only evaluation mode, o4-mini achieved accuracy of only 52.11%, 24.73%, 36.57%, and 29.80% in mathematics, physics, chemistry, and biology, respectively, indicating a significantly higher difficulty level compared to existing benchmarks. More importantly, using MME-SCI's multilingual and fine-grained knowledge attributes, we analyzed existing models' performance in depth and identified their weaknesses in specific domains. The Data and Evaluation Code are available at https://github.com/JCruan519/MME-SCI.

📄 PDF Abstract BibTeX arXiv:2508.13938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation

2025-08-19 · Xian Gao, Jiacheng Ruan, Zongyun Zhang, Jingsheng Gao 외 arxiv

With the rapid growth of academic publications, peer review has become an essential yet time-consuming responsibility within the research community. Large Language Models (LLMs) have increasingly been adopted to assist i…

M4U: Evaluating Multilingual Understanding and Reasoning for Large Multimodal Models

2024-05-24 · Hongyu Wang, Jiayu Xu, Senwei Xie, Ruiping Wang 외

Multilingual multimodal reasoning is a core component in achieving human-level intelligence. However, most existing benchmarks for multilingual multimodal reasoning struggle to differentiate between models of varying per…

Multimodal Reasoning

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

2025-06-01 · Yufei Zhan, Ziheng Wu, Yousong Zhu, Rongkun Xue 외

Despite notable advancements in multimodal reasoning, leading Multimodal Large Language Models (MLLMs) still underperform on vision-centric multimodal reasoning tasks in general scenarios. This shortfall stems from their…

4kMathMathematical ReasoningMultimodal Reasoning

From UAV Imagery to Agronomic Reasoning: A Multimodal LLM Benchmark for Plant Phenotyping

2026-04-10 · Yu Wu, Guangzeng Han, Ibra Niang Niang, Francia Ravelombola 외 arxiv

To improve crop genetics, high-throughput, effective and comprehensive phenotyping is a critical prerequisite. While such tasks were traditionally performed manually, recent advances in multimodal foundation models, espe…

Multimodal Reasoning

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing

2026-04-10 · Aoran Xiao, Shihao Cheng, Yonghao Xu, Yexian Ren 외 arxiv

Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and remote sensing (RS) remains constrained by distinctive challenges: wide…