paper-with-me

홈 › Papers

SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark

2024-02-06 · Zhenwen Liang, Kehan Guo, Gang Liu, Taicheng Guo, Yujun Zhou, Tianyu Yang, Jiajun Jiao, Renjie Pi, Jipeng Zhang, Xiangliang Zhang

The paper introduces SceMQA, a novel benchmark for scientific multimodal question answering at the college entrance level. It addresses a critical educational phase often overlooked in existing benchmarks, spanning high school to pre-college levels. SceMQA focuses on core science subjects including Mathematics, Physics, Chemistry, and Biology. It features a blend of multiple-choice and free-response formats, ensuring a comprehensive evaluation of AI models' abilities. Additionally, our benchmark provides specific knowledge points for each problem and detailed explanations for each answer. SceMQA also uniquely presents problems with identical contexts but varied questions to facilitate a more thorough and accurate assessment of reasoning capabilities. In the experiment, we evaluate both open-source and close-source state-of-the-art Multimodal Large Language Models (MLLMs), across various experimental settings. The results show that further research and development are needed in developing more capable MLLM, as highlighted by only 50% to 60% accuracy achieved by the strongest models. Our benchmark and analysis will be available at https://scemqa.github.io/

📄 PDF Abstract BibTeX arXiv:2402.05138

Code (0)

등록된 구현이 없습니다.

Tasks

Multiple-choiceQuestion Answering

Similar Papers 제목 키워드 기반

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

2024-02-21 · Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 외

Recent advancements have seen Large Language Models (LLMs) and Large Multimodal Models (LMMs) surpassing general human capabilities in various tasks, approaching the proficiency level of human experts across multiple dom…

Logical Fallacies

Leveraging Diverse Lexical Chains to Construct Essays for Chinese College Entrance Examination

2017-11-01 · IJCNLP 2017 11 · Liunian Li, Xiaojun Wan, Jin-Ge Yao, Siming Yan

In this work we study the challenging task of automatically constructing essays for Chinese college entrance examination where the topic is specified in advance. We explore a sentence extraction framework based on divers…

Sentence

Entrance seat allotment system project report.

2025-07-02 · Zenodo 2025 7 · Kamal Acharya

This project Entrance Seat Allotment System is windows application in which students can register with their rank number for the entrance examination and the administrator can allot the seats for the students. Admini…

GAOKAO-MM: A Chinese Human-Level Benchmark for Multimodal Models Evaluation

2024-02-24 · Yi Zong, Xipeng Qiu

The Large Vision-Language Models (LVLMs) have demonstrated great abilities in image perception and language understanding. However, existing multimodal benchmarks focus on primary perception abilities and commonsense kno…

AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

2023-04-13 · Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang 외

Evaluating the general abilities of foundation models to tackle human-level tasks is a vital aspect of their development and application in the pursuit of Artificial General Intelligence (AGI). Traditional benchmarks, wh…

Decision MakingMath