paper-with-me

홈 › Papers

Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams

2025-12-17 · Yiming Cui, Xin Yao, Yuxuan Qin, Xin Li, Shijin Wang, Guoping Hu arxiv

Multimodal scientific reasoning remains a significant challenge for large language models (LLMs), particularly in chemistry, where problem-solving relies on symbolic diagrams, molecular structures, and structured visual data. Here, we systematically evaluate 40 proprietary and open-source multimodal LLMs, including GPT-5, o3, Gemini-2.5-Pro, and Qwen2.5-VL, on a curated benchmark of Olympiad-style chemistry questions drawn from over two decades of U.S. National Chemistry Olympiad (USNCO) exams. These questions require integrated visual and textual reasoning across diverse modalities. We find that many models struggle with modality fusion, where in some cases, removing the image even improves accuracy, indicating misalignment in vision-language integration. Chain-of-Thought prompting consistently enhances both accuracy and visual grounding, as demonstrated through ablation studies and occlusion-based interpretability. Our results reveal critical limitations in the scientific reasoning abilities of current MLLMs, providing actionable strategies for developing more robust and interpretable multimodal systems in chemistry. This work provides a timely benchmark for measuring progress in domain-specific multimodal AI and underscores the need for further advances at the intersection of artificial intelligence and scientific reasoning.

📄 PDF Abstract BibTeX arXiv:2512.14989

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Grounding

Similar Papers 제목 키워드 기반

OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model

2026-04-22 · Qiguang Chen, Chengyu Luan, Jiajun Wu, Qiming Yu 외 arxiv

Large vision-language models (LVLMs) have made substantial advances in reasoning tasks at the Olympiad level. Nevertheless, current Olympiad-level multimodal reasoning benchmarks for these models often emphasize single-i…

Multimodal Reasoning

FrontierScience: Evaluating AI's Ability to Perform Expert-Level Scientific Tasks

2026-01-29 · Miles Wang, Robi Lin, Kat Hu, Joy Jiao 외 arxiv

We introduce FrontierScience, a benchmark evaluating expert-level scientific reasoning in frontier language models. Recent model progress has nearly saturated existing science benchmarks, which often rely on multiple-cho…

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

2024-02-21 · Chaoqun He, Renjie Luo, Yuzhuo Bai, Shengding Hu 외

Recent advancements have seen Large Language Models (LLMs) and Large Multimodal Models (LMMs) surpassing general human capabilities in various tasks, approaching the proficiency level of human experts across multiple dom…

Logical Fallacies

Analysis of Organic Chemistry Tests in 27th Chemistry Olympiad (Provincial Division)

2015-01-01 · Yu-Hao Deng

Based on the basic principle and mechanism of organic chemistry, the organic chemistry tests in 27th chemistry olympiad (Provincial Division) in were analyzed in detail for the reference of instructors and competitors.

ChemLabs on ChemO: A Multi-Agent System for Multimodal Reasoning on IChO 2025

2025-11-20 · Qiang Xu, Shengyuan Bai, Leqing Chen, Zijing Liu 외 arxiv

Olympiad-level benchmarks in mathematics and physics are crucial testbeds for advanced AI reasoning, but chemistry, with its unique multimodal symbolic language, has remained an open challenge. We introduce ChemO, a new …

Multimodal Reasoning