paper-with-me

홈 › Papers

Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

2026-04-30 · Junpeng Ding, Zichen Tang, Haihong E, Mengyuan Ji, Yang Liu, Haolin Tian, Haiyang Sun, Pengqi Sun, Yang Xu, Yichen Liu, Haocheng Gao, Zijie Xi, Ruomeng Jiang, Peizhi Zhao, Rongjin Li, Yuanze Li, Jiacheng Liu, Zhongjun Yang, Jintong Chen, Siying Lin arxiv

We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answering (QA) pairs derived from 1,084 expert-curated images. SPUR features three key innovations: (1) Panel-Level Fine-Grained Perception: evaluating the visual perception of multimodal large language models (MLLMs) across three dimensions (numerical, morphological, and information localization) on six fine-grained panel types; (2) Cross-Panel Relation Understanding: utilizing complex images with an average of 14.3 panels per sample to evaluate MLLMs' ability to decipher intricate cross-panel relations; (3) Expert-Level Reasoning: assessment of qualitative and quantitative reasoning across five experimental paradigms to determine if models can infer conclusions from evidence as human experts do. Comprehensive evaluation of 20 MLLMs and four multimodal Chain-of-Thought (MCoT) methods reveals that current models fall significantly short of the expert-level requirements for scientific image interpretation, underscoring a critical bottleneck in AI for Science (AI4S) research.

📄 PDF Abstract BibTeX arXiv:2604.27604

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

2026-07-17 · Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu 외 arxiv

We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and…

Multimodal ReasoningImage Generation

Spawrious: A Benchmark for Fine Control of Spurious Correlation Biases

2023-03-09 · Aengus Lynch, Gbètondji J-S Dovonon, Jean Kaddour, Ricardo Silva

The problem of spurious correlations (SCs) arises when a classifier relies on non-predictive features that happen to be correlated with the labels in the training data. For example, a classifier may misclassify dog breed…

Image Captioningimage-classificationImage Classification

Causal Decoding for Hallucination-Resistant Multimodal Large Language Models

2026-02-24 · Shiwei Tan, Hengyi Wang, Weiyi Qin, Qi Xu 외 arxiv

Multimodal Large Language Models (MLLMs) deliver detailed responses on vision-language tasks, yet remain susceptible to object hallucination (introducing objects not present in the image), undermining reliability in prac…

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Address Multimodal Hallucination

2025-04-14 · Hao Yin, Guangzong Si, Zilei Wang

Contrastive decoding strategies are widely used to reduce hallucinations in multimodal large language models (MLLMs). These methods work by constructing contrastive samples to induce hallucinations and then suppressing t…

Hallucination

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling

2025-08-30 · Shengyin Sun, Yiming Li, Xing Li, Yingzhao Lian 외 arxiv

Test-time scaling has emerged as a powerful paradigm for enhancing the reasoning capabilities of large language models (LLMs) by allocating additional computational resources during inference. However, this paradigm is i…