paper-with-me

Papers

SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation

2024-05-14 · Jonathan Roberts, Kai Han, Neil Houlsby, Samuel Albanie

Large multimodal models (LMMs) have proven flexible and generalisable across many tasks and fields. Although they have strong potential to aid scientific research, their capabilities in this domain are not well characterised. A key aspect of scientific research is the ability to understand and interpret figures, which serve as a rich, compressed source of complex information. In this work, we present SciFIBench, a scientific figure interpretation benchmark consisting of 2000 questions split between two tasks across 8 categories. The questions are curated from arXiv paper figures and captions, using adversarial filtering to find hard negatives and human verification for quality control. We evaluate 28 LMMs on SciFIBench, finding it to be a challenging benchmark. Finally, we investigate the alignment and reasoning faithfulness of the LMMs on augmented question sets from our benchmark. We release SciFIBench to encourage progress in this domain.

📄 PDF Abstract BibTeX arXiv:2405.08807

Code (1)

jonathan-roberts1/SciFIBench 공식 구현 pytorch

Tasks

BenchmarkingMultiple-choice

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

2024-03-01 · Lei LI, Yuqi Wang, Runxin Xu, Peiyi Wang 외

Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains l…

BenchmarkingMathematical ReasoningQuestion Answering

Every Part Matters: Integrity Verification of Scientific Figures Based on Multimodal Large Language Models

2024-07-26 · Xiang Shi, Jiawei Liu, Yinpeng Liu, Qikai Cheng 외

This paper tackles a key issue in the interpretation of scientific figures: the fine-grained alignment of text and figures. It advances beyond prior research that primarily dealt with straightforward, data-driven visuali…

Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy

2026-05-19 · Aritra Roy, Enrico Grisan, Chiara Gattinoni, John Buckeridge arxiv

Automated extraction of materials composition-property data from scientific literature has advanced considerably with the development of large language model-based pipelines; however, existing frameworks remain limited t…

MuSciClaims: Multimodal Scientific Claim Verification

2025-06-05 · Yash Kumar Lal, Manikanta Bandham, Mohammad Saqib Hasan, Apoorva Kashi 외

Assessing scientific claims requires identifying, extracting, and reasoning with multimodal data expressed in information-rich figures in scientific literature. Despite the large body of work in scientific QA, figure cap…

ArticlesClaim VerificationDiagnosticMultimodal Reasoning

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

2026-08-14 · Jennifer D'Souza, Fahad Ahmed, Cecilia Andrea Bustamante Andrade, Lina Frolova 외 arxiv

Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competit…

Visual Question AnsweringInformation Extraction