paper-with-me

Papers

M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models

2024-11-06 · Chuhan Li, Ziyao Shangguan, Yilun Zhao, Deyuan Li, Yixin Liu, Arman Cohan

Existing benchmarks for evaluating foundation models mainly focus on single-document, text-only tasks. However, they often fail to fully capture the complexity of research workflows, which typically involve interpreting non-textual data and gathering information across multiple documents. To address this gap, we introduce M3SciQA, a multi-modal, multi-document scientific question answering benchmark designed for a more comprehensive evaluation of foundation models. M3SciQA consists of 1,452 expert-annotated questions spanning 70 natural language processing paper clusters, where each cluster represents a primary paper along with all its cited documents, mirroring the workflow of comprehending a single paper by requiring multi-modal and multi-document data. With M3SciQA, we conduct a comprehensive evaluation of 18 foundation models. Our results indicate that current foundation models still significantly underperform compared to human experts in multi-modal information retrieval and in reasoning across multiple scientific documents. Additionally, we explore the implications of these findings for the future advancement of applying foundation models in multi-modal scientific literature analysis.

📄 PDF Abstract BibTeX arXiv:2411.04075

Code (0)

등록된 구현이 없습니다.

Tasks

Information RetrievalQuestion Answering

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Automatic Inter-document Multi-hop Scientific QA Generation

2026-03-15 · Seungmin Lee, Dongha Kim, Yuni Jeon, Junyoung Koh 외 arxiv

Existing automatic scientific question generation studies mainly focus on single-document factoid QA, overlooking the inter-document reasoning crucial for scientific understanding. We present AIM-SciQA, an automated fram…

Machine Reading ComprehensionQuestion Generation

VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering

2025-11-25 · Yuyi Li, Daoyuan Chen, Zhen Wang, Yutong Lu 외 arxiv

Large Vision-Language Models (LVLMs) show promise for scientific applications, yet open-source models still struggle with Scientific Visual Question Answering (SVQA), namely answering questions about figures from scienti…

Visual Question Answering

SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

2024-05-16 · Yuwei Wan, Yixuan Liu, Aswathy Ajith, Clara Grazian 외

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA g…

Open-Ended Question AnsweringQuestion AnsweringScience Question Answering

DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering

2026-01-23 · Haotian Chen, Qingqing Long, Siyu Pu, Xiao Luo 외 arxiv

With the rapid growth of scientific literature, scientific question answering (SciQA) has become increasingly critical for exploring and utilizing scientific knowledge. Retrieval-Augmented Generation (RAG) enhances LLMs …

Question Answering

SciMDR: Advancing Scientific Multimodal Document Reasoning

2026-03-12 · Ziyu Chen, Yilun Zhao, Chengye Wang, Rilyn Han 외 arxiv

Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faithfulness, and realism. To address this challenge, we introduce the synthesize-a…