paper-with-me

Papers

SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

2024-05-16 · Yuwei Wan, Yixuan Liu, Aswathy Ajith, Clara Grazian, Bram Hoex, Wenjie Zhang, Chunyu Kit, Tong Xie, Ian Foster

We introduce SciQAG, a novel framework for automatically generating high-quality science question-answer pairs from a large corpus of scientific literature based on large language models (LLMs). SciQAG consists of a QA generator and a QA evaluator, which work together to extract diverse and research-level questions and answers from scientific papers. Utilizing this framework, we construct a large-scale, high-quality, open-ended science QA dataset containing 188,042 QA pairs extracted from 22,743 scientific papers across 24 scientific domains. We also introduce SciQAG-24D, a new benchmark task designed to evaluate the science question-answering ability of LLMs. Extensive experiments demonstrate that fine-tuning LLMs on the SciQAG dataset significantly improves their performance on both open-ended question answering and scientific tasks. To foster research and collaboration, we make the datasets, models, and evaluation codes publicly available, contributing to the advancement of science question answering and developing more interpretable and reasoning-capable AI systems.

📄 PDF Abstract BibTeX arXiv:2405.09939

Code (1)

masterai-eam/sciqag

Tasks

Open-Ended Question AnsweringQuestion AnsweringScience Question Answering

Similar Papers 제목 키워드 기반

Enhancing Student Learning with LLM-Generated Retrieval Practice Questions: An Empirical Study in Data Science Courses

2025-07-08 · Yuan An, John Liu, Niyam Acharya, Ruhma Hashmi arxiv

Retrieval practice is a well-established pedagogical technique known to significantly enhance student learning and knowledge retention. However, generating high-quality retrieval practice questions is often time-consumin…

Question-Driven Summarization of Answers to Consumer Health Questions

2020-05-18 · Max Savery, Asma Ben Abacha, Soumya Gayen, Dina Demner-Fushman

Automatic summarization of natural language is a widely studied area in computer science, one that is broadly applicable to anyone who routinely needs to understand large quantities of information. For example, in the me…

Medical Question AnsweringQuestion Answering

Automatic question generation for propositional logical equivalences

2024-05-09 · Yicheng Yang, Xinyu Wang, Haoming Yu, Zhiyuan Li

The increase in academic dishonesty cases among college students has raised concern, particularly due to the shift towards online learning caused by the pandemic. We aim to develop and implement a method capable of gener…

AttributeQuestion GenerationQuestion-Generation

A Song of Ice and Fire: Analyzing Textual Autotelic Agents in ScienceWorld

2023-02-10 · Laetitia Teodorescu, Xingdi Yuan, Marc-Alexandre Côté, Pierre-Yves Oudeyer

Building open-ended agents that can autonomously discover a diversity of behaviours is one of the long-standing goals of artificial intelligence. This challenge can be studied in the framework of autotelic RL agents, i.e…

Diversity

Exploiting Reasoning Chains for Multi-hop Science Question Answering

2021-09-07 · Findings (EMNLP) 2021 11 · Weiwen Xu, Yang Deng, Huihui Zhang, Deng Cai 외

We propose a novel Chain Guided Retriever-reader ({\tt CGR}) framework to model the reasoning chain for multi-hop Science Question Answering. Our framework is capable of performing explainable reasoning without the need …

Abstract Meaning RepresentationARCQuestion AnsweringScience Question Answering