paper-with-me

홈 › Papers

ArxEval: Evaluating Retrieval and Generation in Language Models for Scientific Literature

2025-01-17 · Aarush Sinha, Viraj Virk, Dipshikha Chakraborty, P. S. Sreeja

Language Models [LMs] are now playing an increasingly large role in information generation and synthesis; the representation of scientific knowledge in these systems needs to be highly accurate. A prime challenge is hallucination; that is, generating apparently plausible but actually false information, including invented citations and nonexistent research papers. This kind of inaccuracy is dangerous in all the domains that require high levels of factual correctness, such as academia and education. This work presents a pipeline for evaluating the frequency with which language models hallucinate in generating responses in the scientific literature. We propose ArxEval, an evaluation pipeline with two tasks using ArXiv as a repository: Jumbled Titles and Mixed Titles. Our evaluation includes fifteen widely used language models and provides comparative insights into their reliability in handling scientific literature.

📄 PDF Abstract BibTeX arXiv:2501.10483

Code (0)

등록된 구현이 없습니다.

Tasks

HallucinationRetrieval

Similar Papers 제목 키워드 기반

Enhancing Scientific Literature Chatbots with Retrieval-Augmented Generation: A Performance Evaluation of Vector and Graph-Based Systems

2026-02-19 · Hamideh Ghanadian, Amin Kamali, Mohammad Hossein Tekieh arxiv

This paper investigates the enhancement of scientific literature chatbots through retrieval-augmented generation (RAG), with a focus on evaluating vector- and graph-based retrieval systems. The proposed chatbot leverages…

Decision Making

LLM-based Corroborating and Refuting Evidence Retrieval for Scientific Claim Verification

2025-03-11 · Siyuan Wang, James R. Foulds, Md Osman Gani, SHimei Pan

In this paper, we introduce CIBER (Claim Investigation Based on Evidence Retrieval), an extension of the Retrieval-Augmented Generation (RAG) framework designed to identify corroborating and refuting documents as evidenc…

Claim VerificationRAGRetrievalRetrieval-augmented Generation

Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models

2025-09-12 · Ozan Gokdemir, Neil Getty, Robert Underwood, Sandeep Madireddy 외 arxiv

As scientific knowledge grows at an unprecedented pace, evaluation benchmarks must evolve to reflect new discoveries and ensure language models are tested on current, diverse literature. We propose a scalable, modular fr…

Question GenerationDomain Adaptation

GraphMind: Interactive Novelty Assessment System for Accelerating Scientific Discovery

2025-10-17 · Italo Luis da Silva, Hanqi Yan, Lin Gui, Yulan He arxiv

Large Language Models (LLMs) show strong reasoning and text generation capabilities, prompting their use in scientific literature analysis, including novelty assessment. While evaluating novelty of scientific papers is c…

Information RetrievalText Generation

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

2026-05-28 · A. J. Lew, Y. Cao, M. J. Buehler arxiv

Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance …

Semantic Similarity