paper-with-me

홈 › Papers

DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis

2025-08-27 · Liana Patel, Negar Arabzadeh, Harshit Gupta, Ankita Sundar, Ion Stoica, Matei Zaharia, Carlos Guestrin arxiv

The ability to research and synthesize knowledge is central to human expertise and progress. A new class of AI systems--designed for generative research synthesis--aims to automate this process by retrieving information from the live web and producing long-form, cited reports. Yet, evaluating such systems remains an open challenge: existing question-answering benchmarks focus on short, factual answers, while expert-curated datasets risk staleness and data contamination. Neither captures the complexity and evolving nature of real research synthesis tasks. We introduce DeepScholar-bench, a live benchmark and automated evaluation framework for generative research synthesis. DeepScholar-bench draws queries and human-written exemplars from recent, high-quality ArXiv papers and evaluates a real synthesis task: generating a related work section by retrieving, synthesizing, and citing prior work. Our automated framework holistically measures performance across three key dimensions--knowledge synthesis, retrieval quality, and verifiability. To further future work, we also contribute DeepScholar-ref, a simple, open-source reference pipeline, which is implemented on the LOTUS framework and provides a strong baseline. Using DeepScholar-bench, we systematically evaluate prior open-source systems, search agents with strong models, OpenAI's DeepResearch, and DeepScholar-ref. We find DeepScholar-bench is far from saturated: no system surpasses a geometric mean of $31\%$ across all metrics. These results highlight both the difficulty and importance of DeepScholar-bench as a foundation for advancing AI systems capable of generative research synthesis. We make our benchmark code and data available at https://github.com/guestrin-lab/deepscholar-bench.

📄 PDF Abstract BibTeX arXiv:2508.20033

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models

2026-07-02 · Yuanzhi Liu, Shousheng Zhao, Bo Zhou, Kongming Liang 외 arxiv

Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vulnerable to temporal staleness, data contamination, and costly maintenance. We pre…

LiveSecBench: A Dynamic and Event-Driven Safety Benchmark for Chinese Language Model Applications

2025-11-04 · Yudong Li, Peiru Yang, Feng Huang, Zhongliang Yang 외 arxiv

We introduce LiveSecBench, a continuously updated safety benchmark specifically for Chinese-language LLM application scenarios. LiveSecBench constructs a high-quality and unique dataset through a pipeline that combines a…

A Semi-Automated Live Interlingual Communication Workflow Featuring Intralingual Respeaking: Evaluation and Benchmarking

2022-06-01 · LREC 2022 6 · Tomasz Korybski, Elena Davitti, Constantin Orasan, Sabine Braun

In this paper, we present a semi-automated workflow for live interlingual speech-to-text communication which seeks to reduce the shortcomings of existing ASR systems: a human respeaker works with a speaker-dependent spee…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingMachine Translation+4

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

2026-04-08 · Igor Ivanov, David Demitri Africa arxiv

Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the validity of safety and alignment benchmarks. We propose LURE (Live-Usa…

LiveSearchBench: An Automatically Constructed Benchmark for Retrieval and Reasoning over Dynamic Knowledge

2025-11-03 · Heng Zhou, Ao Yu, Yuchen Fan, Jianing Shi 외 arxiv

Evaluating large language models (LLMs) on question answering often relies on static benchmarks that reward memorization and understate the role of retrieval, failing to capture the dynamic nature of world knowledge. We …

Question Answering