paper-with-me

홈 › Papers

SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys

2025-12-02 · Jiahao Zhao, Shuaixing Zhang, Nan Xu, Lei Wang arxiv

LLM-based automatic survey systems are transforming how users acquire information from the web by integrating retrieval, organization, and content synthesis into end-to-end generation pipelines. While recent works focus on developing new generation pipelines, how to evaluate such complex systems remains a significant challenge. To this end, we introduce SurveyEval, a comprehensive benchmark that evaluates automatically generated surveys across three dimensions: overall quality, outline coherence, and reference accuracy. We extend the evaluation across 7 subjects and augment the LLM-as-a-Judge framework with human references to strengthen evaluation-human alignment. Evaluation results show that while general long-text or paper-writing systems tend to produce lower-quality surveys, specialized survey-generation systems are able to deliver substantially higher-quality results. We envision SurveyEval as a scalable testbed to understand and improve automatic survey systems across diverse subjects and evaluation criteria.

📄 PDF Abstract BibTeX arXiv:2512.02763

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DeepSurvey-Bench: Evaluating Academic Value of Automatically Generated Scientific Surveys

2026-01-13 · Guo-Biao Zhang, Ding-Yuan Liu, Da-Yi Wu, Tian Lan 외 arxiv

The rapid development of automated survey generation technology has made it increasingly important to establish a comprehensive benchmark to evaluate the quality of generated surveys. Most existing benchmarks first const…

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

2025-08-21 · Weihang Su, Anzhe Xie, Qingyao Ai, Jianming Long 외 arxiv

The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this process, progress in this area is hindered by …

ARISE: Agentic Rubric-Guided Iterative Survey Engine for Automated Scholarly Paper Generation

2025-11-21 · Zi Wang, Xingqiao Wang, Sangah Lee, Xiaowei Xu arxiv

The rapid expansion of scholarly literature presents significant challenges in synthesizing comprehensive, high-quality academic surveys. Recent advancements in agentic systems offer considerable promise for automating t…

Paper generation

SurveyLens: A Discipline-Aware Benchmark for Automatic Survey Generation

2026-02-11 · Beichen Guo, Zhiyuan Wen, Jia Gu, Haochen Shi 외 arxiv

Automatic Survey Generation (ASG) aims to produce comprehensive literature surveys by retrieving, organizing, and synthesizing academic papers. Despite rapid progress in specialized ASG frameworks and Deep Research agent…

SurveyBench: Can LLM(-Agents) Write Academic Surveys that Align with Reader Needs?

2025-10-03 · Zhaojun Sun, Xuzhou Zhu, Xuanhe Zhou, Xin Tong 외 arxiv

Academic survey writing, which distills vast literature into a coherent and insightful narrative, remains a labor-intensive and intellectually demanding task. While recent approaches, such as general DeepResearch agents …