paper-with-me

홈 › Papers

SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems

2025-08-15 · Beichen Guo, Zhiyuan Wen, Yu Yang, Peng Gao, Ruosong Yang, Jiaxing Shen arxiv

The growing interest in automatic survey generation (ASG), a task that traditionally required considerable time and effort, has been spurred by recent advances in large language models (LLMs). With advancements in retrieval-augmented generation (RAG) and the rising popularity of multi-agent systems (MASs), synthesizing academic surveys using LLMs has become a viable approach, thereby elevating the need for robust evaluation methods in this domain. However, existing evaluation methods suffer from several limitations, including biased metrics, a lack of human preference, and an over-reliance on LLMs-as-judges. To address these challenges, we propose SGSimEval, a comprehensive benchmark for Survey Generation with Similarity-Enhanced Evaluation that evaluates automatic survey generation systems by integrating assessments of the outline, content, and references, and also combines LLM-based scoring with quantitative metrics to provide a multifaceted evaluation framework. In SGSimEval, we also introduce human preference metrics that emphasize both inherent quality and similarity to humans. Extensive experiments reveal that current ASG systems demonstrate human-comparable superiority in outline generation, while showing significant room for improvement in content and reference generation, and our evaluation metrics maintain strong consistency with human assessments.

📄 PDF Abstract BibTeX arXiv:2508.11310

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CharacterEval: A Chinese Benchmark for Role-Playing Conversational Agent Evaluation

2024-01-02 · Quan Tu, Shilong Fan, Zihang Tian, Rui Yan

Recently, the advent of large language models (LLMs) has revolutionized generative agents. Among them, Role-Playing Conversational Agents (RPCAs) attract considerable attention due to their ability to emotionally engage …

MultifacetEval: Multifaceted Evaluation to Probe LLMs in Mastering Medical Knowledge

2024-06-05 · Yuxuan Zhou, Xien Liu, Chen Ning, Ji Wu

Large language models (LLMs) have excelled across domains, also delivering notable performance on the medical evaluation benchmarks, such as MedQA. However, there still exists a significant gap between the reported perfo…

MedQA

Learning a Semantic Calibration Network for Open-Vocabulary Semantic Segmentation

2026-06-06 · Yang Sun, Tao Wang, Anastasia Ioannou, Ge Xu arxiv

Semantic image segmentation assigns a predefined category label to each pixel, has achieved significant progress lately. Open-Vocabulary Segmentation (OVS) extends the segmentation task from a fixed set to an open set, e…

Semantic SegmentationImage Segmentation

Encoding Hierarchical Schema via Concept Flow for Multifaceted Ideology Detection

2024-05-29 · Songtao Liu, Bang Wang, Wei Xiang, Han Xu 외

Multifaceted ideology detection (MID) aims to detect the ideological leanings of texts towards multiple facets. Previous studies on ideology detection mainly focus on one generic facet and ignore label semantics and expl…

Contrastive Learning

MobA: Multifaceted Memory-Enhanced Adaptive Planning for Efficient Mobile Task Automation

2024-10-17 · Zichen Zhu, Hao Tang, Yansi Li, Dingye Liu 외

Existing Multimodal Large Language Model (MLLM)-based agents face significant challenges in handling complex GUI (Graphical User Interface) interactions on devices. These challenges arise from the dynamic and structured …

Decision MakingLanguage ModelingLanguage ModellingLarge Language Model+1