paper-with-me

홈 › Papers

Understanding DeepResearch via Reports

2025-10-09 · Tianyu Fan, Xinyao Niu, Yuxiang Zheng, Fengji Zhang, Chengen Huang, Bei Chen, Junyang Lin, Chao Huang arxiv

DeepResearch agents represent a transformative AI paradigm, conducting expert-level research through sophisticated reasoning and multi-tool integration. However, evaluating these systems remains critically challenging due to open-ended research scenarios and existing benchmarks that focus on isolated capabilities rather than holistic performance. Unlike traditional LLM tasks, DeepResearch systems must synthesize diverse sources, generate insights, and present coherent findings, which are capabilities that resist simple verification. To address this gap, we introduce DeepResearch-ReportEval, a comprehensive framework designed to assess DeepResearch systems through their most representative outputs: research reports. Our approach systematically measures three dimensions: quality, redundancy, and factuality, using an innovative LLM-as-a-Judge methodology achieving strong expert concordance. We contribute a standardized benchmark of 100 curated queries spanning 12 real-world categories, enabling systematic capability comparison. Our evaluation of four leading commercial systems reveals distinct design philosophies and performance trade-offs, establishing foundational insights as DeepResearch evolves from information assistants toward intelligent research partners. Source code and data are available at: https://github.com/HKUDS/DeepResearch-Eval.

📄 PDF Abstract BibTeX arXiv:2510.07861

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

2026-02-03 · Changze Lv, Jie Zhou, Wentao Zhao, Jingwen Xu 외 arxiv

Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward signals. Accordingly, rubric-based evaluation has become a common practi…

Reinforcement Learning

Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework

2025-06-03 · Zhaorui Yang, Bo Pan, Han Wang, Yiyao Wang 외

Visualizations play a crucial part in effective communication of concepts and information. Recent advances in reasoning and retrieval augmented generation have enabled Large Language Models (LLMs) to perform deep researc…

Retrieval-augmented Generation

DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning

2026-06-05 · Lingyong Yan, Can Xu, Yukun Zhao, Wenxuan Li 외 arxiv

Deep Research (DR) has emerged as a new agentic paradigm to tackle complex, open-ended research tasks, demanding systems that can iteratively frame problems, acquire evidence, verify sources, and synthesize long-form rep…

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

2025-06-13 · Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang 외

Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they transform vast amounts of online informatio…

Information RetrievalRetrieval

Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs

2025-09-28 · Shreyas Singh, Kunal Singh, Pradeep Moturi arxiv

Tool-integrated reasoning has emerged as a key focus for enabling agentic applications. Among these, DeepResearch Agents have gained significant attention for their strong performance on complex, open-ended information-s…

Reinforcement LearningInformation Retrieval