paper-with-me

Papers

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

2025-06-13 · Mingxuan Du, Benfeng Xu, Chiwei Zhu, Xiaorui Wang, Zhendong Mao

Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they transform vast amounts of online information into analyst-grade, citation-rich reports--compressing hours of manual desk research into minutes. However, a comprehensive benchmark for systematically evaluating the capabilities of these agents remains absent. To bridge this gap, we present DeepResearch Bench, a benchmark consisting of 100 PhD-level research tasks, each meticulously crafted by domain experts across 22 distinct fields. Evaluating DRAs is inherently complex and labor-intensive. We therefore propose two novel methodologies that achieve strong alignment with human judgment. The first is a reference-based method with adaptive criteria to assess the quality of generated research reports. The other framework is introduced to evaluate DRA's information retrieval and collection capabilities by assessing its effective citation count and overall citation accuracy. We have open-sourced DeepResearch Bench and key components of these frameworks at https://github.com/Ayanami0730/deep_research_bench to accelerate the development of practical LLM-based agents.

📄 PDF Abstract BibTeX arXiv:2506.11763

Code (1)

ayanami0730/deep_research_bench 공식 구현

Tasks

Information RetrievalRetrieval

Similar Papers 제목 키워드 기반

Understanding DeepResearch via Reports

2025-10-09 · Tianyu Fan, Xinyao Niu, Yuxiang Zheng, Fengji Zhang 외 arxiv

DeepResearch agents represent a transformative AI paradigm, conducting expert-level research through sophisticated reasoning and multi-tool integration. However, evaluating these systems remains critically challenging du…

DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent

2026-03-01 · Tongzhou Wu, Yuhao Wang, Xinyu Ma, Xiuqiang He 외 arxiv

Deep-research agents are capable of executing multi-step web exploration, targeted retrieval, and sophisticated question answering. Despite their powerful capabilities, deep-research agents face two critical bottlenecks:…

Reinforcement LearningQuestion Answering

Marco DeepResearch: Unlocking Efficient Deep Research Agents via Verification-Centric Design

2026-03-30 · Bin Zhu, Qianghuai Jia, Tian Lan, Junyang Ren 외 arxiv

Deep research agents autonomously conduct open-ended investigations, integrating complex information retrieval with multi-step reasoning across diverse sources to solve real-world problems. To sustain this capability on …

Information Retrieval

Fathom-DeepResearch: Unlocking Long Horizon Information Retrieval and Synthesis for SLMs

2025-09-28 · Shreyas Singh, Kunal Singh, Pradeep Moturi arxiv

Tool-integrated reasoning has emerged as a key focus for enabling agentic applications. Among these, DeepResearch Agents have gained significant attention for their strong performance on complex, open-ended information-s…

Reinforcement LearningInformation Retrieval

DeepResearch Arena: The First Exam of LLMs' Research Abilities via Seminar-Grounded Tasks

2025-09-01 · Haiyuan Wan, Chen Yang, Junchi Yu, Meiqi Tu 외 arxiv

Deep research agents have attracted growing attention for their potential to orchestrate multi-stage research workflows, spanning literature synthesis, methodological design, and empirical verification. Despite these str…