paper-with-me

Papers

Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports

2025-10-02 · Yang Yao, Yixu Wang, Yuxuan Zhang, Yi Lu, Tianle Gu, Lingyu Li, Dingyi Zhao, Keming Wu, Haozhe Wang, Ping Nie, Yan Teng, Yingchun Wang arxiv

As an embodiment of intelligence evolution toward interconnected architectures, Deep Research Agents (DRAs) systematically exhibit the capabilities in task decomposition, cross-source retrieval, multi-stage reasoning, information integration, and structured output, which markedly enhance performance on complex and open-ended tasks. However, existing benchmarks remain deficient in evaluation dimensions, response format, and scoring mechanisms, limiting their effectiveness in assessing such agents. This paper introduces Dr. Bench, a multidimensional evaluation framework tailored to DRAs and long-form report-style responses. The benchmark comprises 214 expert-curated challenging tasks across 10 broad domains, each accompanied by manually constructed reference bundles to support composite evaluation. This framework incorporates metrics for semantic quality, topical focus, and retrieval trustworthiness, enabling a comprehensive evaluation of long reports generated by DRAs. Extensive experimentation confirms the superior performance of mainstream DRAs over web-search-tool-augmented reasoning models, yet reveals considerable scope for further improvement. This study provides a robust foundation for capability assessment, architectural refinement, and paradigm advancement of DRAs.

📄 PDF Abstract BibTeX arXiv:2510.02190

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities

2025-06-10 · Wendong Bu, Yang Wu, Qifan Yu, Minghe Gao 외

As multimodal large language models (MLLMs) advance, MLLM-based virtual agents have demonstrated remarkable performance. However, existing benchmarks face significant limitations, including uncontrollable task complexity…

Search-Time Contamination in Deep Research Agents: Measuring Performance Inflation in Public Benchmark Evaluation

2026-06-03 · Yongjie Wang, Xinyue Zhang, Kunhong Yao, Zhiwei Zeng 외 arxiv

Public benchmarks enable fair and reproducible evaluation of LLM reasoning, but they become fragile for deep research agents that actively search the web during inference. Such agents may retrieve public benchmark metada…

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents

2026-06-02 · Alex Wang, Georg Meinhardt, Jacob Katz, Joseph H. Kim 외 arxiv

Financial-research answers are decision-relevant only when another analyst can audit how they were produced: which source was chosen, which period and accounting definition were used, which assumptions were made, and how…

Towards Automatic Generation of Questions from Long Answers

2020-04-10 · Shlok Kumar Mishra, Pranav Goel, Abhishek Sharma, Abhyuday Jagannatha 외

Automatic question generation (AQG) has broad applicability in domains such as tutoring systems, conversational agents, healthcare literacy, and information retrieval. Existing efforts at AQG have been limited to short a…

Information RetrievalNatural QuestionsQuestion GenerationQuestion-Generation+2

DREAM: Deep Research Evaluation with Agentic Metrics

2026-02-21 · Elad Ben Avraham, Changhao Li, Ron Dorfman, Roy Ganz 외 arxiv

Deep Research Agents generate analyst-grade reports, yet evaluating them remains challenging due to the absence of a single ground truth and the multidimensional nature of research quality. Recent benchmarks propose dist…