paper-with-me

홈 › Papers

DRBench: A Realistic Benchmark for Enterprise Deep Research

2025-09-30 · Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol, Curtis Fox, Amrutha Varshini Ramesh, Étienne Marcotte, Xing Han Lù, Nicolas Chapados, Spandana Gella, Peter West, Giuseppe Carenini, Christopher Pal, Alexandre Drouin, Issam H. Laradji arxiv

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (for example, "What changes should we make to our product roadmap to ensure compliance with this standard?") that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web. Tasks are generated through a carefully designed synthesis pipeline with human-in-the-loop verification, and agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports. We release 100 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance. We demonstrate the effectiveness of DRBench by evaluating diverse DR agents across open- and closed-source models (such as GPT, Llama, and Qwen) and DR strategies, highlighting their strengths, weaknesses, and the critical path for advancing enterprise deep research. Code and data are available at https://github.com/ServiceNow/drbench.

📄 PDF Abstract BibTeX arXiv:2510.00172

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

One Interaction Is Worth a Thousand Guesses: Benchmarking the Interactive Capabilities of Deep Research Agents

2026-01-10 · Yingchaojie Feng, Qiang Huang, Xiaoya Xie, Zhaorui Yang 외 arxiv

Deep research agents powered by Large Language Models (LLMs) can perform multi-step reasoning, web exploration, and long-form report generation. However, existing systems remain largely autonomous, assuming fully specifi…

IDRBench: Understanding the Capability of Large Language Models on Interdisciplinary Research

2025-07-21 · Yuanhao Shen, Daniel Xavier de Sousa, Ricardo Marçal, Hongyu Guo 외 arxiv

Innovation is a key driving force of human civilization. As the body of knowledge has grown considerably, bridging knowledge across different disciplines, where significant innovation often emerges, has become increasing…

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?

2026-04-10 · Young-Suk Lee, Ramon Fernandez Astudillo, Radu Florian arxiv

Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities in isolation, creating a blind spot in assessing real-world performance. We intro…

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

2026-05-25 · Xinrui Shi, Kai Liu, Ziqing Zhang, Jianze Li 외 arxiv

Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objects, attributes, and relations must be jointly grounded and resolved th…

Characterizing Deep Research: A Benchmark and Formal Definition

2025-08-06 · Abhinav Java, Ashmit Khandelwal, Sukruta Midigeshi, Aaron Halfaker 외 arxiv

Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of \textit{deep research} -- a term also adopted by recent models ta…