paper-with-me

홈 › Papers

DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity

2026-02-12 · Joey Zhong, Hao Zhang, Clare Southern, Jeremy Yang, Thomas Wang, Kate Jung, Shu Zhang, Denis Yarats, Johnny Ho, Jerry Ma arxiv

We present DRACO (Deep Research Accuracy, Completeness, and Objectivity), a benchmark of complex deep research tasks. These tasks, which span 10 domains and draw on information sources from 40 countries, originate from anonymized real-world usage patterns within a large-scale deep research system. Tasks are sampled from a de-identified dataset of Perplexity Deep Research requests, then filtered and augmented to ensure that the tasks are anonymized, open-ended and complex, objectively evaluable, and representative of the broad scope of real-world deep research use cases. Outputs are graded against task-specific rubrics along four dimensions: factual accuracy (accuracy), breadth and depth of analysis (including completeness), presentation quality (including objectivity), and citation quality. DRACO is publicly available at https://hf.co/datasets/perplexity-ai/draco.

📄 PDF Abstract BibTeX arXiv:2602.11685

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training

2026-09-03 · Shubham Gandhi, Saurabh Goyal, Kiran Kate, Yara Rizk hf

Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind setting, where ground-truth success signals ar…

Reinforcement Learning

Dataflow-Guided Retrieval Augmentation for Repository-Level Code Completion

2024-05-30 · Wei Cheng, Yuhan Wu, Wei Hu

Recent years have witnessed the deployment of code language models (LMs) in various code intelligence tasks such as code completion. Yet, it is challenging for pre-trained LMs to generate correct completions in private r…

Code CompletionRetrievaltext similarity

Agentic DraCor and the Art of Docstring Engineering: Evaluating MCP-empowered LLM Usage of the DraCor API

2025-08-19 · Peer Trilcke, Ingo Börner, Henny Sluyter-Gäthje, Daniil Skorinkin 외 arxiv

This paper reports on the implementation and evaluation of a Model Context Protocol (MCP) server for DraCor, enabling Large Language Models (LLM) to autonomously interact with the DraCor API. We conducted experiments foc…

HydraCollab: Adaptive Collaborative-Perception for Distributed Autonomous Systems

2026-06-30 · Luke Chen, Cheng-Ju Wu, David R. Martin, Qilin Ye 외 arxiv

Collaborative-perception enables multi-robot systems to enhance situational awareness by sharing perceptual information. Existing collaborative-perception systems face an inherent trade-off between communication bandwidt…

TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models

2026-07-22 · Mark Schutera arxiv

tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a single-file, drop-in counterpart to Karpathy's tiny_shakespeare. The available Ger…