paper-with-me

홈 › Papers

Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

2026-06-08 · Rishabh Sabharwal, Hongru Wang, Amos Storkey, Jeff Z. Pan arxiv

Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their reports when guided by feedback? To investigate this, we conduct a multi-turn evaluation of DRAs under two feedback settings: self-reflection, in which the agent revises its report without any external diagnostic signal, and process-level feedback, in which the agent receives guidance targeting gaps in its research strategy. To enable process-level feedback, we design Research Gap Inference (RGI), a method that analyzes patterns of satisfied and unsatisfied rubric criteria to infer research-process gaps. Our analysis reveals three key findings: (i) under self-reflection, agents incorporate and regress on rubric criteria at nearly equal rates, yielding negligible net improvement; (ii) a single round of process-level feedback yields substantial gains, raising the normalized score by approximately $8$-$15$ points and yielding a roughly $35$-$40\%$ incorporation rate; (iii) these gains do not compound over subsequent turns, as agents regress on up to $24\%$ of previously satisfied criteria when rewriting the full report to address remaining gaps. Even with targeted guidance, reliable multi-turn improvement remains out of reach for the DRA architectures we evaluate. Our code and results are publicly available at https://github.com/sabharwalrishabh/Multi-Turn-Evaluation-of-DRAs.

📄 PDF Abstract BibTeX arXiv:2606.09748

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EpiBench: Benchmarking Multi-turn Research Workflows for Multimodal Agents

2026-04-07 · Xuan Dong, Huanyang Zheng, Tianhao Niu, Zhe Han 외 arxiv

Scientific research follows multi-turn, multi-step workflows that require proactively searching the literature, consulting figures and tables, and integrating evidence across papers to align experimental settings and sup…

Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision

2026-01-19 · Bingsen Chen, Boyan Li, Ping Nie, Yuyu Zhang 외 arxiv

Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task, which fundamentally diverges from how human researchers iteratively draft and revise reports via self-reflection …

Prompt Engineering

Drift-Bench: Diagnosing Cooperative Breakdowns in LLM Agents under Input Faults via Multi-Turn Interaction

2026-02-02 · Han Bao, Zheyuan Zhang, Pengcheng Jing, Zhengqing Yuan 외 arxiv

As Large Language Models transition to autonomous agents, user inputs frequently violate cooperative assumptions (e.g., implicit intent, missing parameters, false presuppositions, or ambiguous expressions), creating exec…

SDPO: Segment-Level Direct Preference Optimization for Social Agents

2025-01-03 · Aobo Kong, Wentao Ma, Shiwan Zhao, Yongbin Li 외

Social agents powered by large language models (LLMs) can simulate human social behaviors but fall short in handling complex goal-oriented social dialogues. Direct Preference Optimization (DPO) has proven effective in al…

Substance over Style: Evaluating Proactive Conversational Coaching Agents

2025-03-25 · Vidya Srinivas, Xuhai Xu, Xin Liu, Kumar Ayush 외

While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coaching presents unique challenges with initi…