paper-with-me

홈 › Papers

TRACE: Trajectory-Aware Comprehensive Evaluation for Deep Research Agents

2026-02-05 · Yanyu Chen, Jiyue Jiang, Jiahong Liu, Yifei Zhang, Xiao Guo, Irwin King arxiv

The evaluation of Deep Research Agents is a critical challenge, as conventional outcome-based metrics fail to capture the nuances of their complex reasoning. Current evaluation faces two primary challenges: 1) a reliance on singular metrics like Pass@1, creating a "high-score illusion" that ignores the quality, efficiency, and soundness of the reasoning process; and 2) the failure of static benchmarks to quantify crucial attributes like robustness and latent capability. To address these gaps, we introduce TRACE (Trajectory-Aware Comprehensive Evaluation), a framework that holistically assesses the entire problem-solving trajectory. To counter the "high-score illusion", we propose a Hierarchical Trajectory Utility Function that quantifies process efficiency and cognitive quality, including evidence grounding, alongside accuracy. To measure deeper attributes, TRACE introduces a Scaffolded Capability Assessment protocol, quantifying an agent's latent ability by determining the minimum guidance needed for success. Our contributions include the TRACE framework, its novel metrics, and the accompanying DeepResearch-Bench with controllable complexity. Experiments show TRACE delivers a granular ranking that uncovers critical trade-offs between agent accuracy, efficiency, and robustness entirely missed by singular metrics.

📄 PDF Abstract BibTeX arXiv:2602.21230

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Why Your Deep Research Agent Fails? On Hallucination Evaluation in Full Research Trajectory

2026-01-30 · Yuhao Zhan, Tianyu Fan, Linxuan Huang, Zirui Guo 외 arxiv

Diagnosing failure patterns in Deep Research Agents (DRAs) remains a critical challenge. Existing benchmarks predominantly rely on end-to-end evaluation, obscuring intermediate hallucinations that accumulate throughout t…

TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding

2026-02-23 · Fan Yang, Shurong Zheng, Hongyin Zhao, Yufei Zhan 외 arxiv

Recent Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in image understanding and natural language generation. However, current approaches focus predominantly on global image understanding, strug…

Trajectory PredictionScene UnderstandingLogical Reasoning

TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety

2026-05-30 · Zhepei Hong, Lin Wang, Liting Li, Haokai Ma 외 arxiv

Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moderation. Existing turn-level or short-context detectors struggle to re…

SciTrace: Trajectory-Aware Safety Reasoning for Scientific Discovery Agents

2026-06-06 · Tanush Swaminathan, Runmin Jiang, Letian Zhang, Min Xu arxiv

LLM-based scientific agents have shown strong capacity for autonomous research, yet their safety layers remain structurally divorced from core reasoning: they inspect pipeline outputs rather than shaping the deliberation…

Adversarial Robustness

TRACER: Trajectory Risk Aggregation for Critical Episodes in Agentic Reasoning

2026-02-11 · Sina Tayebati, Divake Kumar, Nastaran Darabi, Davide Ettori 외 arxiv

Estimating uncertainty for AI agents in real-world multi-turn tool-using interaction with humans is difficult because failures are often triggered by sparse critical episodes (e.g., looping, incoherent tool use, or user-…

Text Generation