paper-with-me

홈 › Papers

Evaluating Stochasticity in Deep Research Agents

2026-02-26 · Haotian Zhai, Elias Stengel-Eskin, Pratik Patil, Liu Leqi arxiv

Deep Research Agents (DRAs) are promising agentic systems that gather and synthesize information to support research across domains such as financial decision-making, medical analysis, and scientific discovery. Despite recent improvements in research quality (e.g., outcome accuracy when ground truth is available), DRA system design often overlooks a critical barrier to real-world deployment: stochasticity. Under identical queries, repeated executions of DRAs can exhibit substantial variability in terms of research outcome, findings, and citations. In this paper, we formalize the study of stochasticity in DRAs by modeling them as information acquisition Markov Decision Processes. We introduce an evaluation framework that quantifies variance in the system and identify three sources of it: information acquisition, information compression, and inference. Through controlled experiments, we investigate how stochasticity from these modules across different decision steps influences the variance of DRA outputs. Our results show that reducing stochasticity can improve research output quality, with inference and early-stage stochasticity contributing the most to DRA output variance. Based on these findings, we propose strategies for mitigating stochasticity while maintaining output quality via structured output and ensemble-based query generation. Our experiments on DeepSearchQA show that our proposed mitigation methods reduce average stochasticity by 22% while maintaining high research quality.

📄 PDF Abstract BibTeX arXiv:2602.23271

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Revisiting the Arcade Learning Environment: Evaluation Protocols and Open Problems for General Agents

2017-09-18 · Marlos C. Machado, Marc G. Bellemare, Erik Talvitie, Joel Veness 외

The Arcade Learning Environment (ALE) is an evaluation platform that poses the challenge of building AI agents with general competency across dozens of Atari 2600 games. It supports a variety of different problem setting…

Atari Games

Using Subjective Logic to Estimate Uncertainty in Multi-Armed Bandit Problems

2020-08-17 · Fabio Massimo Zennaro, Audun Jøsang

The multi-armed bandit problem is a classical decision-making problem where an agent has to learn an optimal action balancing exploration and exploitation. Properly managing this trade-off requires a correct assessment o…

Decision MakingMulti-Armed Bandits

Evaluating Generalisation in General Video Game Playing

2020-05-22 · Martin Balla, Simon M. Lucas, Diego Perez-Liebana

The General Video Game Artificial Intelligence (GVGAI) competition has been running for several years with various tracks. This paper focuses on the challenge of the GVGAI learning track in which 3 games are selected and…

Reinforcement Learning (RL)

Let's Play Again: Variability of Deep Reinforcement Learning Agents in Atari Environments

2019-04-12 · Kaleigh Clary, Emma Tosch, John Foley, David Jensen

Reproducibility in reinforcement learning is challenging: uncontrolled stochasticity from many sources, such as the learning algorithm, the learned policy, and the environment itself have led researchers to report the pe…

Atari GamesDeep Reinforcement Learningreinforcement-learningReinforcement Learning+1

Practical Considerations for Agentic LLM Systems

2024-12-05 · Chris Sypherd, Vaishak Belle

As the strength of Large Language Models (LLMs) has grown over recent years, so too has interest in their use as the underlying models for autonomous agents. Although LLMs demonstrate emergent abilities and broad experti…